Abstract
Introduction
The purpose of this study is to measure the comparative diagnostic accuracy of telehealth diagnostic examinations for pathologies of the shoulder against an in-person examination. The telehealth examinations were hypothesized to be non-inferior to in-person examinations for accuracy and to demonstrate fair to moderate agreement. This is an expanded study of a data set included in a prior publication.
Methods
Patients underwent in-person standardized clinical examination (SCE) and standardized telehealth examination (STE) during the same visit by two different providers in randomized order. Tests were analyzed for sensitivity, specificity, agreement, and diagnostic accuracy using a nonarthrographic shoulder MRI as a reference standard, and divided into tests for rotator cuff tears (RCTs), glenohumeral arthritis (GHA), and acromioclavicular (AC) joint arthropathy. A pooled diagnostic accuracy was created for SCE and STE and directly compared using a Mann–Whitney U test.
Results
Sixty-two patients, average age of 57.9 years (±11.2), with 60 patients obtaining an MRI, were included in this study. There were no significant differences in the pooled diagnostic accuracy of identifying RCT, GHA, or AC arthropathy between SCE and STE (P = .495, .469, .333, respectively). The highest agreement between SCE and STE was observed for the shoulder shrug test, night pain, and internal rotation limitation for identifying RCT.
Discussion
STE demonstrated non-inferior pooled diagnostic accuracy in comparison to SCE for full-thickness RCT, GHA, and AC joint arthropathy. Secondarily, there was moderate to substantial agreement for selective tests, with a considerable portion ranging from fair to substantial agreement.
Introduction
Shoulder pain is a common musculoskeletal symptom with a worldwide prevalence ranging from 2.5% to 25% that increases with age.1–3 With such high prevalence, research into the utility and accuracy of physical examination in the telehealth/telemedicine context is essential to determine the viability of its use in orthopedic surgery. Especially in the context of a heightened demand for telemedicine after the COVID-19 pandemic with ongoing support from the American Orthopaedic Association, identifying the accuracy of off-site physical examination for the shoulder would aid in generalizing telemedicine's applications.4–6
The use of telehealth examination for shoulder pathology in orthopaedic surgery is a relatively new concept that became more prevalent after the COVID-19 epidemic. Initially, these visits were completed in the postoperative setting after rotator cuff tear (RCT) repair, with more recent studies of diagnostic accuracy and treatment outcomes remaining promising.7–11 As previously published as part of this study, in 2020, Bradley et al. suggested there were no differences in diagnostic accuracy of the shoulder examination for those seen in telemedicine versus in-person examinations with an RCT confirmed on MRI. 12 However, these results were limited by incomplete MRI confirmation of shoulder pathology with interruption of the study by the COVID-19 epidemic, and the authors did not analyze other shoulder derangements. As such, this additional data and analysis contributes to the literature on telehealth examination for detecting RCTs, but also glenohumeral arthritis (GHA), long-head of the biceps tendon (LHBT) pathology, and acromioclavicular arthropathy (ACA).
As telehealth adoption continues, analyzing the diagnostic capabilities of specific subcategories of shoulder pathology is essential to determine the validity of telehealth examination. Therefore, the purpose of this study is to measure the comparative accuracy (pooled) of telehealth diagnostic examinations for pathologies of the shoulder against an in-person examination. The telehealth examinations were hypothesized to be non-inferior to in-person examinations with regard to accuracy in identifying shoulder pathology based upon MRI and hypothesized to have fair to moderate agreement between exams.
Methods
Study design
The study is a randomized case-control design in compliance with Standards for Reporting Diagnostic Accuracy Studies (STARD) and International Committee of Medical Journal Editors (ICMJE) reporting standards and to guide this study in the use of partial previously published data. 13 A significant portion of the data used in this study was previously collected and published prior to completion due to the COVID-19 pandemic as seen in the publication by Bradley et al. 12
Following international review board approval (IRB # Pro00101866, Duke University Health System), all clinic examinations were conducted by sports medicine fellowship-trained orthopedic surgeons. Telehealth examinations were conducted by sports medicine fellows or senior orthopaedic residents. Following informed consent, all patients underwent both in-person standardized clinical examination (SCE) and standardized telehealth examination (STE) during the same examination by two different providers in randomized order.
Eligibility and inclusion/exclusion criteria
Patients were included if they: (a) were 40 years of age or older, (b) presented with shoulder pain, and (c) provided informed consent. Patients were excluded if they: (a) had known contraindications to MRI, (b) had a history of fracture or dislocation on prior advanced imaging, or (c) had a history of shoulder arthroplasty, instability, or prior imaging revealing pathology on the shoulder in question. MRIs were provided for all patients, paid for via a research grant from the health system.
Initial examination
All participants underwent two sets of index tests via both assessment platforms (one SCE and one STE, by two different surgeons) in randomized order. SCE tests included commonly used examination maneuvers targeting diagnosis of RCTs, LHBT pathology, AC joint arthropathy, and GHA with either a high sensitivity, specificity, or both, based upon literature and practice with reasonable diagnostic accuracy.14,15
STE tests were modified SCE examinations, designed by the senior author, with approximately 15 years of experience, and a physical therapist with a PhD, who has specialized in diagnostic accuracy research for 20 years. 12 These modifications were used to create tests that reflected a clinical examination, such that each SCE test had an analogous STE test that could accurately and repeatedly be completed over telemedicine examinations, such that they were transferable to any telehealth setting with a video feed. A description of these testing procedures and modifications for STE examination can be seen in Table 1. Further testing beyond that analyzed in this study was included to provide the standard of care to patients for a shoulder examination.
Shoulder examinations as adapted by Bradley et al. (2021). 12
ER, external rotation; MMT, manual motor testing; IR, internal rotation; PROM, passive range of motion.
Standardized clinical examination
For the SCE examination, one of three fellowship-trained sports medicine orthopedic surgeons performed the standard set of shoulder examination procedures. The standard set of examinations was completed in the same order for patients undergoing both SCE and analogous STE examination. Once all examination procedures were complete, the senior author attained a detailed history and reviewed radiographs to mitigate the risk of bias during examination.
Simulated telehealth examination
Just prior to the STE portion of the study, patients were invited to view a tutorial video of the STE. A research coordinator subsequently used a portable electronic device (Apple iPad; Apple, Cupertino, CA, USA) equipped with a video camera for patient direction and imaging, and a senior orthopedic surgery resident or fellow (PGY4-6) served as the telehealth examiner performing the physician-guided, patient-performed telehealth examination. This was completed from a separate room at a desktop interface (Cisco Webex DX80 or Cisco Webex DX70; Cisco, San Jose, CA, USA), resembling a common environment for telehealth examinations. Both portable electronic device and desktop interface had audiovisual capability such that the patient and examiner could see, hear, and observe each other during the STE portion of the examination.
Senior orthopedic surgery residents and fellows were utilized for the STE in order to limit bias that would otherwise be introduced if the sports fellowship-trained orthopedic surgeon completed both shoulder examinations (STE and SCE). Nine senior residents and fellows participated in the telehealth version of the examination. To limit differences between testers, a script was provided.
Data management
All study data besides MRI results were collected and stored using REDCap electronic data capture tools (Research Electronic Data Capture; Vanderbilt University, Nashville, TN, USA) hosted at Duke University.16,17
MRI
The use of MRI as a reference standard in identifying shoulder pathology in this study is the same as previously published data from Bradley et al. 12 Shoulder MRI examinations were completed for all patients in nonarthrographic fashion. This was done via a 3.0-Tesla MR scanner (Trio TIM; Siemens Healthcare, Erlangen, Germany) using a phased array 8-channel shoulder coil (Invivo), as previously reported.12,14 Axial, oblique sagittal, and oblique coronal fat-suppressed fast spin-echo T2-weighted sequences (slice thickness, 3.0 mm; FOV, 16 cm; TR/TE, 3000/65); axial fat-suppressed fast spin-echo intermediate-weighted sequence (slice thickness, 3.0 mm; FOV, 16 cm; TR/TE, 3000/23); and oblique sagittal T1-weighted sequence (slice thickness, 3.0 mm; FOV, 16 cm; TR/TE, 688/11) were the study parameters. 12 Blinded to the clinical findings, one musculoskeletal radiologist from the institution that completed the study prospectively reviewed the research MRIs. The presence or absence of partial and complete RCTs, as well as tear locations, was recorded. Similarly, the presence or absence of GHA, LHBT pathology, and AC joint arthropathy was noted and characterized as mild, moderate, or severe. These MRI results were considered the reference standard in the diagnosis of shoulder pathology to determine the accuracy of the STE and SCE examinations and serve as the reference for comparison between groups, as mentioned by Bradley et al. 12
Power analysis
To determine the sample size in a noninferiority trial (STE is noninferior to SCE), a noninferiority margin was used to calculate the confidence window around the difference between the treatments, and the acceptability of the difference was subsequently determined. 18 With intention to treat analysis, a 20% difference in pooled diagnostic effectiveness between groups was assumed unacceptable. With the target of 95% power, the projected use of a Mann–Whitney U test for comparison of differences, and an error probability of .05, a projected sample size of 60 was identified to detect differences in accuracy across examination groups for RCTs, instead of diagnostic accuracy of the examinations in general.
Data analysis and statistical considerations
All analysis was completed via SPSS (V26.0; IBM, Armonk, NY, USA) and a publicly available online software calculator from the University of Illinois, Chicago (http://araw.mede.uic.edu/cgi-bin/testcalc.pl). First the values were summarized in tabulated values for age, gender, and diagnoses based on the radiology read.
Agreement between tests using Cohen's Kappa was completed to calculate the chance-corrected agreement between two or more examiners, ranging from 0 (perfect lack of agreement) to +1.0 (perfect agreement). Landis and Koch provided cutoff values for interpretation as follows: <0, no agreement; 0–0.20, slight; 0.21–0.40, fair; 0.41–0.60, moderate; 0.61–0.80, substantial; and 0.81–1, almost perfect agreement. 19
The diagnostic accuracy measures of sensitivity (SN), specificity (SP), positive likelihood ratio (LR+), and negative likelihood ratio (LR−) for each maneuver of the STE and SCE were completed for each individual test in the study. In this study, LR+ of >10.0 and LR− of <0.10 were deemed to have large increases or decreases in the likelihood of the pathology, with LR+ of >5 to 10 and LR− of <0.20 to 0.10 having moderate increases and decreases in the likelihood of the disease. 20 Location of the RCT was not factored into the accuracy of the exam maneuvers targeting the rotator cuff.
Overall diagnostic accuracy (also known as diagnostic accuracy or diagnostic effectiveness) was calculated for both SCE and STE tests and summated to a grand mean (hereby known as pooled diagnostic accuracy. The formula ((True positives (TP) + True negatives (TN)/TP + TN + False Positives (FP) + False Negatives (FN)) was used to assess each test within the pathological spectrum. Values range from 0% to 100% (completely inaccurate to completely accurate). The pooled diagnostic accuracy was compared between SCE and STE using a Mann–Whitney U nonparametric test. Statistical significance was defined a priori as <.05.
Results
From August 2019 to March 2020, 96 consecutive patients met the inclusion and exclusion criteria. Of those, 62 patients [average age of 57.9 years (±11.2), 31 (51.7%) female] were ultimately enrolled, with 60 patients obtaining an MRI (two patients were either unable to due to tolerance or had an unexpected contraindication for imaging).
MRI results
In this cohort, 72% of patients had an MRI-confirmed rotator cuff tear, with 67% of cuff tears involving the supraspinatus, 30% with the infraspinatus, and 23% with the subscapularis. Complete tears were identified in 12% of subjects (7 out of 60) for both the supraspinatus and infraspinatus tendons, while no full-thickness tears were found in the subscapularis tendon (Table 2). Evidence of bursitis was present in only 37% of patients, whereas biceps pathology was revealed on MRI for 65% of patients. Glenohumeral arthritis was only identified in 17% of patients (glenoid or humeral cartilage irregularity), with AC joint arthritis more common at 40% of patients (Table 2).
Prevalence of shoulder pathology identified in this cohort
Diagnostic accuracy for RCT
There was no significant difference in the pooled diagnostic accuracy of identifying a complete rotator cuff tear between SCE and STE examinations (P = .495) (Table 3). The most sensitive exam maneuvers/signs for the SCE examination included the painful arc test and night pain [79.1% (72.3, 87.1), 72.1% (66.1, 80.6), respectively]. Similarly, for the STE examination, painful arc test, Neer's sign, and Night pain showed the highest sensitivity [79.1% (72.3, 87.1), 74.4% (67.0, 82.8), 74.4% (67.6, 82.8), respectively]. External rotation (ER) lag sign, shoulder shrug, drop arm test, Belly-press test, and lift-off sign showed the greatest specificity in both groups, with drop-arm test and ER lag sign approximating 100% specificity [SCE: 100% (83.2, 100), 100% (94.0, 100); STE: 100% (0.88, 100), 94.1% (81.8, 99.7), respectively, for both]. Internal rotation (IR) pain with strength testing specificity was lower for SCE than for STE [31.3% (13.1, 55.1) vs 76.5% (55.0, 91.7), respectively], whereas ER weakness with strength testing specificity was higher for SCE than for STE [94.1% (73.6, 99.7) vs 58.8% (36.5, 79.0), respectively] (Table 3). Exam maneuvers in both the SCE and STE examinations had higher specificity than sensitivity in this cohort.
Diagnostic accuracy of clinical and telehealth testing for full thickness rotator cuff tears
Diagnostic accuracy for other shoulder pathology
There was no significant difference in the pooled diagnostic accuracy between SCE and STE diagnostic accuracy for GHA or AC joint arthropathy (P = .469, .333, respectively) (Table 4). Neither group had an exam with a sensitivity higher than 70% for GHA, with active internal rotation < T12 with 70.0% (37.3, 91.7) sensitivity for SCE [40.0% (14.4, 70.3) for the STE] and crepitus for the STE examination with 70.0% (37.4, 91.7) sensitivity [40.0% (14.4, 70.1) for SCE]. In both the SCE and STE examinations, passive forward flexion <120° showed the highest specificity at 92.7% (87.7, 97.0) and 90.9% (85.8, 95.5), respectively. As for AC joint arthropathy, both examinations had low sensitivities, with AC joint tenderness as the most specific examination maneuver for the SCE [85.4% (76.7, 93.1)] and STE [73.2% (65.3, 82.6)] examination (Table 4). None of the exam maneuvers in the STE or SCE were highly sensitive or specific for disorders of the biceps.
Diagnostic accuracy of clinical and telehealth testing for glenohumeral arthritis, acromioclavicular arthropathy, and biceps disorders
Agreement between SCE and STE
The highest agreement between SCE and STE were observed for the shoulder shrug test, night pain, and IR limitation for identifying RTC [k = 0.57 (P < .01), k = 0.88 (P < .01), k = 0.52 (P < .01), respectively], whereas the lowest agreement was observed in the active to passive flexion limitation, ER weakness with strength testing, lift off sign, and Hawkins–Kennedy tests [k = −0.10 (P = .33), k = −0.07 (P = .56), k = 0.06 (P = .65), k = 0.04 (P = .66), respectively] (Table 5). As for GHA, only the passive forward flexion <120° showed a substantial agreement between exams [k = 0.69 (P < .01)]. AC joint tenderness was the most agreeable exam finding for AC arthropathy between SCE and STE at a kappa value of 0.62 (P < .01). There was only fair agreement with Speeds Test for biceps disorder with a kappa value of 0.49 (P < .01) (Table 5).
Overall accuracy (0–100%) comparisons between clinical and telehealth testing for all conditions
Discussion
The most important finding from this study is the non-inferior pooled diagnostic accuracy for STE in comparison to SCE for full-thickness RCT, GHA, and AC joint arthropathy. Overall, the diagnostic accuracy of individual exams in both the SCE and STE examinations was fairly low, often having either a high sensitivity, specificity, or neither. Secondarily, there was moderate to substantial agreement for selective tests between telehealth and clinical examination, with a considerable portion of examinations for RCT, LHBT pathology, AC joint arthropathy, and GHA ranging from fair to substantial agreement.
Accuracy for SCE and STE
The pooled diagnostic accuracy of diagnosis for RCTs, GHA, and AC joint arthropathy for STE was non-inferior to SCE examination based upon pooled results for each shoulder pathology, except biceps pathology, given there was only one exam testing for its presence. Most individual tests performed in either the SCE or STE were unable to predict the presence of full-thickness RCT, GHA, AC joint arthropathy, or biceps pathology. Although exams with some level of certainty, be it specificity or sensitivity, overlapped between the SCE and STE groups, they were not always the same. Neither setting had a single test that resulted in a high specificity and sensitivity for any shoulder pathology. The lack of diagnostic strength is unsurprising, given extensive prior studies on the effectiveness of these exams in isolation.14,21–23 As such, the overall diagnostic ability of telehealth examination requires further research, where pooled selected exams rather than exams in isolation are analyzed for their collective clinical accuracy and agreement to in-person exams. After all, the traditional in-person examinations are rarely based upon a single exam sign of finding, but rather the gestalt and summation of the patient's presentation across multiple aspects of the exam. Despite this, this study addresses the primary concern in orthopedic telehealth examinations; telehealth exams are, at minimum, non-inferior to their in-person counterpart in diagnostic accuracy with seemingly comparable agreement.
Agreement between SCE and STE
The agreement of SCE and STE was, as expected, highest amongst the more easily performed tests. This includes exams or questions such as night pain for RCT, passive forward flexion limitation for GHA, or AC joint tenderness for AC joint arthropathy, all having at least fair, or better, agreement. As such, the simplest of exams and symptoms likely have the most utility in telehealth examinations. Similarly, more complicated tests, especially those that heavily rely on more in-depth directions or descriptions of an exam finding, had appreciably lower agreement between SCE and STE, even ranging into disagreeable (negative) kappa values. Primary examples of lack of agreement for RTCs include ER weakness with strength testing, active to passive flexion limitation, and sensing crepitus in GHA for patients during STE examinations. These exams may be more suited for in-person examination and may need to be replaced by a simpler exam that provides the same information for the physician.
Clinical versus personal aspects of telemedicine
In this diagnostic accuracy study, the testing environments were carefully controlled (for internal validity), findings were reported in accordance with the STARD guidelines, and the results were anchored to an MRI (a recognized reference standard). 24 Nonetheless, we would be remorse if we did not discuss the additional elements of a well-organized telehealth examination, which, although critical, were not represented in our study. Optimized telehealth examinations require communication and engagement skills to truly understand the whole patient and how the condition has influenced their lifestyle and/or occupation. 25 Because of high heterogeneity and patient and practitioner experience on web-based platforms, intention towards relationship building, a focus on conversational flow, and an understanding of examination results in context of the full picture of morbidity of the patient are essential for each telehealth encounter. 26 Our study provides accuracy data for clinical tests associated with the shoulder, and while these provide validity for their use, it is important to recognize that these do not supersede the whole findings of the examination and encounter-they only contribute to it.
Limitations
There are several identifiable limitations to this study. First, the power analysis was based on the accuracy of detection of RTCs and thus may not be fully powered to assess pooled diagnostic accuracy calculations for other shoulder pathologies. Second, despite higher-level trainees (PGY 4-6) conducting the STE examination in a standardized fashion (use of a script by the trainee and an educational video for the patient), direction during the telehealth may have been communicated differently via Sports Medicine fellowship-trained orthopedic surgeons. Third, full-thickness tears on MRI were used as the gold standard in the determination of accuracy; however, patients' partial-thickness tears may also have abnormal exams. As such, this study may underrepresent RCT pathology seen on exam, due to unmet diagnostic criteria on gold standard MRI. Lastly, overall diagnostic accuracy was a necessary measure used to compare SCE and STE performances, but it has limitations. In scenarios where the prevalence of a condition is very low or very high, diagnostic accuracy can be misleading. Further, tests are often sensitive or specific, but are rarely both. Lower values in either of these areas can reduce the overall accuracy, despite the test's potential ability to rule in or rule out a condition when used in a comprehensive examination platform.
Conclusion
Based upon the findings in this study, STE is non-inferior to SCE accuracy, with fair to substantial agreement for most exam maneuvers at diagnosing RCT, GHA, and AC joint arthropathy. Sensitivity and Specificity were highly variable without a single test that was high in both for STE and SCE examinations. Telehealth examinations for shoulder pathology are adequate for the diagnosis of multiple shoulder pathologies, and future research is warranted on which test or pool of tests is the best predictor in diagnosis.
Footnotes
Acknowledgments
We appreciate the help of Anne Boyd, Theresa Curington, and Elizabeth Pennington, Daniel Le, Duke Institute for Health Innovation, Shilpa Shelton, MHA, Duke Telehealth, Donna Phinney, Evalyn Garrido, Javon Beatty, the residents and fellows involved in administering the telehealth examinations, Dr Nicholas Bonazza, Jonathan Cheah, D. Landry Jarvis, Danica Vance, Peter Casey, Jonathan Peterson, Sean Peterson, Sean Ryan, and John Steele for all of their aid in this study.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Ethical approval and informed consent
International Review Board number Pro00101866, Duke University Health Systems.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Duke Institute for Health Innovation (grant number 2019040104).
Data availability request
Data is available for review upon reasonable request.
Data availability statement
Example text of a Data statement, as provided by the author.
