generalists: 2 heads better than one, and better than specialists

Fw: generalists: 2 heads better than one, and better than specialists
Geoff A. Modest, M.D.
Mon 3/4/2019 7:06 AM
Geoff A. Modest, M.D.
An interesting article just came out finding increased diagnostic accuracy of multiple physician involvement (collective intelligence) vs an individual physician, and generalist groups outperformed individual specialists in that field (see primary care multiple MDs better jamaopen2019 in dropbox, or doi:10.1001/jamanetworkopen.2019.0096 ).

Details:
-- data were from the Human Diagnosis Project (Human Dx) platform, one used by physicians and medical students to practice authoring and diagnosing teaching cases. These cases are from their own clinical practice and identify the intended diagnosis or the differential diagnosis. Respondents generate a ranked differential diagnosis and are notified if they are  correct or not
    -- as of 2019, more than 14,000 users of all specialties from more than 80 countries have made more than 230,000 contributions authoring and diagnosing cases
-- in this study 2069 users solving 1572 cases, from 2014 to 2016
    -- 1228 (59%) were residents or fellows
    -- 431 (21%) were attending physicians
    -- 410 (20%) were medical students
    -- 70% internal medicine (67% general, 4% subspecialty)
    -- 3% surgery
    -- 8% other specialties
    -- location: 91% US, 9% other
-- accuracy in general involved whether any of the top 3 diagnoses from the respondent matched the case author's intended diagnosis or differential diagnosis. The researchers also performed a weighted analysis, giving different weights to the 2nd and later diagnoses in the differential

Results:
-- collective intelligence was associated with increasing diagnostic accuracy:
    -- individual physicians overall: 62.5% (60.1%-64.9%)
        -- individual residents or fellows: 65.5% (63.1%-67.8%)
        -- medical students: 55.8% (53.4%-58.3%), p <0.001 for difference with residents/fellows
        -- attending physicians: 63.9% (61.6%-66.3%), not significantly different from residents/fellows
    -- groups of 9 physicians: 85.6% (83.9%-87.4%)
        -- improvement of 23.0% (14.9%-31.2%), p<0.001
    -- review of their graphs show a large step-up going from an individual to 2 physicians, with a gradual rise from 2 physicians to 9 physicians. This was true when including or excluding medical students or attending physicians
-- assessment by one of 4 specific symptoms, comparing individuals to groups of 9 (specific accuracy numbers not given for 2 of these symptoms):
    -- abdominal pain: about 74% to 92%, improvement of 17.3% (6.4%-28.2%), p=0.002
    -- fever: about 55% to 85%, improvement of 29.8% (3.7%-55.8%), p=0.02
    -- chest pain: about 65% to 88%, improvement of about 13%
    -- shortness of breath: about 58% to 88%, improvement of about 30%
-- comparing groups from 2 to 9 users, to individual specialists in their subspecialty:
    -- groups of 2 users: 77.7% accuracy (70.1%-84.6%)
    -- groups of 9 users: 85.5% accuracy (75.1%-95.9%)
    -- individuals specialists in their own subspecialty: 66.3% accuracy (59.1%-73.5%), p<0.001 vs groups of 2 and 9.
-- Different weighting schemes for diagnoses (e.g. giving differing  emphasis to the 1st diagnosis over others) did not seem to matter much, nor did the number of user diagnoses to construct the collective differential

Commentary:
-- groups consistently outperformed individuals, with increasing diagnostic accuracy with increasing group size (though going from 1 to 2 was the largest increase), was largely independent of the status/experience of the clinicians, and was better than that of specialists as measured by 4 specific diagnoses
-- other studies have highlighted clinical misdiagnosis, which can certainly lead to potential morbidity and mortality, as well as increased testing/radiation exposure/cost/psychological morbidity/medicalization/etc.
    -- most other studies have focused on binary decisions in well-defined tasks, such as reading mammograms or assessing skin lesions, and not the complexity of broader symptoms
--it was quite interesting that the mean individual specialist accuracy was not statistically significant from an individual nonspecialist, and that 2 nonspecialists were statistically more likely to be accurate then an individual specialist. There was a trend towards some improvement going from 2 nonspecialists to 9, though this increase was not statistically significant
-- as a personal anecdote, 20 years ago I started a one-hour case conference session every week at our health center (with CME credit arranged) for our primary care providers. This is an open-ended discussion of any cases that providers want to discuss. It is pretty remarkable that this has continued for 20 years. In addition, my personal experience is that for essentially any case that I personally present, there are useful and important suggestions regarding both the diagnosis and management.
-- limitations of the study include:
    -- the Human Dx users may not be representative of the broader medical community, so the results may not be directly translatable to other clinical situations
    -- Human Dx was not designed to assess collective intelligence vs individual reasoning, so may not be the best tool
    -- the reference standard used in this database seems to be what the referring provider indicated, though it is not entirely clear that this is always accurate
    -- there is also a potentially big difference between seeing a patient with a problem vs reading a brief distillation of the problem on the internet (the actual patient encounter involves much more nuance and more non-verbal cues, and these probably do influence the actual differential in clinical practice); this study does compare apples to apples (all using the computer), but the reality may be that the clinical impression of a single clinician who actually interviews/examines a patient may be very different (likely more accurate) than that of perhaps several people reading the condensed story devoid of these nuances, a story filtered by the authoring clinician vs the raw material of a real patient encounter
    -- and the cases may not represent general practice: i suspect that there are more unusual diagnoses submitted by clinicians than those in run-of-the-mill clinical practice
-- Also, as noted by the researchers, not all differential diagnoses are the same. In many cases, especially for  patients with potentially very serious and urgent diagnoses, a much more limited differential may be essential/appropriate. In the case of non-critical symptoms, a  longer list of potential diagnoses is important in the workup of the patient.
-- It would be useful to know the diagnostic accuracy of the users’ differential diagnoses stratified by the level of certainty that the individual clinician has for their 1st diagnosis, perhaps with some visual analog scale. For example, if the diagnostic accuracy were shown to be much lower in those where there was more uncertainty by an individual clinician (as would be anticipated), it would make sense to focus on these diagnoses for the results of a more collective approach.

So, an interesting study for a few reasons. It does reinforce the importance and utility of broader discussions amongst several clinicians, probably especially so when there is more diagnostic uncertainty. And it would then make sense to have a defined system/structure within the clinical settings to allow for more collective decision-making (which could be face-to-face or virtual). And, the study also showed that specialists are also not so accurate as individuals, which I think is really important for primary care clinicians to recognize. My sense in several different clinical situations is that primary care clinicians often accept specialist diagnoses and recommendations explicitly/unquestioningly, and this study does reinforce that we in primary care should review all recommendations critically and determine their appropriateness for the individual patient.

geoff
If you would like to be on the regular email list for upcoming blogs, please contact me at gmodest@uphams.org

to get access to blogs since 8/15/17:1. go to http://gmodestmedblogs.blogspot.com/ to see them in reverse chronological order
2. click on 3 parallel lines top left, if you want to see blogs by category, then click on "labels" and choose a category​
3. or you can just type in a name in the search box and get all the blogs with that name in them

to access older blogs from the BMJ website, from October 2013 until 8/15/17: go to http://blogs.bmj.com/bmjebmspotlight/category/archive/ 

please feel free to circulate this to others. also, if you send me their emails, i can add them to the list​


Comments

Popular posts from this blog

air pollution and heart disease

resistant hypertension: are diuretics harmful?

Body Roundness Index is better predictor than BMI for clinical problems