pluses and minuses of artificial intelligence in medicine

 As we all know increasingly well, artificial intelligence is gaining prominence in all aspects of our lives, and certainly so in the field of medicine. there are clear benefits from this, but also increasingly clear downsides.


does AI provide accurate information?
-- there are persistent concerns about the quality and accuracy of the information
-- in a very early AI model, i personally found that AI had very well referenced journal articles to back up their recommendations about medical questions. the problem was that the citations were often to very specific articles that actually did not exist. the prevailing explanation was that it was important for people to get an answer, whether right or wrong, in order to accept AI's value
-- one of the founders of Open AI was interviewed on NPR recently. He became disenchanted with AI and the direction it was going, and he took it on himself to assess the accuracy of AI responses continually over time
    -- his findings have continued to show that AI was only about 60% accurate, but 40% of the text involved either some inaccuracies cited in the AI searches and some suggestions were blatantly incorrect
-- a study assessed the accuracy of large language models (LLMs, ie AI) for patient use, such as OpenAI's ChatGPT: artificial intelligence poor accuracy NatMed2025 in dropbox or https://doi.org/10.1038/s41591-025-04074-y
    -- the researchers noted prior studies on clinicians: "although LLMs now achieve strong performances on medical tasks, attempting to support doctors with LLMs in real clinical settings have faced difficulties. On the one hand, LLM scores on medical knowledge benchmarks are now commensurate with passing the US Medical Licensing Exam. LLM-generated clinical documents are rated as equivalent to or better than those written by doctors. On the other hand, excelling at medical tasks in silico does not translate to accurate performance in clinical settings under physician guidance." "In silico" is the computer chip equivalent to the expression "in vitro" in medicine. see the article for the specific references to these comments. the next section below expands on this point.
    -- this patient study enlisted 1298 patients who lack medical expertise to see if AI was able to provide appropriate patient answers outside of the clinical setting. the patients were presented with 10 medical scenarios, with directions for them to pretend that they themselves had encountered this situation at home. Based on the LLM information, they decided whether and how to access professional medical treatment, choosing the best disposition on a five-point scale; their answers were compared to what the 3 physicians who drafted the question concluded. the patients were assigned to be either in a treatment group, accessing AI as much as they felt they needed, or a control group that of people who could use any non-AI assistance they would typically use at home, including internet searches. those in the treatment group were assigned to one of 3 LLMs, either GPT-4o, Llama 3, or Command R+, with their results compared to the control group.
        -- participants using LLMs were significantly less likely than those in the control group to correctly identify at least one medical condition relevant to their scenario, with those in the control group having 1.76 times higher odds of identifying a relevant condition than the aggregate of the participants using the three LLMs. this result was despite the correct suggestions appearing in the LLM-produced list of clinical suggestions. further assessment found that this was mostly because the users failed to supply sufficient information to the LLM; and when the users added extra information later, the additional information was in some cases not central to the case scenario and led to incorrect responses by the LLM
    -- so, the conclusion of this study was that patients relying on AI to make important decisions about what to do with their symptoms, from nothing much to going to the hospital, had a pretty high risk of making the wrong decision, more often so than in the pre-AI era internet control group and potentially with really bad consequences. a real problem as AI is more integrated into our lives
is AI useful in providing insight into the content of the clinician-patient interaction?
-- why does the last section comment that AI (ie LLMs) do well for answering exam questions but not in clinical scenarios. Why may this be true, based on my own clinical, non-AI experiences?
    -- a clinician-patient interaction involves both the actual description the patient has of their problems but also clinician observations of non-verbal patient responses
    -- this information provided may well not fit into a nice "bucket" that the LLMs (AI) requires
        -- for example, a patient may use a term that AI would interpret very differently from a clinician. Perhaps they use the words "abdominal pain" to describe a concern about their abdomen. And a clinician who either knows the patient and family well or deciphers non-verbal cues from the patient may well interpret "abdominal pain" as not really being "pain" but reflect the fact that the patient's parent had abdominal cancer and the patient is really concerned about that. or the word for "pain" in their native language varies some from the English work "pain". or their grandparent, the arbiter of family medical decisions, reconfigures the problem as "pain". the AI system might well search the medical literature for "abdominal pain" and provide an array of possible etiologies that are actually irrelevant to the patient's condition, and this perhaps lead the clinician astray in assessing and treating the patient
        -- this inaccurate information, if acquired directly by the patient as in the above LLM study, could lead to tremendous anxiety in the patient and perhaps family
        -- and, this whole situation is even more difficult in patients with poor medical literacy, poor education, and in those where English is not their primary language, making the use of searchable phrases (eg "abdominal pain") even less likely to be accurate, as noted in the patient study above where not using AI produced more accurate results

what are the downsides of AI-generated information?
-- I have been using Open Evidence many times in the last couple of months to answer clinical questions and for my blogs, with several observations:
    -- i have gotten quite extensive answers to my queries
    -- there is preferential citations of recommendations from published guidelines and reviews but also key articles
    -- the comments are associated with links to the relevant articles cited
    -- BUT:
        -- on checking those links, it is quite common that the citations do NOT really support, or do support somewhat erroneously, the AI summations. i suspect this is related to the detailed AI search done that focuses on key words or combinations of words but does not really reflect the actual context of the question. this is basically the same issue with AI interpreting the clinical questions in the clinician-patient encounter as per above, where the "abdominal pain" patient complaint is misconstrued by AI by reducing the interpretation to a key word or combination of words that does not really reflect the actual content of the patient's problem
-- the concern is that checking the citations may not be done often. for example, there was a recent study finding that 60% of Google searches are not followed by checking the links (https://www.instagram.com/reel/DZw9fPwCFei/). it seems that we clinicians do need to check important links when the AI-generated answers would affect how we diagnose or treat a patient
-- one big issue this brings up is that the medical literature as cited by AI is itself incomplete, limiting our ability as clinicians to figure out the best course for the patient even without AI
    -- studies may not be well designed and actually have sufficient flaws that undercut the accuracy of their conclusions (as i find doing the blogs and listing the often very many limitations of the study being evaluated), and AI does not critique the studies to make sure the conclusions are appropriate and reasonable and does not seem to elucidate study limitations
    -- even randomized controlled trials are often tangential to the results we need to treat a patient. many studies exclude patients with renal failure, or other comorbidities, or those perhaps over 70 years old. i personally see many patients much older than 70yo, and a significant majority of my patients have some degree of renal failure. do the results of the studies really apply to the patient sitting in front of me?? I make my best guess as to whether to use the results for my patient who would not have qualified to be in the study. and these studies are the basis for the AI recommendations
    -- and society guidelines are also fraught: a study of cardiology guidelines in 2008 as compared to 2018 in the US and in 2003 vs 2018 in Europe found that about three quarters of the recommendations were by "expert opinion" and there was no significant change in that number from the earlier to the later time assessments: https://gmodestmedblogs.blogspot.com/2019/04/guidelines-lacking-evidence-based.html.
        -- the guideline committee is typically of physicians who are leaders in their fields, many being very prominent but having little hands-on patient care, and many predominantly do research, both of these issues could raise a significant bias to the "expert opinions"
        -- In my limited experience many years ago, I was as a member of the National Cancer Institute's Initial Technical Evaluation Group's request for proposals (RFP) for the 16-year, multi-center contract to screen 150,000 people for prostate, lung, colorectal, and ovary cancer (the PLCO study). there was a very clear dominance relationship between the members, with deference to the top researchers by the other committee members (which is true for many group dynamics). Not so surprising since these were the leaders in the field. BUT, some of these leaders were important researchers and fund-getters, and may well have been swayed by their own research interests (and the drug company money involved in that) and often had not been in actual clinical practice for a long time
-- i should add, as an aside, that clinical articles in our most highly regarded journals (New England Journal of Medicine, etc) do occasionally have citations that actually have nothing to do with the points being made by the authors (ie, when i read some comments in an article that were unknown to me or were suspect but was pivotal to the authors' argument, it was useful to quickly review the citation to see if it was accurate). it is surprising that the article reviewers in these highly regarded journals do not seem to check the cited important references

where is AI really useful?
-- my guess is that AI is very useful for scanning increasingly large databases that are less involved in the detailed and complex clinician-patient interaction:
-- for example, though not all studies find that AI has a superior interpretation of chest xrays, a study retrospectively assessing contrast-enhanced abdominal CTs found that very small pancreatic cancers (small enough to be potentially treatable, prior to metastases) were picked up in abdominal CT scans and missed by radiologists (https://gmodestmedblogs.blogspot.com/2026/02/ai-detecting-early-pancreatic-cancer.html). this is perhaps related to the cancer nidus being too small to be identified by these trained radiologists? or that the radiologist was focusing on another area of the CT that the referring clinician was concerned about? it should be noted that many comparative radiologist studies do find quite different interpretations of xrays (even, for example, large variations of experienced mammography radiographers in the reading of mammograms, as found in several studies, including https://pmc.ncbi.nlm.nih.gov/articles/PMC2661777/ )...
-- dermatology is another great AI target, given the ability of AI to rapidly review huge numbers of dermatologic photos and diagnoses to match to those of an individual patient. And there are several studies supporting this (https://opendermatologyjournal.com/VOLUME/17/ELOCATOR/e187437222304140/FULLTEXT/, or one on skin cancer diagnoses: https://med.stanford.edu/news/all-news/2024/04/ai-skin-diagnosis.html)

so, AI does not seem to be ready for many of the tasks it is being put to, as noted above
-- OpenEvidence is a really remarkable tool for us clinicians, but i would argue that compelling suggestions be checked by a quick review of the links
-- of course, there will be huge improvements in AI for medicine and everything else over time as the algorithms develop and more information is available for the AI recommendations
-- but AI itself also raises several huge social issues: see https://news.mit.edu/2025/explained-generative-ai-environmental-impact-0117)
    -- its effects on the environment: it needs lots of energy requiring consumption of lots of fossil fuels to make the electricity, and the effects of this massive energy need are increasingly straining our old and somewhat fragile electrical grid. this could potentially lead to problems with the population's needs for electricity as the rush to AI development further tilts this balance
    -- AI development requires lots of water in the process of developing the energy, and huge swaths of the US are already struggling for good water, in part related to the massive prior development of fracking
    -- and a huge increase in income inequality in the US, with the extraordinarily rich getting extraordinarily hugely richer, with not much benefit to those less wealthy to begin with

geoff

-----------------------------------

If you would like to be on the regular email list for upcoming blogs, please contact me at gmodest@bidmc.harvard.edu

to get access to all of the blogs:  go to http://gmodestmedblogs.blogspot.com/ to see the blogs in reverse chronological order


or you can just click on the magnifying glass on top right, then type in a name in the search box and get all the blogs with that name in them


Comments

Popular posts from this blog

air pollution and heart disease

resistant hypertension: are diuretics harmful?

UPDATE: ASCVD risk factor critique