As new AI tools emerge, we must learn to critically evaluate their output while understanding that how we ask a question often determines the accuracy of the answer.
Most physicians are already using artificial intelligence (AI) large language model (LLM) tools to conduct clinical searches and rapidly summarize information, but many of us are struggling to evaluate the best way to use these new tools. Previous articles in this journal have discussed common AI applications and products.1 This article will focus on the following:
- We need to be cautious users of AI tools and learn to evaluate their outputs, lest we become overly confident in this technology or overly dependent on it for clinical decision making and differential development, especially in early career and training.
- We need to learn to refine our inputs (or prompts) to make AI clinical searches more effective and enhance decision making at the point of care, understanding that how we ask a clinical question often determines the accuracy of the answer.
(If you encounter unfamiliar terms, see Table 1 for a glossary.)
KEY POINTS
- Large language model (LLM) artificial intelligence (AI) programs can help with point-of-care clinical decisions but should be validated against trusted sources, including a physician’s own clinical judgment and expertise.
- Better AI prompting leads to better answers. Clinical prompts should be as specific as possible and provide relevant context about the patient and clinician setting.
- It can be useful to instruct LLMs to explain their reasoning, cite specific guidelines, or surface contrary evidence.
WHY WE NEED TO BE CAUTIOUS USERS OF AI
Our patients are using AI tools for their own health-related queries. In clinical practice, this leads to the increased need to rapidly surface in-depth information — from clinical guidelines to treatment options — and apply it at the point of care. AI-powered search is helpful in this regard. It also challenges our existing thought processes and expands our inputs, making us better clinicians.
Despite these advantages, a critical examination of AI/LLM tools is important because they are not always producing the same outputs. Asking each one the same clinical question makes that clear. Variability can also exist within the same tool, with different answers to the same question. As these tools evolve, these differences will likely become more widely understood. Still, physicians must guard against the tendency toward automation bias, meaning a general acceptance of algorithmic outputs and failure to cross-check them, especially as trust increases over time. To address this in my practice, I generally triangulate answers by using three sources: two AI models — a medical-specific one such as OpenEvidence (which is trained on indexed journals) and a general AI tool with academic search such as Perplexity or ChatGPT5 — and then comparing their output to an established conventional resource, such as UpToDate, PubMed, Cochrane, Essential Evidence Plus, or ePocrates.
While physicians should be using AI tools, we should also understand that integration of these models is occurring much faster than our ability to validate their output. Minimum standards for transparency and accuracy have yet to be developed. Therefore, overreliance may increase the risk of accepting errors. The longer a series of questions or dialogue with an LLM is, the greater the likelihood of AI hallucinations — authoritative-sounding answers that are not based in fact.2 This is most likely to occur in citations. In studies of medical errors caused by LLMs, errors are of omission are most common.3 Omission can be worsened if the clinician’s query leaves out critical history, patient context (such as inability to pay for expensive drugs or other social determinants of health), or an important physical exam finding. Additionally, if a topic is new or unusual enough that AI sources do not have current information, the content surfaced may be inaccurate.
Read the full article
Get immediate access, anytime, anywhere.
Choose a single article, issue, or full-access subscription.
Earn up to 5 CME credits per issue.
