Already a member or subscriber? Sign in now

An “EASIER” Framework for Evaluating AI Tools

KEVIN KINDLER, MD
OLGA KRAVCHENKO, PHD
STACY BARTLETT, MD
STEPHANIE BALLARD, PHARMD

FPM. 2026;32(5):21-25.

Author disclosures: no relevant financial relationships.

This content conforms to AAFP criteria for CME.

Deciding which AI tools to implement, and which to avoid, just got EASIER.

Artificial intelligence (AI) is a new paradigm in health care that has the potential to augment or even transform clinical reasoning, but it may also introduce inaccuracies or reinforce biases.1 Physicians need a reliable way to evaluate the clinical usefulness of AI programs, particularly as they continue to develop at a rapid pace.

KEY POINTS

  • Several rubrics exist for health systems to evaluate AI tools, but none are ideal for individual physicians deciding whether to adopt an AI tool for clinical practice.
  • The EASIER method helps physicians evaluate AI tools across six domains: ethics, accuracy, safety, intended use, explainability/transparency, and regulation.
  • Physicians ultimately bear responsibility for the care they provide and need to make informed decisions about implementing AI tools.

Establishing a framework to systematically assess key aspects of AI tools, such as clinical validity, fairness, interpretability, and integration into workflow, will help clinicians make informed decisions and feel more open to their implementation.

There is no framework that neatly fills that role right now. The Coalition for Health AI and the Joint Commission have proposed guidance, but their recommendations are designed for health systems rather than individual clinicians.2,3 Other frameworks46 have some relevance to individual clinicians but require stakeholders in leadership across disciplines and are not practical for physician-directed review. Still other frameworks are limited to addressing domain-specific use cases.711

Not all clinicians practice under a health system's IT umbrella, and those who do may be justifiably wary of using certain AI tools, even if their health system has approved them. Frontline primary care clinicians need our own rubric to evaluate AI tools for implementation in clinical medicine. The following is our attempt to create one.

AN “EASIER” WAY

Just as clinicians rely on the STEPS framework (safety, tolerability, effectiveness, price, and simplicity) to evaluate new medications,12 they can use the EASIER framework to evaluate AI tools in the following areas: ethics, accuracy, safety, intended use, explainability/transparency, and regulation. Each area guides evaluation but is not meant to be prescriptive or provide a specific score or weight. Table 1 provides questions physicians can ask to further explore each area.

TABLE 1. SAMPLE QUESTIONS TO ASK WHEN USING THE “EASIER” FRAMEWORK

Ethics
  1. Patient population: Does this tool account for the diversity of my patient panel, such as race, ethnicity, language, dialect, and cultural background?

  2. Patient experience: Does using this tool negatively impact a patient's willingness to speak openly during the encounter?

  3. User autonomy: Does the tool allow me to remain in control of the final output? Is there a risk of automation complacency?

  4. Bias: Does this tool introduce clinical bias into medical decision making?

Accuracy
  1. Task performance: How often does the tool correctly perform the intended task and appropriately capture relevant medical information? Does the developer publish statistics on accuracy?

  2. Real-world use: Does the tool perform as well in my clinical environment as it claimed to perform in testing? Has the developer published a white paper?

  3. Comparative value: Is this tool noticeably more accurate or efficient than my current process? How does it compare to other products I have used?

Safety
  1. FDA review: Has the FDA completed a review of this tool? If so, what were the findings and what is the assessed risk classification?

  2. Downstream impact: If this tool provides inaccurate output, what is the worst-case effect on patient care? Can that risk be easily mitigated?

  3. Human-in-the-loop: Does the tool provide a clear mechanism for me to review, edit, and verify all content or orders before they are officially signed or implemented?

  4. Data privacy and security: Is there known or potential risk for patient health information being exposed or shared without consent? Is the tool HIPAA-compliant?

Intended use
  1. Use case and scope creep: What specific task(s) will the AI perform? Am I using the tool strictly for its designed purpose, or am I relying on it for tasks beyond its intended scope?

  2. Workflow integration: Does this tool fit in my current clinical workflow, or does it require modifications to my existing processes?

  3. Experience and support: Is the user interface intuitive enough that it reduces my cognitive load rather than adding a new layer of technical frustration? If assistance is needed, is there a support framework in place?

  4. Cost considerations: Is there any direct cost to me or the patient that limits accessibility?

Explainability/transparency
  1. Source tracking: Can I easily trace the AI's output back to the original source data (for example, a conversation transcript or diagnostic variables)?

  2. Model logic: Do I have enough insight into how the tool reached its conclusion to feel comfortable explaining it to a patient?

Regulation
  1. Oversight and governance: Is there evidence of external validation (such as FDA review) or publicly available information that confirms the tool has an established governance structure (internal or external)?

  2. Updates and ongoing education: How will I be notified when the model's underlying logic/features are updated? Are there educational materials describing how those updates might affect my practice?

  3. Feedback loop: Is there a clear way for me to report errors or performance issues directly to the developer or my local technical lead?

Dr. Kindler is assistant professor in the Department of Family Medicine and faculty in the Clinical Informatics Fellowship at the University of Pittsburgh School of Medicine.

Dr. Bartlett and Dr. Kravchenko are assistant professors in the University of Pittsburgh School of Medicine's Department of Family Medicine.

Dr. Ballard is a clinical pharmacist in the Department of Ambulatory Care Pharmacy Services at UPMC Presbyterian-Shadyside and faculty at UPMC Shadyside Family Medicine Residency.

Send comments to fpmedit@aafp.org, or add your comments to the article online.

Author disclosures: no relevant financial relationships.

  1. 1.Feehan M, Owen LA, McKinnon IM, DeAngelis MM. Artificial intelligence, heuristic biases, and the optimization of health outcomes: cautionary optimism. J Clin Med. 2021;10(22):5284.
  2. 2.Joint Commission and Coalition for Health AI (CHAI) release guidance to support responsible AI adoption across U.S. health systems. Joint Commission and CHAI news release. Sept. 17, 2025. Accessed July 8, 2026. https://www.jointcommission.org/en-us/knowledge-library/news/2025-09-jc-and-chai-release-initial-guidance-to-support-responsible-ai-adoption
  3. 3.Responsible AI guidance. CHAI. Accessed July 8, 2026. https://www.chai.org/workgroup/responsible-ai/blueprint-for-trustworthy-ai
  4. 4.Jacob C, Brasier N, Laurenzi E, Heuss S, Cöltekin A, Peter MK. AI for IMPACTS framework for evaluating the long-term real-world impacts of AI-powered clinician tools: systematic review and narrative synthesis. J of Med Internet Res. 2025;27(1):e67485.
  5. 5.Wells BJ, Nguyen HM, McWilliams A, et al. A practical framework for appropriate implementation and review of artificial intelligence (FAIR-AI) in healthcare. Npj Digit Med. 2025;8(1):514.
  6. 6.Di Bidino R, Daugbjerg S, Papavero SC, Haraldsen IH, Cicchetti A, Sacchini D. Health technology assessment framework for artificial intelligence-based technologies. Int J Technol Assess Health Care. 2024;40(1):e61.
  7. 7.Moons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505.
  8. 8.Tejani AS, Klontzas ME, Gatti AA, et al. Checklist for artificial intelligence in medical imaging (CLAIM): 2024 update. Radiol Artif Intell. 2024;6(4):e240300.
  9. 9.Sounderajah V, Guni A, Liu X, et al. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat Med. 2025;31(10):3283-3289.
  10. 10.Liu X, Cruz Rivera S, Moher D, Calvert MJ, Denniston AK; SPIRIT-AI and CONSORT-AI Working Group. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med. 2020;26(9):1364-1374.
  11. 11.Cruz Rivera S, Liu X, Chan A-W, et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med. 2020;26(9):1351-1363.
  12. 12.Wright J. Introducing STEPS. Am Fam Physician. 2003;68(8):1467.
  13. 13.2026 Best in KLAS Awards – Software and Services. KLAS Research. Feb. 4, 2026. Accessed July 7, 2026. https://klasresearch.com/report/2026-best-in-klas-awards-software-and-services/3906
  14. 14.Artificial intelligence-enabled medical devices. U.S. Food and Drug Administration. Reviewed June 16, 2026. Accessed July 7, 2026. https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices
  15. 15.Waldren SE. The promise and pitfalls of AI in primary care. Fam Pract Manag. 2024;31(2):27-31.

Copyright © 2026 by the American Academy of Family Physicians.

This content is owned by the AAFP. A person viewing it online may make one printout of the material and may use that printout only for his or her personal, non-commercial reference. This material may not otherwise be downloaded, copied, printed, stored, transmitted or reproduced in any medium, whether now known or later invented, except as authorized in writing by the AAFP. See permissions for copyright questions and/or permission requests.