Accuracy of Artificial Intelligence Chatbots in Responding to Patient Questions About Anticoagulant Therapy: A Systematic Comparative Evidence Synthesis and Cross-Platform Evaluation Framework

Authors

DOI:

https://doi.org/10.66687/JMRIS

Keywords:

artificial intelligence, direct oral anticoagulants, medication safety

Abstract

Background: Oral anticoagulants are high-risk medicines in which misunderstanding can contribute to bleeding, thrombosis, treatment interruption, or inappropriate self-management. Large language model chatbots can produce immediate conversational answers to patient questions, but the safety of using general medical performance as a proxy for anticoagulation-specific accuracy is uncertain.
Objective: To synthesize high-quality evidence available through 31 December 2023 regarding the accuracy, completeness, reliability, and safety of artificial intelligence chatbots for medical question answering and to translate that evidence into a reproducible cross-platform framework for anticoagulant patient questions.
Methods: A structured evidence synthesis was conducted using Q1 literature in medicine, cardiology, thrombosis, digital health, and artificial intelligence. Direct chatbot studies were combined with anticoagulation guidelines and patient-knowledge research. Because no qualifying pre-2024 Q1 study directly compared multiple public chatbots specifically on anticoagulant patient questions, no model-specific anticoagulation performance values were fabricated.
Results: In 2023 studies, ChatGPT responses were judged appropriate for 21 of 25 cardiovascular prevention questions, were preferred to physician responses in 78.6% of evaluations in a general patient-question study, and achieved a median physician-rated accuracy score of 5.5/6 across 284 medical questions. Med-PaLM achieved 92.6% alignment with scientific consensus on consumer medical questions, yet 5.9% of its answers were still judged potentially harmful. Other studies demonstrated specialty-specific inaccuracies, material response instability, and hallucinated references. Anticoagulant questions are especially sensitive to omissions involving missed doses, bleeding, drug interactions, procedures, renal function, adherence, and the difference between warfarin and direct oral anticoagulants.
Conclusion: General-purpose chatbots showed promising medical question-answering performance by the end of 2023, but the evidence did not support unsupervised anticoagulant counseling or a definitive ranking of public platforms. Cross-platform evaluation should use identical prompts, repeated generations, independent anticoagulation experts, guideline-based reference answers, and separate scoring of accuracy, completeness, readability, actionability, consistency, and potential harm.

References

Steffel J, Collins R, Antz M, Cornu P, Desteghe L, Haeusler KG, et al. 2021 European Heart Rhythm Association Practical Guide on the use of non-vitamin K antagonist oral anticoagulants in patients with atrial fibrillation. Europace. 2021;23(10):1612-1676. doi:10.1093/europace/euab065.

Tomaselli GF, Mahaffey KW, Cuker A, Dobesh PP, Doherty JU, Eikelboom JW, et al. 2020 ACC Expert Consensus Decision Pathway on Management of Bleeding in Patients on Oral Anticoagulants. J Am Coll Cardiol. 2020;76(5):594-622. doi:10.1016/j.jacc.2020.04.053.

Stevens SM, Woller SC, Kreuziger LB, Bounameaux H, Doerschug K, Geersing GJ, et al. Antithrombotic Therapy for VTE Disease: Second Update of the CHEST Guideline and Expert Panel Report. Chest. 2021;160(6):e545-e608. doi:10.1016/j.chest.2021.07.055.

January CT, Wann LS, Calkins H, Chen LY, Cigarroa JE, Cleveland JC Jr, et al. 2019 AHA/ACC/HRS Focused Update of the 2014 Guideline for the Management of Patients With Atrial Fibrillation. Circulation. 2019;140(2):e125-e151. doi:10.1161/CIR.0000000000000665.

Amara W, Larsen TB, Sciaraffia E, Hernández Madrid A, Chen J, Estner H, et al. Patients' attitude and knowledge about oral anticoagulation therapy: results of a self-assessment survey in patients with atrial fibrillation conducted by the European Heart Rhythm Association. Europace. 2016;18(1):151-155. doi:10.1093/europace/euv317.

Sarraju A, Bruemmer D, Van Iterson E, Cho L, Rodriguez F, Laffin L. Appropriateness of Cardiovascular Disease Prevention Recommendations Obtained From a Popular Online Chat-Based Artificial Intelligence Model. JAMA. 2023;329(10):842-844. doi:10.1001/jama.2023.1044.

Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, et al. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern Med. 2023;183(6):589-596. doi:10.1001/jamainternmed.2023.1838.

Goodman RS, Patrinely JR, Stone CA Jr, Zimmerman E, Donald RR, Chang SS, et al. Accuracy and Reliability of Chatbot Responses to Physician Questions. JAMA Netw Open. 2023;6(10):e2336483. doi:10.1001/jamanetworkopen.2023.36483.

Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172-180. doi:10.1038/s41586-023-06291-2.

Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med. 2023;29(8):1930-1940. doi:10.1038/s41591-023-02448-8.

Caranfa JT, Bommakanti NK, Young BK, Zhao PY. Accuracy of Vitreoretinal Disease Information From an Artificial Intelligence Chatbot. JAMA Ophthalmol. 2023;141(9):906-907. doi:10.1001/jamaophthalmol.2023.3314.

Hua HU, Kaakour AH, Rachitskaya A, Srivastava S, Sharma S, Mammo DA. Evaluation and Comparison of Ophthalmic Scientific Abstracts and References by Current Artificial Intelligence Chatbots. JAMA Ophthalmol. 2023;141(9):819-824. doi:10.1001/jamaophthalmol.2023.3119.

Downloads

Published

2024-03-15

Similar Articles

1-10 of 11

You may also start an advanced similarity search for this article.