Large Language Models for Generating Patient Education Materials on Type 2 Diabetes: Comparative Assessment of Readability, Accuracy, Completeness, and Hallucination Risk

Authors

DOI:

https://doi.org/10.66687/JMRIS

Keywords:

large language models, type 2 diabetes, chat gpt, AI

Abstract

Background: Type 2 diabetes requires sustained self-management education covering medication use, nutrition, glucose monitoring, physical activity, complication prevention, and responses to abnormal glucose values. Large language models (LLMs) can generate fluent patient-facing explanations within seconds, but their suitability cannot be judged by factual accuracy alone.
Objective: To synthesize high-quality evidence available through 31 December 2023 concerning the ability of LLMs to generate type 2 diabetes patient education, focusing on readability, factual accuracy, completeness, reproducibility, and hallucination or unsupported-content risk.
Methods: A structured comparative evidence synthesis was conducted using Q1 peer-reviewed literature published no later than 31 December 2023. Diabetes-specific LLM studies were prioritized and supplemented by high-quality general medical evaluations where they directly informed assessment of accuracy, completeness, patient-facing communication, or safety. Heterogeneous tasks and scoring systems precluded meta-analysis; results were therefore synthesized by prespecified quality domain.
Results: Diabetes-specific evaluations showed high but nonuniform factual performance. ChatGPT answered all 24 items of the Diabetes Knowledge Questionnaire correctly in one study, while another diabetes evaluation reported very high expert accuracy scores but a mean Flesch-Kincaid grade level of 13.8, indicating substantial reading burden. In a Turing-test-inspired diabetes study, two of ten ChatGPT answers contained factual errors, including one potentially consequential error concerning exercise and glucose response. An AI dietitian study received favorable professional ratings for 162 of 168 generated responses. Across studies, accuracy frequently coexisted with excessive complexity, omissions, lack of source transparency, and occasional incorrect statements.
Conclusion: LLMs can produce useful first drafts of type 2 diabetes education, but patient-facing deployment requires a multidimensional quality gate. Accuracy, completeness, readability, and hallucination risk should be evaluated separately, with clinician review, guideline verification, version documentation, and post-release monitoring. A highly accurate response is not necessarily understandable, complete, current, or safe.

References

Powers MA, Bardsley JK, Cypress M, Duker P, Funnell MM, Fischl AH, et al. Diabetes self-management education and support in adults with type 2 diabetes: a consensus report of the American Diabetes Association, the Association of Diabetes Care & Education Specialists, the Academy of Nutrition and Dietetics, the American Academy of Family Physicians, the American Academy of PAs, the American Association of Nurse Practitioners, and the American Pharmacists Association. Diabetes Care. 2020;43(7):1636-1649. doi:10.2337/dci20-0023.

Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172-180. doi:10.1038/s41586-023-06291-2.

Lee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N Engl J Med. 2023;388(13):1233-1239. doi:10.1056/NEJMsr2214184.

Sng GGR, Tung JYM, Lim DYZ, Bee YM. Potential and pitfalls of ChatGPT and natural-language artificial intelligence models for diabetes education. Diabetes Care. 2023;46(5):e103-e105. doi:10.2337/dc23-0197.

Nakhleh A, Spitzer S, Shehadeh N. ChatGPT's response to the Diabetes Knowledge Questionnaire: implications for diabetes education. Diabetes Technol Ther. 2023;25(8):571-573. doi:10.1089/dia.2023.0134.

Huang C, Chen L, Huang H, Cai Q, Lin R, Wu X, et al. Evaluate the accuracy of ChatGPT's responses to diabetes questions and misconceptions. J Transl Med. 2023;21(1):502. doi:10.1186/s12967-023-04354-6.

Hulman A, Dollerup OL, Mortensen JF, Fenech ME, Norman K, Støvring H, et al. ChatGPT- versus human-generated answers to frequently asked questions about diabetes: a Turing test-inspired survey among employees of a Danish diabetes center. PLoS One. 2023;18(8):e0290773. doi:10.1371/journal.pone.0290773.

Sun H, Zhang K, Lan W, Gu Q, Jiang G, Yang X, et al. An AI dietitian for type 2 diabetes mellitus management based on large language and image recognition models: preclinical concept validation study. J Med Internet Res. 2023;25:e51300. doi:10.2196/51300.

Goodman RS, Patrinely JR, Stone CA Jr, Zimmerman E, Donald RR, Chang SS, et al. Accuracy and reliability of chatbot responses to physician questions. JAMA Netw Open. 2023;6(10):e2336483. doi:10.1001/jamanetworkopen.2023.36483.

Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183(6):589-596. doi:10.1001/jamainternmed.2023.1838.

Liu S, Wright AP, Patterson BL, Wanderer JP, Turer RW, Nelson SD, et al. Using AI-generated suggestions from ChatGPT to optimize clinical decision support. J Am Med Inform Assoc. 2023;30(7):1237-1245. doi:10.1093/jamia/ocad072.

Arrieta AB, Díaz-Rodríguez N, Del Ser J, Bennetot A, Tabik S, Barbado A, et al. Explainable artificial intelligence (XAI): concepts, taxonomies, opportunities and challenges toward responsible AI. Inf Fusion. 2020;58:82-115. doi:10.1016/j.inffus.2019.12.012.

Downloads

Published

2024-03-15

Similar Articles

11-14 of 14

You may also start an advanced similarity search for this article.