Hamad Hejazi¹*, Abdullah Sheekhuna²*
¹ Newcastle upon Tyne Hospitals NHS Foundation Trust, Newcastle upon Tyne, UK
² Gateshead NHS Foundation Trust, Gateshead, UK
* Both authors have contributed equally to this paper
Thyroid eye disease (TED) is an autoimmune condition and it is the most prevalent extrathyroidal disorder of Grave’s disease and the course of this manifestation can vary from individuals leading to various complications with many treatment options ranging from the use of biologics to surgical interventions (1). The burden of this disorder can be unpredictable therefore it is important that patients understand what the disorder is and what complications to be aware of in order to receive the right treatment which can subsequently improve the outcomes both physically and cosmetically (1).
There is an increased use of artificial intelligence in which patients are seeking information about TED prior to or even replacing consulting with a clinician (2). There are validated tools that are used to evaluate health information such as their quality and readability (4, 5) but there is a lack of evidence regarding their use when applied to evaluating the quality of chat bot generated information regarding patient information for TED. This study evaluates 3 commonly used chatbots and utilises validated tools such as DISCERN (4) and Flesh-Kincaid readability metrics (5) to assess and compare the overall quality of patient information.
Method
We used 12 questions that patients with TED frequently ask about, covering topics such as causes, signs and symptoms, complications and treatment options including teprotumumab (3). The questions were submitted verbatim without any changes to wording into the three chatbots: GPT-4o, Gemini 1.5 Pro and Claude 3.5 Sonnet. We used the 16 item DISCERN tool (4) to assess the quality of the chatbot responses and rated them 1-5 for each item with a higher score indicating better quality health information. Two independent reviewers rated each response a score of one to five and the total score for each response across all 16 items ranged from 16 to 80. The items assessed covered reliability of information (items 1-8) and quality of treatment choice information (items 9-15), as well as an item assessing the overall quality (item 16) (4). The readability of the responses were assessed using two standard tests being Flesch Reading Ease and Flesch-Kincaid Grade Level (5). The results were then compared with the recommended reading level standards set out by the National Institutes of Health and The American Medical Association which recommended a sixth- to eighth-grade reading level for patient information (6). To establish the inter-rater reliability of each response, we utilised the intraclass correlation coefficient to determine the consistency between the scores given by the two raters. We assessed the three chatbots to determine if the differences were meaningful and statistically significant using the Kruskal-Wallis test. Where the Kruskal-Wallis test demonstrated a significant difference; Dunn’s post-hoc testing (bonferroni-corrected) was used to identify which chatbots differed.
The twelve questions used in this study were derived from multiple patient-information websites regarding TED. We identified recurring topics and used them to formulate the 12 questions, encompassing key aspects of TED including aetiology, signs and symptoms, treatment options and psychological impact. The full list of questions is provided in Appendix 1.
Results
The two raters generally scored quite consistently with an inter-rater reliability ICC score of 0.846, with a 95% confidence interval of 0.752 to 0.906. In regards to the quality of health information, GPT-4o appeared to produce the highest quality of information with a median DISCERN total score of 50.5 (IQR, 12.25) and was subsequently followed by Gemini 1.5 Pro which achieved a score of 43.75 (IQR, 16.38) and Claude 3.5 Sonnet with a score of 42.5 (IQR, 8.38). Kruskal-Wallis was applied to these results and the differences were not statistically significant (H=4.94, p=0.085). However a side by side comparison of Chat GPT-4o with Claude Sonnet 3.5 demonstrated a statistically significant difference (p=0.030) but this difference was no longer significant following Bonferroni adjustment (adjusted p=0.090).
Further analysis of the DISCERN items demonstrated that GPT-4o consistently achieved the highest median scores across all items assessed. For information reliability, GPT-4o scored 26.25, compared with 23.5 for both other models. With treatment choice quality, GPT-4o scored 19.75, with a score of 18.0 for Gemini 1.5 Pro and 17.0 for Claude 3.5 sonnet. GPT-4o also achieved the highest global quality rating, with a median score of 3.5, compared with 3.0 for both comparators.
The readability of the three models varied and this was evaluated using the Flesch-Kincaid Grade level model. GPT-4o produced the most readable responses with a grade level of 10.1, compared with 15.5 and 15.8 for Gemini 1.5 and Claude 3.5 Sonnet respectively. Despite GPT-4o having the lowest grade level; all three models exceeded the recommended sixth-to-eight grade reading ability. (6) The Kruskal-Wallis test was applied to the different grades and there was a statistically significant difference between the grade levels (H=23.5 p<0.001).
Chat GPT-4o produced longer responses and the length of responses generated varied between the three chat bots; GPT-4o generated a median word count of 756 compared to a median score count of 582 and 610 for Gemini 1.5 Pro and Claude 3.5 Sonnet respectively despite GPT-4o receiving a lower grade level for its responses.
Discussion
The results demonstrated that AI generated chat bots are capable of generating reasonably good quality information however their responses did not meet the standards for the general patient audience (6). The differences are consequential as poor readability may limit the understanding of complex information making decisions particularly difficult while patients must also cope with the psychological burden of TED (1).
Despite GPT-4o numerically scoring higher across the core elements of the DISCERN assessments, the lack of a significant difference between the models limits our ability to claim that an individual model is superior to another. In contrast, the complexities in readability varied significantly demonstrating that the selection of the model may influence how accessible patient information is. These results show that the developers of AI chatbots should aim to improve the accessibility of health information whilst maintaining a standard of good quality information. Clinicians must also consider that patients may encounter reasonable quality TED information; however due to the complex nature of chatbot responses limiting their understanding, clinicians must signpost to accessible and reliable resources.
Limitations of this study included the small number of independent raters and potentially affecting the robustness of the assessment despite good inter-rater reliability. The use of predetermined questions and single turn responses limited the conversational aspect of chatbot interactions and may not represent the dynamic nature of patient-chatbot interactions. Due to the rapid evolution of AI models and responses being captured at a single point in time, the applicability of these findings may diminish over time indicating that performance may differ in newer versions.
Conclusion
All three AI models were capable of generating reasonable quality patient information and were broadly comparable; however the recommended readability standards were not met (6). While AI chat bots may have a role in educating patients, clinicians should consider them to be an educational resource rather than replacing clinical counselling and patients should be directed towards validated resources.
References
- Bartalena L, Kahaly GJ, Baldeschi L, et al. The 2021 European Group on Graves’ orbitopathy (EUGOGO) clinical practice guidelines for the medical management of Graves’ orbitopathy. Eur J Endocrinol. 2021;185(4).
- Ayers JW, Poliak A, Dredze M, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183(6):589-96.
- Douglas RS, Kahaly GJ, Patel A, et al. Teprotumumab for the treatment of active thyroid eye disease. N Engl J Med. 2020;382(4):341-52.
- Charnock D, Shepperd S, Needham G, Gann R. DISCERN: an instrument for judging the quality of written consumer health information on treatment choices. J Epidemiol Community Health. 1999;53(2):105-11.
- Kincaid JP, Fishburne RP, Rogers RL, Chissom BS. Derivation of new readability formulas for Navy enlisted personnel. Research Branch Report 8-75. Millington, TN: Naval Air Station Memphis; 1975.
- Weiss BD. Health literacy: a manual for clinicians. Chicago, IL: American Medical Association Foundation and American Medical Association; 2003.
Appendix 1 – Thyroid Eye Disease Patient Questions
- What is thyroid eye disease and why does it happen?
- Can I get thyroid eye disease even if my thyroid levels are normal?
- Does smoking make thyroid eye disease worse?
- What are the symptoms of thyroid eye disease I should watch out for?
- When should I seek urgent medical attention for my eyes?
- What tests will my doctor do to diagnose and monitor thyroid eye disease?
- What is the active phase of thyroid eye disease and how long does it last?
- Will my eyes go back to how they looked before?
- What medications are used to treat thyroid eye disease?
- What is teprotumumab and is it available to me in the UK?
- What types of surgery are available for thyroid eye disease and when would I need them?
- How might thyroid eye disease affect my mental health and daily life?
