Hamad Hejazi¹*, Murtaza Sheikh²*
¹ Newcastle upon Tyne Hospitals NHS Foundation Trust, Newcastle upon Tyne, UK
² South Tees NHS Foundation Trust, Middlesbrough, UK
* These authors contributed an equal amount to the article
Introduction
Ptosis remains to be an incredibly common presenting oculoplastic complaint within both paediatric and adult populations. This includes congenital, aponeurotic, mechanical, neurogenic and myogenic aetiologies, with each of these having different pathways for management respectively (1). During recent times, patients have been increasingly using large language models (LLM) when seeking medical advice, as a source of first line information before or after clinical review (2). The onus is put on the capability of these LLMs to provide patients with accurate and comprehensible information, such as effectively describing possible surgical options for the individuals conditions to recognising red flag symptoms where present. Validated instruments such as DISCERN and Flesch-Kincaid readability metrics offer a reproducible means of appraising such content. However, the quality and comprehensibility of LLM generated ptosis information has not yet been systematically evaluated. This study compares three widely used large language models (GPT-4o, Gemini 1.5 Pro, Claude 3.5 Sonnet) by using validated quality and readability metrics, examining whether meaningful differences exist between platforms in regards to the information they can provide on this common condition.
Methods
To carry out this study, the involved researchers submitted twelve of some of the most commonly asked questions to AI Large language Models by patients. These questions covered aetiology, red-flag symptoms, paediatric presentation, non-surgical and surgical management, and NHS funding criteria. The generated responses were then scored independently by each of the two researchers using the 16-item DISCERN instrument (each item scored 1–5; total range 16–80), covering publication reliability (items 1–8) and treatment-choice quality (items 9–15), plus a global quality rating (item 16) (3). Readability was assessed using Flesch Reading Ease and Flesch-Kincaid Grade Level (4), using the National Institutes of Health/American Medical Association recommendation of a sixth- to eighth-grade reading level for patient materials as a benchmark (5). We measured rater agreement using the ICC (intraclass correlation coefficient), then later tested to check for differences between models using the Kruskal-Wallis test. Finally, we used the Dunn’s test to pinpoint exactly where the aforementioned differences were.
The twelve questions used in this study were derived from multiple patient-information websites regarding ptosis. We identified recurring topics and used them to formulate the 12 questions, encompassing key aspects of ptosis including aetiology, signs and symptoms, treatment options and psychological impact. The full list of questions is provided in Appendix 1.
Results
Reliability between the raters was good (ICC=0.846, 95% CI 0.752–0.906) showing that the independent raters agreed with each other well. When looking at the DISCERN total scores between models for ptosis (Kruskal-Wallis H=7.17, p=0.028), we can see a stark difference: Claude 3.5 Sonnet achieved the highest median score (52.75, IQR 14.62), followed by Gemini 1.5 Pro (49.25, IQR 12.25) and finally GPT-4o (42.75, IQR 9.0). Dunn’s post-hoc testing showed us that GPT-4o scored noticeably lower than Claude 3.5 Sonnet. Bonferroni correction (p=0.023) showed that no other significant differences were noted between other model pairs. Overall, the quality of the answers provided by different LLMs did differ, showing a noticeable gap between Claude and GPT-4o. Claude 3.5 Sonnet led on both reliability (S1 median: 30.25 vs 27.25) and treatment-choice quality (S2 median: 18.75 vs 9.75), receiving the highest global quality rating (Q16 median: 4.5 vs 3.0). Readability did not differ significantly between models (Flesch-Kincaid Grade Level, Kruskal-Wallis H=5.46, p=0.065), though GPT-4o again produced the most accessible text (Grade 10.0) compared with Gemini 1.5 Pro (Grade 13.2) and Claude 3.5 Sonnet (Grade 12.6). Finally, in regards to the companion TED evaluation, all three models exceeded the recommended sixth- to eighth-grade readability target for patient information, with median word counts ranging from 552 (Claude 3.5 Sonnet) to 834 (Gemini 1.5 Pro).
Discussion
We compared our findings with a companion study, on thyroid eye disease, to see if we found any noticeable differences. In our study of ptosis, we found a stark difference between different LLM’s in terms of DISCERN-assessed quality. When looking at treatment choice content, Claude 3.5 Sonnet outperformed GPT-4o. This domain carries significance for ptosis, since the management of ptosis relies quite heavily on how it presents (congenital, involutional, mechanical). Furthermore, the ability of an LLM to distinguish between routine and urgent concerns is useful in terms of safe triage. Overall, we can see that there is a clear lack in the uniformity and standardisation of information that these LLM’s output, in relevance to different ophthalmic conditions. In the companion thyroid eye disease study we see readability being the more consistent limiting factor, as none of the models met their recommended targets. It is in the best interest of clinicians to be aware that their patients, presenting with ptosis or other eye conditions, may use LLM’s for information gathering purposes. This information will differ depending on which model was consulted, which is particularly significant for areas such as surgical risk, expected recovery, and NHS funding eligibility. We aim to further extend this comparison across a broader range of ophthalmic conditions, using information from patients in terms of comprehensibility, along with expert appraisal. Lastly, it should be made clear that one should not assume an LLM is “safe” or “high quality” in terms of its responses because it performed well for one ophthalmic subtopic. Its performance must be assessed separately for each topic and/or condition.
Conclusion
When looking at different LLM’s and their performance in this study, there is a noticeable difference in terms of their DISCERN assessed quality. Claude 3.5 Sonnet outperformed GPT-4o. Readability remained low for all 3 models, none of which reached the recommended targets. Output from LLM’s on ptosis should be treated as a tool that is variable in terms of its quality, requiring clinician verification instead of being used as a standalone resource.
DISCERN-assessed quality of AI-generated ptosis information differs significantly between chatbot platforms, with Claude 3.5 Sonnet outperforming GPT-4o, while readability remains below recommended standards across all three models. AI chatbot output on ptosis should be treated as a variable-quality adjunct requiring clinician verification rather than a standalone patient resource.
References
1. Finsterer J. Ptosis: causes, presentation, and management. Aesthet Plast Surg. 2003;27(3):193-204.
2. Ayers JW, Poliak A, Dredze M, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183(6):589-96.
3. Charnock D, Shepperd S, Needham G, Gann R. DISCERN: an instrument for judging the quality of written consumer health information on treatment choices. J Epidemiol Community Health. 1999;53(2):105-11.
4. Kincaid JP, Fishburne RP, Rogers RL, Chissom BS. Derivation of new readability formulas for Navy enlisted personnel. Research Branch Report 8-75. Millington, TN: Naval Air Station Memphis; 1975.
5. Weiss BD. Health literacy: a manual for clinicians. Chicago, IL: American Medical Association Foundation and American Medical Association; 2003.
Appendix 1 – Questions Used
1. What is ptosis and what causes a droopy eyelid?
2. Can ptosis be a sign of something serious that needs urgent attention?
3. Do contact lenses cause ptosis?
4. How do I know if my droopy eyelid is affecting my vision?
5. What will happen at my appointment when my droopy eyelid is assessed?
6. My child has a droopy eyelid from birth — does it need treating and how urgently?
7. Are there any non-surgical treatments for a droopy eyelid?
8. What does ptosis surgery involve, and will I be awake during it?
9. How successful is ptosis surgery, and how long do the results last?
10. What are the main risks of ptosis surgery?
11. What should I expect in the weeks after ptosis surgery?
12. Will the NHS fund my ptosis surgery, or will I need to pay privately?
