Research Article: Evaluating artificial intelligence-generated clinical guidance for patellar dislocation: accuracy, readability and qualitative appraisal of DeepSeek-R1 responses
Abstract:
Large language models (LLMs) are increasingly used to answer patient questions, but the readability and information quality of patellar dislocation guidance remain unclear.
A series of clinical questions covering etiopathogenesis, mechanisms, clinical manifestations, diagnosis, treatment, complications, and postoperative rehabilitation of patellar dislocation were posed to DeepSeek-R1 v2.3.1 and ChatGPT-4o. Readability was assessed using the Flesch Reading Ease score (FRES), Flesch-Kincaid Grade Level (FKGL), Automated Readability Index (ARI), SMOG Index, and Gunning Fog Score (GFS). Three sports medicine surgeons rated responses using DISCERN. Wilcoxon signed-rank tests compared models; Inter-rater reliability (IRR) for DISCERN was assessed with intraclass correlation coefficient (ICC). Additionally, we analysed DeepSeek's responses based on the current literature and therapeutic standards.
Readability favored ChatGPT-4o across all indices (between-model P <?0.05): DeepSeek-R1 vs. ChatGPT-4o medians (Q1–Q3) were ARI 15.62 (13.90–16.60) vs. 12.54 (10.15–13.41), FKGL 14.39 (12.38–15.24) vs. 11.27 (9.44–12.76), GFS 18.25 (15.20–19.43) vs. 14.55 (13.73–16.53), SMOG 11.02 (9.72–12.18) vs. 9.55 (8.14–10.67), and FRES 9.50 (6.75–25.00) vs. 33.00 (19.75–41.50). Both models exceeded the NIH-recommended sixth-grade reading level (all P <?0.05). DISCERN scores were similar [median (Q1–Q3)] for DeepSeek-R1 vs. ChatGPT-4o: 49.17 (40.25–59.17) vs. 50.33 (38.92–58.42), corresponding to “Fair” quality ( P =?0.767). DISCERN agreement was high (single-measure ICC 0.860–0.950).
Although both AI platforms can generate clinically relevant responses of moderate quality, they should be regarded as adjunctive information rather than stand-alone clinical guidance, particularly given the high incidence of patellar instability and developmental characteristics of children and adolescents. Exploratory qualitative assessment of DeepSeek-R1 responses identified strengths in clinical explanation and reasoning, while also revealing limitations in citation reliability and readability. Future work should incorporate multi-turn, multi-generation frameworks, more samples with repeated sampling and longitudinal assessment, enforced citation standards, and expanded evaluation to better quantify both clinical utility and reliability.
Introduction:
Large language models (LLMs) are increasingly used to answer patient questions, but the readability and information quality of patellar dislocation guidance remain unclear.
Read more