Research Article: Performance of large language models in postoperative hip fracture rehabilitation counseling for older adults: a comparative evaluation of safety, accuracy, reliability, readability, and empathy
Abstract:
To compare the performance of five widely used large language models in answering public questions about postoperative rehabilitation after hip fracture in older adults, focusing on safety, accuracy, reliability, readability, empathy, and overall information quality.
This cross-sectional comparative evaluation used a 42-question bank addressing general postoperative rehabilitation after proximal femoral fracture treated with internal fixation or hip arthroplasty. All questions and the standardized safety-oriented prompt were submitted in English, and all responses were returned and analyzed in their original English form without translation. Between May 1 and May 10, 2026, each question was submitted once to five generative artificial intelligence-driven chatbots through their official public interfaces, yielding 210 responses. Three independent raters evaluated each response using DISCERN, EQIP, JAMA benchmark criteria, Global Quality Score, and anchored 5-point accuracy and empathy scales. Safety was classified using prespecified binary criteria and summarized descriptively; readability was assessed using six established indices applied to the original English responses. Paired inferential analyses were used for non-binary outcomes.
Inter-rater agreement was excellent (Fleiss' kappa 0.855; ICC 0.829–0.884). Under the standardized safety-oriented prompt, 199 of 210 responses (94.8%) met the study's prespecified binary safety criterion; the 11 responses that did not meet the criterion were concentrated in three questions and were summarized descriptively. Accuracy ratings were high across models, with no statistically detectable difference ( P = 0.572). Empathy differed across models ( P = 0.003; Kendall's W = 0.096), with Gemini showing the highest median score. DISCERN, EQIP, and Global Quality Score varied across models (all P <0.001), although the GQS effect size was small ( W = 0.169). Readability was suboptimal: only 2 of 210 responses (1.0%) met the commonly recommended sixth-grade threshold by FKGL.
Under a standardized safety-oriented prompt, most responses met the study's prespecified binary safety criterion and received high accuracy ratings; however, models differed in empathy, reliability, information quality, and readability. These prompt-conditioned findings do not establish unguided LLM responses as generally safe. LLMs may serve only as supervised adjunctive educational tools, with outputs framed as complementary to procedure-specific professional care.
No summary available.