Research Article: Development and expert evaluation of a culturally adapted large language model for Malay clinical history-taking simulation: mixed methods study
Abstract:
Clinical history-taking is a core medical competency, but conventional teaching may be limited by patient availability, variable clinical exposure, and restricted opportunities for repeated practice. Large language models (LLMs) may support simulated practice; however, their clinical, linguistic, and cultural suitability requires careful evaluation in multilingual settings.
This study aimed to develop and expert-evaluate a culturally and linguistically adapted English–Malay clinical dialogue corpus and to descriptively benchmark a fine-tuned LLM framework for clinical history-taking simulation in Malaysian undergraduate medical education.
A multiphase mixed-methods development and evaluation study was conducted across nine internal medicine systems. Of 63,250 AI-generated dialogue entries, each comprising a simulated doctor inquiry and corresponding patient response, 59,184 (93.6%) underwent consultant-level subject matter expert review using an author-developed holistic 5-point rubric informed by clinical relevance, accuracy, clarity and coherence, completeness, and human-likeness. Free-text comments were analysed inductively. After SME evaluation and clinical-quality and mechanical filtering, 58,886 entries across 2,497 simulated patient profiles were retained. The profiles were partitioned into 1,997 training profiles, 250 validation profiles, and 250 held-out test profiles. Four model configurations were descriptively benchmarked using the same held-out profiles and an automated large language model judge. This benchmark was intended for internal model comparison and did not constitute independent external or human clinical validation of the post-fine-tuning model outputs.
Mean subject matter expert scores were 4.52/5 for simulated doctor inquiries and 4.58/5 for simulated patient responses. Qualitative analysis identified recurring limitations involving question structure, incomplete probing, repetition, linguistic accuracy, over-explanation, and simulated-patient realism. In descriptive benchmarking, the fine-tuned model achieved the highest observed mean score among the four configurations evaluated.
Developing a consultant-reviewed, culturally adapted English–Malay clinical dialogue corpus and using it to fine-tune an LLM for simulated history-taking was feasible. The fine-tuned model achieved the highest observed score in the internal automated benchmark, but this finding does not establish statistical, clinical, or educational superiority. The post-fine-tuning model outputs were not independently evaluated by clinicians in the present study. The framework may offer a scalable, contextually relevant supplement to clinical communication training, but blinded human evaluation of post-fine-tuning outputs, targeted safety assessment, external validation, and learner-outcome evaluation are required before routine student-facing implementation.
Introduction:
Clinical history-taking is a core medical competency, but conventional teaching may be limited by patient availability, variable clinical exposure, and restricted opportunities for repeated practice. Large language models (LLMs) may support simulated practice; however, their clinical, linguistic, and cultural suitability requires careful evaluation in multilingual settings.
Read more