Generative artificial intelligence and large language models in sports medicine: a scoping review of applications, accuracy, and ethical implications.
Dergaa I, Dergaa MA, Razzak M, Ceylan Hİ, Stefanica V, Muntean RI et al. · Frontiers in public health · 2026
Background: Generative artificial intelligence (GenAI), particularly large language models (LLMs) such as ChatGPT, GPT-3.5, and GPT-4, is rapidly being integrated into sports medicine practice. These tools are increasingly used by health professionals, coaches, and athletes for training prescription, rehabilitation, nutrition, mental health support, injury prevention, and academic writing. However, their adoption has outpaced the development of robust evidence regarding their clinical utility, accuracy, and safety, and no comprehensive synthesis of this emerging field currently exists.
Aim: This scoping review aimed to map the current evidence on GenAI and LLM applications in sports medicine and athlete health, evaluate their accuracy and hallucination risks across domains, and synthesise reported ethical and governance concerns.
Methods: The review followed PRISMA-ScR guidelines and Joanna Briggs Institute methodology. Six databases (PubMed/MEDLINE, Scopus, Web of Science, SPORTDiscus, IEEE Xplore, and CINAHL) were searched for studies published between January 2022 and March 2026. Eligibility criteria were defined using the Population-Concept-Context framework. Two independent reviewers conducted screening, full-text assessment, and data extraction, with strong inter-rater agreement ( κ = 0.82).
Results: Of 1,847 records identified, 32 studies were included. Applications were classified into seven domains: training and exercise prescription ( n = 5), nutrition ( n = 4), rehabilitation ( n = 3), mental health ( n = 3), clinical decision support ( n = 5), academic writing ( n = 6), and ethics/governance ( n = 6). LLM accuracy varied substantially: in a single validation study, content validity ratios for sleep recommendations ranged from 0.33 (GPT-3.5) to 0.67 (GPT-4), while only GPT-4 achieved acceptable validity for jet lag guidance (CVR = 0.68). In one bibliometric analysis, AI-generated text in sports medicine journals increased from 2.38% (early 2023) to 6.25% (late 2024). Hallucination risk was rated critical for general-
Purpose: chatbots but substantially reduced in retrieval-augmented systems.
Conclusion: GenAI shows promise as a supervised decision-support tool in sports medicine, but current evidence does not support unsupervised clinical use. Key challenges include hallucination risks, a lack of sport-specific validation datasets, and insufficient ethical and governance frameworks. Addressing these gaps is essential before widespread integration into athlete health and public health contexts.