Generative Artificial Intelligence in Hip and Knee Arthroplasty: A Systematic Review of Emerging Clinical Applications in Patient Communication and Education, Documentation, and Decision Support.
Background
Generative artificial intelligence (AI), including large language models (LLMs), has been increasingly explored in orthopedic surgery; however, its application within total hip and knee arthroplasty (THA/TKA) has not been clearly characterized. Therefore, we performed a systematic review to further evaluate generative AI in THA/TKA across 3 clinical domains.
Methods
A PubMed and Embase systematic literature review was performed on July 9, 2025, in accordance with the Preffered Reporting Items for Systematic Reviews and Meta-Analyses guidelines. Included studies evaluated generative AI use in THA/TKA and addressed 1 of our 3 domains. Excluded studies used nongenerative AI, involved populations not undergoing THA/TKA, or were non-English or non-peer-reviewed. Quality metrics that were assessed included blinded clinician ratings, readability scores, DISCERN scores, and diagnostic accuracy measures. The heterogeneity of the included studies led to a narrative synthesis, and no formal risk of bias was conducted.
Results
Of the 91 articles retrieved, 23 met the inclusion criteria. ChatGPT versions 3.5 or 4 were assessed across all studies, and 3 studies included Google Gemini, Claude 3 Opus, and DeepSeek. Among studies in patient communication and education (n = 19), blinded clinician ratings showed that ChatGPT-generated responses to frequently asked questions (FAQ's) were as accurate, clear, and complete as surgeon-written responses. Regarding documentation (n = 2), LLMs demonstrated 97.5% to 100% accuracy in identifying operative report information and created patient consent documents with improved readability scores (grade level 12.6 vs. 16.8) and completeness (2.4/3 vs. 1.8/3). Among studies assessing decision support (n = 2), ChatGPT demonstrated high sensitivity when predicting surgical candidacy but low specificity when predicting surgical outcomes.
Conclusion
LLMs performed well in several aspects of patient communication and education with support from high clinician ratings, accuracy, and readability scores. However, evidence for documentation and decision support remained limited. Moreover, proper implementation that considers clinical workflows could increase clinician efficiency and positively affect patient experience.
Conflict of interest statement
Disclosure: The Disclosure of Potential Conflicts of Interest forms are provided with the online version of the article (https://links.lww.com/JBJSOA/B341).

Comments
Sign in to join the conversation.
Sign in to commentNo comments yet. Be the first to share your thoughts.