Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automating Case-Based Learning in Obstetrics and Gynecology: Validation of a Locally Deployed Large Language Model
Xueyan Hu1,2, Xiaoju Luo1, Yuan Zhao1
1Department of Obstetrics and Gynecology, Wenzhou People's Hospital, Wenzhou, Zhejiang, People's Republic of China.
Background:
Standardized residency training in obstetrics and gynecology (OB/GYN) relies heavily on Case-Based Learning (CBL). Yet, manually curating high-quality teaching cases remains a labor-intensive burden.
Objective:
To develop and evaluate a methodological framework using a locally deployed large language model (LLM) to automate the extraction and construction of structured CBL databases.
Methods:
We employed a locally deployed Qwen3-8B-Instruct model to analyze 3678 OB/GYN PubMed case abstracts, extracting primary diagnosis, core teaching points, clinical pitfalls, and difficulty levels. Two senior OB/GYN educators established a ground-truth dataset from 100 randomly selected cases to quantify extraction accuracy and AI hallucination rates. Teaching utility was assessed via a 5-point Likert scale, calculating inter-rater reliability using intraclass correlation coefficients (ICC).
Results:
Our automated framework effectively categorized the 3678 cases into diverse subspecialties, with Reproductive Endocrinology and Obstetrics (28.10%) and Andrology and Male Infertility (25.20%) representing the largest cohorts. The LLM demonstrated high precision in extracting objective clinical data, achieving a 96.0% accuracy rate for primary diagnoses and 92.0% for core teaching points. Notably, the overall incidence of AI hallucinations remained low at 3.0%. In pedagogical evaluations, the model earned high scores for diagnostic accuracy (Mean = 4.48/5.0, p < 0.05), showing strong expert consensus (ICC = 0.92). However, the model's performance faltered when addressing subjective clinical nuances; inter-rater reliability dropped significantly regarding the utility of clinical pitfalls (ICC = 0.65) and the appropriateness of difficulty levels (ICC = 0.54).
Conclusion:
This study demonstrates that LLMs offer an efficient, scalable framework for building large-scale CBL databases. While highly accurate in extracting objective facts, AI still lacks the "tacit knowledge" and empirical intuition of senior clinicians. Thus, human oversight remains indispensable to validate content and mitigate clinical risks. Future research must evaluate how these databases directly impact resident competency and learning outcomes.