Related Experiment Video
Updated: May 20, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Leveraging computerized adaptive testing for cost-effective evaluation of large language models in medical
Tianpeng Zheng1,2, Zhehan Jiang3,4, Jiayi Liu5
1Institute of Medical Education, Health Science Center, Peking University, Beijing, China.
NPJ Digital Medicine
|May 18, 2026
Summary
A new computerized adaptive testing (CAT) framework significantly reduces evaluation time and cost for large language models (LLMs) in healthcare. This psychometrically rigorous method enables efficient, scalable assessment of medical knowledge.
Area of Science:
- Artificial Intelligence in Medicine
- Psychometrics
- Natural Language Processing
Background:
- Large language models (LLMs) are increasingly used in healthcare.
- Current LLM evaluation methods use static benchmarks that are costly and prone to contamination.
- Existing benchmarks lack precise measurement properties for detailed performance tracking.
Purpose of the Study:
- To develop and validate a computerized adaptive testing (CAT) framework for assessing standardized medical knowledge in LLMs.
- To enable scalable and psychometrically rigorous evaluation of LLMs in healthcare.
- To improve the efficiency and cost-effectiveness of LLM performance tracking.
Main Methods:
- A two-phase study using Monte Carlo simulations and empirical evaluation of 38 LLMs.
- Development of a CAT framework based on item response theory.
- Assessment of standardized medical knowledge in LLMs.
Main Results:
- The CAT protocol achieved near-perfect correlation (r=0.988) with full-bank results using only 1.3% of items.
- Evaluation time decreased from 6.85 hours to 8.4 minutes per model.
- Token usage dropped from 1.77 million to 0.03 million, and costs decreased from ~$1,475 to under $5 per model.
- Model rankings were fully preserved (Spearman's ρ=1.0).
Conclusions:
- The developed CAT framework provides a scalable, psychometrically rigorous, and cost-effective method for evaluating LLMs in healthcare.
- This adaptive methodology is suitable for pre-screening and continuous monitoring of foundational knowledge in LLMs.
- The CAT framework complements, but does not replace, real-world clinical validation and safety studies.