MedCalc-Bench: Evaluating Large Language Models for Medical Calculations.

Nikhil Khandekar1, Qiao Jin1, Guangzhi Xiong2

  • 1National Library of Medicine, National Institutes of Health.

Arxiv
|October 1, 2025
PubMed
Summary

This study introduces MedCalc-Bench, a new dataset for evaluating large language models (LLMs) in medical calculations. Current LLMs struggle with quantitative reasoning, highlighting a gap for clinical applications.