Related Experiment Videos

How to benchmark medical AI agents

Silas Ruhrberg Estévez1,2,3, Dyke Ferber1, Mihaela van der Schaar2,3

  • 1Else Kroener Fresenius Center for Digital Health, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology, Dresden, Germany.

Plos Medicine
|July 9, 2026
PubMed
Summary

Medical artificial intelligence is moving towards multimodal large language models for clinical workflows. New benchmarks are needed to evaluate reasoning, safety, and resource use, not just final outcomes.