Related Experiment Video
Updated: Sep 13, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
LongHealth: A Question Answering Benchmark with Long Clinical Documents
Lisa Adams1, Felix Busch1,2, Tianyu Han3
1Institute for Diagnostic and Interventional Radiology, School of Medicine and Health, TUM University Hospital, Technical University of Munich (TUM), Ismaninger Str. 22, Munich, 81675 Germany.
None:
Recent advancements in large language models (LLMs) offer potential benefits in healthcare, particularly in processing extensive patient records. However, existing benchmarks do not fully assess LLMs' capability in handling real-world, lengthy clinical data. We present the LongHealth benchmark, comprising 20 detailed fictional patient cases across various diseases, with each case containing 5090 to 6754 words. The benchmark challenges LLMs with 400 multiple-choice questions in three categories: information extraction, negation, and sorting, challenging LLMs to extract and interpret information from large clinical documents. We evaluated eleven open-source LLMs with a minimum of 16,000 tokens and also included OpenAI's proprietary and cost-efficient Generative Pre-trained Transformers-3.5 Turbo for comparison. The highest accuracy was observed for Mistral-Small-24B-Instruct-2501 and Llama-4-Scout-17B-16E-Instruct, particularly in tasks focused on information retrieval from single and multiple patient documents. However, all models struggled significantly in tasks requiring the identification of missing information, highlighting a critical area for improvement in clinical data. In conclusion, while LLMs show considerable potential for processing long clinical documents, their current accuracy levels are insufficient for reliable clinical use, especially in scenarios requiring the identification of missing information. The LongHealth benchmark provides a more realistic assessment of LLMs in a healthcare setting and highlights the need for further model refinement for safe and effective clinical application. We make the benchmark and evaluation code publicly available.
Supplementary Information:
The online version contains supplementary material available at 10.1007/s41666-025-00204-w.
More Related Videos
07:26Executing Complexity-Increasing Queries in Relational MySQL and NoSQL MongoDB and EXist Size-Growing ISO/EN 13606 Standardized EHR Databases
Published on: March 19, 2018
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities
Longitudinal Studies
Longitudinal Research
Methods of Documentation II: POMR
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Purpose of Health Records I
Here's a breakdown of how health records serve these purposes: