Related Experiment Video
Updated: Sep 13, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
LongHealth: A Question Answering Benchmark with Long Clinical Documents.
Lisa Adams1, Felix Busch1,2, Tianyu Han3
1Institute for Diagnostic and Interventional Radiology, School of Medicine and Health, TUM University Hospital, Technical University of Munich (TUM), Ismaninger Str. 22, Munich, 81675 Germany.
Large language models (LLMs) show promise for processing lengthy clinical data, but current accuracy is insufficient for reliable healthcare use. The new LongHealth benchmark reveals significant struggles in identifying missing patient information.
Area of Science:
- Artificial Intelligence in Healthcare
- Natural Language Processing for Clinical Data
- Medical Informatics
Background:
- Large language models (LLMs) offer potential for analyzing extensive patient records.
- Existing benchmarks inadequately assess LLM performance on real-world, lengthy clinical data.
Purpose of the Study:
- To introduce the LongHealth benchmark for evaluating LLMs on long clinical documents.
- To assess the capabilities of various LLMs in processing and interpreting complex patient case files.
Main Methods:
- Developed the LongHealth benchmark with 20 detailed fictional patient cases (5090-6754 words each).
- Created 400 multiple-choice questions across information extraction, negation, and sorting tasks.
- Evaluated eleven open-source LLMs and GPT-3.5 Turbo on the benchmark.
Main Results:
- Mistral-Small-24B-Instruct-2501 and Llama-4-Scout-17B-16E-Instruct achieved the highest accuracy, particularly in information retrieval.
- All evaluated models demonstrated significant difficulties in identifying missing information within clinical documents.
- Current LLM accuracy is insufficient for safe and reliable clinical application.
Conclusions:
- LLMs have potential for processing long clinical documents but require further refinement for healthcare.
- The LongHealth benchmark offers a realistic assessment of LLM performance in clinical settings.
- Future research should focus on improving LLM accuracy in identifying missing clinical data.
More Related Videos
07:26Executing Complexity-Increasing Queries in Relational MySQL and NoSQL MongoDB and EXist Size-Growing ISO/EN 13606 Standardized EHR Databases
Published on: March 19, 2018
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities
Longitudinal Studies
Longitudinal Research
Methods of Documentation II: POMR
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Purpose of Health Records I
Here's a breakdown of how health records serve these purposes: