Related Experiment Video
Updated: May 11, 2026

Reconstruction of 3-Dimensional Histology Volume and its Application to Study Mouse Mammary Glands
Published on: July 26, 2014
Preclinical HistoBench: A Pilot Benchmark Dataset for Evaluating Large Language Models on Preclinical
Avan Kader1, Marie-Luise H H Ranner-Hafferl2, Felix Reuter1
1Department of Diagnostic and Interventional Radiology, Technical University of Munich, Ismaninger Str. 22, 81675 Munich, Germany.
This study introduces a benchmark dataset for evaluating large language models (LLMs) in preclinical histopathology. LLM performance varies significantly, showing sensitivity to class imbalance and potential as research screening tools.
Area of Science:
- * Preclinical histopathology
- * Artificial intelligence in pathology
- * Large language model (LLM) evaluation
Background:
- * Lack of standardized benchmarks for assessing LLMs in preclinical histopathology.
- * Need for multi-dimensional classification capabilities in AI models for pathology.
- * Development of a pilot dataset for LLM evaluation in this domain.
Purpose of the Study:
- * To create and present a benchmark dataset for evaluating LLM performance on histological samples.
- * To assess LLMs on multi-dimensional classification tasks including species, organ, staining, and preparation type.
- * To address the need for standardized evaluation metrics in AI-driven histopathology.
Main Methods:
- * Evaluation of three LLMs (GPT-4.1, GPT-4o-mini, Llama 3.2) on 378 preclinical histological samples.
- * Four classification dimensions: species (mouse, rabbit, rat), organ, staining method, and preparation type (frozen vs. paraffin-embedded).
- * Performance assessment using sensitivity, specificity, and confusion matrix analysis.
Main Results:
- * Substantial variation in LLM performance across tasks and high sensitivity to class imbalance.
- * GPT-4.1 showed balanced preparation type classification; Llama 3.2 struggled with paraffin samples.
- * Llama 3.2 identified all species but had poor mouse recognition; GPT-4.1 excelled in mouse identification.
- * Llama 3.2 demonstrated high staining classification performance; GPT-4o-mini achieved perfect H&E recognition.
Conclusions:
- * Current LLMs exhibit variable performance in histological classification, highly sensitive to class imbalance.
- * LLMs are not suitable for standalone diagnostic use in histopathology.
- * LLMs show potential as screening tools in research settings with human oversight.
More Related Videos
09:06Whole-brain Segmentation and Change-point Analysis of Anatomical Brain MRI—Application in Premanifest Huntington's Disease
Published on: June 9, 2018
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Genetic Lingo
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...