Related Experiment Video
Updated: Aug 8, 2026

Guidelines and Experience Using Imaging Biomarker Explorer IBEX for Radiomics
Published on: January 8, 2018
Auto-METRICS: LLM-assisted scientific quality control for radiomics research
José Guilherme de Almeida1, Nickolas Papanikolaou2
1Champalimaud Foundation, Lisbon, Portugal.
Purpose:
The quality of radiomics research is critical for reliable clinical translation, yet methodological flaws remain prevalent. This study evaluates whether large language models (LLMs) can reliably assess radiomics methodological quality using the METhodological RadiomICs Score (METRICS).
Methods:
We compared a commercial cloud-based LLM (Gemini Flash 2.0) METRICS assessments for 46 articles with those of radiologists using two reproducibility studies (ADA2025 and K2025, with 6 radiologist groups and 3 radiologists, respectively, with varying degrees of experience). Cohen's kappa (κ) and METRICS Pearson's correlation (PC), and error rates between LLMs and human raters were evaluated. Prompt clarifications to METRICS were suggested to improve human-LLM agreement. Twenty four privacy-preserving open LLMs were compared with Gemini Flash 2.0.
Results:
In ADA2025, the commercial LLM achieved inter-rater agreements with human raters comparable to those between human raters (average κ = 0.48 vs. average κ = 0.48, respectively, Wilcoxon rank-sum test p = 0.41), leading to similar correlation values in METRICS scoring (average PC = 0.62 vs. average PC = 0.56, Wilcoxon rank-sum test p = 0.11). This was confirmed with K2025 (mean human-LLM κ = 0.58 vs. human-human κ = 0.57, Wilcoxon rank-sum test p = 0.28), with no evidence for correlation differences (PC = 0.68 vs. PC = 0.51, respectively, Wilcoxon rank-sum test p = 0.55). Phi4-Reasoning, an open model which can be run locally, performed comparably to Gemini Flash 2.0 (median ranking = 1 vs. median ranking = 3, respectively, across all raters).
Conclusion:
LLMs can assist in standardized radiomics quality assessment. Open privacy-preserving models can offer comparable performance to commercial cloud-based LLMs, suggesting their utility in supporting human raters for evaluating radiomics research integrity.
More Related Videos
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
13:01Industrialized, Artificial Intelligence-guided Laser Microdissection for Microscaled Proteomic Analysis of the Tumor Microenvironment
Published on: June 3, 2022
Related Concept Videos
Instrument Calibration
Analytical Balance Calibration
An analytical balance measures mass and requires regular calibration to...
Controlled-Current Coulometry: Overview
Control Systems
At the heart...
Electronic Distance Measuring Instruments
Leveling Equipment