Related Experiment Video
Updated: Jun 27, 2026

10:17
Guidelines and Experience Using Imaging Biomarker Explorer IBEX for Radiomics
Published on: January 8, 2018
13.6K
Evaluating methodological quality in radiomics research using large language models: Added value of METRICS-E3
1Department of Radiology, Uskudar State Hospital, Istanbul, Turkey.
European Journal of Radiology
|November 22, 2025
Summary
The METRICS-E3 resource enhanced a large language model's (LLM) ability to assess radiomics research quality, improving agreement with human experts. This LLM-assisted appraisal shows promise as a scalable tool for pre-screening and auditing.
Area of Science:
- Medical Imaging and Artificial Intelligence
- Radiomics Research Methodology
- Natural Language Processing in Science
Background:
- Assessing the methodological quality of radiomics research is crucial for reliable clinical translation.
- Large language models (LLMs) show potential for automating research appraisal, but their accuracy needs improvement.
- The METhodological RadiomICs Score (METRICS) is a tool for evaluating radiomics study quality.
Purpose of the Study:
- To evaluate if the METRICS-E3 (Explanation and Elaboration with Examples) resource enhances LLM performance in appraising radiomics research quality using the METRICS framework.
- To compare LLM performance with and without METRICS-E3 against human consensus and other LLM benchmarks.
Main Methods:
- A meta-research study assessed 48 radiomics articles using the METRICS tool with two GPT-5 pipelines: baseline (GPT-5) and enhanced with METRICS-E3 (GPT-5 E3).
- Each article was evaluated three times per pipeline, with results aggregated using majority voting.
- GPT-5 and GPT-5 E3 results were compared against previously reported GPT-4o data and a reference human consensus.
Main Results:
- While GPT-4o achieved the highest median METRICS score (79.50%), GPT-5 E3 showed improved concordance with human ratings (Kendall's τ increased from 0.474 to 0.626) compared to baseline GPT-5.
- Agreement, measured by intra-class correlation coefficient, increased from 0.539 (GPT-5) to 0.793 (GPT-5 E3).
- GPT-5 E3 demonstrated improved agreement on 83.3% of items and 80% of conditions assessed by METRICS.
Conclusions:
- The METRICS-E3 resource significantly improved GPT-5's alignment with human consensus and reliability in assessing radiomics methodological quality.
- LLM-assisted appraisal using METRICS and METRICS-E3 can serve as a scalable aid for pre-screening or auditing under expert supervision.
- Further validation across diverse datasets, human readers, and LLM architectures is necessary to confirm these findings.
Related Concept Videos
Methods to Assess Microbial Communities
Microbial communities, comprising bacteria, archaea, and eukaryotic microorganisms, inhabit diverse ecosystems and play crucial roles in environmental and biological processes. Their diversity is defined by three main parameters: species richness (the number of distinct species), species abundance (the relative quantity of each species), and species evenness (how uniformly individual species are distributed in various locations). These factors together shape the structure and ecological balance...
Methods of Medium Optimization
Optimizing growth media enhances microbial proliferation and maximizes product yield. Statistical experimental design methodologies provide structured and reproducible approaches, offering progressively higher levels of robustness and efficiency.The One-Factor-at-a-Time (OFAT) MethodThe One-Factor-at-a-Time (OFAT) method involves adjusting a single variable while keeping all others constant. However, it cannot detect interactions between variables, often leading to suboptimal outcomes when...

