Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Reporting checklist for foundation and large language models in medical research (REFINE): an international consensus guideline.

Diagnostic and interventional radiology (Ankara, Turkey)·2026
Same author

Soft Annotations versus Pixel-Based Segmentation Masks of Prostate Anatomies: The Effect of Annotation Type on Radiomics.

Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference·2025
Same author

Scalable Clinical Annotation with Location Evidence (SCALE).

Computers in biology and medicine·2025
Same author

Self-supervised learning leads to improved performance in biparametric prostate MRI classification.

Computers in biology and medicine·2025
Same author

A Review of Methods for Trustworthy AI in Medical Imaging: The FUTURE-AI Guidelines.

IEEE journal of biomedical and health informatics·2025
Same author

Telomere attrition becomes an instrument for clonal selection in aging hematopoiesis and leukemogenesis.

Nature genetics·2025

Related Experiment Video

Updated: Sep 11, 2025

Guidelines and Experience Using Imaging Biomarker Explorer IBEX for Radiomics
10:17

Guidelines and Experience Using Imaging Biomarker Explorer IBEX for Radiomics

Published on: January 8, 2018

13.3K

Auto-METRICS: LLM-assisted scientific quality control for radiomics research.

José Guilherme de Almeida1, Nickolas Papanikolaou2

  • 1Champalimaud Foundation, Lisbon, Portugal.

European Journal of Radiology
|August 14, 2025
PubMed
Summary

Large language models (LLMs) show promise in assessing radiomics research quality, achieving agreement comparable to human radiologists. Privacy-preserving open LLMs perform similarly to commercial models, supporting research integrity.

Keywords:
Artificial intelligenceLarge language modelsMETRICSRadiomics

More Related Videos

Industrialized, Artificial Intelligence-guided Laser Microdissection for Microscaled Proteomic Analysis of the Tumor Microenvironment
13:01

Industrialized, Artificial Intelligence-guided Laser Microdissection for Microscaled Proteomic Analysis of the Tumor Microenvironment

Published on: June 3, 2022

3.9K
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.9K

Related Experiment Videos

Last Updated: Sep 11, 2025

Guidelines and Experience Using Imaging Biomarker Explorer IBEX for Radiomics
10:17

Guidelines and Experience Using Imaging Biomarker Explorer IBEX for Radiomics

Published on: January 8, 2018

13.3K
Industrialized, Artificial Intelligence-guided Laser Microdissection for Microscaled Proteomic Analysis of the Tumor Microenvironment
13:01

Industrialized, Artificial Intelligence-guided Laser Microdissection for Microscaled Proteomic Analysis of the Tumor Microenvironment

Published on: June 3, 2022

3.9K
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.9K

Area of Science:

  • Medical Imaging
  • Artificial Intelligence
  • Radiomics

Background:

  • Radiomics research quality is crucial for clinical translation but often suffers from methodological flaws.
  • Standardized assessment tools like the METhodological RadiomICs Score (METRICS) are needed to ensure research integrity.

Purpose of the Study:

  • To evaluate the reliability of large language models (LLMs) in assessing radiomics methodological quality using the METRICS framework.
  • To compare LLM performance against human radiologists in a reproducibility setting.

Main Methods:

  • A commercial LLM (Gemini Flash 2.0) and 24 open LLMs were used to assess the METRICS score for 46 radiomics articles.
  • Assessments were compared with those of radiologists in two reproducibility studies (ADA2025 and K2025).
  • Inter-rater reliability (Cohen's kappa) and scoring correlation (Pearson's correlation) were analyzed.

Main Results:

  • The commercial LLM demonstrated inter-rater agreement with human radiologists comparable to human-human agreement in both studies (average kappa ≈ 0.48-0.58).
  • METRICS scoring correlation between the LLM and human raters was also comparable to human-human correlation (average PC ≈ 0.62-0.68).
  • An open, privacy-preserving LLM (Phi4-Reasoning) performed comparably to the commercial LLM.

Conclusions:

  • LLMs can effectively assist in the standardized quality assessment of radiomics research.
  • Open-source, privacy-preserving LLMs offer a viable alternative to commercial models for supporting radiomics research integrity evaluation.