Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Oncogene inactivation-induced senescence facilitates tumor relapse.

Nature communications·2026
Same author

Deep Learning-Based Analysis of Gene Expression Data and Gene-Related Information in Pediatric Surgical Oncology: A Scoping Review.

Cancer medicine·2026
Same author

Altered cholesterol immunometabolism activates the macrophage NLRP3-inflammasome in lung fibrosis.

American journal of respiratory cell and molecular biology·2026
Same author

Multiplexed biomarkers dynamically detect heterogeneous residual neuroblastoma cell clone activity in the bone marrow niche.

Cancer letters·2026
Same author

Flexynesis: A deep learning toolkit for bulk multi-omics data integration for precision oncology and beyond.

Nature communications·2025
Same author

Evaluating large language model workflows in clinical decision support for triage and referral and diagnosis.

NPJ digital medicine·2025

Related Experiment Video

Updated: May 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

462

Enhancing biomarker based oncology trial matching using large language models.

Nour Alkhoury1, Maqsood Shaik1, Ricardo Wurmus1

  • 1Berlin Institute for Medical Systems Biology (BIMSB), Max Delbrück Center for Molecular Medicine, Berlin, Germany.

NPJ Digital Medicine
|May 5, 2025
PubMed
Summary

Open-source language models effectively extract genomic biomarkers from clinical trial data, outperforming closed-source alternatives. Fine-tuning further improves their ability to structure this crucial information for cancer drug development.

More Related Videos

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
07:41

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases

Published on: May 17, 2019

8.8K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

607

Related Experiment Videos

Last Updated: May 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

462
Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
07:41

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases

Published on: May 17, 2019

8.8K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

607

Area of Science:

  • Oncology
  • Bioinformatics
  • Natural Language Processing

Background:

  • Clinical trial enrollment relies on patient eligibility criteria often found in unstructured text.
  • Genomic biomarkers are vital for precision medicine and targeted cancer therapies.
  • Matching patients to clinical trials requires efficient extraction of eligibility information.

Purpose of the Study:

  • To explore strategies for extracting genetic biomarkers from oncology clinical trial descriptions.
  • To structure unstructured clinical trial data for improved patient matching.
  • To evaluate the performance of large language models (LLMs) in this task.

Main Methods:

  • Utilized open-source and closed-source large language models (LLMs) to process clinical trial study descriptions.
  • Focused on extracting and structuring genomic biomarker information from eligibility criteria.
  • Compared out-of-the-box LLM performance with fine-tuned models.

Main Results:

  • Open-source LLMs demonstrated effectiveness in capturing complex logical expressions and structuring genomic biomarkers.
  • Out-of-the-box open-source models outperformed closed-source models like GPT-4.
  • Fine-tuning open-source models with additional data led to significant performance enhancements.

Conclusions:

  • Open-source LLMs are a viable solution for structuring unstructured clinical trial data, specifically for genomic biomarkers.
  • These models aid in identifying suitable patients for precision oncology trials.
  • Further development and fine-tuning can optimize LLM performance for clinical trial data extraction.