Related Experiment Video
Updated: Jan 12, 2026

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
Reasoning Models for Text Mining in Oncology: A Comparison Between o1 Preview, GPT-4o, and GPT-5 at Different
Paul Windisch1,2, Fabio Dennstädt2, Julia Weyrich1,2
1Department of Radiation Oncology, Cantonal Hospital Winterthur, Winterthur, Switzerland.
New AI models, including GPT-5 with enhanced reasoning, show superior performance in oncology text mining. These advanced language models accurately extract patient eligibility criteria from clinical trial abstracts, potentially setting a new standard for medical research.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing
- Oncology Research
Background:
- Chain-of-thought prompting enables large language models (LLMs) to generate intermediate reasoning steps for complex problem-solving.
- Recent advancements in LLMs like OpenAI's o1 preview and GPT-5 incorporate internal chain-of-thought capabilities, claiming improved performance on complex reasoning tasks.
Purpose of the Study:
- To evaluate the performance of GPT-4o, o1 preview, and GPT-5 in text mining specifically within the field of oncology.
- To assess the accuracy of these LLMs in classifying clinical trial eligibility based on abstracts.
Main Methods:
- A dataset of 600 clinical trials from high-impact medical journals was curated.
- Trials were classified based on the inclusion of patients with localized and/or metastatic disease.
- GPT-4o, o1 preview, and GPT-5 were tasked with classifying trial eligibility using abstracts, with GPT-5 tested at varying reasoning effort settings.
Main Results:
- GPT-4o achieved F1 scores of 0.80 for localized disease and 0.97 for metastatic disease.
- o1 preview demonstrated higher performance with F1 scores of 0.91 for localized and 0.99 for metastatic disease.
- GPT-5, with increased reasoning effort, reached F1 scores of up to 0.94 for localized and 0.99 for metastatic disease, outperforming the other models.
Conclusions:
- o1 preview surpassed GPT-4o in accurately identifying trial eligibility for patients with localized and/or metastatic disease from abstracts.
- GPT-5, particularly at high reasoning settings, demonstrated superior performance compared to both GPT-4o and o1 preview.
- The findings suggest that advanced reasoning models have the potential to become the new standard for text mining applications in medicine.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Genomics
Tumor Progression
Colon cancer is one of the best-documented examples of tumor progression. Early mutation in the APC gene in colon cells causes a small growth on the colon wall called a polyp. With time, this polyp grows into a benign, pre-cancerous tumor. Further...
Deductive Reasoning
For example, a researcher can deduce specific predictions...