Related Experiment Video
Updated: Jan 12, 2026

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
Reasoning Models for Text Mining in Oncology: A Comparison Between o1 Preview, GPT-4o, and GPT-5 at Different
Paul Windisch1,2, Fabio Dennstädt2, Julia Weyrich1,2
1Department of Radiation Oncology, Cantonal Hospital Winterthur, Winterthur, Switzerland.
Purpose:
Chain-of-thought prompting is a method to make large language models generate intermediate reasoning steps when solving a complex problem. OpenAI's o1 preview and GPT-5 have been trained to create such a chain of thought internally before giving a response and have been claimed to surpass various benchmarks requiring complex reasoning. The purpose of this study was to evaluate their performance in text mining in oncology.
Methods:
Six hundred trials from high-impact medical journals were classified depending on whether they allowed for the inclusion of patients with localized and/or metastatic disease. GPT-4o, o1 preview, and GPT-5 at different reasoning effort settings were instructed to do the same classification based on the publications' abstracts.
Results:
For predicting whether patients with localized disease were enrolled, GPT-4o and o1 preview achieved F1 scores of 0.80 (0.76-0.83) and 0.91 (0.89-0.94), respectively. For predicting whether patients with metastatic disease were enrolled, GPT-4o and o1 preview achieved F1 scores of 0.97 (0.95-0.98) and 0.99 (0.99-1.00), respectively. For GPT-5, the F1 scores for predicting the eligibility of patients with localized disease increased from 0.84 to 0.93 and 0.94 with increased reasoning effort. F1 scores for metastatic disease were 0.97, 0.99, and 0.99.
Conclusion:
o1 preview outperformed GPT-4o in extracting if people with localized and/or metastatic disease were eligible for a trial from its abstract. GPT-5 at high reasoning effort settings outperformed both GPT-4o and o1 preview, supporting the notion that reasoning models could become the new standard for text mining in medicine.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Genomics
Tumor Progression
Colon cancer is one of the best-documented examples of tumor progression. Early mutation in the APC gene in colon cells causes a small growth on the colon wall called a polyp. With time, this polyp grows into a benign, pre-cancerous tumor. Further...
Deductive Reasoning
For example, a researcher can deduce specific predictions...