Related Experiment Video
Updated: Jan 9, 2026

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
Systemic Anticancer Therapy Timelines Extraction From Electronic Medical Records Text: Algorithm Development and
Jiarui Yao1, Eli Goldner1, Harry Hochheiser2
1Computational Health Informatics Program, Boston Children's Hospital, Harvard Medical School, 401 Park Drive, Boston, MA, 02115, United States, 1 7813545014.
Background:
The systemic treatment of cancer typically requires the use of multiple anticancer agents in combination or sequentially. Clinical narrative texts often contain extensive descriptions of the temporal sequencing of systemic anticancer therapy (SACT), setting up an important task that may be amenable to automated extraction of SACT timelines.
Objective:
We aimed to explore automatic methods for extracting patient-level SACT timelines from clinical narratives in the electronic medical records (EMRs).
Methods:
We used two datasets from two institutions: (1) a colorectal cancer (CRC) dataset including the entire EMR of the 199 patients in the THYME (Temporal Histories of Your Medical Event) dataset and (2) the 2024 ChemoTimelines shared task dataset including 149 patients with ovarian cancer, breast cancer, and melanoma. We explored finetuning smaller language models trained to attend to events and time expressions, and few-shot prompting of large language models (LLMs). Evaluation used the 2024 ChemoTimelines shared task configuration-Subtask1 involving the construction of SACT timelines from manually annotated SACT event and time expression mentions provided as input in addition to the patient's notes and Subtask2 requiring extraction of SACT timelines directly from the patient's notes.
Results:
Our task-specific finetuned EntityBERT model achieved 93% F1-score, outperforming the best results in Subtask1 of the 2024 ChemoTimelines shared task (90%). It ranked second in Subtask2. LLM (LLaMA2, LLaMA3.1, and Mixtral) performance lagged the task-specific finetuned model performance for both the THYME and shared task datasets. On the shared task datasets, the best LLM performance was 77% macro F1-score, 16% points lower than the task-specific finetuned system (Subtask1).
Conclusions:
In this paper, we explored approaches for patient-level timeline extraction through the SACT timeline extraction task. Our results and analysis add to the knowledge of extracting treatment timelines from EMR clinical narratives using language modeling methods.
Insights
Automated extraction of systemic anticancer therapy (SACT) timelines from electronic medical records (EMRs) is crucial. A finetuned EntityBERT model achieved 93% F1-score, outperforming large language models for SACT timeline extraction.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Bioinformatics
Background:
- Systemic anticancer therapy (SACT) often involves complex drug combinations and sequences.
- Clinical narratives in electronic medical records (EMRs) contain detailed SACT timelines.
- Automated extraction of these timelines is a significant challenge.
Purpose of the Study:
- To explore automatic methods for extracting patient-level SACT timelines from clinical narratives in EMRs.
- To compare the performance of finetuned language models and large language models (LLMs) for this task.
Main Methods:
- Utilized two datasets: THYME (colorectal cancer) and ChemoTimelines shared task (ovarian, breast cancer, melanoma).
- Explored finetuning smaller language models (EntityBERT) and few-shot prompting of LLMs (LLaMA, Mixtral).
- Evaluated performance on Subtask1 (timeline construction from annotated input) and Subtask2 (direct extraction from notes).
Main Results:
- The finetuned EntityBERT model achieved a 93% F1-score, surpassing the best Subtask1 result (90%) in the ChemoTimelines shared task.
- EntityBERT ranked second in Subtask2.
- LLMs (LLaMA2, LLaMA3.1, Mixtral) underperformed the finetuned model, with the best LLM achieving 77% macro F1-score on shared task datasets (Subtask1).
Conclusions:
- Task-specific finetuning of language models, like EntityBERT, is highly effective for extracting SACT timelines from clinical narratives.
- This approach outperforms general-purpose LLMs for this specialized task.
- The findings contribute to advancing automated treatment timeline extraction from EMRs.
More Related Videos
11:18Generation of Comprehensive Thoracic Oncology Database - Tool for Translational Research
Published on: January 22, 2011
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Related Concept Videos
Cancer Survival Analysis
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Targeted Cancer Therapies
There are several types of targeted therapies against...