Related Experiment Video
Updated: Jan 23, 2026

A Fine Motor Task to Study Joint Kinematics in a Preclinical Model of Neurodegenerative Disease
Published on: June 13, 2025
BioTriplex: a full-text annotated corpus for fine-tuning language models in gene-disease relation extraction tasks
Charlotte Collins1,2, Panagiotis Fytas1, İlknur Karadeniz1,3
1Language Technology Laboratory, Theoretical and Applied Linguistics, Faculty of Modern and Medieval Languages and Linguistics, University of Cambridge, Cambridge CB3 9DA, United Kingdom.
We created BioTriplex, a new dataset for training AI models on gene-disease relationships. Fine-tuned AI models using BioTriplex show improved accuracy in extracting complex gene-disease connections from biomedical texts.
Area of Science:
- Biomedical informatics
- Computational biology
- Natural Language Processing
Background:
- Machine learning for biomedical text analysis requires specialized models due to technical terminology and semantic variations.
- Existing large language models (LLMs) struggle with the nuances of biomedical literature.
- There is a need for annotated full-text datasets to fine-tune LLMs for specific biomedical applications, particularly gene-disease relationship extraction.
Purpose of the Study:
- To introduce BioTriplex, a novel annotated corpus of full-length biomedical research articles.
- To develop and evaluate a fine-tuned language model for extracting gene-disease relationships.
- To demonstrate the effectiveness of specialized datasets in improving LLM performance for biomedical tasks.
Main Methods:
- Developed BioTriplex, a corpus of 100 full-length biomedical articles manually annotated with diseases, genes, and 21 subtypes of gene-disease relationships.
- Utilized BioTriplex to fine-tune the LLaMA 3.1 8B language model for gene-disease relation extraction.
- Compared the performance of the fine-tuned model against zero-shot and few-shot approaches within the same architecture and against other state-of-the-art LLMs (GPT-4, Claude Sonnet 3.7).
Main Results:
- The BioTriplex-trained LLaMA 3.1 8B model significantly outperformed baseline and other advanced LLMs in gene-disease relation extraction.
- The fine-tuned model demonstrated superior accuracy and a broader, more granular classification of gene-disease relationship types.
- BioTriplex proved to be a valuable resource for enhancing LLM capabilities in biomedical information extraction.
Conclusions:
- BioTriplex is an effective resource for fine-tuning language models for biomedical applications.
- Specialized datasets are crucial for improving the performance of LLMs in complex scientific domains like gene-disease relationship extraction.
- The fine-tuned LLaMA 3.1 8B model represents a significant advancement in automated extraction of gene-disease relationships.
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Genome Annotation and Assembly
Fineness of Cement
Direct...
Fineness Modulus
Consider performing sieve analysis on sand through a set of ASTM sieves. The weight of aggregate retained in each sieve and pan placed at the bottom is recorded, as given in Column B of Table 1.
To determine the fineness modulus of...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

