Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

15.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.7K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.8K
3.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

GlyGen: a knowledgebase linking glycan data with protein and gene data to reveal novel biological connections.

Research square·2026
Same author

Knowledge Graph-Driven AI in Biohealth: From Biomedical Discovery to Health Risk Prediction.

Delaware journal of public health·2026
Same author

Enhanced Adverse-Event Detection and Drug-Event Relation Extraction from Clinical Notes.

medRxiv : the preprint server for health sciences·2026
Same author

A network-centric approach reveals novel pathways impacted by Prader-Willi Syndrome.

PloS one·2026
Same author

The Common Fund Data Ecosystem (CFDE).

bioRxiv : the preprint server for biology·2026
Same author

Desiderata for a biomedical knowledge network: opportunities, challenges and future directions.

Bioinformatics advances·2026

Related Experiment Video

Updated: Apr 19, 2026

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.5K

Pretraining effective T5 generative models for clinical and biomedical applications.

Saad Althabiti1,2,3, Chuming Chen4,5, Sultan Alrowili6

  • 1Department of Health Informatics, King Saud bin Abdulaziz University for Health Sciences, Riyadh, Saudi Arabia.

Plos One
|April 17, 2026
PubMed
Summary

Selecting the right pretraining corpus and vocabulary is crucial for optimizing T5 language models in clinical and biomedical natural language processing (NLP). Domain-specific data significantly boosts performance on targeted tasks.

Related Experiment Videos

Last Updated: Apr 19, 2026

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.5K

Area of Science:

  • Natural Language Processing
  • Computational Linguistics
  • Medical Informatics

Background:

  • T5 language models show promise in various domains.
  • Optimizing models for specialized fields like clinical and biomedical text requires careful consideration of training data and vocabulary.
  • Previous research has not comprehensively evaluated the impact of specific corpus and vocabulary choices on T5 performance in these domains.

Purpose of the Study:

  • To investigate how corpus selection and vocabulary design affect the performance of T5 language models in clinical and biomedical domains.
  • To quantify the impact of pretraining data and tokenization choices on downstream task performance.
  • To compare the effectiveness of domain-specific versus general corpora and vocabularies.

Main Methods:

  • Introduced five T5-EHR models pretrained from scratch with varied clinical and biomedical corpora and vocabularies.
  • Evaluated model performance across diverse clinical and biomedical tasks.
  • Analyzed the influence of pretraining data composition and vocabulary tokenization on downstream results.

Main Results:

  • Models pretrained exclusively on clinical data outperformed others on clinical tasks, with limited gains from adding biomedical data.
  • Clinical-specific vocabularies significantly improved performance on clinical language tasks compared to general biomedical vocabularies.
  • T5 generative models demonstrated competitive performance against state-of-the-art discriminative models on biomedical benchmarks.

Conclusions:

  • Aligning pretraining corpus and vocabulary with the target domain is essential for optimal T5 model performance in clinical and biomedical NLP.
  • Task-specific data selection is critical for maximizing model effectiveness.
  • T5 models exhibit strong generalization capabilities within the biomedical domain.