Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Survival Tree01:19

Survival Tree

514
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
514

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Evaluating the role of pretraining dataset size and diversity on single-cell foundation model performance.

Nature methods·2026
Same author

Gemcitabine and nab-paclitaxel with or without the VDR agonist paricalcitol for metastatic pancreatic cancer: a randomized, multiarm, run-in phase trial.

Nature cancer·2026
Same author

Evidence-based classification of genes implicated in craniosynostosis disorders using the ClinGen curation framework.

Genetics in medicine : official journal of the American College of Medical Genetics·2026
Same author

Tackling the complexity of cancer with generative models.

Cell·2026
Same author

Unlocking <i>KRAS</i>: Navigating Its Molecular Biology and Treatment Landscape Among Gastrointestinal Malignancies.

Current oncology (Toronto, Ont.)·2026
Same author

Scalable nonparametric clustering with unified marker gene selection for single-cell RNA-seq data.

Cell reports methods·2026

Related Experiment Video

Updated: May 6, 2026

Discrimination and Characterization of Heterocellular Populations Using Quantitative Imaging Techniques
09:48

Discrimination and Characterization of Heterocellular Populations Using Quantitative Imaging Techniques

Published on: June 30, 2017

7.4K

Consequences of training data composition for deep learning models in single-cell biology.

Ajay Nadig1,2,3, Akshaya Thoutam4, Madeline Hughes4

  • 1Harvard Medical School, Boston, MA, USA.

Biorxiv : the Preprint Server for Biology
|March 10, 2025
PubMed
Summary

Foundation models for single-cell transcriptomics need diverse training data for better performance. Optimizing datasets improves generalization to new cell types and disease states.

More Related Videos

Analyzing Mitochondrial Morphology Through Simulation Supervised Learning
12:06

Analyzing Mitochondrial Morphology Through Simulation Supervised Learning

Published on: March 3, 2023

3.9K
Author Spotlight: Enhancing PSC-to-Functional Cell Differentiation Using ML Models Based on Live-Cell Bright-Field Imaging
11:38

Author Spotlight: Enhancing PSC-to-Functional Cell Differentiation Using ML Models Based on Live-Cell Bright-Field Imaging

Published on: October 4, 2024

457

Related Experiment Videos

Last Updated: May 6, 2026

Discrimination and Characterization of Heterocellular Populations Using Quantitative Imaging Techniques
09:48

Discrimination and Characterization of Heterocellular Populations Using Quantitative Imaging Techniques

Published on: June 30, 2017

7.4K
Analyzing Mitochondrial Morphology Through Simulation Supervised Learning
12:06

Analyzing Mitochondrial Morphology Through Simulation Supervised Learning

Published on: March 3, 2023

3.9K
Author Spotlight: Enhancing PSC-to-Functional Cell Differentiation Using ML Models Based on Live-Cell Bright-Field Imaging
11:38

Author Spotlight: Enhancing PSC-to-Functional Cell Differentiation Using ML Models Based on Live-Cell Bright-Field Imaging

Published on: October 4, 2024

457

Area of Science:

  • Computational biology
  • Genomics
  • Machine learning

Background:

  • Foundation models offer potential for single-cell transcriptomics analyses, especially with sparse data.
  • Current single-cell foundation models train on large corpora, often overlooking training data composition's impact on performance.
  • Large language model research highlights the critical role of training data composition in shaping model behavior.

Purpose of the Study:

  • To systematically investigate how training dataset composition affects deep learning models for single-cell transcriptomics.
  • To evaluate model generalization to unseen cell types and disease states.
  • To identify strategies for optimizing future single-cell foundation models.

Main Methods:

  • Focused on human hematopoiesis as a model system.
  • Included diverse cell types from adult and developing tissues, disease states, and perturbation atlases.
  • Systematically varied training dataset composition to assess model behavior.

Main Results:

  • Models demonstrated poor generalization to unseen cell types.
  • Incorporating malignant cells into training data did not consistently improve modeling of unseen malignant cells.
  • Including embryonic stem cell differentiation data enhanced performance on out-of-distribution tasks.

Conclusions:

  • Training data diversity is crucial for effective single-cell foundation models.
  • Current training strategies may lead to poor generalization.
  • Future models should prioritize curated, diverse datasets for improved performance and broader applicability.