Related Experiment Video
Updated: Feb 12, 2026

Author Spotlight: Emerging Technologies and Advanced Tools for Decoding Metabolomics Data Analysis
Published on: November 10, 2023
Data-driven discovery of digital twins in biomedical research
Clémence Métayer1, Annabelle Ballesta1, Julien Martinelli2
1Inserm U1331, Institut Curie, PSL Research University, CBIO-Center for Computational Biology, Mines Paris, Cancer Systems Pharmacology team, Saint-Cloud 92210, France.
Automated methods are crucial for creating biological digital twins from complex data. Sparse regression shows promise, but hybrid approaches combining mechanistic models and deep learning are key for future advancements.
Area of Science:
- Systems Biology
- Computational Biology
- Biomedical Informatics
Background:
- High-throughput biological data enables digital twin development for personalized medicine.
- Current digital twin creation relies on manual data integration, necessitating automated methods.
- Biology presents unique challenges for data-driven discovery compared to physics.
Purpose of the Study:
- To review and evaluate methodologies for automated inference of biological digital twins from time-series data.
- To assess algorithms against key biological and methodological challenges.
- To propose future directions for developing robust digital twin inference methods.
Main Methods:
- Systematic review of 177 methodologies for inferring digital twins.
- Evaluation of algorithms based on eight challenges: data integration, multiple conditions, prior knowledge, latent variables, high dimensionality, unobserved variable derivatives, library design, and uncertainty quantification.
- Comparison of sparse regression, symbolic regression, deep learning, and large language models.
Main Results:
- Sparse regression generally outperformed symbolic regression, especially with Bayesian frameworks.
- Deep learning and large language models show potential for integrating prior knowledge but require reliability improvements.
- No single method currently addresses all challenges effectively.
Conclusions:
- Hybrid and modular frameworks are essential for advancing digital twin development.
- Future progress requires combining mechanistic grounding (chemical reaction networks), Bayesian uncertainty quantification, and deep learning's generative capabilities.
- Standardized benchmarks are needed to evaluate methods across all identified challenges.
Related Concept Videos
Drug Discovery: Overview
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
ATP Driven Pumps I: An Overview
There are four main types of ATP-driven pumps - P-type, V-type, F-type, and ABC transporter. All these pumps are of varying complexities and...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Data Reporting and Recording
Xylem and Transpiration-driven Transport of Resources

