Related Experiment Video
Updated: Aug 25, 2025

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Conformal prediction under feedback covariate shift for biomolecular design.
Clara Fannjiang1, Stephen Bates2, Anastasios N Angelopoulos1
1Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, CA 94720.
This study introduces a new method for quantifying prediction uncertainty in machine learning, crucial for iterative data collection and model training. It ensures reliable model performance even when test data selection depends on training data, aiding in protein design and algorithm selection.
Area of Science:
- Machine Learning
- Computational Biology
- Biotechnology
Background:
- Iterative machine learning protocols are common in scientific discovery, involving data collection, model training, and using model outputs to guide further data acquisition.
- In protein design, regression models predict sequence fitness, but validating sequences is costly, necessitating accurate uncertainty quantification for model predictions.
- A key challenge is distribution shift: test data (designed sequences) are statistically dependent on training data, complicating error estimation.
Purpose of the Study:
- To develop a method for constructing confidence sets that accurately quantify prediction uncertainty in iterative machine learning settings.
- To address the challenge of statistically dependent training and test data distributions common in design applications.
- To provide finite-sample guarantees for confidence sets applicable to any regression model, irrespective of test-time input distribution selection.
Main Methods:
- Introduced a novel method for constructing confidence sets for regression model predictions.
- Developed a technique to account for the statistical dependence between training and test data distributions.
- Ensured finite-sample guarantees for the proposed confidence sets, applicable across various regression models and input selection strategies.
Main Results:
- Demonstrated the ability of the method to quantify uncertainty in predicted protein sequence fitness using real datasets.
- Showcased how the confidence sets provide reliable uncertainty estimates even when test data is chosen based on model outputs.
- Validated the finite-sample guarantees of the confidence sets across different regression models and input selection schemes.
Conclusions:
- The developed method effectively quantifies prediction uncertainty in iterative machine learning, particularly in data-dependent settings like protein design.
- This uncertainty quantification enables informed selection of machine learning algorithms that balance predictive accuracy with confidence.
- The approach offers a robust framework for reliable decision-making in costly experimental validation scenarios, such as designing novel proteins.
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Predicting Reaction Outcomes
Behavioral Genetics and Its Designs
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
Improving Translational Accuracy
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Physiological Pharmacokinetic Models: Assumption with Protein Binding

