Related Experiment Videos
Unbiased descriptor and parameter selection confirms the potential of proteochemometric modelling
Eva Freyhult1, Peteris Prusis, Maris Lapinsh
1The Linnaeus Centre for Bioinformatics, Uppsala University, Box 598, S-751 24 Uppsala, Sweden. Eva.Freyhult@lcb.uu.se
BMC Bioinformatics
|March 12, 2005
Summary
Proteochemometrics models predict protein function from interaction data. A double cross-validation loop provides unbiased performance estimates, revealing that single models can be misleading, especially with small datasets.
Area of Science:
- Computational chemistry
- Bioinformatics
- Cheminformatics
Background:
- Proteochemometrics predicts protein function using interaction data, bypassing 3D structures.
- Existing models offer insights into biomolecular interactions but lack rigorous statistical evaluation.
- Variable selection in current methods often uses all data, similar to gene expression analysis.
Purpose of the Study:
- To implement a methodology for unbiased evaluation of proteochemometric model predictive power.
- To assess the performance of proteochemometric models on large datasets.
- To address limitations in current statistical evaluation and interpretation of these models.
Main Methods:
- Developed and applied a double cross-validation (CV) loop procedure.
- Estimated expected performance of proteochemometric design methods.
- Utilized two large-scale proteochemometric datasets for evaluation.
Main Results:
- Unbiased performance estimates (P2) confirm useful predictive power of well-designed single proteochemometric models.
- Standard cross-validation may yield models with limited performance.
- Commercial software can produce misleading performance estimates; chemical interpretation of single models is uncertain with small datasets.
Conclusions:
- The double CV loop provides unbiased estimates, identifying non-predictive proteochemometric models.
- Chemical interpretations should be based on multiple models from the double CV loop, not single models.
- This approach enhances the reliability and interpretability of proteochemometric modeling.