Related Experiment Video
Updated: May 26, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Leveraging domain information to restructure biological prediction.
Xiaofei Nan1, Gang Fu, Zhengdong Zhao
1Department of Computer and Information Science, University of Mississippi, USA.
This study introduces a new algorithm to identify discrete attributes that simplify machine learning tasks. The method uses conditional entropy to select attributes, improving prediction performance by partitioning complex problems into simpler ones.
Area of Science:
- Machine Learning
- Data Science
- Computational Statistics
Background:
- Incorporating domain knowledge into prediction models is crucial but challenging.
- Discrete or categorical attributes can partition problem domains into simpler sub-problems.
- Identifying attributes that effectively simplify learning tasks is a key goal.
Purpose of the Study:
- Develop an algorithm to identify discrete/categorical attributes that maximally simplify supervised learning tasks.
- Propose a metric to rank attributes based on their potential to reduce classification uncertainty.
- Enhance prediction performance by restructuring learning problems.
Main Methods:
- Restructuring supervised learning problems using attribute-based partitions.
- Proposing a conditional entropy metric to quantify attribute effectiveness.
- Approximating solutions using expected minimum conditional entropy with random projections to manage computational cost.
Main Results:
- The proposed metric effectively ranks attributes for problem simplification.
- Tested on artificial, cheminformatics, and gene expression datasets.
- Classifiers built on restructured problems consistently outperform those on the original problem.
Conclusions:
- The conditional entropy metric successfully identifies effective partitions for classification problems.
- This approach enhances prediction performance by simplifying the learning task.
- The method provides a computationally feasible way to leverage domain knowledge encoded in discrete attributes.
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Genome Annotation and Assembly
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Evolutionary Relationships through Genome Comparisons
