Related Experiment Video
Updated: Jul 10, 2026

JUMPn: A Streamlined Application for Protein Co-Expression Clustering and Network Analysis in Proteomics
Published on: October 19, 2021
Emergent unsupervised clustering paradigms with potential application to bioinformatics.
David J Miller1, Yue Wang, George Kesidis
1Dept of Electrical Engineering, Pennsylvania State University, University Park, PA 16802, USA. djmiller@engr.psu.edu
Machine learning, including data clustering, is increasingly used in molecular biology for analyzing DNA microarray expression data. This review highlights key challenges and recent methods for advanced bioinformatics analysis, such as unsupervised feature selection and semisupervised learning.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Machine learning techniques, particularly data clustering, are widely applied in molecular biology.
- These methods analyze DNA microarray expression data to group genes or samples, aiding in understanding gene function, co-regulation, and disease pathways.
Purpose of the Study:
- To identify challenging, yet relevant, machine learning problems in bioinformatics that have been underexplored.
- To review recent methods addressing these specific challenges and discuss their suitability for bioinformatics applications.
Main Methods:
- Review of existing literature on machine learning applications in bioinformatics.
- Identification and discussion of advanced machine learning techniques relevant to biological data analysis.
Main Results:
- Several key areas requiring further investigation in bioinformatics machine learning were identified: unsupervised clustering with unsupervised feature selection, semisupervised learning, handling confounding variables in learning, and assessing clustering solution stability.
- Recent methods addressing these challenges show promise for common bioinformatics scenarios.
Conclusions:
- Advanced machine learning techniques, particularly those addressing unsupervised feature selection, semisupervised learning, confounding variables, and stability, are crucial for future bioinformatics research.
- These methods offer powerful solutions for complex biological data analysis and discovery.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Applications of Molecular Taxonomy
Modern Molecular Taxonomy
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
