Related Experiment Video
Updated: Jan 30, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Prediction using step-wise L1, L2 regularization and feature selection for small data sets with large number of
Ozgur Demir-Kavuk1, Mayumi Kamada, Tatsuya Akutsu
1Institute of Chemistry and Biochemistry, Freie Universität Berlin, Fabeckstrasse 36A, 14195 Berlin, Germany.
This study introduces a two-step machine learning approach to improve biological predictions. It effectively selects relevant molecular features from large datasets, enabling accurate modeling even with limited training data.
Area of Science:
- Computational Biology
- Cheminformatics
- Machine Learning
Background:
- Machine learning (ML) is widely used for biological predictions, but limited training data and numerous molecular descriptors pose challenges.
- Developing accurate predictive models requires careful feature selection when dealing with complex biological data.
Purpose of the Study:
- To develop and evaluate a robust two-step ML method for biological regression tasks with limited training data and extensive features.
- To address the challenge of high-dimensional feature spaces in biological prediction problems.
Main Methods:
- A two-step feature selection and model training procedure was applied to four biological regression datasets.
- The first stage utilized L1 regularization for optimized feature selection, reducing thousands of descriptors to approximately 50.
- The second stage employed L2 regularization on selected features, coupled with a soft loss function to mitigate outlier influence.
Main Results:
- The method successfully reduced over 6,000 features to around 50 in the initial stage.
- The two-step approach demonstrated improved prediction performance across all four CoEPrA regression tasks.
- The use of L2 regularization in the second stage further enhanced predictive accuracy.
Conclusions:
- The proposed two-step ML method is effective for biological prediction problems with limited sample sizes and high-dimensional descriptor spaces.
- This approach offers a valuable tool for drug discovery and other areas requiring accurate biological predictions from complex datasets.
More Related Videos
Related Concept Videos
Pericarditis II: Clinical Features and Diagnostic Tests
Endocarditis II: Clinical Features of Infective Endocarditis
COPD: Pathogenesis and Clinical Features
The primary cause for the onset of COPD is cigarette smoking and exposure to air pollution. These hazardous factors initiate a chain reaction within the lungs, resulting in chronic inflammation, damage to the airways, and a...
Special Features of Adaptive Immunity
The primary cell types involved in adaptive immunity are T cells and B cells. Each type has a unique role in defending the body against pathogens. T cells are responsible for cell-mediated immunity. They identify and eliminate infected cells directly,...
Esophageal Strictures-II: Clinical Features and Management
Healthcare providers should gather a comprehensive medical history and conduct a physical examination for diagnosis. If esophageal stricture is...
Esophageal Varices-II: Clinical Features and Management
In the initial assessment, a thorough review of the patient's medical history is vital to identify risk factors such as liver disease, alcohol...

