Related Experiment Video
Updated: Dec 20, 2025

09:33
Author Spotlight: Finding New Therapeutic Targets for Malignant Peripheral Nerve Sheath Tumor Through Genome-Scale shRNA Screens
Published on: August 25, 2023
1.5K
parSMURF, a high-performance computing tool for the genome-wide detection of pathogenic variants.
Alessandro Petrini1, Marco Mesiti1, Max Schubach2,3
1Università degli Studi di Milano, AnacletoLab - Dipartimento di Informatica, via Giovanni Celoria 18, 20135 Milano, Italy.
Gigascience
|May 24, 2020
Summary
parSMURF is a novel computational biology tool that effectively handles large, imbalanced genomic datasets. This parallel machine learning method significantly speeds up the prediction of rare pathogenic variants in genomic medicine.
Area of Science:
- Genomic Medicine
- Computational Biology
- Machine Learning
Background:
- Genomic medicine and computational biology face challenges with big data and imbalanced datasets.
- Positive examples (e.g., pathogenic variants) are often a small fraction of the total data.
- Classical methods struggle to identify rare pathogenic variants and manage large genomic datasets.
Purpose of the Study:
- To develop a method that addresses big and imbalanced genomic data challenges.
- To improve the prediction of rare pathogenic variants in genomic medicine.
- To accelerate computational analysis of genomic data.
Main Methods:
- parSMURF utilizes a hyper-ensemble approach with oversampling and undersampling techniques.
- Parallel computational techniques (MPI, OpenMP) are employed for big data management and speed.
- Bayesian optimization facilitates efficient hyper-parameter tuning.
Main Results:
- parSMURF successfully manages big genomic data by partitioning it across computing nodes.
- Achieved state-of-the-art results on synthetic and real-world genomic datasets.
- Demonstrated an 80-fold speed-up compared to sequential versions.
Conclusions:
- parSMURF is a scalable parallel machine learning tool for big and imbalanced genomic data.
- Its parallelization enables efficient fitting of complex genomic problems.
- Available in C++ OpenMP and C++ MPI/OpenMP hybrid versions for diverse computing environments.
Keywords:
GWASMendelian diseasesensemble methodshigh-performance computinghigh-performance computing tool for genomic medicinemachine learning for genomic medicinemachine learning for imbalanced genomic dataparallel machine learning tool for big dataparallel machine learning tool for imbalanced dataprediction of deleterious or pathogenic variants
