Related Experiment Video
Updated: Jan 2, 2026

08:27
Large-Scale Multi-Omics Genome-Wide Association Studies Mo-GWAS: Guidelines for Sample Preparation and Normalization
Published on: July 27, 2021
4.7K
TSLRF: Two-Stage Algorithm Based on Least Angle Regression and Random Forest in genome-wide association studies.
Jiali Sun1, Qingtai Wu1, Dafeng Shen1
1College of Science, Nanjing Agricultural University, Nanjing, 210095, China.
Scientific Reports
|December 4, 2019
Summary
This study introduces a new two-stage algorithm for genome-wide association analysis, improving the detection of trait-related single-nucleotide polymorphisms (SNPs) and quantitative trait nucleotides (QTNs) with higher accuracy and efficiency.
Area of Science:
- Genomics
- Bioinformatics
- Statistical Genetics
Background:
- Genome-Wide Association Studies (GWAS) face challenges with high-dimensional genetic data.
- Traditional methods struggle with massive datasets, while existing machine learning approaches have limitations like overfitting and low accuracy.
- Accurate detection of single-nucleotide polymorphisms (SNPs) linked to traits is crucial for genetic research.
Purpose of the Study:
- To develop a novel, robust algorithm for identifying trait-associated SNPs and quantitative trait nucleotides (QTNs) in large genetic datasets.
- To overcome the limitations of existing machine learning methods in GWAS, such as poor generalization and low detection accuracy.
- To enhance the efficiency and accuracy of genetic analyses for complex traits.
Main Methods:
- Proposed a two-stage algorithm combining Least Angle Regression (LARS) and Random Forest (RF).
- The algorithm first controls for population structure and polygenic effects, then uses LARS for initial SNP selection.
- Random Forest is employed for further analysis of selected SNPs to detect QTNs.
Main Results:
- The TSLRF algorithm demonstrated superior detection ability for QTNs in simulation experiments and real data analyses compared to existing methods.
- Achieved improved model fitting, reduced calculation time, and a significant distinction between QTNs and other SNPs.
- In Arabidopsis, the method identified 60 confirmed trait-related genes and multiple gene clusters associated with flowering traits, outperforming other approaches.
Conclusions:
- The proposed two-stage algorithm (TSLRF) offers a powerful and efficient solution for GWAS, enhancing QTN detection accuracy.
- This method effectively addresses challenges posed by high-dimensional genetic data and improves upon existing machine learning techniques.
- TSLRF provides a significant advancement in identifying genes and genetic variations associated with complex traits in plants and potentially other organisms.
Related Concept Videos
Genome-wide Association Studies-GWAS
15.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.2K
Evolutionary Relationships through Genome Comparisons
6.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.8K
Genome Annotation and Assembly
20.4K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.4K

