Machine Learning Study of SNPs in Noncoding Regions to Predict Non-small Cell Lung Cancer Susceptibility
1Health Management Center, General Practice Medical Center, West China Hospital, Sichuan University, Chengdu, Sichuan 610041, China; Institute of Respiratory Healthy, West China Hospital, Sichuan University, Chengdu, Sichuan 610041, China.
This study identified 17 new genetic loci linked to non-small cell lung cancer (NSCLC) risk in a Chinese population. Machine learning accurately predicted NSCLC using genetic and clinical data, aiding early diagnosis.
Area of Science:
- Genetics
- Oncology
- Bioinformatics
Background:
- Non-small cell lung cancer (NSCLC) is the most prevalent lung cancer subtype.
- Both environmental and genetic factors influence lung cancer susceptibility.
- Understanding genetic risk factors is crucial for early detection and prevention.
Purpose of the Study:
- To conduct a genome-wide association study (GWAS) to identify novel genetic variants associated with NSCLC risk in a Chinese population.
- To develop a predictive model for NSCLC susceptibility using genetic and clinical data with machine learning.
- To explore the application of machine learning in precision medicine for NSCLC.
Main Methods:
- Genome-wide association study (GWAS) of 287 NSCLC patients and 467 healthy controls using Illumina Genome-Wide Asian Screening Array Chip (712,095 SNPs).
- Logistic regression modeling to identify significant single nucleotide polymorphism (SNP) loci associated with NSCLC risk.
- Machine learning algorithms applied to a combined dataset of identified SNPs, previously reported SNPs, and clinical covariates (smoking status, age, CT screening, sex).
Main Results:
- GWAS identified 17 novel noncoding region SNP loci associated with NSCLC risk, with top SNPs (rs80040741, rs9568547, rs6010259) achieving stringent p-values (<3.02e-6).
- Two top SNPs, rs80040741 and rs6010259, were intronic variants within MUC3A and MLC1 genes, respectively.
- A machine learning model integrating genetic and clinical data achieved 86% accuracy in distinguishing NSCLC patients from healthy controls.
Conclusions:
- This study identified novel genetic loci contributing to NSCLC risk in the Chinese population.
- The integration of genetic and clinical data with machine learning offers a promising approach for NSCLC early diagnosis.
- Findings enhance the understanding of machine learning applications in precision medicine for cancer susceptibility.
More Related Videos
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
07:59Author Spotlight: Advancements in Molecular Biomarker Testing for Non-Squamous Non-Small Cell Lung Cancer
Published on: September 8, 2023
Related Concept Videos
Single Nucleotide Polymorphisms-SNPs
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
lncRNA - Long Non-coding RNAs
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
