A Bayesian genomic selection approach incorporating prior feature ordering and population structures with application

Xiaotian Dai1, Xuewen Lu1, Thierry Chekouo2

  • 1Department of Mathematics and Statistics, University of Calgary, Calgary, Canada.

Insights

This study introduces a new Bayesian method to identify genetic variants linked to coronary artery disease (CAD). The approach improves accuracy by considering variant order and population structure for better disease risk prediction.

Area of Science:

  • Genetics
  • Biostatistics
  • Cardiovascular Disease Research

Background:

  • Coronary artery disease (CAD) is a leading cause of death, with significant genetic influences in both sexes.
  • Identifying specific genetic variants associated with CAD is crucial for understanding disease mechanisms and improving risk prediction.

Purpose of the Study:

  • To propose a novel Bayesian variable selection framework for identifying genetic variants associated with coronary artery disease (CAD) status.
  • To develop an innovative prior that accounts for the ordering structure of genetic variants, improving selection accuracy.
  • To incorporate population structure by fitting separate regressions for different subject groups, enhancing disease risk reflection.

Main Methods:

  • Developed a Bayesian variable selection framework incorporating an innovative prior for genetic variant inclusion probabilities, considering their ordering.
  • Implemented a method to group subjects based on population structure and fit separate regression models.
  • Utilized a Markov random field-inspired prior to borrow strength across regression models, enhancing model performance.

Main Results:

  • The proposed framework demonstrated improved variable selection and prediction performance in simulation studies.
  • The method was successfully applied to the CATHeterization GENetics (CG) dataset for binary CAD status.
  • The novel approach effectively identifies important genetic variants and accounts for population heterogeneity.

Conclusions:

  • The novel Bayesian framework offers improved accuracy in identifying genetic variants associated with coronary artery disease.
  • Accounting for variant ordering and population structure enhances the precision of genetic association studies for CAD.
  • This approach provides a robust tool for genetic research in cardiovascular diseases.

Related Concept Videos

Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
13.6K
Behavioral Genetics and Its Designs01:23

Behavioral Genetics and Its Designs

Behavior genetics explores how genetic inheritance influences human behavior. It focuses on how genes, passed from parents to offspring, contribute to the development of behavioral traits and tendencies. This branch of genetics seeks to understand the complex interplay between inherited genetic factors and environmental influences in shaping our behaviors.
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
418
Coronary Artery Disease I: Introduction01:30

Coronary Artery Disease I: Introduction

Coronary Artery Disease (CAD): An Overview with Scientific InsightsCoronary Artery Disease (CAD), often referred to as C-A-D, is a prevalent blood vessel disorder classified under the broader category of atherosclerosis. Atherosclerosis is a pathological process characterized by the hardening and narrowing of arteries due to the accumulation of atherosclerotic plaques. These plaques are composed of cholesterol, fatty substances, inflammatory cells, calcium, and fibrin, reducing blood flow to...
28
Frequency-dependent Selection01:21

Frequency-dependent Selection

When the fitness of a trait is influenced by how common it is (i.e., its frequency) relative to different traits within a population, this is referred to as frequency-dependent selection. Frequency-dependent selection may occur between species or within a single species. This type of selection can either be positive—with more common phenotypes having higher fitness—or negative, with rarer phenotypes conferring increased fitness.
22.1K
Genetic Variation01:25

Genetic Variation

Genetic variation is the diversity in DNA sequences found among individuals of the same species. This diversity is crucial for a species' survival because it helps organisms adapt to environmental changes. Genetic variation begins with fertilization, where an egg and sperm cell merge. Each of these cells carries 23 chromosomes, up to 46 in the fertilized egg. Chromosomes are long DNA strands that contain genes, the basic units of heredity.
Genes exist in different versions called alleles,...
331
Pedigree Analysis01:35

Pedigree Analysis

Overview
84.5K