Related Experiment Video
Updated: Sep 20, 2025

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
Precise estimation of in-depth relatedness in biobank-scale datasets using deepKin
Qi-Xin Zhang1, Dovini Jayasinghe2, Zhe Zhang3
1Institute of Bioinformatics, Zhejiang University, Hangzhou, Zhejiang 310058, China; Center for Laboratory Medicine, Department of Genetic and Genomic Medicine, and Clinical Research Institute, Zhejiang Provincial People's Hospital, People's Hospital of Hangzhou Medical College, Hangzhou, Zhejiang 310014, China; Australian Centre for Precision Health and UniSA Allied Health and Human Performance, University of South Australia, Adelaide, SA 5000, Australia.
DeepKin accurately estimates genetic relatedness in large biobanks by accounting for sampling variance. This method identifies over 212,000 relative pairs in the UK Biobank, revealing geographic patterns and improving distant relative detection.
Area of Science:
- Genetics
- Bioinformatics
- Statistical Genetics
Background:
- Accurate estimation of genetic relatedness is crucial for large-scale genetic studies and biobanks.
- Traditional methods often rely on fixed thresholds, which may not be optimal for diverse datasets.
- Understanding relatedness is key to controlling for population structure and identifying familial relationships.
Purpose of the Study:
- To introduce deepKin, a novel method-of-moments framework for accurate relatedness estimation in biobank-scale genetic studies.
- To develop a method that accounts for sampling variance to enable robust statistical inference and classification of relatedness.
- To provide tools for determining data-specific significance thresholds and estimating statistical power for detecting distant relatives.
Main Methods:
- Developed deepKin, a method-of-moments framework incorporating sampling variance.
- Implemented data-specific significance threshold computation.
- Determined the minimum effective number of markers and estimated statistical power.
- Validated through simulations and application to UK Biobank data.
Main Results:
- deepKin accurately infers unrelated pairs and relatives by leveraging sampling variance.
- Analysis of the UK Biobank 3K Oxford subset showed increased power for distant relative detection with larger effective marker sets.
- In the UK Biobank White British subset, deepKin identified over 212,000 significant relative pairs across six degrees.
- Geographic patterns of relatedness were revealed across 19 UK Biobank assessment centers.
Conclusions:
- deepKin offers a statistically rigorous approach to relatedness estimation in large genetic datasets.
- The method enhances the power to detect distant relatives and provides insights into population structure and familial relationships.
- An R package (deepKin) is available for broader application in genetic research.
More Related Videos
09:47Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024