Related Experiment Video
Updated: Aug 27, 2026

Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
P-KNN: joint calibration of multiple pathogenicity prediction tools streamlines variant classification
1Center for Human Genetics & Genomics, New York University Grossman School of Medicine, NY; Department of Neurology, National Cheng Kung University Hospital, College of Medicine, National Cheng Kung University, Tainan, Taiwan; Department of Genomic Medicine, National Cheng Kung University Hospital, College of Medicine, National Cheng Kung University, Tainan, Taiwan.
Purpose:
Clinical guidelines for interpreting genetic variants in the context of Mendelian disease require converting the outputs of pathogenicity prediction tools into well-calibrated probabilities. However, the existing calibration method is only valid when pre-committing to one tool, preventing clinical laboratories from using multiple tools with complementary strengths. To lift this restriction, we introduce Pathogenicity K-Nearest Neighbors (P-KNN), a flexible method that jointly calibrates any set of tools.
Methods:
P-KNN represents each variant in a multidimensional space defined by tool scores and estimates the probability of pathogenicity based on the proportion of pathogenic neighbors. We compared P-KNN against standard single-tool calibration of multiple predictors and meta-predictors at four historical time points.
Results:
P-KNN outperforms standard calibration of single tools and meta-predictors in two aspects: i) overall evidence strength and ii) alignment of the calibrated probabilities with true pathogenicity frequencies. Additionally, the evidence from P-KNN keeps improving with the addition of newer tools. It also correctly integrates correlated computational and experimental evidence that is overestimated by existing protocols.
Conclusion:
P-KNN provides robust joint calibration for any set of pathogenicity prediction tools, thereby alleviating the constraint of pre-committing to a single predictor while enhancing statistical rigor and diagnostic yield. P-KNN is available via command line (https://github.com/Brandes-Lab/P-KNN) and precomputed scores (https://huggingface.co/datasets/brandeslab/P-KNN).
More Related Videos
Related Concept Videos
Single Nucleotide Polymorphisms-SNPs
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Evolutionary Relationships through Genome Comparisons
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Modern Molecular Taxonomy
Multiple Allele Traits

