Related Experiment Video
Updated: Aug 27, 2026

Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information
Published on: August 15, 2019
OtoVCE: a mechanism-aware language-model evidence layer for hereditary hearing-loss variant interpretation
Shaopei Ye1, Lan Wang2, Peng Chen3
1Faculty of Biology, University of Barcelona, Barcelona, Spain.
Background:
Missense variants in genes implicated in hereditary hearing loss are frequently returned to clinicians as variants of uncertain significance. In silico predictors and rule-based ACMG/AMP frameworks each capture only a fraction of the required evidence, and neither directly accesses case-level and functional evidence reported in the published literature.
Methods:
We developed OtoVCE, a four-stage framework that places a large language model inside a calibrated ACMG/AMP rule engine as a structured evidence extractor rather than an end-to-end classifier. Rule-encodable evidence-six in silico missense predictors, gnomAD allele frequencies, UniProt domain annotations, and a hearing-loss-specific protein language model with a pathogenicity head and a mechanism head-is aggregated by the ClinGen Hearing Loss VCEP rule set. For variants that remain uncertain, OtoVCE retrieves the literature from five sources and prompts the language model using the Brnich functional evidence rubric, with the predicted disease mechanism as context, returning PS3, BS3, and PS4 strength assignments with PubMed-verified citations. Rule-based and language-model-derived strengths are then combined into a posterior probability mapped to the five ACMG/AMP classes.
Results:
On a held-out post-2024 ClinVar cohort (N = 1,885), OtoVCE achieved an area under the receiver-operating curve (AUC) of 0.997 and a sensitivity of 93.6% at a 1% false-positive rate. Performance was preserved across an OTOF gene-leave-out cohort (AUC = 0.993), the external Deafness Variation Database (AUC = 0.917), and the protein-language-model training cohort itself, with model-derived rules disabled (AUC = 0.976), excluding training-set memorization. A paired A/B comparison (N = 1,930) attributed the literature evidence contribution to the mechanism cue: 7.1-fold more cited identifiers, 11.6-fold more quantitative evidence, and diagnostic strength scores (paired-bootstrap ΔAUC +0.18). OtoVCE identified 211 of 2,389 uncertain variants as one piece of supporting evidence short of likely pathogenic at 98.1% precision, agreed with ClinGen-VCEP curation at Cohen's κ = 0.57, and verified 100% of 3,185 cited PubMed identifiers.
Conclusion:
A literature-grounded, mechanism-aware language-model evidence layer can serve as a complementary component within a calibrated ACMG/AMP workflow for clinical variant interpretation in hereditary hearing loss.
Related Concept Videos
The Cochlea
Pleiotropy
Genetic Lingo

