Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Pedigree Analysis01:35

Pedigree Analysis

88.9K
Overview
88.9K
ER Retrieval Pathway01:45

ER Retrieval Pathway

4.7K
In the secretory pathway, vesicles transport proteins from one cellular compartment to another in forward transport to deliver the protein to its correct location. Occasionally, misfolded proteins and incorrect proteins escape their original compartments, and a retrieval pathway is used to return the escaped proteins to their original compartment.
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
4.7K
Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

15.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
15.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Optimizing protein tokenization: reduced amino acid alphabets for efficient and accurate protein language models.

Bioinformatics (Oxford, England)·2026
Same author

Genetic effects put into context.

Science (New York, N.Y.)·2026
Same author

SECmeres outperform extracellular vesicles as potential blood RNA biomarkers for Alzheimer's disease.

Nature communications·2026
Same author

Combinatorial effects of gene dosage, polygenic background and environment on complex traits.

medRxiv : the preprint server for health sciences·2026
Same author

GROMTools: scalable individual-level GReX imputation for mega-biobank-scale cohorts.

medRxiv : the preprint server for health sciences·2026
Same author

The Proteo-Transcriptome of Extracellular Vesicles and Particles Is Largely Preserved After Cryopreservation.

Journal of extracellular biology·2026

Related Experiment Video

Updated: Jan 18, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

9.2K

Phecoder: semantic retrieval for auditing and expanding ICD-based phenotypes in EHR biobanks.

Jamie J R Bennett1,2,3,4,5,6, Simone Tomasi1,2,3,4,5,6,7, Sonali Gupta1,2,3,4,5,6

  • 1Department of Psychiatry, Icahn School of Medicine at Mount Sinai, New York, NY, USA.

Medrxiv : the Preprint Server for Health Sciences
|January 16, 2026
PubMed
Summary

Phecoder, an AI tool, automates the identification of relevant diagnostic codes for electronic health record research. This improves the accuracy and completeness of patient cohorts for studies like genome-wide association studies.

Keywords:
EHR-linked biobanksICD codeselectronic health recordsmachine learningphecodesphenotypingsemantic retrievaltext embeddings

More Related Videos

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
07:35

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports

Published on: October 13, 2023

2.1K
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.4K

Related Experiment Videos

Last Updated: Jan 18, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

9.2K
A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
07:35

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports

Published on: October 13, 2023

2.1K
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.4K

Area of Science:

  • Computational biology
  • Biomedical informatics
  • Genomics

Background:

  • Electronic health record (EHR)-based phenotyping is crucial for genome-wide association studies (GWAS).
  • Current methods rely on manually curated ICD code lists (e.g., Phecodes), which are labor-intensive, subjective, and may miss relevant codes, reducing study power.
  • Advances in text embedding models offer a path to automate and standardize ICD-based phenotype construction.

Purpose of the Study:

  • To develop and evaluate Phecoder, an ensemble of text embedding models for automated ICD-based phenotyping.
  • To compare Phecoder's performance against existing PhecodeX phenotypes.
  • To assess the impact of Phecoder on cohort size and diversity.

Main Methods:

  • Developed Phecoder, an ensemble of pre-trained text embedding models to rank ICD codes by similarity to free-text descriptions.
  • Evaluated nine embedding models and unsupervised ensemble rank-fusion methods against 1,125 PhecodeX phenotypes.
  • Assessed retrieval performance using recall and average precision at top-100 (R@100, AP@100).
  • Conducted expert clinical review for six neuropsychiatric phenotypes and compared cohort sizes in the Million Veteran Program (MVP).

Main Results:

  • Phecoder, particularly with ensemble rank-fusion, significantly improved retrieval performance (e.g., +3% R@100, +8% AP@100) over individual models.
  • Expert review confirmed Phecoder identified additional clinically relevant ICD codes not present in PhecodeX.
  • Phecoder demonstrated substantial potential case expansion (median 200%, up to 2000% for specific disorders), increasing cohort completeness across demographic groups.

Conclusions:

  • Phecoder offers an automated, objective, and efficient approach to ICD-based phenotyping, addressing limitations of manual curation.
  • The framework enhances the identification of relevant diagnostic codes, leading to more comprehensive and reproducible EHR research.
  • Phecoder's applicability to future ICD code versions and its potential to improve cohort completeness across diverse populations highlight its value in biomedical research.