Related Experiment Video
Updated: May 26, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
475
A large language model framework for literature-based disease-gene association prediction
Peng-Hsuan Li1, Yih-Yun Sun1, Hsueh-Fen Juan1,2,3,4
1Taiwan AI Labs, 6F., No. 70, Sec. 1, Chengde Road, Datong Dist., Taipei 10355, Taiwan.
Briefings in Bioinformatics
|February 25, 2025
Summary
Large Language Models (LLMs) can now understand biomedical literature for precision medicine. A new method, LORE, extracts gene-disease relationships with 90% precision, aiding therapeutic target discovery.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Precision Medicine
Background:
- The exponential growth of biomedical literature necessitates advanced tools for knowledge extraction.
- Current Large Language Model (LLM) approaches struggle with reliability, verifiability, and scalability in analyzing complex biological relationships.
Purpose of the Study:
- To develop a novel, unsupervised methodology for automated biomedical literature understanding using LLMs.
- To overcome limitations in current LLM applications for extracting complex biological relationships and advancing precision medicine.
Main Methods:
- Proposed LORE, a two-stage reading methodology utilizing LLMs.
- Modeled biomedical literature as a knowledge graph of verifiable factual statements and semantic embeddings.
- Applied LORE to PubMed abstracts for gene-disease relationship extraction.
Main Results:
- LORE effectively captured essential gene pathogenicity information from PubMed abstracts.
- Achieved 90% mean average precision in identifying relevant genes across 2097 diseases by modeling latent pathogenic flow with ClinVar supervision.
- Demonstrated scalability and reproducibility in biomedical literature analysis.
Conclusions:
- LORE offers a scalable and reproducible approach for leveraging LLMs in biomedical literature analysis.
- This methodology enhances the identification of therapeutic targets by efficiently extracting disease-gene relationships.
- Paves the way for improved precision medicine through automated knowledge discovery.
Related Concept Videos
Genome-wide Association Studies-GWAS
12.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.3K
Pleiotropy
39.4K
Pleiotropy is the phenomenon in which a single gene impacts multiple, seemingly unrelated phenotypic traits. For example, defects in the SOX10 gene cause Waardenburg Syndrome Type 4, or WS4, which can cause defects in pigmentation, hearing impairments, and an absence of intestinal contractions necessary for elimination. This diversity of phenotypes results from the expression pattern of SOX10 in early embryonic and fetal development. SOX10 is found in neural crest cells that form melanocytes,...
39.4K
Genomics
35.6K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
35.6K
Single Nucleotide Polymorphisms-SNPs
13.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
13.8K
Translation
14.4K
Translation is the process of synthesizing proteins from the genetic information carried by messenger RNA (mRNA). Following transcription, it constitutes the final step in the expression of genes. This process is carried out by ribosomes, complexes of protein and specialized RNA molecules. Ribosomes, transfer RNA (tRNA), and other proteins produce a chain of amino acids—the polypeptide—as the end product of translation.
Translation Produces the Building Blocks of Life
Proteins are...
Translation Produces the Building Blocks of Life
Proteins are...
14.4K
Incomplete Dominance
20.9K
Gregor Mendel's work (1822 - 1884) was primarily focused on pea plants. Through his initial experiments, he determined that every gene in a diploid cell has two variants called alleles inherited from each parent. He suggested that amongst these two alleles, one allele is dominant in character and the other recessive. The combination of alleles determines the phenotype of a gene in an organism.
20.9K

