Related Experiment Video
Updated: Jan 12, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.0K
Scalable scientific interest profiling using large language models
Yilun Liang1, Gongbo Zhang2, Edward Sun3
1Tandon School of Engineering, New York University, Brooklyn, NY, USA.
Journal of Biomedical Informatics
|November 2, 2025
Summary
Large Language Models (LLMs) can automate scientific interest profiling. Profiles generated using Medical Subject Headings (MeSH) terms are more readable, though human-written profiles offer more novel concepts.
Area of Science:
- Biomedical informatics
- Artificial intelligence in research
Background:
- Scientific research profiles are crucial for talent discovery and collaboration.
- Existing profiles are often outdated, necessitating automated and scalable solutions.
- Large Language Models (LLMs) offer a potential solution for dynamic profile generation.
Purpose of the Study:
- To design and evaluate LLM-based methods for generating scientific interest profiles.
- To compare machine-generated profiles with researchers' self-summarized interests.
- To assess the performance of profiles generated from PubMed abstracts versus Medical Subject Headings (MeSH) terms.
Main Methods:
- Two LLM-based methods were developed: one summarizing researcher abstracts and another using MeSH terms.
- GPT-4o-mini was used to generate summaries for 595 researchers from Columbia University Irving Medical Center.
- Automated metrics (ROUGE-L, BLEU, METEOR, BERTScore, KL Divergence) and manual evaluations were employed for comparison.
Main Results:
- Automated metrics showed low lexical overlap but moderate semantic similarity (BERTScore F1: ~0.55) between machine-generated and human-written profiles.
- Manually paraphrased summaries achieved higher similarity (F1: 0.851).
- MeSH-based profiles demonstrated superior readability (93.44% favorable ratings) and were preferred in 67.86% of manual reviews, despite differences in keyword usage and factual accuracy compared to human-written profiles.
Conclusions:
- LLMs show promise for scalable automation of scientific interest profiling.
- MeSH-based LLM-generated profiles offer better readability than abstract-based ones.
- While LLMs can generate semantically similar profiles, human-written summaries tend to introduce more novel concepts.
Related Concept Videos
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K
Ribosome Profiling
4.0K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
4.0K
Scaling
545
In designing and analyzing filters, resonant circuits, or circuit analysis at large, working with standard element values like 1 ohm, 1 henry, or 1 farad can be convenient before scaling these values to more realistic figures. This approach is widely utilized by not employing realistic element values in numerous examples and problems; it simplifies mastering circuit analysis through convenient component values. The complexity of calculations is thereby reduced, with the understanding that...
545
Language and Cognition
702
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
702
Aggregates Classification
962
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
962

