Related Experiment Video
Updated: Apr 3, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Natural language querying of biological databases with large language models.
Vladimir A Makarov1, Oleg Stroganov2, Laura I Furlong3
1Pistoia Alliance, Wakefield, MA 01880, USA.
Drug Discovery Today
|April 1, 2026
Summary
Querying biological databases with natural language is improved by using multiple large language models (LLMs) that collaborate and interact with users. This approach balances accuracy and flexibility for effective biological data mining.
Area of Science:
- Bioinformatics
- Computational Biology
- Artificial Intelligence in Life Sciences
Background:
- Biological knowledge bases are vast and complex.
- Accessing information often requires structured query languages.
- Natural language querying offers a more intuitive interface.
Purpose of the Study:
- To systematically assess current methods for natural language querying of biological data using large language models (LLMs).
- To identify optimal strategies for balancing accuracy and flexibility in LLM-based biological data mining.
- To propose requirements for effective benchmarks and future applications of LLMs in this domain.
Main Methods:
- Systematic assessment of natural language querying techniques.
- Evaluation of multiple large language model (LLM) agents.
- Analysis of agent interaction and human-in-the-loop feedback.
Main Results:
- Multiple interacting LLM agents achieve the best balance between accuracy and flexibility.
- Agent-to-agent challenges and human interaction enhance query performance.
- Current benchmarks are insufficient for robust assessment of natural language data mining systems.
Conclusions:
- Collaborative LLM agent systems offer a promising approach for natural language querying of biological knowledge bases.
- Development of standardized benchmarks is crucial for advancing the field.
- Future applications should focus on structured agent interactions and human oversight for optimal biological data mining.
Keywords:
artificial intelligencedrug discoverylarge language modellife sciencesmachine learningpharmaceuticalMore Related Videos
Related Concept Videos
Genetic Lingo
117.9K
Overview
117.9K
lncRNA - Long Non-coding RNAs
10.2K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
10.2K

