Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automated extraction of toxicological mechanism of action from PubMed literature using Large Language Models
Arnaud Molle1, Nelly D Saenen2, Karen Smeets2
1European Food Safety Authority, Via Carlo Magno 1A, Parma, 43126, Italy.
Abstract:
Risk assessment activities at the European Food Safety Authority face mounting challenges from an increasing volume of compounds requiring evaluation and exponential growth in scientific literature. To address these challenges, toxicologists started grouping compounds based on their mechanism of action in the body, allowing study comparisons similar to current risk assessment strategies. The LLMs-rev pipeline was created to retrieve potentially relevant papers from PubMed, assess their relevance, and extract the mechanism of action (MoA) information from accessible publications using large language models. Applied to 121 compounds, the system processed over 400,000 papers, identifying 30,250 as containing relevant information and extracting specific MoA quotations from 4500 open-access publications. The automated approach demonstrated processing speeds exceeding 8000 papers per hour, dramatically outpacing conventional manual screening methods that typically assess 100-120 papers per hour. While the system proved particularly effective for compounds for which abundant literature was available, human expertise remained essential for interpreting complex MoAs that required contextual data analysis. The optimized prompts minimized model hallucinations by restricting outputs to direct quotations from verified sources. The methodology demonstrates considerable potential for accelerating risk assessment workflows while maintaining scientific rigor through the complementary use of human oversight.
