Related Experiment Video
Updated: Jun 10, 2026

09:22
Ultrafast Lignin Extraction from Unusual Mediterranean Lignocellulosic Residues
Published on: March 9, 2021
Librarian of Alexandria: A Modular Chemical Data Extraction Pipeline to Compare LLM Performance
Morgan Grougan1, Janya Subasinghe1, Mark A Hix1
1Department of Chemistry, Wayne State University, Detroit, Michigan 48202, United States.
Journal of Chemical Information and Modeling
|June 9, 2026
Summary
Creating large chemistry datasets for machine learning is now easier with the Librarian of Alexandria (LoA) framework. This tool uses large-language models (LLMs) to automatically extract data from scientific papers, achieving high accuracy.
Area of Science:
- Computational Chemistry
- Data Science
- Machine Learning
Background:
- Manual creation of large experimental chemistry datasets is time-consuming and challenging, especially for specialized properties.
- Scarcity of comprehensive datasets hinders the development of predictive machine learning models in chemistry.
Purpose of the Study:
- To introduce the Librarian of Alexandria (LoA), an open-source framework for evaluating large-language models (LLMs) in generating large chemistry datasets.
- To provide a modular and updatable testing environment for LLMs performing data extraction from scientific literature.
Main Methods:
- LoA utilizes LLMs to assess research paper relevance and extract data, allowing for independent or identical model selection for each task.
- The framework supports easy integration of new LLMs and automates the retrieval of open-access research papers from chemical journals.
- Performance of various LLMs was compared for relevance assessment and data extraction capabilities.
Main Results:
- The LoA framework facilitates the automated generation of substantial chemistry datasets with minimal scientific effort.
- The best-performing LLMs demonstrated approximately 95% accuracy in data extraction, suitable for training predictive machine learning models.
- LoA offers a robust testing ground for developing advanced LLM-driven data generation tools.
Conclusions:
- LoA significantly streamlines the process of creating large-scale chemistry datasets for machine learning applications.
- The framework's modular design and high accuracy pave the way for more efficient AI-driven scientific discovery in chemistry.