Related Experiment Video
Updated: Jun 10, 2026

Ultrafast Lignin Extraction from Unusual Mediterranean Lignocellulosic Residues
Published on: March 9, 2021
Librarian of Alexandria: A Modular Chemical Data Extraction Pipeline to Compare LLM Performance
Morgan Grougan1, Janya Subasinghe1, Mark A Hix1
1Department of Chemistry, Wayne State University, Detroit, Michigan48202, United States.
Abstract:
Dataset creation is a critical component of predictive machine learning technology. In the field of chemistry, large experimental datasets are scarce, especially for niche chemical properties and topics, and their creation is cumbersome and time-consuming when performed manually. Here, we present Librarian of Alexandria (LoA), an open-source and lightweight framework for testing large-language models (LLMs) on the task of generating large datasets via direct extraction from scientific literature. LoA is available on GitHub along with example inputs, outputs, Docker container, and a Colab-friendly Jupyter Notebook. LoA has the chosen LLM(s) check the relevance of a research paper and perform data extraction. Two separate or identical LLMs may be independently and modularly specified by the end-user for these separate tasks. LoA can be easily updated via simplified user incorporation of the latest available LLMs. We compare several LLMs for both relevance and extraction functions and automate the collection of research papers for several popular chemical journals providing open access. LoA provides a much-needed testing environment which can be used to develop LLMs capable of assembling enormous chemical datasets with minimal effort on the part of the scientist. The best models we were able to test achieved ∼95% accuracy, which is sufficient for the purposes of training predictive machine learning models.