Related Experiment Video
Updated: Mar 27, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Utilizing large language models to construct a dataset of Württemberg's 19th-century fauna from historical records
Maximilian C Teich1, Belen Escobari2, Malte Rehbein1
1Chair of Computational Humanities, University of Passau, Passau, Germany.
Abstract:
Constructing datasets on past biodiversity from historical sources is crucial for understanding long-term ecological changes. Typically, compiling such datasets relies on prior knowledge of the sources' composition and requires considerable manual effort. To overcome these challenges, we implement an automated approach based on prompted large language models (LLMs) to detect mentions of species in texts from 19th-century Württemberg and link these mentions to identifiers in the GBIF database. Based on our evaluation, we find that LLMs can reliably identify species in the texts with high recall (92.6%) and precision (95.3%), while providing estimates of the correct species identifier with considerable accuracy (83.0%). As our approach is easily scalable and adaptable to other contexts and languages, it offers a promising way to advance dataset generation from historical material using limited resources.
More Related Videos
Related Concept Videos
The Fossil Record
What is Evolutionary History?
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Habitat Fragmentation
Evolutionary Relationships through Genome Comparisons
The Tree of Life - Bacteria, Archaea, Eukaryotes

