Related Experiment Video
Updated: Jan 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Information Extraction of Doctoral Theses Using Two Different Large Language Models vs Health Services Researchers:
Jonas Cittadino1, Pia Traulsen1, Teresa Schmahl1
1Institute of Family Medicine, University Hospital Schleswig-Holstein, Maria-Goeppert-Straße 9a, Lübeck, 23538, Germany, 49 451 3101 ext 8001.
Large language models (LLMs) like GPT-4o and Gemini-1.5-Flash can efficiently extract information and generate abstracts from historical doctoral theses. LLM-generated abstracts are comparable to human-written ones and significantly faster to produce.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Archival Science
Background:
- The Archive of German-Language General Practice (ADAM) contains approximately 500 paper-based doctoral theses from 1965 to present.
- No systematic information extraction (IE) has been performed on these historical medical documents.
- Large language models (LLMs) show potential for IE in medical texts, but concerns about hallucinations and use with non-recent documents exist.
Purpose of the Study:
- To assess the efficacy of LLMs (GPT-4o and Gemini-1.5-Flash) in extracting information from paper-based doctoral theses in the ADAM archive.
- To evaluate the quality of LLM-generated abstracts compared to human-generated abstracts.
Main Methods:
- Random selection of 10 doctoral theses (1965-2022) for analysis.
- Utilized two LLM pipelines (OpenAI's GPT-4o and Google's Gemini-1.5-Flash) for information extraction and abstract generation.
- Compared LLM-generated abstracts with a human-generated abstract using blinded raters and bidirectional encoder representations from transformers (BERT) scores.
Main Results:
- Key dissertation characteristics (institute, title, author, year) were extracted for all 10 theses.
- Abstracts were successfully generated by GPT-4o for 9/10 theses and by Gemini-1.5-Flash for all 10 theses.
- LLM-generated abstracts showed moderate-to-high semantic similarity (BERT F1-scores: GPT-4o 0.72, Gemini 0.71) and were created 24-36 times faster than human abstracts.
- Raters found no significant difference in quality between LLM-generated and human-generated abstracts (P=.44).
Conclusions:
- LLMs (GPT-4o and Gemini-1.5-Flash) are effective tools for information extraction and abstract generation from historical doctoral theses.
- LLM-generated abstracts are semantically similar to human-generated ones, faster to produce, and do not lose information during translation.
- This study demonstrates the feasibility of using LLMs in a standard workflow to improve searchability of historical family medicine literature.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019