Related Experiment Video
Updated: Aug 6, 2026

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Benchmarking a Local Schema-Constrained Large Language Model Pipeline for Abstract Screening and Evidence Mapping
1Psychiatry, Università degli Studi di Enna Kore, Enna, ITA.
Cureus
|July 22, 2026
Summary
A novel, fully local pipeline using open-weight large language models (LLMs) effectively supports scalable and auditable biomedical literature screening and evidence mapping. This approach addresses governance and reproducibility challenges posed by proprietary cloud APIs in evidence synthesis.
Area of Science:
- Biomedical Informatics
- Artificial Intelligence in Healthcare
- Systematic Review Methodology
Background:
- The exponential growth of biomedical literature challenges manual evidence synthesis.
- Current large language model (LLM) workflows often rely on proprietary cloud APIs, hindering governance, reproducibility, and scalability.
- There is a need for local, open-weight LLM solutions for efficient literature screening.
Purpose of the Study:
- To develop and evaluate a fully local, open-weight, schema-constrained pipeline for title/abstract-based scoping workflows.
- To benchmark the pipeline's performance against established systematic reviews using precision, recall, and F1 scores.
- To assess the pipeline's capability for structured abstract extraction and evidence mapping.
Main Methods:
- Developed a local pipeline using gpt-oss-20b via Ollama on an Apple M1 Max.
- Combined deterministic metadata filtering, LLM-assisted screening, and structured abstract extraction.
- Benchmarked performance against three systematic reviews, reporting standard and audit-adjusted metrics.
Main Results:
- Achieved 100.0% recall and 98.9% F1 score in the ketamine/neuroimaging benchmark after audit adjustment.
- Demonstrated strong performance in clozapine reviews (F1 scores of 76.0% and 83.6%), with limitations noted for non-informative abstracts.
- Successfully recovered audited metadata fields without errors and generated thematically concordant evidence maps.
Conclusions:
- A fully local LLM pipeline can effectively support scalable, auditable abstract-based scoping and evidence mapping.
- The findings, while promising, require confirmation in larger, independently adjudicated evaluations.
- Human audit and expert full-text synthesis remain crucial for non-informative abstracts and mechanistic precision.
