Related Experiment Video
Updated: Aug 12, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
A local, privacy-oriented multi-agent LLM framework for framework-grounded manuscript editing: A proof-of-concept
Alon Gorenshtein1, Rohan Bhansali1, Brandon Westover1
1Department of Neurology, Beth Israel Deaconess Medical Center, Boston, MA, USA; Harvard Medical School, Boston, MA, USA.
Objectives:
Manuscript preparation is a bottleneck in publishing, and cloud-based AI tools raise confidentiality concerns for clinical researchers. We developed the Paper Analysis Tool (PAT), a free, multi-agent framework that audits manuscripts with a local open-weight language model.
Materials And Methods:
PAT comprises 31 components (one deterministic text-metrics module and 30 language-model agents) run locally. For six manuscripts (three in-house neurology, three external preprints), suggestions from the PAT pipeline, the same local model, and a frontier model (Claude Sonnet 5), each with one generic prompt, were pooled, source-stripped, shuffled, and scored blind by two co-authors for actionability and usefulness; text suggestions mapped to 15 grounded domains, figures to five figure-quality domains.
Results:
Orchestrated by PAT, a local open-weight 27B model covered 11.3 of 15 useful domains, versus 5.5 for the same model with one generic prompt and 9.3 for the frontier benchmark, exceeding the same model in every paper and after matching (8.7 versus 5.5). Coverage and per-suggestion usefulness were distinct axes: the frontier was more useful (81 % versus 29 %) and led at the clearly-useful threshold (7.7 versus 5.4) and after matching (9.3 versus 6.4). Inter-rater agreement was 90.5 % and 91.7 % (Cohen kappa 0.81, 0.82). By deterministic Phase-0 metrics (rater-independent), PAT's rewrites cut passive-voice sentences from 51 % to 6 %, long sentences by 73 %, and word count by 22 %.
Discussion:
In this proof-of-concept, PAT's structured audit reliably broadened a local open-weight model's coverage of the 15 quality-relevant domains, exceeding the same one-prompt model in every paper and after matching, with inference kept local. Per-suggestion usefulness is bounded by the base model: the local model trailed the frontier and PAT's suggestions still need author judgment. As a model-agnostic framework, PAT can pair this breadth with a stronger base model to raise usefulness, without claiming frontier-equivalent per-suggestion quality.
Conclusion:
Local, privacy-preserving multi-agent review is feasible and broadens the domains a base model addresses.
Related Concept Videos
Feedback Inhibition
Proofreading
Mismatch Repair
MicroRNAs
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on the metal...