Related Experiment Video
Updated: Sep 28, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Core open science practices in major medical journals: development of an automatized tool based on a Large Language
Constant Vinatier1, Guillaume Freyermuth2, Margaux Millour2
1Univ Rennes, CHU Rennes, Inserm, EHESP, Irset (Institut de recherche en santé, environnement et travail) - UMR_S 1085, Rennes, France.
Background:
Open science practices are increasingly promoted to improve the transparency and reproducibility of biomedical research, and their large-scale monitoring has been identified as a priority by initiatives such as the UNESCO Recommendation on Open Science. Existing automated tools cover only narrow subsets of practices. We developed and validated an automated tool, based on locally run open-weights large language models (LLMs), to extract a core set of 13 open science practices from published medical research articles.
Methods:
We built a validation database of 303 research articles published in 2020-2021 in 10 major general medical journals, stratified across randomized controlled trials (RCTs), meta-analyses, and other designs. Duplicate manual extraction with expert adjudication served as the reference standard. The pipeline combined document conversion (GROBID), BM25-based chunk retrieval, and classification by three Llama models (3-8B, 3-70B, 3.3-70B) run entirely locally, with and without supplementary materials. Diagnostic performance (sensitivity, specificity, predictive values, F1) was computed against the reference standard.
Results:
Usable outputs were obtained for 301/303 articles. Performance was heterogeneous across criteria rather than across models: practices reported in standardized statements (conflicts of interest, funding, author contributions, data sharing, pre-registration, reporting guidelines) were extracted reliably (F1 ≥ 0.75 for at least one model), whereas rare or inconsistently reported practices (code sharing, statistical analysis plan, protocol sharing, ORCID, open access status, preprints) remained difficult regardless of model size. Study-type classification exceeded 0.90 on all metrics. Larger models improved several criteria at a 6.6-7.5× computational cost, and adding supplementary materials did not improve performance, mainly due to conversion failures.
Conclusions:
A fully local, open-weights LLM pipeline can reliably monitor a substantial subset of open science practices. Document conversion and layout heterogeneity, rather than language understanding, constitute the main barrier to scaling automated open science monitoring.
Related Concept Videos
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Improving Translational Accuracy
Improving Translational Accuracy
Language and Cognition
Statistical Software for Data Analysis and Clinical Trials
Introduction to Language of Pathophysiology l
