Related Experiment Video
Updated: Dec 23, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Data-driven classification of the certainty of scholarly assertions
Mario Prieto1, Helena Deus2, Anita de Waard3
1Departamento de Biotecnología-Biología Vegetal, Escuela Técnica Superior de Ingeniería Agronómica, Alimentaria y de Biosistemas, Centro de Biotecnología y Genómica de Plantas, Universidad Politécnica de Madrid (UPM)- Instituto Nacional de Investigación y Tecnología Agraria y Alimentaria (INIA), Pozuelo de Alarcon, Madrid, Spain.
Scholars
Area of Science:
- Linguistics
- Computer Science
- Scholarly Communication
Background:
- Scholarly writing uses grammatical structures to indicate certainty.
- Existing certainty categorization systems lack objective validation regarding reader interpretation.
- Automated text analysis often misses subtle linguistic cues of certainty.
Purpose of the Study:
- To objectively test the validity of different scholarly certainty classification systems.
- To determine how researchers categorize assertions based on certainty.
- To develop an automated method for detecting and representing certainty in scholarly text.
Main Methods:
- Administered questionnaires to researchers to classify scholarly assertions.
- Utilized three distinct certainty classification systems for analysis.
- Developed and validated a machine learning model for automated certainty detection.
Main Results:
- Identified three distinct categories of certainty along a spectrum from high to low.
- Achieved high accuracy in automated certainty detection: 89.2% (author-annotated) and 82.2% (publicly-annotated).
- Demonstrated the feasibility of embedding certainty as metadata in machine-accessible formats like Nanopublications.
Conclusions:
- Reader interpretation of scholarly certainty can be objectively categorized.
- Machine learning models can accurately detect and classify linguistic certainty cues.
- Integrating certainty metadata into text-mining enhances scholarly information extraction.
Related Concept Videos
Uncertainty: Confidence Intervals
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Confidence Coefficient
Classification of Systems-II
Propagation of Uncertainty from Systematic Error

