Related Experiment Video
Updated: Aug 22, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Performance of Two AI Approaches in ASReview Compared With Manual Screening for Dementia Care Literature Screening:
Dirk Steijger1,2,3, Stella Thissen4, Sil Aarts2,3
1Department of Psychiatry and Neuropsychology, Mental Health and Neuroscience Research Institute, Faculty of Health, Medicine and Life Sciences, Maastricht University, Dr Tanslaan, 12, Maastricht, 6229ET, The Netherlands, 31 612365497.
Background:
Literature reviews rely on rigorous title and abstract screening by researchers, which is time-consuming. AI-assisted literature screening tools have been proposed to improve efficiency by prioritizing titles and abstracts with the highest likelihood of meeting the inclusion criteria, thereby reducing the need to screen all records.
Objective:
This study aims to evaluate the performance of two AI-assisted screening approaches in ASReview (version 1.3; Department of Methodology and Statistics, Utrecht University) compared with manual title and abstract screening in a previously completed and published scoping review on how AI can support the quality of life in people with dementia.
Methods:
This study used a dataset of 4690 titles and abstracts from a published scoping review. The manual screening decisions of the scoping review served as the reference standard. Both ASReview approaches were applied by the same author who conducted the majority of the original manual title and abstract screening. Approach A used a simpler model with minimal prior input, whereas approach B used a more advanced model with a larger training set. Both ASReview approaches applied predefined stopping rules: (1) more than 10% of the dataset to be screened; and (2) 50 consecutive irrelevant titles and abstracts. Performance was evaluated in terms of sensitivity, specificity, precision, accuracy, and screening time. 95% CIs were calculated for sensitivity, specificity, precision, and accuracy. Agreement between manual and ASReview approaches was assessed using the Cohen κ, and differences in how manual and both ASReview approaches classified titles and abstracts were examined using the McNemar test. Performance and agreement were calculated at two levels: (1) after title and abstract screening and (2) after full-text inclusion.
Results:
Manual screening identified 283 titles and abstracts for full-text review and resulted in 30 final included studies, requiring 19 hours. Of the 4690 titles and abstracts, approach A screened 830 (17.7%) in 4.3 hours and retrieved 16 of the 30 (sensitivity 0.53, 95% CI 0.36-0.70) final included studies, whereas approach B screened 798 (17.0%) in 5.5 hours and retrieved 21 of the 30 (sensitivity 0.70, 95% CI 0.52-0.83) final included studies. Although both ASReview approaches showed high specificity and accuracy, these metrics should be interpreted cautiously because the dataset was highly imbalanced and contained relatively few relevant titles and abstracts. McNemar tests showed significant directional imbalance (P<.001): ASReview missed more manually selected titles and abstracts at level 1, whereas at level 2, ASReview more often labeled titles and abstracts not included in the final review as relevant.
Conclusions:
ASReview can support workload reduction in title and abstract screening, but the evaluated ASReview approaches did not retrieve all final included studies from the original dementia care scoping review. These findings suggest that the evaluated ASReview configurations may be insufficient for reviews in which near-complete retrieval of relevant evidence is required.