Related Experiment Video
Updated: Apr 21, 2026

RBDT: A Computerized Task System based in Transposition for the Continuous Analysis of Relational Behavior Dynamics in Humans
Published on: July 17, 2021
CLEAR: Comparative Letter Examination and Analysis for Red Flags
Jaclyn Wiggins1, Melissa Jerdonek Sacco2, Elizabeth Bradley3
1is an Assistant Professor of Pediatrics, Department of Pediatrics, University of Virginia School of Medicine, Charlottesville, Virginia, USA.
Background:
Narrative letters of recommendation (LORs) remain a central element in fellowship selection. Programs screen these letters for language signifying a struggling learner or professionalism concerns, also known as "red flags," when determining interview offers. However, thorough screening is time consuming for program directors.
Objective:
To compare the speed and consistency of Microsoft Copilot to human reviewers in screening LORs for red flags.
Methods:
A retrospective analysis was conducted using de-identified LORs submitted during the 2024-2025 neonatal-perinatal medicine fellowship application cycle at a single fellowship site. Two reviewers independently screened each letter for predefined red flags. Disagreements were resolved by consensus or third-party adjudication. A rule-based natural language processing (NLP) model, refined through prompt adjustments, screened the same letters. Time to completion and red flag detection were compared.
Results:
A total of 195 LORs were reviewed. Following adjudication, red flags were confirmed in 21 letters. The NLP model flagged 16 letters and showed 76% (16 out of 21) agreement with the final adjudicated review. It processed all letters in 25 minutes, compared to the 554 minutes required by human reviewers. The model reliably identified terms "solid" and "good" with sentence-level context and showed consistency across the dataset, while humans varied more in detection, particularly with vague or indirect phrasing.
Conclusions:
A rule-based NLP model offers an efficient and consistent method for initial LOR screening.
Related Concept Videos
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Sign Test for Matched Pairs
To conduct the sign test, we first calculate the differences in...
Assessment of the Cardiovascular System II: Inspection
Head and Neck
Introduction to the Sign Test
Detection of Gross Error: The Q Test
Critical Region, Critical Values and Significance Level
In hypothesis testing, a sample statistic is converted to a test statistic using z, t, or chi-square distribution. A critical region is an area under the curve in probability distributions demarcated by the critical value. When the test statistic falls in this region, it suggests that the null hypothesis must be rejected. As this region contains all those values of the...

