Related Experiment Videos
Semantic Similarity Search Approach to Extract Exemplars of Stigmatizing and Positive Language in Obstetric Clinical
Jihye Kim Scroggins1, Ismael Ibrahim Hulchafo2, Veronica Barcelona2
1School of Nursing, University of North Carolina at Chapel Hill, 120 Medical Drive, Chapel Hill, NC, 27514, United States, 1 (919) 966-4260.
Background:
Natural language processing can extract meaningful information from clinical notes. However, human annotation is time-consuming and costly, and scarce data poses a challenge.
Objective:
This study aimed to explore a semantic similarity search approach to extract exemplars of stigmatizing and positive language in obstetric clinical notes.
Methods:
We used electronic health record data from labor and birth admissions at 2 hospitals in the United States from 2017 to 2019. We used a semantic similarity search approach, which used 200 randomly selected true exemplars, stratified by language categories, as queries to search for similar exemplar candidates. We extracted the top 5 candidates with the highest cosine similarities, which were assessed for accuracy.
Results:
We retrieved 1000 candidates. An average precision of 0.69 was achieved when candidates with cosine similarity thresholds of 0.75 or higher were included, at which point 68.8% (64/93) of exemplar candidates accurately represented true cases. At the 0.75 threshold, the proportion of true cases was higher for preferred language (41/56, 73.2%) and unilateral/authoritarian decisions (5/7, 71.4%). The proportion of true cases was lower for difficult patients (2/7, 28.6%) and marginalized identities (3/9, 33.3%).
Conclusions:
The semantic similarity search approach shows promise in efficiently extracting exemplars while reducing the annotation burden, laying the groundwork for future applications in other domains.