Identification of misspelled words without a comprehensive dictionary using prevalence analysis

Alexander Turchin1, Julia T Chu, Maria Shubina

  • 1Partners HealthCare, Boston, MA, USA.

Summary

An algorithm accurately identifies medical misspellings by analyzing word prevalence in documents. This method enhances information retrieval and corrects errors in clinical notes with high precision.

Related Concept Videos

Prevalence and Incidence01:08

Prevalence and Incidence

In statistical epidemiology and health sciences, two essential metrics—prevalence and incidence—are fundamental for understanding disease dynamics within a population. These measures enable public health officials, epidemiologists, and researchers to assess the burden of diseases, allocate resources effectively, and design impactful public health policies and interventions.
Prevalence indicates the proportion of individuals in a population who have a specific disease or health condition at a...
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
Statistical Package for the Social Sciences (SPSS)01:22

Statistical Package for the Social Sciences (SPSS)

The Statistical Package for the Social Sciences, or SPSS, is a data management and analysis software suite. Developed by SPSS Inc. in 1968 and acquired by IBM in 2009, this tool was initially designed for social science data analysis, evolving to serve a wider range of disciplines. It was later renamed to Statistical Product and Service Solutions.
SPSS streamlines the process from data preparation to analysis and reporting. It is characterized by its user-friendly interface, which conceals...
Predicting Products: Substitution vs. Elimination02:52

Predicting Products: Substitution vs. Elimination

When a nucleophile and an alkyl halide react, nucleophilic substitution and β-elimination reactions compete to generate products.
The following factors can influence the mechanisms competing against each other: