Related Experiment Video
Updated: Dec 17, 2025

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Automated Spelling Correction for Clinical Text Mining in Russian.
Ksenia Balabaeva1, Anastasia Funkner1, Sergey Kovalchuk1
1ITMO University, Saint Petersburg, Russia.
This study introduces a Russian clinical text spell checker using machine learning and string distance algorithms. The developed tool achieves high precision, aiding medical text mining by addressing common errors and linguistic complexities.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Medical Informatics
Background:
- Clinical text presents unique challenges for spell checking due to specialized vocabulary and potential for errors.
- Accurate spell checking is crucial for reliable medical text mining and analysis.
- Existing tools may not adequately address the nuances of Russian clinical language.
Purpose of the Study:
- To develop and evaluate a novel spell checker module specifically designed for Russian clinical text.
- To integrate this spell checker into a broader medical text mining tool.
- To improve the accuracy of automated analysis of clinical documents.
Main Methods:
- Combination of string distance measure algorithms (e.g., Levenshtein, Damerau-Levenshtein) for initial error detection.
- Application of machine learning embedding methods (e.g., Word2Vec, FastText) to understand contextual relevance and word meaning.
- Development of a hybrid approach leveraging both algorithmic and machine learning techniques.
Main Results:
- Achieved an overall precision of 0.86 for the spell checker module.
- Demonstrated high lexical precision of 0.975, indicating accurate identification of correct terms.
- Reported an error precision of 0.74, showing effectiveness in identifying and correcting misspellings.
- The spell checker is a component of a larger medical text mining tool.
Conclusions:
- The developed spell checker module significantly enhances the accuracy of processing Russian clinical text.
- The hybrid approach combining string distance and machine learning is effective for clinical spell correction.
- This advancement supports more reliable medical text mining, including detection of misspelling, negation, experiencer, and temporality.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
12:49Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019
Related Concept Videos
Proofreading
Proofreading
Errors During Replication are Corrected by the DNA Polymerase...
Improving Translational Accuracy
Improving Translational Accuracy
Types of Errors: Detection and Minimization
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.