Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Quality Assurance01:19

Quality Assurance

155
Quality assurance is the overarching term used to describe the activities employed to ensure the proper performance of a system. These activities can be classified into three categories: quality control, quality assessment, and internal corrective measures. Typically, these activities work cyclically: quality control is performed before and during the analysis, while quality assessment occurs during and after the investigation. Internal corrective measures are implemented based on the findings...
155
Data Validation01:15

Data Validation

183
Method validation is a crucial process in analytical chemistry designed to confirm that a given method consistently produces reliable and high-quality results. This process is essential when a method is applied to different sample matrices or when procedural modifications are made, ensuring that the results meet acceptable standards across various applications.
Key parameters for method validation include:
183
Protein Folding Quality Check in the RER01:29

Protein Folding Quality Check in the RER

3.7K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
3.7K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

1.6K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Qualitative Analysis01:10

Qualitative Analysis

297
Qualitative analysis is the process of identifying elements, ions, or compounds in an unknown sample. It is the first and most fundamental type of analysis based on the hierarchy of analytical goals. This hierarchy is significant as it provides a structured approach to scientific research, with qualitative analysis serving as the initial step, providing essential information before moving on to quantitative or other forms of analysis.
There are two main approaches to qualitative analysis:...
297
Review and Preview01:13

Review and Preview

9.0K
Data are individual items of information obtained from a population or sample. Data may be classified as qualitative (categorical), quantitative continuous, or quantitative discrete. Because it is not practical to measure the entire population in a study, researchers use samples to represent the population. A random sample is a representative group from the population chosen by using a method that gives each individual in the population an equal chance of being included in the sample. Random...
9.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Prevalence and risk factors of hypertension in rural Bangladesh: A population-based cross-sectional study.

PloS one·2026
Same author

Health selection at out- and return migration in Manitoba, Canada. A retrospective matched cohort register study.

Social science & medicine (1982)·2026
Same author

Impact of grassroots development of interprofessional team-based practices: Retrospective matched cohort study using ICES data.

Canadian family physician Medecin de famille canadien·2026
Same author

Implementing the screening for poverty and related social determinants and intervening to improve knowledge of and links to resources (SPARK) in primary care clinics across Canada.

Family practice·2026
Same author

Toward a family medicine capability framework for Canada: From competence to capability.

Canadian family physician Medecin de famille canadien·2026
Same author

Timeliness, continuity or travel time: results of a choice-based analysis exploring public preferences for access to primary care in Canada.

BMC primary care·2026

Related Experiment Video

Updated: Jul 17, 2025

Optimization for Sequencing and Analysis of Degraded FFPE-RNA Samples
07:30

Optimization for Sequencing and Analysis of Degraded FFPE-RNA Samples

Published on: June 8, 2020

12.1K

A scoping review of preprocessing methods for unstructured text data to assess data quality.

Marcello Nesca1,2, Alan Katz1,2,3, Carson K Leung4

  • 1Department of Community Health Sciences, University of Manitoba, Winnipeg, MB, Canada.

International Journal of Population Data Science
|September 6, 2023
PubMed
Summary

Natural language processing (NLP) methods preprocess unstructured text data (UTD) for research. This review examines NLP preprocessing techniques and their impact on data quality, particularly for electronic medical record (EMR) data.

Keywords:
data qualitynatural language processingreview

More Related Videos

Author Spotlight: Investigating the Role of Repetitive DNA Misregulation in Cancer Initiation and Immunotherapy Resistance
04:58

Author Spotlight: Investigating the Role of Repetitive DNA Misregulation in Cancer Initiation and Immunotherapy Resistance

Published on: December 13, 2024

2.5K
Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
10:41

Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms

Published on: May 9, 2017

9.3K

Related Experiment Videos

Last Updated: Jul 17, 2025

Optimization for Sequencing and Analysis of Degraded FFPE-RNA Samples
07:30

Optimization for Sequencing and Analysis of Degraded FFPE-RNA Samples

Published on: June 8, 2020

12.1K
Author Spotlight: Investigating the Role of Repetitive DNA Misregulation in Cancer Initiation and Immunotherapy Resistance
04:58

Author Spotlight: Investigating the Role of Repetitive DNA Misregulation in Cancer Initiation and Immunotherapy Resistance

Published on: December 13, 2024

2.5K
Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
10:41

Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms

Published on: May 9, 2017

9.3K

Area of Science:

  • Computer Science
  • Health Informatics
  • Data Science

Background:

  • Unstructured text data (UTD) is increasingly prevalent in research databases, including electronic medical records (EMRs).
  • The quality of UTD significantly impacts its research utility.
  • Natural Language Processing (NLP) techniques are commonly used to preprocess UTD for analysis, but methods vary and can affect data quality.

Purpose of the Study:

  • To systematically document current research and practices concerning NLP preprocessing methods for UTD.
  • To identify methods used to describe or improve the quality of UTD, with a focus on EMR data.

Main Methods:

  • A scoping review of peer-reviewed studies published between December 2002 and January 2021.
  • Searches were conducted across Scopus, Web of Science, ProQuest, and EBSCOhost.
  • Data extracted included article characteristics, data types, preprocessing methods, and data quality aspects, analyzed via narrative synthesis.

Main Results:

  • 41 articles were included, with over 50% published from 2016-2021; nearly 20% were in health science journals.
  • Common NLP preprocessing methods included stop word removal, punctuation/number removal, tokenization, and part-of-speech tagging.
  • EMR data quality concerns involved misspellings, de-identification, word variability, noise, annotation quality, and abbreviation ambiguity.

Conclusions:

  • Various NLP techniques exist for UTD preprocessing, with specific adaptations for EMR data.
  • Data quality dimensions for UTD share similarities with structured data.
  • Few general-purpose data quality measures exist for UTD, primarily focusing on noise measurement.