Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

What Are Outliers?01:12

What Are Outliers?

5.1K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
5.1K
Outliers and Influential Points01:08

Outliers and Influential Points

6.2K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
6.2K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

4.0K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
4.0K
How Data are Classified: Numerical Data00:59

How Data are Classified: Numerical Data

37.7K
Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
37.7K
How Data are Classified: Categorical Data01:11

How Data are Classified: Categorical Data

44.2K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
44.2K
Data Reporting and Recording01:24

Data Reporting and Recording

5.4K
Reporting and recording are crucial in data documentation. The timely, thorough, and accurate documentation of facts is essential when recording patient data. Failure to record findings during an assessment or interpretation of a problem will result in loss of information and make the patient document unreliable. The reader is left with general impressions if the information is not specific. A recording is documenting data of the individual's health information in a traceable, secure, and...
5.4K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Association between multi-intensity olfactory test performance and the Montreal Cognitive Assessment score after adjustment for gray matter volume.

Scientific reports·2026
Same author

Axial Length Growth Reference Curves and LMS Parameters for Japanese Children and Adolescents Aged 4-20 Years: The TMM BirThree Cohort Study.

Ophthalmic & physiological optics : the journal of the British College of Ophthalmic Opticians (Optometrists)·2026
Same author

Maternal Cardiovascular Health During Pregnancy and Offspring Developmental Delay.

JAMA network open·2026
Same author

Bidirectional associations of screen time with externalizing and internalizing behaviors in children: a random intercept cross-lagged panel model.

American journal of epidemiology·2026
Same author

Knowledge of Colorectal Cancer Risk and Cancer Screening, with Colorectal Cancer Screening Attendance: A Nationwide Study in Japan (INFORM Study, 2020).

Asian Pacific journal of cancer prevention : APJCP·2026
Same author

Response to: Comment on: Association of subjective and objective physical activity with home hypertension.

Hypertension research : official journal of the Japanese Society of Hypertension·2026

Related Experiment Video

Updated: Jan 28, 2026

Author Spotlight: Developing Efficient Cryopreservation and Biobanking Technologies for Global Reef Restoration
05:25

Author Spotlight: Developing Efficient Cryopreservation and Biobanking Technologies for Global Reef Restoration

Published on: June 7, 2024

1.3K

Outlier detection for questionnaire data in biobanks.

Rieko Sakurai1,2, Masao Ueki1,2, Satoshi Makino1,3

  • 1Statistical Genetics Team, RIKEN Center for Advanced Intelligence Project, Tokyo, Japan.

International Journal of Epidemiology
|March 9, 2019
PubMed
Summary

Biobanks need efficient outlier detection for omics and epidemiologic data. We developed kurPCA and RAMP, unsupervised machine learning methods that effectively identify data errors, reducing manual effort in large-scale biobanks.

Keywords:
Outlier detectionanomaly detectionkurtosisprincipal component analysisregression adjustment

More Related Videos

Biobank for Translational Medicine: Standard Operating Procedures for Optimal Sample Management
08:01

Biobank for Translational Medicine: Standard Operating Procedures for Optimal Sample Management

Published on: November 30, 2022

5.6K
Creation and Maintenance of a Living Biobank - How We Do It
13:08

Creation and Maintenance of a Living Biobank - How We Do It

Published on: April 10, 2021

8.1K

Related Experiment Videos

Last Updated: Jan 28, 2026

Author Spotlight: Developing Efficient Cryopreservation and Biobanking Technologies for Global Reef Restoration
05:25

Author Spotlight: Developing Efficient Cryopreservation and Biobanking Technologies for Global Reef Restoration

Published on: June 7, 2024

1.3K
Biobank for Translational Medicine: Standard Operating Procedures for Optimal Sample Management
08:01

Biobank for Translational Medicine: Standard Operating Procedures for Optimal Sample Management

Published on: November 30, 2022

5.6K
Creation and Maintenance of a Living Biobank - How We Do It
13:08

Creation and Maintenance of a Living Biobank - How We Do It

Published on: April 10, 2021

8.1K

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Epidemiology

Background:

  • Biobanks collect vast amounts of omics and epidemiologic data.
  • Data cleaning is labor-intensive due to data complexity and scale.
  • Efficient outlier detection methods are crucial for data integrity.

Purpose of the Study:

  • To develop an unsupervised machine learning method for outlier detection in biobank data.
  • To improve outlier detection accuracy through a novel regression adjustment approach.
  • To reduce manual effort in data cleaning processes.

Main Methods:

  • Developed kurPCA (kurtosis-based Principal Component Analysis) for unsupervised outlier detection.
  • Proposed RAMP (Regression Adjustment for data by Missing Patterns) to enhance detection.
  • Applied and validated methods on large-scale biobank epidemiological data.

Main Results:

  • The combination of kurPCA and RAMP effectively identified known errors and inconsistencies.
  • Methods demonstrated good performance in both simulations and real-world biobank data.
  • Successful application in the Tohoku Medical Megabank Organization (Japan).

Conclusions:

  • The proposed kurPCA and RAMP methods are effective for outlier detection in biobanks.
  • These methods significantly reduce manual labor in data cleaning.
  • The approaches are versatile and applicable to various practical data analysis scenarios.