Jove
Visualize
Contact Us

Related Concept Videos

What Are Outliers?01:12

What Are Outliers?

4.6K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
4.6K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

2.8K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.8K
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

6.5K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.5K
Outliers and Influential Points01:08

Outliers and Influential Points

4.9K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.9K
Modified Boxplots00:57

Modified Boxplots

10.5K
A standard box and whisker plot informs us about the spread of the data in a given sample. One can identify the minimum value, maximum value, first quartile value, second quartile or median value, and third quartile.
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
10.5K
Cluster Sampling Method01:20

Cluster Sampling Method

13.3K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
13.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

How the Outliers Influence the Quality of Clustering?

Entropy (Basel, Switzerland)·2022
Same author

Exploration of Outliers in If-Then Rule-Based Knowledge Bases.

Entropy (Basel, Switzerland)·2020
See all related articles
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Oct 25, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.1K

Qualitative Data Clustering to Detect Outliers.

Agnieszka Nowak-Brzezińska1, Weronika Łazarz1

  • 1Institute of Computer Science, Faculty of Science and Technology, University of Silesia, Bankowa 12, 40-007 Katowice, Poland.

Entropy (Basel, Switzerland)
|August 6, 2021
PubMed
Summary

This study evaluates three different computational methods for identifying unusual data points, known as outliers, within datasets composed entirely of non-numerical, categorical information. By comparing the K-modes, STIRR, and ROCK algorithms, the researchers determine how effectively each approach isolates anomalies. The results highlight how different algorithmic settings and dataset characteristics influence the detection of these rare observations.

Keywords:
K-modesROCKSTIRRdata clusteringoutliers detectionqualitative datacategorical variablesK-modes algorithmSTIRR algorithmROCK algorithmdata mining

Frequently Asked Questions

More Related Videos

A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
08:33

A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences

Published on: September 4, 2019

7.2K
ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
05:12

ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data

Published on: January 16, 2019

11.6K

Related Experiment Videos

Last Updated: Oct 25, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.1K
A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
08:33

A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences

Published on: September 4, 2019

7.2K
ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
05:12

ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data

Published on: January 16, 2019

11.6K

Area of Science:

  • Data science and qualitative data clustering methodologies
  • Statistical analysis within computational intelligence

Background:

Identifying anomalies remains a persistent challenge across statistics, machine learning, and data mining fields. Most existing strategies prioritize large numerical datasets, leaving a significant gap in handling non-numerical information. Researchers have developed numerous techniques to isolate unusual behaviors from common observations. However, these tools often struggle when applied to categorical variables rather than quantitative values. This discrepancy creates a clear limitation in current analytical frameworks for non-numerical data. That uncertainty drove the need to explore specialized approaches for qualitative variable sets. Prior research has shown that standard methods are frequently ill-suited for these specific data types. No prior work had resolved the comparative effectiveness of categorical clustering for outlier identification until now.

Purpose Of The Study:

This study aims to compare three specific categorical data clustering algorithms for their ability to detect outliers. The researchers address the problem of identifying unusual behavior in datasets containing only non-numerical variables. While many solutions exist for quantitative sets, fewer options are available for qualitative data. This gap motivated the team to investigate how different methods perform under varying conditions. They sought to determine if these algorithms identify anomalies similarly or if they produce divergent results. The authors also examined how individual parameters and set characteristics influence the detection process. By analyzing these factors, the study provides insight into the reliability of clustering techniques for categorical information. This investigation clarifies the dependencies between algorithmic settings and the identification of rare data points.

Main Methods:

The authors conducted a comparative analysis of three distinct categorical clustering algorithms. They utilized the K-modes approach, derived from the classic K-means framework, alongside the STIRR and ROCK techniques. The team performed experiments using several datasets that varied in object counts and variable numbers. This review approach involved adjusting specific algorithmic parameters to observe changes in output. They systematically examined how these settings influenced the formation of clusters. The researchers also evaluated the sensitivity of each method to the number of categories present. By testing these variables, the study assessed the consistency of outlier identification across different configurations. This structured evaluation provided a clear basis for comparing the performance of the selected computational tools.

Main Results:

The study reveals that the three algorithms produce different outcomes when identifying unusual observations in categorical sets. Key findings from the literature suggest that the number of tuples and variables significantly affects detection rates. The researchers observed that K-modes, STIRR, and ROCK do not detect outliers in a uniform manner. Algorithmic sensitivity varies greatly depending on the specific parameters chosen by the user. The analysis shows that the distribution of categories within the data influences the final cluster assignments. Each method demonstrates unique strengths and weaknesses when processing non-numerical information. The results indicate that the complexity of the dataset directly impacts the stability of the detected anomalies. These findings confirm that the choice of clustering tool is a critical factor for accurate outlier isolation.

Conclusions:

The researchers demonstrate that the three evaluated algorithms exhibit distinct behaviors when isolating anomalies in categorical sets. These findings suggest that the choice of algorithm significantly impacts the identification of unusual observations. The study highlights how varying input parameters alters the resulting cluster structures and outlier counts. Authors propose that dataset size and the number of categories influence the performance of these tools. The analysis indicates that no single algorithm consistently outperforms the others across all tested conditions. These results provide a framework for selecting appropriate methods based on specific data characteristics. The authors suggest that future applications should carefully calibrate parameters to improve detection accuracy. This synthesis clarifies the operational differences between K-modes, STIRR, and ROCK in qualitative environments.

The researchers propose that these algorithms identify anomalies by partitioning datasets into distinct groups. K-modes, STIRR, and ROCK utilize different mathematical logic to isolate rare observations compared to the majority of data points. Each method produces unique cluster assignments based on categorical variable distributions.

The study evaluates the K-modes algorithm, which adapts MacQueen's K-means approach, alongside the STIRR and ROCK methods. These three techniques represent different computational strategies for grouping non-numerical information to reveal hidden patterns or unusual entries.

The authors indicate that testing across multiple datasets is necessary to account for variations in object counts and variable types. This approach ensures that the performance of each algorithm is measured against diverse structural complexities found in real-world qualitative data.

The researchers utilize qualitative variables to assess how different algorithms handle non-numerical attributes. This data type is central to the study, as it allows for the evaluation of clustering performance in environments where traditional quantitative metrics are unavailable.

The study measures the sensitivity of each algorithm to changes in input parameters and dataset dimensions. By observing how these factors alter outlier detection, the authors quantify the reliability of each method under varying conditions.

The authors suggest that practitioners must consider the specific characteristics of their dataset, such as the number of categories, when choosing an outlier detection method. This implication emphasizes that algorithmic selection is highly dependent on the underlying structure of the qualitative data.