Related Experiment Video
Updated: Oct 11, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
On Using Classification Datasets to Evaluate Graph Outlier Detection: Peculiar Observations and New Insights
1Heinz College Information Systems & Public Policy, Carnegie Mellon University, Pittsburgh, Pennsylvania, USA.
Repurposing graph classification datasets for graph-level outlier detection (GLOD) causes performance flips. Model performance drastically changes based on which class is down-sampled, highlighting issues with current evaluation methods.
Area of Science:
- Data Mining
- Machine Learning
- Graph Analytics
Background:
- Outlier detection models are often evaluated using repurposed classification datasets.
- Graph-level outlier detection (GLOD) is an understudied area with significant real-world potential.
- Current practices repurpose graph classification datasets for GLOD by down-sampling one class to create outlier samples.
Purpose of the Study:
- To identify and analyze an issue with repurposing graph classification datasets for GLOD.
- To investigate the causes of performance variations in GLOD models.
- To question the appropriateness of current GLOD evaluation methodologies.
Main Methods:
- Investigated the impact of down-sampling different classes on ROC-AUC performance for GLOD.
- Analyzed graph embedding spaces generated by propagation-based models.
- Examined various graph embedding methods and downstream outlier detectors.
Main Results:
- A significant "performance flip" was observed, where model performance drastically changes (from high to worse-than-random) depending on the down-sampled class.
- Performance gaps were amplified by propagation in certain models.
- Disparity in within-class densities and overlapping class supports in embedding spaces were identified as key factors.
- The performance flip issue persists across different embedding methods, though the specific down-sampled version yielding higher performance may vary.
Conclusions:
- The identified performance flip issue raises concerns about the validity of averaging GLOD performance across different down-sampled dataset versions.
- There is a need to develop improved graph embedding methods to mitigate the observed performance flip.
- Further research is required to establish robust evaluation protocols for GLOD.
More Related Videos
Related Concept Videos
Outliers and Influential Points
Quantifying and Rejecting Outliers: The Grubbs Test
What Are Outliers?
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
Detection of Gross Error: The Q Test
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:

