Related Experiment Video
Updated: May 11, 2026

Knowing What Counts: Unbiased Stereology in the Non-human Primate Brain
Published on: May 14, 2009
The depth problem: identifying the most representative units in a data group
Itziar Irigoien1, Francesc Mestres, Concepción Arenas
1Department of Computation Science and Artificial Intelligence, University of the Basque Country, Donostia, Spain. itziar.irigoien@ehu.es
Abstract:
This paper presents a solution to the problem of how to identify the units in groups or clusters that have the greatest degree of centrality and best characterize each group. This problem frequently arises in the classification of data such as types of tumor, gene expression profiles or general biomedical data. It is particularly important in the common context that many units do not properly belong to any cluster. Furthermore, in gene expression data classification, good identification of the most central units in a cluster enables recognition of the most important samples in a particular pathological process. We propose a new depth function that allows us to identify central units. As our approach is based on a measure of distance or dissimilarity between any pair of units, it can be applied to any kind of multivariate data (continuous, binary or multiattribute data). Therefore, it is very valuable in many biomedical applications, which usually involve noncontinuous data, such as clinical, pathological, or biological data sources. We validate the approach using artificial examples and apply it to empirical data. The results show the good performance of our statistical approach.
Related Concept Videos
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Problem Solving: Dimensional Analysis
What is Central Tendency?
The central tendency is the most conventionally used data characteristic. It is a...
Detection of Gross Error: The Q Test
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...

