Clustering on hierarchical heterogeneous data with prior pairwise relationships
Wei Han1,2, Sanguo Zhang1,2, Hailong Gao3
1School of Mathematical Sciences, University of Chinese Academy of Sciences, Beijing, China.
BMC Bioinformatics
|January 23, 2024
Summary
This study introduces a novel hierarchical clustering framework for heterogeneous data, improving cancer subtype discovery. Incorporating prior sample relationships enhances clustering accuracy and reveals essential biological insights.
Area of Science:
- Statistics
- Bioinformatics
- Computational Biology
Background:
- Traditional clustering methods overlook feature differences, limiting applications in complex biological data.
- Unequal feature treatment in cancer data analysis hinders accurate diagnosis and effective anti-cancer therapies.
- Heterogeneity in biological data and cancer itself necessitates advanced clustering approaches.
Purpose of the Study:
- To propose a hierarchical clustering framework for heterogeneous data incorporating prior pairwise relationships.
- To characterize feature differences and identify hierarchical structures through rough and refined clustering.
- To enhance cancer subtyping and provide deeper biological insights beyond existing methods.
Main Methods:
- Developed a clustering framework for hierarchical heterogeneous data.
- Employed rough clustering for initial groupings and refined clustering for subtype identification.
- Integrated prior pairwise relationships of samples to improve clustering performance.
Main Results:
- The refined clustering successfully identified distinct cancer subtypes, offering deeper insights.
- The framework demonstrated flexibility in incorporating prior information, boosting clustering accuracy.
- Statistical consistency properties, including parameter estimation and structure determination, were rigorously established.
Conclusions:
- The proposed method outperforms existing approaches in simulation studies, especially with added prior information.
- Hierarchical clustering revealed necessary and reasonable structures in lung adenocarcinoma analysis using imaging and omics data.
- This approach offers a more nuanced understanding of cancer heterogeneity and potential therapeutic targets.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
How Data are Classified: Categorical Data
32.8K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
32.8K
Phylogeny
44.1K
Phylogeny is concerned with the evolutionary diversification of organisms or groups of organisms. A group of organisms with a name is called a taxon (singular). Taxa (plural) can span different levels of the evolutionary hierarchy. For instance, the group containing all birds is a taxon (comprising the class Aves), and the group of all species of daisies (the genus Bellis) is a taxon. Phylogenies can likewise include just one genus (i.e., depict species relationships) or span an entire kingdom.
44.1K
Phylogenetic Trees
45.3K
Phylogenetic trees come in many forms. It matters in which sequence the organisms are arranged from the bottom to the top of the tree, but the branches can rotate at their nodes without altering the information. The lines connecting individual nodes can be straight, angled, or even curved.
45.3K
Test for Homogeneity
2.0K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.0K
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K


