在具有先前对对关系的层次异质数据上进行聚类
Wei Han1,2, Sanguo Zhang1,2, Hailong Gao3
1School of Mathematical Sciences, University of Chinese Academy of Sciences, Beijing, China.
BMC bioinformatics
|January 23, 2024
概括
这项研究为异质数据引入了一种新的层次聚类框架,改善了癌症亚型的发现. 纳入先前的样本关系可以提高聚类的准确性,并揭示重要的生物学见解.
科学领域:
- 统计 统计 统计 统计
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 传统的聚类方法忽视了特征差异,限制了复杂生物数据中的应用.
- 在癌症数据分析中不平等的特征治疗阻碍了准确的诊断和有效的抗癌疗法.
- 生物数据和癌症本身的异质性需要先进的集群方法.
研究的目的:
- 为包含先前对对关系的异质数据提出一个层次的集群框架.
- 通过粗略和精细的集群来描述特征差异和识别层次结构.
- 为了提高癌症亚型,并提供比现有方法更深入的生物学见解.
主要方法:
- 开发了一个对层次异质数据的集群框架.
- 采用初始分组的粗略分类和精细分类用于亚型识别.
- 集成样本的先前对对关系,以提高聚类性能.
主要成果:
- 精细的聚类成功识别了不同的癌症亚型,提供了更深入的见解.
- 该框架在整合先前信息方面表现出灵活性,提高了聚类准确度.
- 包括参数估计和结构确定在内的统计一致性属性被严格确定.
结论:
- 拟议的方法优于模拟研究中的现有方法,特别是添加了先前信息.
- 层次聚类揭示了在使用成像和OMIC数据进行肺腺癌分析时必要和合理的结构.
- 这种方法提供了对癌症异质性和潜在治疗点的更细致的理解.
相关概念视频
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
How Data are Classified: Categorical Data
32.8K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
32.8K
Phylogeny
44.1K
Phylogeny is concerned with the evolutionary diversification of organisms or groups of organisms. A group of organisms with a name is called a taxon (singular). Taxa (plural) can span different levels of the evolutionary hierarchy. For instance, the group containing all birds is a taxon (comprising the class Aves), and the group of all species of daisies (the genus Bellis) is a taxon. Phylogenies can likewise include just one genus (i.e., depict species relationships) or span an entire kingdom.
44.1K
Phylogenetic Trees
45.3K
Phylogenetic trees come in many forms. It matters in which sequence the organisms are arranged from the bottom to the top of the tree, but the branches can rotate at their nodes without altering the information. The lines connecting individual nodes can be straight, angled, or even curved.
45.3K
Test for Homogeneity
2.0K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.0K
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K


