Related Experiment Video
Updated: Feb 17, 2026

10:12
Droplet Barcoding-Based Single Cell Transcriptomics of Adult Mammalian Tissues
Published on: January 10, 2019
19.1K
Multiple-cumulative probabilities used to cluster and visualize transcriptomes
Xingang Jia1,2, Yisu Liu3, Qiuhong Han4
1School of Mathematics Southeast University Nanjing China.
FEBS Open Bio
|December 12, 2017
Summary
This study introduces a novel gene expression analysis method using Pearson
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Gene expression data analysis is crucial for biological discovery.
- Clustering and visualization are key techniques for interpreting high-dimensional gene expression data.
- Existing methods face challenges with high-dimensional data and clear cluster visualization.
Purpose of the Study:
- To develop an improved method for clustering and visualizing gene expression data.
- To address the challenge of high-dimensional Pearson's correlation coefficient of multiple-cumulative probabilities (PCC-MCP) data.
- To enhance the understanding of relationships between gene expression clusters.
Main Methods:
- Utilized Pearson's correlation coefficient of multiple-cumulative probabilities (PCC-MCP) to define gene expression similarity.
- Employed icc-cluster, an iterative clustering algorithm, with PCC-MCP for gene grouping.
- Applied t-statistic stochastic neighbor embedding (t-SNE) of KC-data for optimal cluster mapping (t-SNE-MCP-O maps).
Main Results:
- Demonstrated clear advantages of icc-cluster with PCC-MCP over conventional clustering methods across multiple transcriptome datasets.
- Showcased t-SNE-MCP-O maps providing clear projecting boundaries for PCC-MCP clusters.
- Facilitated easy visualization and understanding of relationships between gene expression clusters.
Conclusions:
- The combination of icc-cluster and PCC-MCP offers a superior approach for gene expression data analysis.
- t-SNE-MCP-O maps effectively visualize complex gene expression patterns and cluster relationships.
- This method enhances biological knowledge discovery from transcriptome data.
Related Concept Videos
RNA-seq
12.2K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.2K
Probability Histograms
13.3K
A probability histogram is a visual representation of a probability distribution. Similar a typical histogram, the probability histogram consists of contiguous (adjoining) boxes. It has both a horizontal axis and a vertical axis. The horizontal axis is labeled with what the data represents. The vertical axis is labeled with probability. Each rectangular bar in the histogram is 1 unit wide, which suggests that the area under each bar equals the probability, P(x), where x is 1, 2, 3, and so on.
13.3K
Cluster Sampling Method
14.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
14.9K
Cumulative Frequency Distribution
8.7K
A cumulative frequency distribution is another type of frequency distribution. Instead of reporting how many data values fall in some classes, it reports how many data values are contained in either that class or any class to its left. Technically, it means the sum of frequencies of the class and all the classes below it in a frequency distribution. A cumulative frequency is calculated by adding the frequency of each class lower than the corresponding class interval or category. In general, a...
8.7K
Biostatistics: Overview
908
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
908
DNA Microarrays
21.3K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
21.3K

