Automated subset identification and characterization pipeline for multidimensional flow and mass cytometry data
Stephen Meehan1, Gleb A Kolyagin2, David Parks1
11Department of Genetics, Stanford University School of Medicine, Stanford, CA 94305 USA.
Communications Biology
|June 27, 2019
Summary
Researchers developed an automated pipeline for analyzing high-dimensional data, improving cluster identification and visualization. This tool enhances data analysis accuracy and reproducibility across scientific institutions.
Area of Science:
- Computational biology
- Data science
- Bioinformatics
Background:
- Automated clustering methods are crucial for analyzing high-dimensional datasets.
- Current automated methods lack intuitive visualization and cluster matching, hindering adoption.
- User-guided clustering is time-consuming and less standardized for complex datasets.
Purpose of the Study:
- To develop a fully automated pipeline for subset identification and characterization (SIC).
- To provide robust cluster matching and intuitive data visualization for high-dimensional data.
- To overcome the limitations of existing automated clustering techniques.
Main Methods:
- Developed a novel automated subset identification and characterization (SIC) pipeline.
- Integrated robust cluster matching algorithms.
- Implemented intuitive two-dimensional data visualization techniques for high-dimensional datasets.
- Ensured the method is safe from the curse of dimensionality.
Main Results:
- The SIC pipeline automatically generates intuitive 2D representations of high-dimensional data.
- Achieved robust cluster matching and enhanced data interpretability.
- Demonstrated a method that is safe from the curse of dimensionality.
Conclusions:
- The developed SIC pipeline offers a more accurate, standardized, and faster approach to data analysis.
- This tool facilitates reproducible research and the establishment of new gold standard practices.
- Enhances the interpretation of fully automatically generated clusters in flow/mass cytometry and other data types.
Related Concept Videos
Flow Cytometry
15.9K
The development of flow cytometry techniques began in 1934 with initial attempts by Andrew Moldavan, a bacteriologist who counted the cells in a flowing capillary system. Moldavan pumped cells through a capillary tube focused under a microscope for visualization. The invention of photometry allowed the measurement of differentially-stained cells, and Louis Kamentsky developed the first multiparameter flow cytometer in 1965 to identify and count the cancer cells in cervical tissue specimens.
In...
In...
15.9K
Peptide Identification Using Tandem Mass Spectrometry
8.2K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
8.2K
Cluster Sampling Method
14.1K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
14.1K
Vesicular Tubular Clusters
3.1K
After budding out from the ER membrane, some COPII vesicles lose their coat and fuse with one another to form larger vesicles and interconnected tubules called vesicular tubular clusters or VTCs. These clusters constitute a compartment at the ER-Golgi interface known as ERGIC (Endoplasmic Reticulum Golgi Intermediate Compartment). The ERGIC is a mobile membrane-bound cargo transport system that sorts proteins secreted from ER and delivers them to the Golgi.
With the help of motor proteins such...
With the help of motor proteins such...
3.1K
Atomic Mass
69.7K
Atoms — and the protons, neutrons, and electrons that compose them — are extremely small. For example, a carbon atom weighs less than 2 × 10−23 g. When describing the properties of tiny objects such as atoms, we use appropriately small units of measure, such as the atomic mass unit (amu). The amu was originally defined based on hydrogen, the lightest element, then later in terms of oxygen. Since 1961, it has been defined with regard to the most abundant isotope of carbon, atoms of which...
69.7K
Molar Mass
86.2K
The identity of a substance is defined not only by the types of atoms or ions it contains but by the quantity of each type of atom or ion. For example, water, H2O, and hydrogen peroxide, H2O2, are alike in that their respective molecules are composed of hydrogen and oxygen atoms. However, because a hydrogen peroxide molecule contains two oxygen atoms, as opposed to the water molecule, which has only one, the two substances exhibit very different properties.
86.2K


