Related Experiment Video
Updated: May 23, 2026

Visualization and Quantification of High-Dimensional Cytometry Data using Cytofast and the Upstream Clustering Methods FlowSOM and Cytosplore
Published on: December 12, 2019
Multivariate analysis of flow cytometric data using decision trees
Svenja Simon1, Reinhard Guthke, Thomas Kamradt
1Research Group Systems Biology/Bioinformatics, Leibniz Institute for Natural Product Research and Infection Biology - Hans Knöll Institute Jena, Germany.
Insights
Machine learning, specifically induction of decision trees, can analyze complex flow cytometry data. This method reveals cytokine expression patterns, including dependencies on expression intensity, advancing immunological research.
Area of Science:
- Immunology
- Computational Biology
- Data Science
Background:
- Understanding host immune responses to pathogens is crucial.
- Flow cytometry is a key tool in immunology for single-cell analysis.
- High-dimensional flow cytometry data requires advanced statistical methods.
Purpose of the Study:
- To evaluate the effectiveness of the "induction of decision trees" machine learning method for analyzing flow cytometry data.
- To investigate cytokine co-expression patterns and dependencies in immune cells.
- To explore novel analytical approaches for complex immunological datasets.
Main Methods:
- Utilized supervised machine learning, specifically induction of decision trees.
- Analyzed intracellular cytokine staining data for six cytokines.
- Employed stratified fivefold cross-validation and quality criteria for tree selection.
Main Results:
- Decision trees identified known cytokine co-expression patterns.
- Revealed that cytokine expression depends not only on co-expression but also on expression intensity.
- Successfully applied induction of decision trees to high-dimensional flow cytometry data.
Conclusions:
- Induction of decision trees is a feasible method for analyzing complex flow cytometry data.
- This approach can uncover intricate patterns in cytokine expression.
- The method provides new insights into host immune system responses.
Abstract:
Characterization of the response of the host immune system is important in understanding the bidirectional interactions between the host and microbial pathogens. For research on the host site, flow cytometry has become one of the major tools in immunology. Advances in technology and reagents allow now the simultaneous assessment of multiple markers on a single cell level generating multidimensional data sets that require multivariate statistical analysis. We explored the explanatory power of the supervised machine learning method called "induction of decision trees" in flow cytometric data. In order to examine whether the production of a certain cytokine is depended on other cytokines, datasets from intracellular staining for six cytokines with complex patterns of co-expression were analyzed by induction of decision trees. After weighting the data according to their class probabilities, we created a total of 13,392 different decision trees for each given cytokine with different parameter settings. For a more realistic estimation of the decision trees' quality, we used stratified fivefold cross validation and chose the "best" tree according to a combination of different quality criteria. While some of the decision trees reflected previously known co-expression patterns, we found that the expression of some cytokines was not only dependent on the co-expression of others per se, but was also dependent on the intensity of expression. Thus, for the first time we successfully used induction of decision trees for the analysis of high dimensional flow cytometric data and demonstrated the feasibility of this method to reveal structural patterns in such data sets.

