Multivariate analysis of flow cytometric data using decision trees

Svenja Simon1, Reinhard Guthke, Thomas Kamradt

  • 1Research Group Systems Biology/Bioinformatics, Leibniz Institute for Natural Product Research and Infection Biology - Hans Knöll Institute Jena, Germany.

Insights

Machine learning, specifically induction of decision trees, can analyze complex flow cytometry data. This method reveals cytokine expression patterns, including dependencies on expression intensity, advancing immunological research.

Area of Science:

  • Immunology
  • Computational Biology
  • Data Science

Background:

  • Understanding host immune responses to pathogens is crucial.
  • Flow cytometry is a key tool in immunology for single-cell analysis.
  • High-dimensional flow cytometry data requires advanced statistical methods.

Purpose of the Study:

  • To evaluate the effectiveness of the "induction of decision trees" machine learning method for analyzing flow cytometry data.
  • To investigate cytokine co-expression patterns and dependencies in immune cells.
  • To explore novel analytical approaches for complex immunological datasets.

Main Methods:

  • Utilized supervised machine learning, specifically induction of decision trees.
  • Analyzed intracellular cytokine staining data for six cytokines.
  • Employed stratified fivefold cross-validation and quality criteria for tree selection.

Main Results:

  • Decision trees identified known cytokine co-expression patterns.
  • Revealed that cytokine expression depends not only on co-expression but also on expression intensity.
  • Successfully applied induction of decision trees to high-dimensional flow cytometry data.

Conclusions:

  • Induction of decision trees is a feasible method for analyzing complex flow cytometry data.
  • This approach can uncover intricate patterns in cytokine expression.
  • The method provides new insights into host immune system responses.