Related Experiment Video
Updated: Mar 9, 2026

Flow Cytometric Characterization of Murine B Cell Development
Published on: January 22, 2021
Leveraging Kappa-Lambda Signatures in a Multistage Machine Learning Pipeline for B-Cell Lymphoma Detection by Flow
Iris Zhang1, Sulov Chalise2, Mikhail Roshal2
1Department of Biostatistics, School of Global Public Health, New York University, New York, New York.
Insights
This study introduces a machine learning pipeline for B-cell lymphoma detection using flow cytometry. Integrating immunoglobulin light chain signatures significantly improves diagnostic accuracy and reproducibility.
Area of Science:
- Hematology
- Computational Biology
- Immunology
Background:
- Manual interpretation of flow cytometry data for B-cell lymphoma diagnosis is subjective and time-consuming.
- Existing computational methods often fail to incorporate crucial biological principles like immunoglobulin light chain restriction.
Purpose of the Study:
- To develop a biologically informed machine learning pipeline for accurate and reproducible B-cell lymphoma detection.
- To integrate immunoglobulin kappa (IGK) and lambda (IGL) signatures into an automated analysis.
Main Methods:
- A three-stage XGBoost machine learning pipeline was developed using 21 immunophenotypic markers on over 15 million single-cell events.
- The pipeline sequentially classified light chain expression, cell phenotypes, and sample-level predictions, incorporating IGK/IGL signatures.
Main Results:
- The IGK/IGL classifier achieved 88.0% test accuracy (AUC 0.957), and cell-level classification reached 92.9% accuracy (AUC 0.983).
- Sample-level classification achieved 94.7% accuracy (AUC 0.976), with IGK/IGL enrichment being the most informative feature.
Conclusions:
- Incorporating biologically grounded features like IGK/IGL signatures enhances automated flow cytometry analysis accuracy and interpretability.
- This approach provides a scalable, reproducible, and clinically relevant alternative to manual review for B-cell lymphoma diagnosis.
Abstract:
Flow cytometry immunophenotyping is essential for diagnosing B-cell lymphomas, but manual interpretation of high-dimensional data remains subjective, time-consuming, and prone to interoperator variability. Previous computational approaches often overlook clinically relevant principles, such as Ig light chain restriction. To address this gap, a biologically informed, three-stage machine learning pipeline that integrates Ig κ (IGK) and Ig λ (IGL) signatures to improve B-cell lymphoma detection was developed. A total of 200 peripheral blood samples (100 normal, 100 abnormal) were analyzed, comprising >15 million single-cell events characterized by 21 immunophenotypic markers. Three XGBoost models were trained sequentially: the first classified light chain expression (IGK, IGL, or nuisance), the second identified cell phenotypes using marker intensities and IGK/IGL-based neighborhood enrichment, and the third produced sample-level predictions based on aggregated cell features. The IGK/IGL classifier achieved 88.0% test accuracy [area under the receiver operating characteristic curve (AUC), 0.957], whereas the cell-level classification reached 92.9% accuracy (AUC, 0.983), with IGK/IGL enrichment as the most informative feature. Similarly, sample-level classification achieved 94.7% accuracy (AUC, 0.976), with improved performance when IGK/IGL enrichment was included. These findings demonstrate that incorporating biologically grounded features enhances both the accuracy and interpretability of automated flow cytometry analysis. This approach offers a scalable, reproducible, and clinically aligned alternative to the manual review of flow cytometry data for B-cell lymphomas.

