Related Experiment Video
Updated: Aug 15, 2026

Identification of Mouse and Human Antibody Repertoires by Next-Generation Sequencing
Published on: March 15, 2019
LM-QASAS: reference-free identification of antigen-specific sequences from the BCR repertoire using antibody language
Genki Masuda1, Yohei Funakoshi2, Shunsuke Iizumi1
1Department of Computer Science, School of Computing, Institute of Science Tokyo, Yokohama, Kanagawa, Japan.
None:
The B-cell receptor (BCR) repertoire serves as a historical record of immunological events. However, deciphering antigen-specific sequences from this vast dataset remains a challenge, particularly for novel pathogens where prior knowledge is absent. While time-course analysis methods such as QASAS have proven effective for tracking immune responses, they rely on existing antibody databases, limiting their applicability to emerging diseases. To overcome this limitation, we introduce LM-QASAS, a reference-free computational framework that integrates antibody language models (AbLMs) with longitudinal repertoire dynamics. By mapping sequences into a high-dimensional semantic embedding space, LM-QASAS identifies clusters of functionally convergent sequences that are semantically similar and exhibit transient expansion upon immune stimulation. In a SARS-CoV-2 cohort, candidate sequences extracted by LM-QASAS were enriched for overlap with the CoV-AbDab neutralizing-antibody database relative to a random baseline. Leave-one-out cross-validation showed that these reference-free databases could reconstruct longitudinal immune-response dynamics in unseen individuals without external references. Conversely, the method showed limited sensitivity in an influenza vaccine cohort, indicating that the approach is most effective under conditions of robust, synchronized clonal expansion, such as those induced by mRNA vaccination. LM-QASAS provides a rapid, reference-free approach for monitoring humoral immunity against emerging threats.
