Related Experiment Video
Updated: May 31, 2026

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions
Published on: January 26, 2024
Unifying pKa and Protonation Prediction with Sequence-Based Deep Learning.
Charlotte Infante1, Jieyu Lu1, Xiaolin Pan1
1Department of Chemistry, New York University, New York, New York 10003, United States.
Accurate prediction of acidity constants (pKa) is crucial for understanding molecular behavior. This study introduces T5pKa, a novel deep learning model that uses sequence-based methods and a curated dataset to predict microscopic pKa values, overcoming limitations of existing approaches.
Area of Science:
- Computational Chemistry
- Machine Learning in Drug Discovery
- Chemical Informatics
Background:
- Accurate prediction of acidity constants (pKa) is vital for understanding molecular properties like solubility and binding affinity.
- Scarcity of experimental microscopic pKa data and inconsistent terminology in existing datasets impede the development of reliable prediction models.
- While graph-based neural networks dominate, sequence-based deep learning for pKa prediction remains an underexplored area.
Purpose of the Study:
- To introduce a curated dataset, pKaCHU, containing 9000 experimentally derived microscopic pKa entries with ionization-state annotations.
- To develop and present T5pKa, a novel text-based transformer model for small-molecule pKa prediction utilizing sequence-based deep learning.
- To leverage a unified multitasking framework for both microstate enumeration and microscopic pKa prediction.
Main Methods:
- Developed pKaCHU, a comprehensive dataset by combining, honing, and updating experimental microscopic pKa data.
- Built T5pKa, a sequence-to-sequence model based on the T5Chem architecture, to predict molecular protonation/deprotonation.
- Employed multitask learning within T5pKa to enumerate microstates and a separate regression model for pKa value prediction.
Main Results:
- T5pKa successfully predicts microscopic pKa values by treating protonation/deprotonation as a language modeling task.
- The model demonstrates performance comparable to existing state-of-the-art pKa prediction tools across benchmark datasets.
- T5pKa offers a unified framework for microstate enumeration and pKa prediction, enhancing model development efficiency.
Conclusions:
- T5pKa represents a significant advancement in small-molecule pKa prediction using sequence-based deep learning.
- The pKaCHU dataset provides a valuable resource for training and benchmarking pKa prediction models.
- This approach offers a promising direction for improving the accuracy and efficiency of predicting key molecular properties.
Related Concept Videos
Predicting Products: SN1 vs. SN2
With increased substitution on the alkyl halide,...
Basicity of Aliphatic Amines
To measure the basicity of amines, two conventions are generally used. The first defines Kb as the basicity constant for the deprotonation reaction of water by the amine, as presented in Figure 1. Conventionally, lower Kb indicates higher...
Acid and Bases: Ka, pKa, and Relative Strengths
Relative Strengths of Conjugate Acid-Base Pairs
Protein Organization
The primary structure of a protein is its amino acid sequence.
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...

