Related Experiment Video
Updated: Aug 9, 2026

Improving Small RNA-seq: Less Bias and Better Detection of 2'-O-Methyl RNAs
Published on: September 16, 2019
CABA-Bind: Confounder-aligned backdoor adjustment for debiased RNA-ligand binding prediction
Xiaoqing Wang1, Xingyu Liu1, Zhiwei Zhang1
1Academy of Artificial Intelligence, Beijing Institute of Petrochemical Technology, Beijing 102617, China.
Context:
RNA-ligand molecular recognition is important for RNA-targeted drug discovery and candidate small-molecule prioritization. However, RNA-ligand binding datasets often contain biased associations between RNA sequences and binding labels, which may cause models to rely on sequence-driven cues rather than ligand-dependent binding features. Such sequence-bias-driven reliance can reduce model reliability in challenging prediction scenarios, especially when evaluating unseen RNA targets, structurally dissimilar ligands, or hard decoy molecules.
Methods:
We propose CABA-Bind, a causal debiasing framework for RNA-ligand binding prediction. CABA-Bind encodes RNA sequences and ligand SMILES using RNA-FM and ChemBERTa, constructs confounder centers by K-means clustering to represent recurrent RNA prior patterns, and combines RNA-Confounder Alignment with backdoor-adjusted prediction to reduce the influence of RNA sequence-driven bias. The model was evaluated on the Robin and Biosensor datasets using four data-splitting strategies, prior-dependency metrics, ablation studies, hard decoy ranking, and structure-guided interpretation. CABA-Bind reduced the Score-prior |ρ| by 40.8% compared with the baseline model, achieved an MRR of 0.65 in hard decoy evaluation, and provided computational evidence suggesting that model-highlighted RNA regions around G17 and A53/A54 may contribute to ligand-associated recognition. These results suggest that CABA-Bind improves the reliability and ligand-specific interpretability of RNA-ligand molecular recognition modeling.
Related Concept Videos
Leaky Scanning
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...

