Related Experiment Video
Updated: Apr 22, 2026

PAR-CliP - A Method to Identify Transcriptome-wide the Binding Sites of RNA Binding Proteins
Published on: July 2, 2010
KANBind as a diagnostic probe for DNA-binding protein prediction: A prevalence-calibrated reality check under strict
Qipeng Wen1, Shaohua Jiang1, Yiwen Zhang1
1College of Information Science and Engineering, Hunan Normal University, Changsha, P. R. China.
Abstract:
Deep learning reports over 90% DNA-binding protein (DBP) prediction performance on common benchmarks, but these results are usually obtained on balanced test sets and may not translate to proteome-wide scans with extreme class imbalance. Here, we use KANBind as a diagnostic probe to stress-test sequence-based DBP prediction under strict homology control and realistic prevalence. Evaluated on the homology-controlled HBTD benchmark with prevalence-calibrated reporting, KANBind achieves a calibrated precision of 0.0558 at a realistic bacterial prevalence ([Formula: see text]), implying an expected false discovery rate (FDR) of 94.42%. In a proteome-scale scan, this corresponds to approximately 95 false positives per 100 predicted DBPs. Interpretability analysis indicates that predictions are driven mainly by coarse physicochemical cues such as electrostatics, which may be necessary for DNA binding but are insufficient to determine DBP function. Together, these results suggest that apparent benchmark gains can be dominated by homology leakage and evaluation on balanced sets rather than by generalizable functional rules, motivating stress-test benchmarks with strict homology control and realistic negative backgrounds.
Related Concept Videos
Labeling DNA Probes
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
Single-Strand DNA Binding Proteins
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...

