Related Experiment Video
Updated: Jan 8, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
AttentionScore: A Target-Specific, Bias-Aware Scoring Function for Structure-Based Virtual Screening: A Case Study on
Muhammad Junaid1,2, Muhammad Zeeshan3, Wenjin Li1
1Institute for Advanced Study, Shenzhen University, Shenzhen 518060, China.
Abstract:
Target-specific scoring functions offer a promising route to improve structure-based virtual screening beyond generic, bias-prone scoring schemes. Here, we introduce AttentionScore, a deep learning-based scoring function for METTL3 that integrates ligand-only and protein-ligand interaction information within a single end-to-end architecture. AttentionScore combines multihead attention-based encoders, joint autoencoder-style latent representations, and a multiscale fusion module that couples PLEC interaction fingerprints with ligand fingerprints (Avalon/ECFP4), enabling the model to exploit both interaction patterns and ligand chemotypes. To obtain a bias-aware evaluation, we construct a similarity-constrained (SC) split using ECFP4 Tanimoto clustering (τ = 0.50), define a harder extrapolative test set (Set 2) containing only molecules with Tc ≤ 0.50 to all training compounds, and systematically compare against scaffold-based and random splits. Cross-set similarity diagnostics with ECFP4 and PLEC fingerprints, together with threshold-sensitivity analyses, confirm that the SC protocol minimizes analogue leakage relative to conventional splits. On the SC test set (Set 1), AttentionScore (PLEC + Avalon) achieves PR-AUC = 0.9609, Precision = 0.9698, Recall = 0.7277, F1 = 0.8057, and MCC = 0.7388, and it maintains strong performance on the stricter Set 2 while consistently outperforming generic scoring functions and machine-learning baselines under matched conditions. Statistical analysis using paired Wilcoxon tests, bootstrap confidence intervals, and effect sizes supports that these gains are robust rather than split-specific artifacts. All data, code, pretrained models, and a Streamlit-based graphical interface for nonexpert use are publicly available, providing a transparent and accessible framework for bias-aware, target-specific virtual screening.

