Related Experiment Video
Updated: Aug 28, 2026

A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
Explainable Multi-Isoform QSAR, PubChem Concordance, and Applicability-Domain-Guided Prioritization of Selective
Alaa M Elsayad1, Khaled A Elsayad2
1Biomedical Group, Department of Electrical Engineering, College of Engineering, Prince Sattam Bin Abdulaziz University, Wadi Alddawasir 11991, Saudi Arabia.
Abstract:
Background/Objectives: Human carbonic anhydrase (CA) isoforms are clinically relevant zinc metalloenzymes, but their highly conserved catalytic Zn2+ site makes isoform-selective inhibitor design challenging. This study developed an explainable multi-isoform QSAR workflow for prioritizing selective inhibitors of human CA I, CA II, CA IX, and CA XII, integrating PubChem experimental concordance and applicability-domain (AD) analysis to distinguish model-supported candidates from exploratory extrapolative hypotheses. Methods: Curated ChEMBL Ki datasets were standardized to pKi and modeled in KNIME using 4185 two-dimensional descriptors, Random Forest-based feature selection, and H2O.ai AutoML stacked ensembles. A matched 3200-compound four-isoform matrix supported direct selectivity profiling. Potent ChEMBL inhibitors (Ki < 10 nM) seeded 90% PubChem similarity expansion; analogues were scored across all four isoform-specific models, while PubChem records were filtered to retain only human isoform-specific Ki evidence. AD was calibrated using Morgan fingerprint similarity, descriptor-space distance, and leverage analysis. Candidates were prioritized by predicted potency, selectivity, SAR plausibility, SwissADME profile, structural alerts, and AD membership. Results: Models achieved held-out R2 values of 0.727, 0.719, 0.652, and 0.607 for CA I, CA II, CA IX, and CA XII, respectively. PubChem concordance identified 1098 prioritized compounds with assay records, including 847 with direct Ki evidence and 428 with complete four-isoform coverage; same-target agreement showed Pearson r > 0.86 and linear-fit R2 > 0.74. Test-set error rose from very-high-AD to low/out-of-domain classes. Conclusions: The workflow integrates QSAR prediction, PubChem concordance, explainable SAR, ADME triage, and AD-based reliability assessment. CA II and CA XII candidates showed the strongest support, CA IX showed intermediate support, and CA I candidates require cautious interpretation as low-domain exploratory hypotheses.
