Related Experiment Video
Updated: Mar 27, 2026

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
Published on: June 21, 2018
A heterogeneous graph neural network for candidate gene prediction in endometriosis
Zelia Soo1, Hua Lin2, Mengjia Wu1
1Australian Artificial Intelligence Institute, Faculty of Engineering and Information Technology, University of Technology Sydney, 61 Broadway, Ultimo, 2007, NSW, Australia.
Objective:
Precision medicine applications depend on elucidating the underlying molecular mechanisms of disease. However, many disorders, such as endometriosis, remain poorly characterised genetically due to data scarcity, positive-unlabelled (PU) imbalance, and the heterogeneous structure of biomedical knowledge. This study aims to develop HetBio-CLiP, a novel heterogeneous graph-based contrastive learning (CL) methodology that improves candidate gene prioritisation across heterogeneous biomedical graphs by explicitly addressing positive-unlabelled learning constraints and structural heterogeneity.
Methods:
HetBio-CLiP (Heterogeneous Biomedical graph Contrastive Learning with Interactions and Positive-Unlabelled learning) integrates multi-relational genomic, variant, and clinical data from real-world patient cohorts. To address the challenges of data scarcity and class imbalance, the graph neural network (GNN) methodology combines heterogeneous graph CL with PU learning. Furthermore, the model incorporates an interpretable GNNShap explainer to provide transparency at both the feature and edge levels, to assess the biological relevance of the predictions.
Results:
HetBio-CLiP achieved superior performance across key evaluation metrics, attaining the highest Area Under the Curve (AUC) of 0.9489 ± 0.04 and Area Under the Precision-Recall Curve (AUPR) of 0.9401 ± 0.05. This outperformed state-of-the-art baselines such as M-SAGEGraph (AUC 0.9149 ± 0.04) and GAT (AUC 0.9138 ± 0.07). In terms of ranking candidate genes, the model achieved the highest Precision@42 (0.6933 ± 0.06) and TP@42 (29.1 ± 2.4) and second-highest Normalised Discounted Cumulative Gain (NDCG@42) (0.827 ± 0.06). Ablation studies confirmed that the integration of PU and contrastive learning was essential for these performance gains. Biologically, the model successfully prioritised known pathophysiology drivers, with top-ranked validated gene predictions including WNT4, GREB1, and ESR1.
Conclusion:
HetBio-CLiP presents an effective and interpretable methodology for uncovering candidate gene-disease relationships in data-scarce environments. By combining CL with PU-training, the proposed method addresses critical limitations of current gene prioritisation models and produces biologically meaningful predictions. The model demonstrates a strong balance between ranking accuracy and recall, identifying broader sets of plausible candidates than existing methods. This approach benefits researchers, clinicians, and downstream genomic studies by offering an improved foundation for identifying novel disease-associated genes and guiding future experimental investigation in endometriosis and other complex diseases.
Related Concept Videos
End Point Prediction: Gran Plot
For potentiometric titration, the Gran plot is created by plotting...
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...

