Related Experiment Videos
From Biomedical Datasets to Fairness-Aware Recommendations: An Integrated Data Orchestration Pipeline for Binary
Marta Alberola1, Pedro Copado1, Alfredo Vellido1
1Intelligent Data Science and Artificial Intelligence (IDEAI-UPC) Research Centre, Universitat Politècnica de Catalunya - Barcelona Tech (UPC), Barcelona 08034, Spain.
Abstract:
Many problems in biomedicine can be posed as binary classification. When they are addressed using artificial intelligence methods, though, average performance alone does not show whether a dataset is artificial intelligence ready, whether the endpoint is clinically valid, or whether errors are unevenly distributed across patient subgroups. This article presents the Fairness-Aware Data Orchestration Pipeline (FADOP), a reusable workflow that analyzes biomedical datasets, trains baseline binary classifiers, audits subgroup error patterns, tests mitigation strategies, and generates a documented recommendation. Such a pipeline is intended for systematic evaluation before clinical translation, not as an automatic deployment tool. Two publicly available case studies illustrate its use: the HIV-related ACTG175 dataset was repurposed from a treatment-comparison trial into a 1-year baseline mortality-prediction task, with death by day 365 as the positive class rather than the cid AIDS/failure composite endpoint; then, a stroke-risk dataset was analyzed as direct event prediction. The case studies show how the same workflow can generate cohort, performance, fairness, mitigation, and recommendation evidence across different rare-event clinical datasets.