Related Experiment Video
Updated: Mar 31, 2026

08:03
Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
2.9K
Privacy-preserving verification of preprocessing in federated learning for genomic data.
Wenbiao Li1, Anisa Halimi2, Jaideep Vaidya3
1Case Western Reserve University, Cleveland, OH, United States.
JAMIA Open
|March 30, 2026
Summary
Federated genomic studies can now verify identical data preprocessing using differentially private LIME explanations. This method ensures data integrity without revealing sensitive genetic information.
Area of Science:
- Genomics
- Computational Biology
- Privacy-Preserving Technologies
Background:
- Federated learning (FL) enables collaborative genomic studies without centralizing raw data.
- Ensuring consistent data preprocessing across sites is crucial for reliable results.
- Protecting participant privacy is paramount in genomic data analysis.
Purpose of the Study:
- To develop and validate a method for verifying identical preprocessing pipelines in federated genomic studies.
- To ensure data integrity and reproducibility without compromising raw genotype privacy.
Main Methods:
- Institutions applied local differential privacy (LDP) to perturb genomic data slices.
- A RandomForest classifier was trained locally, and LIME explanation vectors were transmitted.
- A central server simulated preprocessing combinations and trained a classifier to detect site configurations.
Main Results:
- Centralized simulations achieved 80% accuracy in verifying preprocessing across 15 configurations.
- Membership-inference attack power was maintained below 0.05 at ε=3.
- Distributed FL experiments reached 70% accuracy in binary compatibility detection.
Conclusions:
- A single differentially private explanation vector serves as an auditable preprocessing fingerprint.
- The framework demonstrates feasible automated preprocessing verification in federated genomic consortia.
- This approach maintains participant privacy while ensuring data standardization.
Related Concept Videos
Improving Translational Accuracy
15.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.5K
Improving Translational Accuracy
3.8K
3.8K
Censoring Survival Data
649
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
649