Related Experiment Video
Updated: Mar 21, 2026

07:24
A Precision Medicine Tool for Measurement and Monitoring of Hemoglobin S in Sickle Cell Disease Patients Receiving Transfusion Therapy
2.0K
Novel use of a hash-based tokenization for sharing sickle cell disease data without sharing protected health
Najibah Galadanci1, Gerhard Stefan Hellemann2, Samuel David Washko3
1Lifespan Comprehensive Sickle Cell Center, University of Alabama at Birmingham, Birmingham, AL.
Blood Advances
|March 20, 2026
Summary
Researchers developed a secure method to link sickle cell disease (SCD) registries, creating a unified dataset. This approach overcomes data fragmentation, enabling better understanding of SCD progression and treatment development.
Area of Science:
- Biomedical Informatics
- Hematology
- Data Science
Background:
- Sickle cell disease (SCD) care improvements are hindered by fragmented, uncoordinated registries.
- Existing data aggregation is limited by poor interoperability and inconsistent data elements.
- Optimizing current data resources is crucial to avoid data loss and advance SCD research.
Purpose of the Study:
- To develop a privacy-preserving method for securely linking multiple SCD data collection efforts.
- To create a unified, longitudinal dataset for comprehensive SCD research.
Main Methods:
- Leveraged IRB-approved access to three major US SCD registries.
- Generated and hashed identity tokens using SHA-256 for secure, privacy-preserving linkage.
- Utilized deterministic matching of hashed tokens to identify unique individuals across datasets.
Main Results:
- Successfully linked data from the Sickle Cell Data Collection project, ASH Research Collaborative Data Hub, and GRNDaD.
- Identified 1,080 unique individuals across at least two of the three registries.
- Demonstrated the first privacy-preserving linkage of multiple SCD registries.
Conclusions:
- Secure data integration enhances interoperability for SCD research.
- This method enables richer longitudinal analyses for advancing SCD treatment.
- Optimized data linkage is vital for understanding lifelong SCD progression.
Related Concept Videos
Single Nucleotide Polymorphisms-SNPs
19.7K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
19.7K
Genome-wide Association Studies-GWAS
16.5K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
16.5K

