Related Experiment Video
Updated: Mar 30, 2026

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
NHash: Randomized N-Gram Hashing for Distributed Generation of Validatable Unique Study Identifiers in Multicenter
Guo-Qiang Zhang1, Shiqiang Tao, Guangming Xing
1Institute of Biomedical Informatics, University of Kentucky, Lexington, KY, United States.
Researchers developed Randomized N-gram Hashing (NHash) for unique, error-tolerant study identifiers in multicenter research. This method ensures data integrity and privacy across distributed networks, successfully linking deidentified clinical data.
Area of Science:
- Biomedical Informatics
- Data Management
- Clinical Research
Background:
- Unique study identifiers are crucial for linking research data while protecting patient privacy.
- Traditional identifiers face challenges in large, multicenter studies due to scale and potential for errors.
- Error tolerance (validatability) is a key requirement for study identifiers to ensure data integrity.
Purpose of the Study:
- Introduce Randomized N-gram Hashing (NHash), a novel method for generating unique study identifiers.
- Enable distributed and validatable identifier generation for multicenter research.
- Ensure identifiers are pseudonymous, collision-resistant, and error-tolerant.
Main Methods:
- NHash employs a two-phase approach: randomized N-gram hashing and subsequent encryption.
- Phase 1 generates an intermediate string using N-gram hashes derived from randomized inputs.
- Phase 2 encrypts the intermediate string and concatenates it with a random number for the final identifier.
Main Results:
- Experiments with synthesized data showed a negligible probability of identifier collision with NHash.
- NHash was successfully implemented for the Center for SUDEP Research (CSR) multicenter collaboration.
- The CSR involves 14 institutions focused on understanding sudden unexpected death in epilepsy (SUDEP).
Conclusions:
- The CSR Data Repository effectively utilized NHash to link deidentified multimodal clinical data.
- NHash met all objectives for generating unique, distributed, and validatable study identifiers.
- The method supports secure data linkage in complex, multicenter research environments.
More Related Videos
Related Concept Videos
Randomized Experiments
Simple randomization
Simple...
Bioequivalence Experimental Study Designs: Completely Randomized and Randomized Block Designs
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...

