NHash: Randomized N-Gram Hashing for Distributed Generation of Validatable Unique Study Identifiers in Multicenter

Guo-Qiang Zhang1, Shiqiang Tao, Guangming Xing

  • 1Institute of Biomedical Informatics, University of Kentucky, Lexington, KY, United States.

JMIR Medical Informatics
|November 12, 2015
PubMed
Summary

Researchers developed Randomized N-gram Hashing (NHash) for unique, error-tolerant study identifiers in multicenter research. This method ensures data integrity and privacy across distributed networks, successfully linking deidentified clinical data.

Related Concept Videos

Randomized Experiments01:13

Randomized Experiments

The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
9.3K
Bioequivalence Experimental Study Designs: Completely Randomized and Randomized Block Designs01:20

Bioequivalence Experimental Study Designs: Completely Randomized and Randomized Block Designs

Bioequivalence experimental study designs are crucial methodologies used in evaluating and comparing the bioavailability of different drug products. These designs are categorized into various types: completely randomized, randomized block, repeated measures, cross and carry-over, and Latin square designs.Completely randomized designs involve randomly allocating treatments to all subjects participating in the experiment. This allocation is achieved by assigning unique random numbers to subjects...
368
Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
16.6K