Related Experiment Video
Updated: Apr 9, 2026

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
MapReduce Based Personalized Locality Sensitive Hashing for Similarity Joins on Large Scale Data
1School of Information Science and Technology, Xiamen University, Xiamen 361005, China ; Shenzhen Research Institute of Xiamen University, Shenzhen 518058, China.
Personalized Locality Sensitive Hashing (PLSH) offers tailored control over false positives and negatives for high-dimensional data similarity joins. This novel banding scheme improves efficiency and accuracy in large-scale data processing.
Area of Science:
- Computer Science
- Data Mining
- Machine Learning
Background:
- Locality Sensitive Hashing (LSH) is vital for efficient similarity joins in high-dimensional datasets.
- LSH performance hinges on managing false positives and false negatives.
- Reducing false positives is critical, while some applications require balancing both error types.
Purpose of the Study:
- To introduce Personalized Locality Sensitive Hashing (PLSH) for fine-grained control over LSH error rates.
- To develop a novel banding scheme within PLSH to tailor false positives, false negatives, and their sum.
- To enable efficient, large-scale similarity joins using a parallel MapReduce implementation of PLSH.
Main Methods:
- Developed a novel banding scheme for Personalized Locality Sensitive Hashing (PLSH).
- Implemented PLSH in parallel using the MapReduce framework for scalability.
- Conducted experiments on both real and simulated high-dimensional datasets.
Main Results:
- PLSH effectively allows users to adjust the trade-off between false positives and false negatives.
- The parallel MapReduce implementation demonstrates significant efficiency for large-scale data.
- Experimental results confirm PLSH's superiority over existing state-of-the-art methods.
Conclusions:
- PLSH provides a flexible and effective approach to similarity joins for high-dimensional data.
- The technique offers improved control over approximation errors compared to standard LSH.
- PLSH is a scalable solution for handling massive datasets in similarity search applications.
Related Concept Videos
Causes of Similarity-Dissimilarity Effect
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Wilcoxon Signed-Ranks Test for Matched Pairs
Maximum Size of Aggregate
Weighted Mean
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...

