Automatic (near-) duplicate content document detection in a cancer registry.

Tapio Niemi1, Jean Pierre Ghobril1, Gautier Defossez1

  • 1Centre for Primary Care and Public Health (Unisanté), University of Lausanne, Lausanne, Switzerland.

Summary

This study developed an efficient method to identify duplicate and near-duplicate medical documents using Simhash and Smith-Waterman algorithms. The system successfully detects these documents in large datasets, improving data management and research accuracy.