Related Experiment Video
Updated: May 31, 2025

06:53
Detection and Monitoring of Tumor Associated Circulating DNA in Patient Biofluids
Published on: June 8, 2019
8.6K
Automatic (near-) duplicate content document detection in a cancer registry.
Tapio Niemi1, Jean Pierre Ghobril1, Gautier Defossez1
1Centre for Primary Care and Public Health (Unisanté), University of Lausanne, Lausanne, Switzerland.
International Journal of Medical Informatics
|January 22, 2025
Summary
This study developed an efficient method to identify duplicate and near-duplicate medical documents using Simhash and Smith-Waterman algorithms. The system successfully detects these documents in large datasets, improving data management and research accuracy.
Area of Science:
- Medical Informatics
- Data Management
- Computational Biology
Background:
- Duplicate and near-duplicate medical documents pose challenges in data management, clinical use, and research.
- Large volumes of multisourced documents in cancer registries complicate processing.
- Pairwise comparison is infeasible for identifying near-duplicates in extensive datasets.
Purpose of the Study:
- To develop and validate an efficient system for identifying full-duplicates and near-duplicates in multisourced medical documents.
- To address the challenges of document processing in population-based cancer registries.
- To reduce the false positive rate in near-duplicate detection without compromising sensitivity.
Main Methods:
- Implemented a system combining Simhash (Locality Sensitive Hashing) and Smith-Waterman text alignment.
- Validated the method on 3042 manually verified document pairs.
- Achieved high performance metrics: AUC 0.96, sensitivity 0.92, specificity 1.00, PPV 1.00, NPV 0.91.
Main Results:
- Applied to 224,398 medical documents, identifying 5.5% duplicates and 0.17-0.24% near-duplicates.
- Most near-duplicates shared patient and transmitter information.
- Differences were primarily administrative (83%) rather than medical (2%).
Conclusions:
- The developed method efficiently identifies full-duplicates and most near-duplicates in large medical document sets.
- The system demonstrates high accuracy and can be applied across various domains.
- Potential improvements and broader applications of the method are discussed.
More Related Videos
Related Concept Videos
Combination Therapies and Personalized Medicine
4.8K
Combining two or more treatment methods increases the life span of cancer patients while reducing damage to vital organs or tissue from the overuse of a single treatment. Combination therapy also targets different cancer-inducing pathways, thus reducing the chances of developing resistance to treatment.
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
4.8K
Comparing Copy Number Variations and SNPs
17.1K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.1K

