Related Experiment Video
Updated: Feb 28, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Document retrieval on repetitive string collections
Travis Gagie1, Aleksi Hartikainen2, Kalle Karhu3
1CeBiB - Center of Biotechnology and Bioengineering, School of Computer Science and Telecommunications, Diego Portales University, Santiago, Chile.
This study introduces novel indexing techniques for repetitive string collections, significantly reducing space usage and enabling efficient document retrieval operations like listing, top-k retrieval, and counting.
Area of Science:
- Computer Science
- Information Retrieval
- Data Compression
Background:
- Fast-growing string collections are often repetitive, with many similar documents.
- Existing document retrieval techniques are less developed for generic and repetitive string collections.
- Exploiting repetitiveness can drastically reduce space usage in large collections.
Purpose of the Study:
- To develop efficient indexing methods for repetitive string collections.
- To enable fast document retrieval operations on these compressed collections.
- To address the challenges posed by the increasing size and repetitiveness of data.
Main Methods:
- Development of two novel techniques: interleaved Longest Common Prefixes (LCPs) and precomputed document lists.
- Creation of highly compressed indexes for document listing, top-k retrieval, and document counting.
- Adaptation of classical data structures for compressibility on repetitive data.
Main Results:
- Achieved highly compressed indexes for efficient document retrieval.
- Demonstrated the compressibility of classical data structures on repetitive data.
- Successfully combined developed tools to solve ranked multi-term queries.
Conclusions:
- Novel indexing techniques significantly improve space efficiency and retrieval performance for repetitive string collections.
- The developed methods provide effective solutions for document listing, top-k retrieval, and counting.
- The approach offers a robust framework for handling large, repetitive datasets in information retrieval.
More Related Videos
07:26Executing Complexity-Increasing Queries in Relational MySQL and NoSQL MongoDB and EXist Size-Growing ISO/EN 13606 Standardized EHR Databases
Published on: March 19, 2018
07:49Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
Related Concept Videos
Data Collection I
Data Collection III
The principles to begin the physical assessment include conducting a comprehensive or problem-related history in a quiet, well-lit room, emphasizing privacy and comfort for the...
Data Collection II
Data Collection by Survey
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Methods of Documentation I: Source-Oriented Records
In an SOR, each discipline involved in patient care maintains a separate medical record section. This record-keeping method enables easy tracking of patient progress and ensures healthcare staff have access to up-to-date information.
Key Attributes include the following: