Related Experiment Videos
Arabic text preprocessing for dynamic searchable symmetric encryption: An empirical evaluation and preprocessor
1Department of Information Technology, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia.
Abstract:
Dynamic Searchable Symmetric Encryption (DSSE) enables keyword search over encrypted data without revealing plaintext to the server. Arabic morphological richness - where a single root generates dozens of surface forms - creates substantial challenges for encrypted search: unprocessed vocabularies inflate encrypted-index size and transmission cost, and limit retrieval recall by failing to match morphological variants. This paper presents the first empirically grounded preprocessor selection framework for Arabic DSSE, derived from a systematic evaluation of four Arabic preprocessing strategies - the Khoja stemmer, ISRI stemmer, Lucene Arabic Analyzer, and Farasa segmenter - against a normalization-only baseline within the Incidence Matrix DSSE (IM-DSSE) scheme on the Khaleej Arabic news corpus. We evaluate vocabulary size, search latency, search quality, and retrieval breadth on an annotated 1,400-document corpus using 56 benchmark queries, and assess scalability on cloud infrastructure across corpus sizes up to 45,500 documents. Stemming reduces vocabulary by up to 83% relative to the normalization-only baseline, proportionally reducing encrypted-index size and transmission cost. We find that per-query search latency is driven primarily by result-set size rather than vocabulary size: the normalization-only baseline is fastest per query because it matches the fewest documents, while preprocessors that broaden retrieval incur higher latency. At scale, all five configurations - including the normalization-only baseline - operate successfully to 45,500 documents, and the result-set-size effect on latency persists. In terms of search quality, light stemming and morphological segmentation achieve the best trade-off between retrieval precision and recall, while root-based stemmers sacrifice precision without commensurate recall gains. Based on these findings, the framework provides actionable, evidence-based guidelines mapping deployment requirements - memory scalability, search latency, search quality, and retrieval breadth - to concrete preprocessor choices for Arabic DSSE system designers.
Related Concept Videos
Automatic Processing and Automatic Social Behavior
Pre-mRNA Processing: RNA Splicing