Related Experiment Videos
The adversarial game between detection and evasion: A survey of anti-detection techniques for machine-generated texts
1Department of Information and Communication Sciences, Faculty of Science & Technology, Sophia University, Chiyoda-ku, 102-8554, Tokyo, Japan.
Abstract:
With the explosive growth of large language models (LLMs), research on machine-generated text detection (MGTD) has also proliferated. Alongside these developments, a wide range of attack algorithms targeting MGTD systems have emerged. While previous studies have surveyed detection techniques, few have examined the dynamic interplay between attack and defense. Following PRISMA 2020, this paper systematically synthesizes 27 studies of attacks against MGTD and the available evidence on corresponding defenses. We categorize existing research into four major types of evasion strategies: watermark attacks, paraphrasing attacks, prompt-based attacks, and adversarial-text attacks, and summarize the available defense evidence. Furthermore, to better understand the practical implications of these methods, we compile the reported performance results of attack and defense techniques across different detectors. Finally, we highlight the current challenges in this area and outline potential future research directions. A companion repository containing the categorized literature, paper links, and available code, data, and project repositories is provided at https://github.com/AIGC1999/A-Survey-of-Anti-Detection-Techniques-for-Machine-Generated-Texts.