Auditing unauthorized training data from AI generated content using information isotopes
Tao Qi1,2, Jinhua Yin2, Dongqi Cai3
1State Key Laboratory of Networking and Switching Technology, School of Computer Science, Beijing University of Posts and Telecommunications, Beijing, China.
Abstract:
The rapid growth of AI systems has been fueled by large-scale human data, intensifying concerns over the unauthorized use of intellectual property and privacy-sensitive content during model training. Auditing such misuse is particularly challenging since mainstream AI services operate as black boxes, exposing only generated outputs while concealing their training and inference processes. In this work, inspired by chemical isotope tracing, we introduce the concept of information isotopes to trace training data within opaque AI systems. We propose an information-isotope tracing framework that selectively marks target data elements and detects their propagation in model outputs, providing concrete evidence of data utilization under black-box access. Experiments on thirteen AI models across six datasets demonstrate that our method distinguishes training from non-training data with up to 99% accuracy and strong statistical significance (p < 0.01) using approximately 4,000 words of evidence. An open-source tool is released to support practical data rights protection.
More Related Videos
07:57Sampling and Pretreatment of Tooth Enamel Carbonate for Stable Carbon and Oxygen Isotope Analysis
Published on: August 15, 2018
12:47Workflow Based on the Combination of Isotopic Tracer Experiments to Investigate Microbial Metabolism of Multiple Nutrient Sources
Published on: January 22, 2018
Related Concept Videos
Mass Spectrometry: Isotope Effect
Isotopes and Radioisotopes
An isotope containing...
Isotopes
An element's atomic mass, or weight,...
Inductively Coupled Plasma-Mass Spectrometry (ICP-MS): Interferences
Non-equilibrium in the Cell
Atomic Emission Spectroscopy: Lab
