CurveMark:確率曲線とダイナミック・セマンティック・ウォーターマークによるAI生成テキストの検出
Yuhan Zhang1, Xingxiang Jiang2,3, Hua Sun2,3
1School of Computer Science and Artificial Intelligence, Nanjing University of Finance and Economics, Nanjing 210023, China.
Entropy (Basel, Switzerland)
|August 28, 2025
まとめ
人工知能が作成したテキストを 検出するのは難しいのです CurveMarkは確率曲線とウォーターマークを使用して,テキストの品質を維持し,攻撃に抵抗しながら,大規模な言語モデル (LLM) のコンテンツを正確に識別します.
科学分野:
- 人工知能
- 自然言語処理
- 情報セキュリティ
背景:
- 大型言語モデル (LLM) は,高度なAIテキスト生成により,重要なコンテンツ認証課題を提示します.
- 既存の検出方法は 限られた情報収集と 悪いトレードオフと 敵対的な脆弱性で苦戦しています
研究 の 目的:
- LLMからAIで生成されたテキストを検出するための堅固なフレームワークを開発する.
- 既存の検出方法の限界を克服し,モデルに関する事前の知識の必要性を克服する.
主な方法:
- 確率曲線解析とダイナミック・セマンティック・ウォーターマークを組み合わせた2チャネル検出フレームワークであるCurveMarkを導入しました.
- テキストソースと観測可能な特徴の間の相互情報を最大化するために情報理論的原理を活用した.
- モデルアグノスティックな統計推論のためのベイジアンマルチ仮説検出フレームワークを組み込みました.
主要な成果:
- 複数のデータセットとLLMアーキテクチャで95.4%の検出精度を達成しました.
- 質の劣化が最小で,難解度が1. 3未満であることが証明された.
- 72~94%の情報を保持した.
結論:
- CurveMarkはAIで生成されたテキスト検出に 高い精度で堅牢なソリューションを提供します
- このフレームワークは,検出の精度とテキストの品質と攻撃に対する回復力を効果的にバランスします.
- このアプローチは,洗練されたLLMの時代にコンテンツの認証を進めます.
さらに関連する動画
関連する概念動画
Non-equilibrium in the Cell
4.8K
An important concept in studying metabolism and energy is that of chemical equilibrium. Most chemical reactions are reversible. They can proceed in both directions, releasing energy into their environment in one direction, and absorbing it from the environment in the other direction. The same is true for the chemical reactions involved in cell metabolism, such as the breaking down and building up of proteins into and from individual amino acids, respectively. Reactants within a closed system...
4.8K
Detection of Black Holes
2.3K
Although black holes were theoretically postulated in the 1920s, they remained outside the domain of observational astronomy until the 1970s.
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
2.3K
Difference from Background: Limit of Detection
7.1K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
7.1K
Detection of Gross Error: The Q Test
6.4K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.4K


