Related Experiment Video
Updated: Apr 26, 2026

09:09
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
1.1K
Research on ChatGPT generated text detection model based on phonetic feature extraction and semantic features.
Hao Chen1, Huancheng Chen1, Boyu Hu1
1School of Computer and Information Engineering, Tianjin Chengjian University, Tianjin, 300384, China.
Scientific Reports
|April 24, 2026
Summary
Detecting AI-generated text is hard, especially when rewritten. This study introduces a new framework combining deep learning with surface features, improving accuracy and robustness for identifying machine-written content.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Computational Linguistics
Background:
- The proliferation of large language models (LLMs) like ChatGPT presents significant challenges in distinguishing AI-generated text from human writing.
- Current detection methods, primarily based on semantic analysis, lack robustness against paraphrased or rewritten AI content.
Purpose of the Study:
- To develop an integrated framework for robust AI-generated text detection.
- To enhance detection capabilities by combining deep contextual semantics with auxiliary surface-level features.
Main Methods:
- Utilized a RoBERTa encoder for deep contextual semantic embeddings.
- Incorporated a convolutional neural network (CNN) for multi-scale representation aggregation.
- Integrated surface-level features: pronunciation cues, structural, lexical, and readability descriptors.
Main Results:
- The proposed RoBERTa-CNN framework significantly outperformed existing baselines on the HAGTC and ChatGPT abstract datasets.
- Achieved superior accuracy and F1 scores in detecting AI-generated text, including rewritten content.
- Ablation studies confirmed the performance gains from fusing multiple feature types.
Conclusions:
- Combining contextual semantic embeddings with auxiliary surface features offers a practical and effective approach for AI-generated text detection.
- The integrated framework demonstrates enhanced robustness, particularly for detecting paraphrased or rewritten AI content.
- Feature-level fusion is a key strategy for advancing the accuracy and reliability of AI text detection systems.

