A model-independent redundancy measure for human versus ChatGPT authorship discrimination using a Bayesian

Silvia Bozza1,2, Claude-Alain Roten3, Antoine Jover3,4

  • 1Ca' Foscari University of Venice, Department of Economics, Venice, 30121, Italy. silvia.bozza@unive.it.

Scientific Reports
|November 6, 2023
PubMed
Summary

Detecting AI-generated text is crucial. This study introduces a novel, AI model-independent method using a Bayes factor to analyze text redundancy, effectively distinguishing human-written content from artificial intelligence (AI) outputs like ChatGPT.

Related Concept Videos

Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Probability Laws01:49

Probability Laws

Overview
40.9K
The Representativeness Heuristic02:13

The Representativeness Heuristic

The representative heuristic describes a biased way of thinking, in which you unintentionally stereotype someone or something. For example, you may assume that your professors spend their free time reading books and engaging in intellectual conversation, because the idea of them spending their time playing volleyball or visiting an amusement park does not fit in with your stereotypes of professors.
15.8K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data01:16

Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data

Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
135
Stereotype Content Model02:16

Stereotype Content Model

The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.7K
Multiple Allele Traits01:49

Multiple Allele Traits

The Concept of Multiple Allelism
34.3K