Related Experiment Video
Updated: May 1, 2026

05:21
Computerized Adaptive Testing System of Functional Assessment of Stroke
Published on: January 7, 2019
5.8K
Reviewer Experience Detecting and Judging Human Versus Artificial Intelligence Content: The Stroke Journal Essay
Gisele S Silva1, Rohan Khera2,3,4, Lee H Schwamm3,5,6
1Hospital Israelita Albert Einstein and Departamento de Neurologia e Neurocirurgia, Universidade Federal de São Paulo, Brazil (G.S.S.).
Stroke
|September 3, 2024
Summary
Human and artificial intelligence (AI) large language models (LLMs) generated scientific essays. Reviewers struggled to distinguish AI from human authors, rating AI essays higher for composition but showing bias against AI-generated content.
Area of Science:
- Neurology
- Artificial Intelligence
- Scientific Publishing
Background:
- Large language models (LLMs) generate human-like text and images.
- The ability of LLMs to produce persuasive scientific essays for peer review is not well understood.
- Assessing AI's impact on scientific authorship and quality perception is crucial.
Purpose of the Study:
- To evaluate human and AI-generated scientific essays in a blinded competition.
- To measure reviewer perceptions of essay quality, persuasiveness, and authorship.
- To determine if AI authorship can be accurately identified by expert reviewers.
Main Methods:
- A 2024 essay contest featured human authors and 4 distinct LLMs on controversial stroke care topics.
- 38 Stroke Editorial Board members (vascular neurologists) blinded to author identity rated 34 essays.
- Reviewers assessed essays for quality, persuasiveness, best in topic, and author type (human vs. AI).
Main Results:
- Human and AI essays received similar overall ratings, but AI essays scored higher for composition quality.
- Reviewers accurately identified author type only 50% of the time; prior LLM experience improved accuracy.
- Persuasiveness was linked to reviewers assigning AI as the author type (aOR, 1.53; P=0.01).
Conclusions:
- Expert reviewers had difficulty distinguishing human from AI-generated scientific essays.
- A bias against AI-generated essays was observed, particularly for 'best in topic' ratings.
- Journals should educate reviewers on AI in scientific writing and establish clear AI authorship policies.
Related Concept Videos
Stereotype Content Model
13.1K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.1K
Non-equilibrium in the Cell
4.0K
An important concept in studying metabolism and energy is that of chemical equilibrium. Most chemical reactions are reversible. They can proceed in both directions, releasing energy into their environment in one direction, and absorbing it from the environment in the other direction. The same is true for the chemical reactions involved in cell metabolism, such as the breaking down and building up of proteins into and from individual amino acids, respectively. Reactants within a closed system...
4.0K
Intelligence
15.4K
The term "intelligence" is complex because it refers to both behavior and individuals, and its interpretation varies across cultures. European Americans tend to link intelligence with reasoning and cognitive skills, while in Kenya, it is tied to responsible participation in family and social life. In Uganda, intelligence is seen as the ability to know the right actions and carry them out effectively, while the Iatmul people of Papua New Guinea associate it with the capacity to remember...
15.4K

