Related Experiment Video
Updated: Aug 6, 2026

Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
Published on: September 5, 2019
Much Ado about Prompting: LLM classification of text messages from experiments
Can Çelebi1, Stefan P Penczynski2
1Vienna Center for Experimental Economics (VCEE), University of Vienna, Oskar-Morgenstern-Platz 1, Vienna, Austria.
Researchers found that detailed, human-readable codebooks significantly improve large language model (LLM) text classification accuracy. Focusing on prompt content, not just engineering, is key for better LLM performance.
Area of Science:
- Natural Language Processing
- Artificial Intelligence
- Behavioral Economics
Background:
- Large language models (LLMs) are increasingly used for text classification.
- Existing codebooks for human annotators are a potential resource for LLM prompting.
- The effectiveness of using these codebooks as prompts requires investigation.
Purpose of the Study:
- To evaluate how well human-annotator codebooks can be repurposed as prompts for LLMs.
- To determine the impact of prompt information level, formatting, framing, and wording on LLM classification accuracy.
- To compare performance across different LLM architectures (proprietary vs. open-weight).
Main Methods:
- Utilized three experimental economics codebooks for promise classification and strategic thinking tasks.
- Systematically varied prompt information levels (low to high).
- Manipulated prompt formatting, framing, and wording at extreme information levels.
- Tested across two proprietary and two open-weight LLMs.
Main Results:
- Codebooks as prompts achieved 82-88% agreement with human annotators.
- Model choice was critical for recognition-heavy tasks; prompt detail was key for learning-heavy tasks.
- Prompt content, not model reasoning, drove performance improvements.
- Larger models better utilized detailed prompts and were robust to variations; smaller models were sensitive to prompt presentation.
Conclusions:
- Prioritizing detailed, human-readable prompt content is more effective than complex prompt engineering techniques.
- Codebooks prepared for human annotators serve as effective prompts for LLMs.
- Future LLM development should focus on optimizing prompt content quality and detail.
Related Concept Videos
Statistical Significance
Routes of Persuasion
Group Design
Non-Verbal Cues
What is an Experiment?
Automatic Processing and Automatic Social Behavior