Related Experiment Video
Updated: Sep 19, 2025

05:54
Eye-tracking to Distinguish Comprehension-based and Oculomotor-based Regressive Eye Movements During Reading
Published on: October 18, 2018
6.3K
EMTeC: A corpus of eye movements on machine-generated texts
Lena S Bolliger1, Patrick Haller2, Isabelle C R Cretton2
1Department of Computational Linguistics, University of Zurich, Andreasstrasse 15, Zurich, 8050, Switzerland. bolliger@cl.uzh.ch.
Behavior Research Methods
|June 3, 2025
Summary
This study introduces the Eye movements on Machine-generated Texts Corpus (EMTeC), detailing eye-tracking data from human reading of AI-generated content. The corpus enables research into reading behaviors and AI model interpretability.
Area of Science:
- Cognitive Science
- Computational Linguistics
- Human-Computer Interaction
Background:
- Understanding human reading behavior is crucial for AI development.
- Existing corpora lack naturalistic eye-tracking data on machine-generated text.
Purpose of the Study:
- Introduce the Eye movements on Machine-generated Texts Corpus (EMTeC).
- Provide a resource for studying human reading of AI-generated content.
- Facilitate research on AI model interpretability and reading behavior.
Main Methods:
- Collected eye-tracking data from 107 participants reading texts generated by three large language models.
- Utilized five decoding strategies and six text-type categories.
- Included pre-processed eye movement data, model internals, and linguistic annotations.
Main Results:
- The EMTeC corpus offers raw and processed eye-tracking data, including fixation sequences and reading measures.
- It contains language model internals (transition/attention scores, hidden states) and linguistic annotations.
- The corpus includes corrected fixation sequences accounting for vertical calibration drift.
Conclusions:
- EMTeC supports diverse research, including reading behavior analysis on AI text and decoding strategies.
- It aids in developing new data processing algorithms and enhancing AI model interpretability.
- The corpus is valuable for assessing predictive models of human reading times.

