Comparison of Structural Parsers and Neural Language Models as Surprisal Estimators
Byung-Doh Oh1, Christian Clark1, William Schuler1
1Department of Linguistics, The Ohio State University, Columbus, OH, United States.
This study introduces a left-corner parser for sentence processing, finding it predicts human reading times better than large neural models. This suggests linguistic structure is key for understanding processing costs.
Area of Science:
- Computational linguistics
- Cognitive science
- Psycholinguistics
Background:
- Expectation-based theories link sentence processing difficulty to contextual predictability.
- Surprisal quantifies predictability but lacks a clear cognitive basis.
- Approximating the human probability model in language comprehension remains an open challenge.
Purpose of the Study:
- To develop and evaluate a computational model of sentence processing.
- To compare the predictive power of different models (left-corner parser, structural parsers, neural language models) against human processing data.
- To investigate how linguistic generalizations influence the prediction of processing costs.
Main Methods:
- Developed an incremental left-corner parser incorporating syntactic and semantic information.
- Evaluated surprisal estimates from various parsers and deep neural language models.
- Compared model predictions against self-paced reading, eye-tracking, and fMRI data.
Main Results:
- The proposed left-corner parser showed comparable or superior fits to reading and eye-tracking data compared to large neural language models.
- A negative correlation was observed between parameter count and model fit for Transformer-based language models.
- The left-corner model's strong linguistic generalizations better predicted human processing costs.
Conclusions:
- Computational models incorporating linguistic abstractions may better approximate human sentence processing than large-scale neural models.
- Humanlike processing costs are potentially better predicted by models emphasizing linguistic structure over sheer data volume.
- Large neural models might rely more on lexical patterns than structural linguistic generalizations.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Language and Cognition
