Related Experiment Video
Updated: Jun 14, 2026

Advancing Dyslexia Assessment in Children Through Computerized Testing
Published on: August 16, 2024
Morphosyntactic annotation of CHILDES transcripts
Kenji Sagae1, Eric Davis, Alon Lavie
1Institute for Creative Technologies, University of Southern California, CA 90292, USA. sagae@usc.edu
Abstract:
Corpora of child language are essential for research in child language acquisition and psycholinguistics. Linguistic annotation of the corpora provides researchers with better means for exploring the development of grammatical constructions and their usage. We describe a project whose goal is to annotate the English section of the CHILDES database with grammatical relations in the form of labeled dependency structures. We have produced a corpus of over 18,800 utterances (approximately 65,000 words) with manually curated gold-standard grammatical relation annotations. Using this corpus, we have developed a highly accurate data-driven parser for the English CHILDES data, which we used to automatically annotate the remainder of the English section of CHILDES. We have also extended the parser to Spanish, and are currently working on supporting more languages. The parser and the manually and automatically annotated data are freely available for research purposes.
Related Concept Videos
Polytene Chromosomes
Polytene Chromosomes
Additional Subnuclear Structures
The nucleus contains many membrane-less subnuclear organelles or nuclear bodies, such as nucleoli, Cajal bodies, speckles, paraspeckles, etc. These nuclear...
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Chromatin Structure and RNA Splicing
The chromatin structure, especially...
