Machine-Scored Syntax: Comparison of the CLAN Automatic Scoring Program to Manual Scoring
Jenny A Roberts1, Evelyn P Altenberg1, Madison Hunter2
1Department of Speech-Language-Hearing Sciences, Hofstra University, Hempstead, NY.
Language, Speech, and Hearing Services in Schools
|March 19, 2020
Summary
Machine scoring of child language syntax using the Computerized Language ANalysis (CLAN) tools is less accurate than manual scoring. Researchers should report accuracy measures to understand CLAN
Area of Science:
- Linguistics
- Computational Linguistics
- Developmental Psychology
Background:
- The Index of Productive Syntax (IPSyn) is a measure of grammatical development.
- Automatic scoring tools like Computerized Language ANalysis (CLAN) aim to streamline IPSyn analysis.
- Accuracy of automated linguistic analysis tools is crucial for research and clinical applications.
Purpose of the Study:
- To evaluate the accuracy of machine-scored Index of Productive Syntax (IPSyn) using CLAN tools.
- To compare machine scoring results against manual scoring benchmarks.
- To introduce and utilize novel metrics for assessing automated syntactic analysis.
Main Methods:
- Comparison of machine-scored and manually-scored IPSyn data from 20 transcripts (10 children, 30 and 42 months).
- Analysis using traditional metrics (absolute point difference, point-to-point accuracy) and new metrics (Machine Item Accuracy - MIA, Cascade Failure Rate).
- Examination of differences in total scores, subscale scores (Noun Phrase, Verb Phrase, Question/Negation, Sentence Structures), and individual structures.
Main Results:
- Machine scoring showed a mean absolute point difference of 3.65 and 72.6% point-to-point agreement with manual scoring.
- Machine Item Accuracy (MIA) was 74.9%, with significantly more erroneous items than missed items.
- Noun Phrase and Verb Phrase subscales demonstrated higher accuracy than Question/Negation and Sentence Structures subscales.
Conclusions:
- The CLAN program's automatic scoring of IPSyn demonstrated notable inaccuracies compared to manual scoring.
- Recommendations for CLAN improvement include addressing second exemplar violations and implementing cascaded credit.
- Researchers and clinicians should routinely report detailed accuracy metrics, including MIA, to understand the limitations of machine-scored syntax.
Related Concept Videos
Multiple Comparison Tests
4.3K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.3K
Automatic Processing and Automatic Social Behavior
169
Automatic processing refers to the cognitive operations that occur without conscious intent or awareness, playing a fundamental role in shaping social cognition and behavior. These processes enable individuals to navigate complex social environments efficiently by relying on mental shortcuts and pre-existing knowledge structures known as schemas. One of the most influential mechanisms underlying automatic processing is priming, which subtly activates mental representations through exposure to...
169
Introduction to z Scores
981
A z score (or standardized value) is measured in units of the standard deviation. It indicates how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a zero z score. It is important to note that the mean of the z scores is zero, and the standard deviation is one.
z scores...
z scores...
981
Introduction to z Scores
10.8K
A z score (or standardized value) is measured in units of the standard deviation. It tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a zero z score. It is important to note that the mean of the z scores is zero, and the standard deviation is one.
z scores...
z scores...
10.8K
z Scores and Unusual Values
10.9K
The z score is one of the three measures of relative standing. It describes the location of a value in a dataset relative to the mean. z scores are obtained after the standardization of the values in a dataset. The z score for the mean is 0.
This score indicates how far a value is from the mean in terms of standard deviation. For example, if a data value has a z score of +1, the researcher can infer that the particular data value is one standard deviation above the mean. If another data...
This score indicates how far a value is from the mean in terms of standard deviation. For example, if a data value has a z score of +1, the researcher can infer that the particular data value is one standard deviation above the mean. If another data...
10.9K
z Scores and Area Under the Curve
17.8K
z scores are the standardized values obtained after converting a normal distribution into a standard normal distribution. A z score is measured in units of the standard deviation. The z score tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a z score of...
17.8K


