Related Experiment Video
Updated: Apr 2, 2026

Evaluation of Commercial-Off-The-Shelf Wrist Wearables to Estimate Stress on Students
Published on: June 16, 2018
Artificial test-takers as transformed controls: measuring SAT difficulty drift and student performance.
Vikram K Suresh1, Saannidhya Rawat1
1University of Cincinnati, Cincinnati, OH, United States.
Artificial test-takers reveal a significant decline in SAT Math difficulty. After adjusting for this, student performance dropped by 34 points, with varied impacts across racial groups, highlighting issues with standardized test comparability.
Area of Science:
- Educational Measurement
- Artificial Intelligence in Education
- Psychometrics
Background:
- Standardized test scores are crucial for tracking student performance and informing educational policy.
- Interpreting score trends is challenging due to changes in exam content over time.
- Existing methods for equating test difficulty may be opaque or incomplete.
Purpose of the Study:
- To introduce an artificial test-taker framework using a fixed large language model (LLM) as a benchmark.
- To measure SAT Math difficulty drift over time.
- To construct difficulty-adjusted measures of student performance.
Main Methods:
- Developed a longitudinal SAT Math item bank (2007-2023).
- Generated bootstrapped SAT forms matching yearly blueprints.
- Administered items to GPT-4 with fixed parameters to create counterfactuals and difficulty benchmarks.
Main Results:
- Identified a statistically significant decline in SAT Math difficulty (-0.21σ relative to 2012).
- Adjusted student performance shows a 34-point decline (Average Difference in Scores) from 2012-2023.
- Observed non-uniform performance declines across different racial groups.
Conclusions:
- Artificial test-takers offer a scalable, protocol-invariant method for auditing longitudinal test comparability.
- Evolving SAT Math content can obscure underlying performance declines and mask subgroup trends.
- AI-driven transformed-control designs can benchmark educational outcomes and differentiate performance changes from instrument changes.
Related Concept Videos
Reliability and Validity
Comparing Experimental Results: Student's t-Test
Binet's Contribution to Measures of Intelligence
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...

