Are Large-Scale Test Scores Comparable for At-Home Versus Test Center Testing?
Katherine E Castellano1, Matthew S Johnson1, Rene Lawless1
1ETS, NJ, USA.
Applied Psychological Measurement
|August 21, 2024
Summary
Remote proctored assessments, or at-home testing, showed no significant score differences compared to traditional test center exams. This finding is crucial for understanding the validity of remote testing in admissions.
Area of Science:
- Educational Measurement
- Psychometrics
- Remote Assessment Technologies
Background:
- The COVID-19 pandemic accelerated the adoption of remote-proctored (at-home) assessments.
- At-home testing environments lack standardization in setting, device, and supervision compared to test centers.
- Understanding score comparability between remote and traditional testing is vital for assessment validity.
Purpose of the Study:
- To investigate score comparability between remote-proctored and test center administrations of a large-scale admissions test.
- To determine if environmental and procedural differences impact test scores in remote assessment settings.
Main Methods:
- Employed both a randomized controlled trial and an observational study design.
- Compared test scores from at-home assessments with those from traditional test centers.
- Utilized statistical adjustments for baseline characteristic differences in sample composition.
Main Results:
- No statistically significant differences were found between at-home and test center scores.
- Adjustments for sample composition confirmed score comparability across testing modalities.
- Findings suggest remote proctoring does not inherently alter admissions test outcomes.
Conclusions:
- Remote-proctored assessments demonstrate score equivalence to traditional test center administrations.
- The study supports the validity and reliability of at-home testing for large-scale admissions.
- These findings have implications for the future of standardized testing and educational accessibility.
Related Concept Videos
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Comparing Experimental Results: Student's t-Test
1.5K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
1.5K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Measures of Intelligence
7.1K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
7.1K
One-Way ANOVA: Equal Sample Sizes
3.2K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.2K
Ratio Level of Measurement
17.5K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated....
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated....
17.5K


