When analyzing treatment effects on test scores, researchers face many choices and
competing guidance for scoring tests and modeling results. This study examines the
impact of scoring choices through simulation and an empirical application. Results
show that estimates from multiple methods applied to the same data will vary because
two-step models using sum or factor scores provide attenuated standardized treatment
effects compared to latent variable models. This bias dominates any other differences
between models or features of the data generating process, such as the use of scoring
weights. An errors-in-variables (EIV) correction removes the bias from two-step models.
An empirical application to data from a randomized controlled trial demonstrates the
sensitivity of the results to model selection. This study shows that the psychometric
principles most consequential in causal inference are related to attenuation bias rather
than optimal scoring weights.
Gilbert, Joshua B.. (). How Measurement Affects Causal Inference: Attenuation Bias is (Usually) More Important Than Scoring Weights. (EdWorkingPaper: -766). Retrieved from
Annenberg Institute at Brown University: https://doi.org/10.26300/4hah-6s55
Head or assistant principal evaluation systems are widely deployed but rarely examined for validity, and AI scoring tools have entered K–12 workflows ahead of the evidence.
Blocked cluster randomized trials (CRTs) are widely used to evaluate educational interventions. In such trials, researchers face critical choices about how to define and estimate the average treatment effect (ATE).