Survey of Attendance Practices
Category: Student Well-Being
Head or assistant principal evaluation systems are widely deployed but rarely examined for validity, and AI scoring tools have entered K–12 workflows ahead of the evidence. We examine human supervisor ratings, Graded Response Model (GRM) rescaling, and large language model (LLM) scoring of 1,385 yearly leader-authored, standards-based, retrospective accounts of school improvement efforts and validate each against five school outcomes. Some structural school characteristics predict human-based and LLM scores, but often in different directions. As predictors, human and GRM scores often converge but predict few outcomes, though GRM may recover aspects of instructional leadership captured by student achievement scores. LLM scores diverge and uniquely predict staffing stability—teacher retention—a signal that survives adjustment for surface text features. Higher human-based scores predict worse staffing stability in the highest-racial-minority and highest-FRPL enrollment schools, while LLM scores show no such inversion. Scoring method choice may be equity-consequential and capture different aspects of school-improvement leadership.