Head or assistant principal evaluation systems are widely deployed but rarely examined for validity, and AI scoring tools have entered K–12 workflows ahead of the evidence. We examine human supervisor ratings, Graded Response Model (GRM) rescaling, and large language model (LLM) scoring of 1,385… more →