Large language models (LLMs) can make tutoring more scalable, but only if students use them to reason through mistakes rather than avoid effort. We study this question in a randomized field experiment with more than 6,000 middle-school students in Hamilton County Schools using NUMI, a research-… more →