Abstract: Experimental studies of generative AI in education mostly report average effects, yet field evidence shows that AI can narrow attainment gaps or widen them. We argue that the direction depends on the judgement burden, the expertise a learner must supply to screen AI output before learning from it, and that expert verification before release moves this burden from students to an accountable tutor. We test the argument in a two-cohort difference-in-differences design in which one half of a compulsory firstyear university economics course received AI-generated podcasts, FAQs and quiz-based study guides, produced with a source-grounded model and checked by a named graduate teaching assistant (170 students; 340 examination marks). Access was associated with a 2.34-mark advantage on a 50-mark component. The share of marks below the upper-second classification boundary fell by 24.7 percentage points relative to the counterfactual, effects were significant at every threshold from 23 to 31 marks and at none above, and roughly three-quarters of the average originated in the bottom quintile. The threshold estimate is robust to removing the lowest-scoring students from the pre-intervention cohort; the average effect is not. Interviews and feedback from 36 students indicate that the verification label gave students a reason to engage with AI-generated material without ending their scrutiny of it. Evaluations of AI learning resources that report only mean effects cannot detect whether the students the resources are meant to help are the ones who gain.
Read the original article:
