Students given GPT-4o earned substantially higher grades on a short marketing assignment, but the experiment did not test whether they learned more. The distinction highlights a growing problem for education: AI can improve the product that teachers grade without proving that a student developed the underlying skill.

Researchers randomly divided 1,053 first-year Bocconi University students into four groups. They received either no intervention, a lesson in causal reasoning, access to GPT-4o, or both. Each student wrote up to 180 words of recommendations for the university merchandise shop.

GPT-4o raised scores by nearly one point on a five-point scale. The assisted answers offered about two additional ideas on average, followed more coherent logic and aligned more closely with three experts’ recommendations. The researchers attributed the remaining advantage to stronger content, not greater student knowledge.

The causal-reasoning lesson produced a different result. It did not improve conventional scores and slightly reduced them on average, but students more often explained why a proposal should work, identified conditions under which it could fail and generated ideas unlike their classmates’. Combining the lesson with GPT-4o did not add another grade boost. The study therefore measures better submitted work, not durable learning or independent performance after AI access is removed.