The skills that earn top grades are the ones AI can fake best
Key Points
- An experiment at Bocconi University with over 1,000 freshmen shows that GPT-4o significantly boosted grades on a business assignment.
- A short lesson on causal reasoning didn't improve traditional scores but pushed students toward more diverse and unusual solutions.
- Whether AI use also improved actual learning remains unanswered. There was no follow-up test without ChatGPT.
ChatGPT significantly improved student work, while a teaching intervention encouraged more diverse and unusual solutions. Whether AI also improves learning remains an open question.
What makes a good student paper, and which parts of that can AI improve? A randomized experiment with 1,053 freshmen at Bocconi University found that GPT-4o helped students earn significantly better grades on a business assignment. A short lesson on causal reasoning didn't raise traditional scores, but it pushed students toward more diverse solutions and deeper thinking about causes, assumptions, and how their proposals would actually work.
GPT-4o delivered a major boost to graded performance
In November 2025, 13 sections of an introductory management course were randomly split into four groups: control, causal reasoning lesson, GPT-4o access, or both. Students had to write marketing recommendations for the university's merchandise shop in up to 180 words. The lesson covered coherent causal logic, falsifiability, and how a proposed action might lead to a desired outcome.
Students with GPT-4o scored nearly a full point higher on a 1-to-5 scale. Their answers contained about two more ideas on average, were more logically coherent, and more closely matched the recommendations of three subject-matter experts. Even after the researchers controlled for argumentation quality, number of ideas, idea diversity, and text properties, a measurable GPT-4o advantage remained. The authors attribute this to higher content quality, not greater student knowledge.
The causal reasoning lesson didn't raise traditional scores. Work from those students actually scored slightly worse on average, but they more often explained why a proposed action should work and under what conditions it might fail. They also generated more diverse ideas that diverged from what their peers wrote.
Combining the lesson with GPT-4o didn't add any further boost to traditional scores. On causal reasoning markers, the two approaches complemented each other, and the greater idea diversity from the lesson group held up.

The grading rubric favored conventional answers
More ideas, more coherent arguments, and greater idea diversity within a single answer all correlated with higher scores. But stronger falsifiability, more detailed explanations of how proposed actions would work, and greater divergence from other students' ideas correlated with lower scores.
That doesn't mean originality was penalized across the board. For this assignment, the traditional score mainly rewarded well-structured answers that stayed within the expected solution space. The authors conclude that diversity and originality need to be explicitly built into grading criteria if they're supposed to count.
But current grading systems measure polish, structure, and completeness, but not learning and understanding. That makes it easy to use AI as a cheating tool.
What the study doesn't show
There was no follow-up test where students had to demonstrate what they actually understood or retained without ChatGPT. The authors explicitly acknowledge that it remains unclear whether the GPT advantage came from knowledge students actually gained or simply reflected better output with AI assistance. The study shows improved graded performance on this assignment, not improved learning.
The experiment only looked at freshmen at a single university working on a narrow marketing task. Randomization happened across 13 class sections rather than individually among all 1,053 students.
The main performance score came from human graders who didn't know which group each text belonged to. Several additional measures like causal reasoning and idea diversity were evaluated using models from OpenAI and Anthropic.
OpenAI provides the technology being studied and was involved in the research. Several authors work at OpenAI or were employed there during the study.
Other studies ask what sticks after the AI goes away
Research so far clearly suggests that what matters isn't whether students use AI but whether it supports their own thinking or replaces it. When it replaces it, students suffer.
A study covering more than 500,000 US college grades found that top grades after ChatGPT's launch increased most in writing- and programming-heavy courses with a large homework component. Controlled experiments found that after AI was taken away, participants performed worse than a control group. The drop was steepest among those who had mainly used GPT to get direct answers.
A 30-month study of more than 26,000 students in China showed a similar pattern over a longer period. Homework got better and faster with AI, but exam scores dropped. On later entrance exams, results were 18 to 24 percent lower over the long term. Students who spent roughly the same amount of time on homework as non-users despite having AI access didn't show comparable declines.
Despite this growing body of evidence, OpenAI frames the results mainly as a challenge for grading systems: if AI can produce polished, expert-like work, the final product alone says less about what students actually understand. The authors argue for "rewarding students for producing work that reflects originality, reasoning, and consideration of multiple approaches, not just the most conventional or polished answers." That may be true, but changing the grading system likely means changing the entire education system along with it.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.