ChatGPT significantly improved student work, while a teaching intervention encouraged more diverse and unusual solutions. Whether AI also improves learning remains an open question.

What makes a good student paper, and which parts of that can AI improve? A randomized experiment with 1,053 freshmen at Bocconi University found that GPT-4o helped students earn significantly better grades on a business assignment. A short lesson on causal reasoning didn't raise traditional scores, but it pushed students toward more diverse solutions and deeper thinking about causes, assumptions, and how their proposals would actually work.

GPT-4o delivered a major boost to graded performance

In November 2025, 13 sections of an introductory management course were randomly split into four groups: control, causal reasoning lesson, GPT-4o access, or both. Students had to write marketing recommendations for the university's merchandise shop in up to 180 words. The lesson covered coherent causal logic, falsifiability, and how a proposed action might lead to a desired outcome.

Students with GPT-4o scored nearly a full point higher on a 1-to-5 scale. Their answers contained about two more ideas on average, were more logically coherent, and more closely matched the recommendations of three subject-matter experts. Even after the researchers controlled for argumentation quality, number of ideas, idea diversity, and text properties, a measurable GPT-4o advantage remained. The authors attribute this to higher content quality, not greater student knowledge.