GPT-4o Boosts Business Student Grades in New Study
A study of 1,053 university students found that OpenAI's GPT-4o significantly boosted grades on a business assignment, raising questions about whether AI improves actual learning.

An experiment conducted at Bocconi University involving 1,053 freshmen revealed that using OpenAI's GPT-4o model significantly improved student performance on a business assignment. In the November 2025 study, 13 sections of an introductory management course were split into four groups: a control group, a group given a lesson on causal reasoning, a group with GPT-4o access, and a group with both. Students were tasked with writing marketing recommendations of up to 180 words for the university's merchandise shop. Those with access to GPT-4o scored nearly a full point higher on a 1-to-5 grading scale, generating an average of two more ideas per response and aligning closely with the recommendations of three subject-matter experts.
Interestingly, the short lesson on causal reasoning did not improve traditional scores, and actually resulted in slightly lower average grades. However, it did encourage students to produce more diverse ideas and better explain the conditions under which their proposals might fail. When combined with GPT-4o, the lesson did not further boost traditional grades, but the benefits of increased idea diversity remained. The researchers, some of whom are affiliated with OpenAI, noted that the grading rubric heavily favored conventional, well-structured answers, which penalized the highly original and falsifiable proposals generated by the lesson-only group.
For educational practitioners and AI developers, the study highlights a critical gap between graded performance and actual comprehension. Because there was no follow-up test without ChatGPT, it remains unclear if the students actually learned anything. This aligns with broader research, including a study of 500,000 U.S. college grades and a 30-month study of 26,000 students in China, which showed that while AI improves homework quality, exam scores can drop by 18 to 24 percent when the technology is removed. Educators must therefore redesign grading criteria to reward original reasoning rather than the polished, conventional outputs that AI can easily replicate.
This is our own summary of reporting by The Decoder



