ChatGPT improves answer quality, critical-thinking training boosts originality, study finds

ChatGPT improves answer quality, critical-thinking training boosts originality, study finds

Researchers at Bocconi University, in collaboration with OpenAI Economic Research, ran a randomized experiment with more than 1,000 first-year undergraduates. Students worked on a real-world business case: developing marketing recommendations for the university's own merchandise store. They were randomly assigned by class period into one of four groups, receiving access to ChatGPT (GPT-4o), training in causal reasoning, both, or neither.

Causal reasoning, the form of critical thinking the training targeted, means linking cause and effect and explaining why a proposed solution might or might not work. The training itself was unrelated to AI: students learned the concepts through a game, examples, questions and feedback.

Submissions were evaluated two ways. Trained human graders scored each one on a five-point rubric that measured how well the recommendations addressed two standard marketing goals, increasing awareness of the store and increasing its use. Separately, researchers ran automated text analysis on the same submissions to measure the number and variety of ideas, evidence of causal reasoning, and how closely each submission resembled work from three experts.

Students with ChatGPT access scored almost a full point higher on the five-point rubric than those without it. Their submissions contained more ideas, followed clearer logic, and read more like the expert-written recommendations. The source stresses these students still had to decide what to ask ChatGPT, evaluate its answers, and choose what to include, rather than simply submitting AI output unchanged.

Students who completed the causal-reasoning training explained more clearly why their proposed ideas might succeed or fail, but this did not translate into higher rubric scores, since the rubric measured only the two marketing goals and not reasoning quality. Text analysis, however, showed that as a group these students produced a wider range of ideas that were more distinct from each other than their peers' ideas were, an effect the rubric could not capture.

Students who received both ChatGPT access and the training matched the idea variety of the training-only group, matched the rubric scores and idea counts of the ChatGPT-only group, and additionally showed stronger logical coherence and more evidence of seeking explanations and questioning assumptions. Across the four groups, this combined group registered gains on the widest range of measures.

Because the design was randomized rather than observational, the researchers could separate the effect of ChatGPT access from the effect of the causal-reasoning training and from their combination. The piece frames ChatGPT access and critical-thinking training as complementary rather than as a choice between letting students use AI or teaching them to think, and argues that because AI can help students produce polished, expert-like answers, evaluating only the final answer reveals less about what a student actually understands, so assignments and rubrics may need to change to reward originality and reasoning as well as polish.

Key facts

  • In a randomized trial, researchers at Bocconi University, working with OpenAI Economic Research, assigned more than 1,000 first-year students by class period into four groups: ChatGPT (GPT-4o) access, causal-reasoning training, both, or neither, for a real assignment writing marketing recommendations for the university's merchandise store.
  • Students with ChatGPT access scored almost a full point higher on a five-point human-graded rubric, and their work had more ideas, clearer logic and closer resemblance to submissions from three experts; researchers note these students still chose what to ask and what to keep, rather than submitting AI output directly.
  • Causal-reasoning training did not raise rubric scores, since the rubric measured only two marketing goals, but automated text analysis found the trained group produced a wider range of more distinct ideas than their peers, plus clearer explanations of why an idea might work or fail.
  • Students who got both ChatGPT access and training matched the idea variety of the training-only group and the rubric scores of the ChatGPT-only group, plus showed stronger logical coherence, registering gains across the widest range of measures of the four groups.
  • The article frames AI access and critical-thinking training as complementary rather than competing, and argues that because AI can make novice work look polished, assignments and rubrics may need to change to reward originality and reasoning, not just a well-structured final answer.

Why it matters

Generative AI tools are already good enough to make a novice's work look polished and professional, which risks leaving a traditional grading rubric blind to whether a student actually worked through a problem or just prompted well. This is a randomized experiment, not an observational survey, so it can isolate what ChatGPT access contributes, what critical-thinking training contributes, and what happens when a student gets both. The answer reframes the usual choice in education debates between letting students use AI or teaching them to think for themselves: the study finds the two produce different, complementary gains rather than substituting for each other.

Who it affects

The direct subjects are the more than 1,000 first-year undergraduates at Bocconi University who took the assignment. The findings speak most directly to instructors and institutions designing courses and grading rubrics where students already have access to tools like ChatGPT, since a rubric built only around whether an answer hits the brief can miss whether a student generated an original idea. It also matters to OpenAI, whose Economic Research group collaborated on the study as part of a wider effort to measure ChatGPT's effects in education.

How to use it

There is no product or setting to adopt here; the piece describes a completed research design rather than a tool or a feature. Its practical takeaway for educators is methodological: pairing a standard grading rubric with a separate measure of idea variety, as the researchers did through automated text analysis, can surface effects such as the critical-thinking group's wider range of distinct ideas that a rubric focused only on the final answer will not detect.

How solid is it

The core strength is the randomized design: students were assigned to the four groups by class period rather than choosing their own group, which is what lets the researchers attribute the differences to ChatGPT access and training rather than to which students were already stronger. Results were also checked two separate ways, trained human graders scoring against a five-point rubric and an independent automated text analysis of idea count, variety and similarity to expert answers, and the two approaches point the same direction. Against that: the source names no individual researcher, only the two institutions; gives no date for when the experiment ran or was published; does not say how many students were in each of the four groups; and reports none of the results with a confidence interval, p-value or other significance test, so a reader cannot check from this account alone how far the differences exceed normal variation.

Risks and caveats

This was one assignment: a marketing-recommendation case for a university store, graded against a rubric built around two specific marketing goals. How far ChatGPT's quality gains and the training's originality gains generalize to other subjects, assignment types or grading schemes is untested here. The piece is also published on OpenAI's own site, which is a reason to read its framing of AI and critical-thinking training as complementary as the company's preferred takeaway, even where the underlying university-run experiment looks methodologically sound. And the caution the study itself raises, that AI-polished answers can score well on a conventional rubric while revealing little about original thinking, applies as much to how this study's own rubric-based scores should be read as to any other classroom.