Data Colada finds evidence of tampering in Ariely and Wertenbroch's 2002 procrastination study

Data Colada finds evidence of tampering in Ariely and Wertenbroch's 2002 procrastination study

The blog Data Colada has published a post presenting evidence that the data behind Ariely and Wertenbroch's 2002 paper "Procrastination, Deadlines, and Performance: Self-Control by Precommitment" were tampered with. The paper, published in Psychological Science and cited more than 2,100 times on Google Scholar, reported that people who were assigned separate, evenly spaced deadlines for three tasks performed far better than people who set their own deadlines or faced a single deadline for all tasks on the last day. The trigger for the post was a new replication study by Kyle Hyndman and Alberto Bisin, forthcoming in Psychological Science, which failed to reproduce Study 2 of the original paper. Hyndman had received the original data files by email in 2006, files whose properties would later indicate they had last been saved by Dan Ariely, and passed them to the Data Colada authors in 2023, a week after Francesca Gino sued the Data Colada authors for $25 million. According to a footnote in the replication paper, the authors asked Ariely in October 2024 for permission to include an analysis of the Study 2 file in their paper; Ariely refused, arguing the file they had might not be the real data, and provided no further data. That prompted Data Colada to analyze the original files themselves and lay out four red flags. First, the effect size is implausibly large: participants with evenly spaced deadlines corrected an average of 136.1 errors versus 71.1 for those with a single last-day deadline, a Cohen's d of 2.5 and a correlation of r = .79 between condition and corrections, an effect the authors say is bigger than the gender gap in height (d about 1.8) and larger than a documented manipulation check (d = 1.49). Second, in the Last Day Deadline condition, 18 of the 20 participants had a "corrections twin", another participant who logged the exact same number of corrections on all three separate tasks, with twinned subject IDs exactly 10 apart; no such twins appeared in the other two conditions. Third, measures that should correlate strongly with each other do not in the original data: five subjective ratings correlate at only -.29 to +.18 in the original files versus +.63 to +.92 in the new replication, performance across the three near-identical proofreading tasks correlates at +.03 to +.27 versus +.74 to +.90 in the replication, and self-reported minutes spent on each task correlate at +.05 to +.17 versus +.79 to +.95 in the replication. Fourth, self-reported minutes in the original data show almost no rounding: only 11.7% of values were round numbers, versus 85% in the replication, where people tend to say "20 minutes" rather than "17". The post states that, to the authors' knowledge, co-author Klaus Wertenbroch never had access to the study data. It also notes that after the authors shared their findings with Ariely and Wertenbroch, the two contacted Psychological Science to request the article's retraction, a process described as ongoing at the time of writing.

Key facts

  • Data Colada concludes the data in Studies 1 and 2 of Ariely and Wertenbroch's 2002 procrastination paper, cited more than 2,100 times, were tampered with.
  • The trigger was a new replication study (Kyle Hyndman and Alberto Bisin, forthcoming in Psychological Science) that failed to reproduce Study 2; Dan Ariely refused the replicators' October 2024 request to include an analysis of the original data file.
  • The reported effect is implausibly large: 136.1 average corrections with evenly spaced deadlines versus 71.1 with a single last-day deadline, Cohen's d = 2.5.
  • 18 of 20 participants in the Last Day Deadline condition have a 'corrections twin' with identical scores on all three tasks and subject IDs exactly 10 apart; measures that should correlate strongly (liking and interest, performance across tasks, self-reported minutes) show near-zero correlation in the original data but strong correlation in the replication.
  • Only 11.7% of self-reported minutes in the original data were round numbers, versus 85% in the replication; Ariely and Wertenbroch have since asked Psychological Science to retract the article.

Why it matters

The 2002 paper has been assigned reading in economics and psychology courses for two decades and accumulated more than 2,100 citations on Google Scholar. Data Colada's analysis, presented as evidence rather than assertion, argues that a foundational, widely taught finding about deadlines and self-control rests on data that do not behave the way real experimental data behave.

Who it affects

It directly concerns the original paper's authors, Dan Ariely and Klaus Wertenbroch, and the replication authors, Kyle Hyndman and Alberto Bisin, whose forthcoming Psychological Science paper failed to reproduce Study 2. It also touches researchers and instructors who have cited or taught the original findings, and Psychological Science, which is processing a retraction request from Ariely and Wertenbroch.

How to use it

Data Colada states its ResearchBox contains the data and code behind the analysis, so the claims can be checked directly rather than taken on faith. Readers who want to evaluate the case can compare the four flagged patterns (effect size, duplicate score 'twins', missing correlations, and unrounded time estimates) against that released material.

How solid is it

The analysis rests on original data files that Hyndman received by email in 2006 from what the file metadata attributes to Dan Ariely, and later passed to Data Colada in 2023. The post reproduces all nine means and standard errors from the original paper's figure and six other reported means, then lays out four independent statistical anomalies: an effect size (d = 2.5) larger than well-documented manipulation checks; matched 'twin' participants confined to one condition with a fixed ID offset; correlations near zero where the replication data show strong positive correlations; and self-reported time estimates that are almost never rounded, unlike the replication's. Data Colada frames its tampering conclusion as based entirely on these posted analyses and invites readers to draw their own conclusions from the released data and code.

Risks and caveats

Dan Ariely disputed that the file Data Colada analyzed is the genuine Study 2 data and did not supply an alternative; the post does not report what he offers as an explanation for the anomalies. The captured text cuts off mid-discussion of the fourth red flag, so further detail on that point, and anything the post covers after it, is not available here. No date is given for the post's publication or for the forthcoming replication paper, and the retraction request's timeline and outcome are described only as ongoing.

“We conclude that the data in Studies 1 and 2 were tampered with.”

— Data Colada