Aaron Swartz was prosecuted for JSTOR scraping, Meta wasn't for AI training, blog argues

Aaron Swartz was prosecuted for JSTOR scraping, Meta wasn't for AI training, blog argues

A blog post revisits the case of Aaron Swartz, described as a co-creator of the RSS protocol among other contributions, who was prosecuted after downloading about 70 gigabytes of academic articles from JSTOR. The author says the charges were pursued so aggressively that Swartz faced exposure of up to 35 years in prison, a $1 million fine, and asset forfeiture (no dollar figure is given for the forfeiture), brought, in the post's words, "to be made an example of." Swartz died by suicide before the case reached trial; the post frames this as the legal system having effectively destroyed him rather than a choice he made freely, saying he "felt the need to take his own life rather than deal with the court circus and impending financial ruin."

The post then turns to Meta, sarcastically calling it "Facebook (oh I'm sorry Meta)" before saying the company has torrented 80 terabytes of books to train its AI models. According to the author, Meta faces "virtually no consequences other than a court case" and will most likely get only "some sort of financial slap on the wrist," while its AI products keep generating revenue. No court, case name, jurisdiction or date is given for this litigation, and the post reports no ruling or settlement; the prediction of a light penalty is the author's own expectation, not a stated result.

The author draws a direct contrast between the two: Swartz's downloading, the post says, served "the dissemination and archival of knowledge," while Meta's copying, in the author's words, powers "proprietary plagiarism code" that the post says harms the environment and enriches an already very wealthy person. The piece closes on a personal note, with the author writing they never met Swartz but feel angry on his behalf, unsure what to do with that anger "other than get more radicalized."

The post is written in the first person throughout and carries no byline; it appears on the personal blog blog.curiousquail.com. It supplies no court records, case numbers or dates for either matter, and does not say whether the 35 year and $1 million figures were an actual sentence and fine or only the maximum exposure once charged; no ruling or settlement is reported for the Meta case either. Submitted to Hacker News, the piece drew more than 1,100 points and 265 comments within about eight hours.

Key facts

  • Aaron Swartz, described as a co-creator of the RSS protocol, was prosecuted over downloading about 70 gigabytes of academic articles from JSTOR.
  • The post says the charges carried exposure of up to 35 years in prison, a $1 million fine and asset forfeiture (no amount given for the forfeiture), brought, in its words, "to be made an example of."
  • Swartz died by suicide before the case reached trial; the post frames this as the legal system effectively destroying him rather than an outcome he chose freely.
  • The same post says Meta has torrented 80 terabytes of books to train its AI models and faces "virtually no consequences other than a court case," which the author expects will end in only "some sort of financial slap on the wrist."
  • No court, case name, date or ruling is given for either matter; the piece is a first person opinion post with no byline, not a reported news account.

Why it matters

This piece revives an argument that keeps resurfacing whenever a well-funded AI company's use of copyrighted or pirated training material draws attention: an individual accused of unauthorized downloading can face severe criminal exposure, while a large company faces a civil case the author predicts will end comparatively lightly. Whether or not the specific figures hold up, the underlying question, whether copyright and computer-access laws land more harshly on individuals than on companies training AI models, remains part of the wider debate over how AI systems get their training data.

Who it affects

This resonates most with people who track how data-access prosecutions and AI copyright disputes are handled: on one side, those sympathetic to Swartz's legacy and wary of harsh computer-crime enforcement; on the other, authors and publishers concerned about their work being used without permission to train AI, and Meta itself, described in the post as facing a court case over the practice.

How to use it

There is no new filing, dataset or link here to act on, only an argument, so treat the piece as opinion rather than investigation. Anyone who wants the underlying record rather than the framing should look up the primary case materials for the prosecution of Aaron Swartz and for whatever litigation Meta faces over its use of torrented books, since this post gives no case citations, links or dates for either. Read the 70 gigabytes, 80 terabytes, 35 year and $1 million figures as the author's own numbers, not confirmed courtroom facts.

How solid is it

This is a personal, unsigned opinion post, not a reported investigation: it carries no byline, and gives no court documents, case numbers or dates for either the Swartz prosecution or the Meta litigation. Whether the 35 years and $1 million reflect an actual sentence and fine or only the maximum exposure once charged is left unstated, and no ruling or settlement is reported for the Meta case, only the author's own guess at what it will produce. The comparison itself, a criminal prosecution against a person set against a civil case against a company, is a familiar one in arguments about data scraping and AI training, which makes this a recognizable argument rather than a fresh disclosure.

Risks and caveats

The main risk is treating a criminal prosecution, brought by government prosecutors under computer-crime statutes, as directly comparable to a civil court case, typically brought by private plaintiffs over copyright, since the two rest on different laws, different standards of proof and different possible penalties. The specific numbers, the prison and fine exposure, the asset forfeiture with no amount given, and the 80 terabyte claim about Meta, all rest on the author's own account with no supporting citation, and the predicted "slap on the wrist" for Meta is the author's expectation, not a reported result. The piece includes no response from Meta.

“Swartz' use case was the dissemination and archival of knowledge; Meta's use case is powering up their proprietary plagiarism code.”

— the post on blog.curiousquail.com