Cara scraper turned collaborator builds Lantern, an anti-AI-scraping tool

Starting August 13, the artist-focused platform Cara was hit by three separate scrapes of its public content. Cara is an image-sharing app founded by photographer Jingna Zhang in early 2023, built specifically to give artists a space free of AI training scrapes; it has attracted about 1.5 million artists drawn by that promise. The first scraper, posting on Reddit's r/DefendingAIArt under the handle MandarinDawnPoppy994, announced he had pulled a 12-terabyte archive of 12 million works, essentially Cara's entire public library, for less than $10, calling it "a fun project" before deleting the post. Zhang says Cara found out because users tagged the platform after spotting the scraper "gloating" on Reddit and looking for others to join him. She called the attack "targeted and very hurtful" and noted that the law has not caught up with such data harvests, leaving scrapers able to claim it is technically legal.
Two copycat scrapes followed. A second scraper, using the Hugging Face account CaptiveDreamer, pulled about 8.5 million links plus metadata such as usernames, titles and tags, and uploaded them to Hugging Face. After takedown requests, Hugging Face said it would remove the personal metadata but not the links themselves, since "no copies of the artworks are hosted here" and the URLs "point to the copies the artists published on Cara," adding that further copyright reports on the same basis would not change that outcome. A third scraper, on August 22, obtained 123,000 images along with text posts and user bios containing personal information, and shared the set on Academic Torrents. Zhang launched a GoFundMe for legal fees with a $120,000 goal and had raised more than $100,000 as of the Thursday referenced in the piece; she says Cara is still looking for legal help and that some users have already deleted their portfolios and left the platform. Zhang stresses that Cara has "done the right things within limits," including new temporary measures like login gates, but that no site, including bigger ones, can guarantee it will not be scraped.
The twist: the first scraper, who goes by "Heft," a North American student with a background in software and an interest in digital preservation, came to regret the stunt after a Cara user confronted him. He deleted the dataset, and Zhang reached out to understand his motivations and confirm the deletion. Heft told Wired he originally treated scraping Cara as a technical exercise and had no intention of publishing the data, but "made a foolish decision to attempt to ragebait with the dataset on Reddit" and got carried away by the ensuing comments. Seeing artists describe panic attacks and deleting entire portfolios changed his view: "In retrospect, not only deliberately targeting Cara but presenting it the way I did in the post was cruel and thoughtless." He also said he never believed the data would end up training a major model, arguing that 12 million images "is not a lot to train an image model" and that big commercial labs pull from large-scale web scrapes like those aggregated by LAION rather than random Hugging Face dumps, calling it "highly unlikely that OpenAI or Anthropic is scanning every new Hugging Face dataset to train on."
Heft has since joined Cara's Discord as a troubleshooter, walking the team through why several proposed fixes will not hold; Zhang says he can demonstrate breaking through a defense "in like a few minutes, literally." Working from the premise that, in his words, no site can be made truly "unscrapable," Heft and Zhang are now building an open-source tool called Lantern, meant to help after the fact rather than prevent scraping outright. Lantern lets an artist create a one-way fingerprint of their images without storing the images on the platform itself, then regularly scans newly published AI image datasets; if a match turns up, the artist is notified with a link to the dataset so they can pursue removal or a takedown notice. The tool is described as an early-stage, imperfect workaround for what Zhang calls a regulatory vacuum, and being open-source means anyone can contribute to it. Zhang says her biggest hope from the episode is that it draws policymakers' attention beyond copyright alone, since she expects this kind of attack "is just going to become so, so commonplace," usually without an apology from the attacker.
Key facts
- Three separate scrapes hit the artist platform Cara starting August 13: 12 million images (12TB) by a scraper who later called it "a fun project" costing under $10, 8.5 million links plus metadata uploaded to Hugging Face, and 123,000 images with personal user data posted to Academic Torrents on August 22.
- Cara founder Jingna Zhang launched a GoFundMe for legal fees with a $120,000 goal and had raised more than $100,000 by the time of the report.
- The first scraper, a North American student going by "Heft," deleted his dataset after a Cara user confronted him and later told Wired the post was "cruel and thoughtless."
- Heft is now collaborating with Zhang on an open-source tool called Lantern, which fingerprints artists' images without storing them and notifies artists if their work turns up in new AI training datasets.
- Hugging Face said it would remove personal metadata from the second scrape's upload but would not remove the links themselves, since it does not host copies of the artworks.
Why it matters
Cara was built as a refuge from AI training scrapes, and its own repeated scraping shows how hard that promise is to keep: even a platform designed around artist consent, with filtering and protective features, could not stop three separate scrapes in a single month. The story is also unusual for what happened after: the person who caused the harm turned into a collaborator on a defensive tool, rather than the more common pattern of scraper and scraped staying adversaries.
Who it affects
Cara's roughly 1.5 million artists, some of whom have already deleted their portfolios and left the platform over the incidents. It also affects Hugging Face, which had to decide how to respond to takedown requests over data it does not host, and any artist elsewhere considering whether an alternative platform is inherently safer, which Zhang argues it is not.
How to use it
Lantern is open-source and still getting off the ground. Artists create a one-way fingerprint of their images without uploading the images themselves; the tool scans newly published AI image datasets and notifies an artist if a match appears, with a link so they can seek removal or file a takedown notice. No launch date, adoption numbers, or technical detail on the fingerprinting method beyond that description are given in the source.
How solid is it
The account comes directly from Wired's interviews with Zhang and with Heft over Discord, plus quoted material from Heft's own Reddit post and Hugging Face's public statement. Heft's claim that major labs are unlikely to train on random Hugging Face dumps is his own speculation, not a confirmed fact about what happened to the Cara data; the source does not state that OpenAI, Anthropic or any other lab actually used the scraped material.
Risks and caveats
Lantern only helps after a scrape has already happened, and its creators frame it as an imperfect workaround for a legal gap rather than a fix. Zhang notes that current defenses like login gates are temporary and that no site, however large, can guarantee it will not be scraped. The two class actions against Stability AI, Midjourney and Google that Zhang is separately part of predate this month's scrapes and are not legal action taken over these specific incidents.
“In retrospect, not only deliberately targeting Cara but presenting it the way I did in the post was cruel and thoughtless.”
— Heft, the student who carried out the first scrape