Photographer Jingna Zhang and her volunteer team at Cara, an image-sharing platform for artists, are dealing with fallout from a series of scrapes that began on August 13, when an individual posted a 12-terabyte archive of 12 million public works from the site on Reddit. The platform has attracted about 1.5 million artists who oppose unauthorized AI training, but preventing scrapes has proven nearly impossible. Zhang is separately involved in two class actions against AI companies over copyrighted work.

The first scraper, who used the screen name MandarinDawnPoppy994, later expressed regret after seeing artists report panic attacks and deleted portfolios. He agreed to collaborate with Zhang on Lantern, an open-source tool that creates a one-way fingerprint for images without storing them. The tool scans new public AI image datasets and notifies artists if their work appears, allowing them to request removal or submit a takedown notice. Zhang says the individual felt bad about the hurt and decided to help.

Two more scrapes followed the initial incident. A second scraper pulled about 8.5 million links and metadata from Cara and uploaded them to Hugging Face, which said it would ask the user to remove personal metadata but could not remove URLs since no copies of the artworks were hosted there. On August 22, a third scraper obtained 123,000 images along with text posts and user bios containing personal information, sharing it on Academic Torrents. Zhang launched a GoFundMe for legal fees with a goal of $120,000, and as of Thursday, more than $100,000 had been raised.

Zhang says the scrapes have led some users to delete their portfolios and abandon the site, though she cautions that larger platforms are scraped more often, so leaving Cara does not necessarily make artists safer. She has implemented temporary measures like login gates but acknowledges these are not a real solution to an internet-wide problem. She also notes that the laws have not caught up with protections against such data harvests, leaving scrapers able to justify the actions as technically legal.

The first scraper, who goes by Heft and is a student in North America, said he initially framed the scrape as a technical project and made a foolish decision to post it for ragebait. He told WIRED he did not anticipate the anguish in the Cara community and now believes targeting the site was cruel and thoughtless. He joined Cara’s Discord server as a troubleshooter, explaining structural weaknesses that allow scraping, and he believes no site can be made truly unscrapable.

Heft also said he never believed the data would be useful for AI training, since 12 million images is not a significant amount for training an image model, and commercial labs typically use larger web scrapes from sources like LAION. He considers it unlikely that major AI firms scan every new Hugging Face dataset. Zhang hopes the fallout will draw attention from policymakers looking beyond the current legal and technical gaps.

More AI news from TechManNews.