Argues the contrast is stark and morally indefensible: Swartz downloaded JSTOR articles he had legitimate MIT-affiliated access to and faced 35 years plus $1M in fines under the CFAA, while Meta torrented 81.7TB of openly pirated books from LibGen and got a fair-use ruling. The post frames this as evidence that computer-access law protects corporate scale while criminalizing individual idealism.
By submitting the post and driving it to 1,385 points, speckx and the HN community elevated the argument that the same conduct — unauthorized bulk downloading — draws felony charges when done by a researcher and a fair-use blessing when done by a trillion-dollar company. The community's engagement signals broad agreement that this asymmetry is the defining scandal.
Contends the CFAA has historically been elastic enough to prosecute a simple curl script, and while Van Buren v. United States (2021) narrowed the statute, it still hangs over individual researchers and hobbyists. The editorial frames the statute itself — not just uneven enforcement — as the underlying failure that made Swartz's prosecution possible.
Points out that the Kadrey v. Meta ruling turned on plaintiffs failing to prove market harm, not on whether Meta's conduct — including internal Slack messages showing engineers hiding seeding activity from corporate IPs — was legitimate. The editorial argues this creates a de facto safe harbor: if your infringement is large enough to feed a foundation model, fair use absorbs it.
A blog post by developer Curious Quail hit the top of Hacker News with 1,385 points, dragging a decade-old wound back into the timeline. The argument is simple and hard to dismiss: Aaron Swartz — co-creator of RSS, co-founder of Reddit, architect of Creative Commons infrastructure — was indicted in 2011 on 13 felony counts under the Computer Fraud and Abuse Act for downloading academic articles from JSTOR through an MIT network closet. He faced up to 35 years in prison and $1 million in fines. He took his own life in January 2013, at 26, while the case was still pending.
Meanwhile, court filings from the *Kadrey v. Meta* class action revealed in early 2025 that Meta torrented roughly 81.7 terabytes of pirated books from LibGen, Anna's Archive, and Z-Library to train the Llama family of models. Internal Slack messages showed engineers explicitly worrying about the optics of seeding pirated material from corporate IP addresses and taking steps to hide it. In June 2025, a federal judge ruled that Meta's use qualified as fair use — largely because the plaintiffs failed to prove market harm — and dismissed the core copyright claims. No executive was charged. No felony indictment. No 35-year exposure.
The asymmetry is the story: Swartz downloaded articles he had legal access to as an MIT-affiliated researcher and was prosecuted as a felon; Meta downloaded pirated books it had no license for and got a fair-use ruling.
The easy read is 'laws are for the little guys,' and that's true but not sufficient. The more precise read is that the U.S. legal system has quietly bifurcated computer-access law into two regimes: one for individuals, and one for corporations at scale.
Under the CFAA — the statute Swartz was charged with — 'exceeding authorized access' has historically been elastic enough to prosecute a curl script. The Supreme Court narrowed this somewhat in *Van Buren v. United States* (2021), but the statute still hangs over anyone doing systematic data collection without a lawyer. Meanwhile, the Ninth Circuit's *hiQ v. LinkedIn* decision effectively blessed scraping of public data at commercial scale. And now *Kadrey v. Meta* adds a third layer: even scraping copyrighted material without a license can survive if you can argue transformative use for AI training.
The legal doctrine isn't the whole picture, though. Prosecutors have finite resources and infinite discretion. Carmen Ortiz's office pursued Swartz with what her critics called disproportionate zeal — reportedly because they wanted a CFAA scalp to deter 'hacktivism.' No U.S. Attorney is going to indict Mark Zuckerberg for torrenting LibGen, because the case would be politically catastrophic, technically complex, and would require going up against Meta's legal war chest. Selective enforcement isn't a bug in the system; it's how the system is calibrated.
The uncomfortable truth for practitioners is that 'is this legal?' has become a question with two answers depending on how much runway you have. A solo developer scraping Reddit for a side project lives in a different legal universe than a well-funded startup doing the exact same thing with a term sheet in hand. The former can be ruined by a cease-and-desist. The latter can absorb the lawsuit as a line item.
Community reaction on Hacker News zeroed in on this framing. The top comments weren't about Swartz specifically — they were about the pattern. One thread traced how the *hiQ*, *Van Buren*, and *Kadrey* rulings together create a de facto regime where the legal risk of scraping is now inversely proportional to your ability to pay for legal defense. That's not a rule of law. It's a rule of resources.
If you're building anything that touches scraped or crawled data — an AI product, a search index, a research tool, a price aggregator — the operational reality has shifted in ways worth being explicit about.
First, robots.txt is now legally load-bearing in ways it wasn't five years ago. Ignoring it used to be a norms violation. Post-*hiQ* and post-*Kadrey*, courts are increasingly using compliance (or non-compliance) with published access controls as a signal of good faith. If you're scraping, respect robots.txt not because it's polite but because it's the cheapest insurance policy you can buy. The same logic applies to rate limits, ToS acknowledgments, and any technical measure a site publishes.
Second, the 'transformative use' defense is real but narrow. *Kadrey* worked for Meta partly because Llama's outputs don't reproduce the training books verbatim in any commercially meaningful way. If your product regurgitates its training data — a code assistant that emits GPL code verbatim, an image model that reproduces watermarks — you don't get the same shield. Build with output filtering and provenance tracking from day one, not as a v2 feature.
Third, keep your internal Slack clean. The single most damaging evidence in Kadrey wasn't the scraping itself — it was the internal messages showing Meta employees knew they were pirating and tried to hide the seeding. Intent matters in fair use analysis. Discovery is brutal. If your team is joking about 'liberating' datasets in a channel that will be subpoenaed someday, you're building future exhibits for the plaintiffs.
The Swartz case will keep resurfacing because the wound hasn't healed — and because every new corporate-scale scraping ruling makes the disproportionality more visible, not less. The realistic path forward isn't waiting for Congress to reform the CFAA (they've had 13 years and produced nothing); it's building in a way that assumes the legal ground beneath scraping will keep shifting. Assume public data can become private, assume fair use rulings can be reversed on appeal, assume that the same act that gets Meta a favorable ruling could get you a lawsuit you can't afford to fight. Swartz's legacy on the technical side — RSS, Creative Commons, Markdown — is embedded in the web. His legacy on the legal side is a warning that the CFAA is still loaded, still pointed at individuals, and still selectively fired.
I don't like talking about this, but first hand knowledge is rarer by the day, and there are entire organizations profiting off this mythology. It's pissing me off. Aaron is not a data point to build stupid metaphors around. He was a bright and broken child.Aaron attracted influential and
He wasn’t prosecuted for scraping. He trespassed into a room with a router, plugged his laptop into it, downloaded papers as quickly as possible, and then rotated his MAC address to dodge the bans that the admin was trying to place on him. That’s very different from downloading a webpage on the open
I don't think it matters much for the argument, which is valid (or not) regardless of whether you get the precise facts about the Swartz prosecution right, but Swartz was not facing 35 years. That's the statutory maximum sentence you'd get if you ignored the sentencing guidelines and
I recently came to the conclusion that it was never about copyright. It's about corporate control, about punishing contempt for business model.Aaron Swartz was punished because he disrespected a business model. All the kids sued by the MAFIAA were punished because they disrespected a business m
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
The part that still bothers me so much about the US vs Swartz case is that JSTOR didn't pursue civil litigation against Aaron. It was the US government that pursued him.There was little for the government to lose in the case. In a case vs Meta, at the scale it has reached, it could have wide ra