On June 18, Andy Baio published the forensic breakdown on Waxy.org: a marketing agency had taken The Dictionary of Obscure Sorrows—John Koenig’s decade-long, 311-word labor of love, a book with an ISBN, a publisher, and a real author who gets royalties—and republished it verbatim on a slick new website, complete with an AI word generator and an invitation to “give your sorrows a voice.” The site lifted Koenig’s entire 800-word foreword, every definition, every etymology, every short essay. It even slapped the book on Amazon, as if it were a product the agency had any right to sell.
The immediate reaction was to file this under “AI plagiarism,” the latest entry in a growing ledger of generative-AI horror stories. That framing is tidy. It is also wrong in a way that lets the actual culprits off the hook.
The AI Part Is a Distraction
Let’s be clear about what the agency actually used AI for. According to Baio’s report, the site employed AI to generate new synthetic neologisms—things like “glintwist” and “solomber”—and to produce accompanying illustrations. The core act of theft, though, was not algorithmic. It was a copy-paste job. Someone at the agency took a PDF or a Kindle file and dumped it into a website template. The AI-generated add-ons were decorative, a thin veneer of “innovation” painted over the brute fact of content misappropriation.
This matters because the AI framing lets people believe the problem is a novel one—that machine learning has created some unprecedented legal gray zone. It hasn’t. Republishing an author’s entire work without permission has been copyright infringement since long before anyone worried about transformer architectures. The agency didn’t stumble into this because Midjourney’s terms of service are unclear. It did it because, in its world, a book is just raw material that happens to come with a cover.
The Real Offender Is the Marketing Stack
What kind of outfit does this? Baio traced the site to a marketing agency—the kind that pitches clients on “content amplification” and “digital brand experiences.” These agencies sit in the middle of a supply chain that has spent a decade normalizing the idea that anything on the internet is, at some level, repurposable. They scrape product descriptions for SEO landing pages. They spin up “review” sites that are really affiliate-link farms. They build “fan experiences” for IP they don’t own, then claim fair use when the cease-and-desist arrives.
A former brand strategist at one of these shops, speaking on condition of anonymity from a WeWork in Santa Monica, put it plainly: “When your entire business model is ‘making content work harder,’ an author’s book stops looking like a finished work and starts looking like underperforming inventory.”
This is not an AI problem. It is a business-model problem that predates ChatGPT by a solid decade. The agency’s decision to republish Koenig’s book was, in its own logic, no different from republishing a manufacturer’s product catalog—something you do because the content exists, the brand needs “assets,” and no one has explicitly told you to stop.
What the Publisher Didn’t Do
The strangest part of the story is the silence. Simon & Schuster published The Dictionary of Obscure Sorrows in 2021. It has a legal department. Where was it? The agency’s site had been live and indexed, by Baio’s account, for some time before the MetaFilter user spotted it. A book with an active Amazon listing and a major publisher behind it had its entire text republished on a third-party domain, and the publisher’s automated enforcement systems—the kind that scan for pirated PDFs on torrent sites—appear not to have noticed.
That’s because those systems are tuned to find torrents and unauthorized uploads to file-sharing platforms, not polished, marketing-agency-owned websites that look like legitimate promotional partners. The publisher’s infrastructure is optimized to catch the piracy it expects—individuals sharing files—not the piracy that arrives in a blazer and a pitch deck.
The Lesson That Won’t Be Learned
The Obscure Sorrows theft will get a few days of outrage, a DMCA takedown, and maybe a terse statement from the agency about “a contractor who failed to follow our content sourcing guidelines.” It will then be filed under “AI bad” and forgotten. That’s the comfortable outcome for everyone involved, because it means the real lesson—that the internet’s marketing middlemen have spent years building a machine that treats creative work as an unowned input—doesn’t have to be confronted.
If you want to be angry at something, don’t be angry at the word generator. Be angry at the business culture that decided a book with an author, a publisher, and a copyright notice was just another piece of content to optimize for click-through. That culture was here before the first chatbot went live, and it will be here after the last one is forgotten.