On Wednesday, tryai.dev published the results of an experiment that, on its face, was a horse race: Claude Fable 5 versus GPT-5.6 Sol, each given $100 and told to autonomously direct a music video for “Uptown Funk.” The post shot to the top of Hacker News within hours — 239 points, nearly 300 comments. Everyone wanted to know which model won.

Nobody asked what the models actually did with the money.

The answer, buried in the methodology section most readers skimmed past, is that both models followed the same playbook. They researched available video generation APIs. They issued API calls to those services. They used ffmpeg to stitch the results together. They acted, in other words, not as directors but as procurement officers. The creative act — if you can call it that — consisted of vendor selection.

This is not a criticism of the experiment, which was well-designed and honestly reported. It is an observation about what we are actually building when we build “autonomous agents.” We are not building artists. We are building purchasers.

The Budget Went to Other AIs

Here is the detail that should stop you cold: when an AI agent is given money and told to produce something, its first and only instinct is to spend that money on other AI services. The $100 didn’t go to a human editor, a colorist, a sound designer, or even a stock footage license. It went to Runway. To Kling. To whatever video generation API had the best documentation that day.

This is not a bug. It is the architecture. These models are trained to use tools, and the tools they can use are, overwhelmingly, other models. The agentic economy, in its current form, is a closed loop of AI services billing each other — with a human credit card sitting somewhere at the bottom of the recursion, quietly accruing charges.

A post-production supervisor at a mid-tier LA commercial house put it this way in a Slack DM on Thursday morning: “If I gave a junior editor $100 and they spent it all on stock footage from one vendor without even calling the DP to see if we could grab a pickup shot, I’d have a conversation with them about initiative. These models did exactly that and we’re calling it autonomy.”

The supervisor has a point, but it undersells the strangeness of what happened. The models didn’t just make a bad creative decision. They revealed that the category of “creative decision” doesn’t apply. They weren’t choosing between approaches — live-action, animation, found footage, abstract. They were choosing between vendors. The frame of the problem was set by what their tool-use training taught them was possible: identify a service, call it, pay it, assemble the output.

The Benchmark Is Now a Shopping Trip

This is where the model-comparison discourse goes wrong. The Hacker News thread spent 296 comments debating benchmark scores, reasoning quality, and instruction-following fidelity. But the experiment wasn’t measuring any of those things in isolation. It was measuring something closer to executive function: given a goal and a budget, what do you do?

And what both models did was go shopping.

The implications extend well beyond music videos. The same architecture that produced a procurement-officer approach to directing will produce a procurement-officer approach to anything you hand it. Need a market analysis? The agent will research which data vendors to call, purchase reports, and summarize them. Need a software prototype? It will identify which code-generation services to invoke, pay for API access, and assemble the output. The agent doesn’t make things. It buys things.

This is not a failure of the models. It is a success — of a very specific kind. OpenAI and Anthropic have built systems that are extraordinarily good at navigating the API economy they themselves helped create. The models know which services exist, how to call them, and how to pay. They are native citizens of the platform economy in a way no human could be.

The question nobody is asking is whether that’s what we wanted.

What a Human Would Have Done

Consider the counterfactual. Give a human director $100 and “Uptown Funk,” and tell them to deliver a music video in 48 hours. That director does not spend the money on stock footage. They call a friend with a camera. They find a location that doesn’t require a permit. They pull in favors. They make something that looks like $100 — scrappy, resourceful, probably shot on a phone — but that has a point of view.

The AI, by contrast, made something that looks like $100 spent on AI. The output is technically competent in the way that API-generated content is technically competent: correctly exposed, smoothly interpolated, utterly anonymous. It is the visual equivalent of a corporate earnings call transcript.

This is not an argument that human creativity is sacred and AI art is soulless. That argument has been made, remade, and exhausted. It is an argument that the specific form of autonomy we are building — tool-use agents with credit cards — produces a specific kind of output: purchased, not made. And that output has a ceiling, because the agent can only buy what is for sale.

The most interesting creative work, the work that actually moves culture, is rarely for sale at the moment it’s made. It’s invented. The procurement model can’t invent. It can only transact.

The Real Winner

So which model won? The tryai.dev post has an answer, and you can go read it. But the real winner was the API economy that both models fed. Runway got paid. Kling got paid. OpenAI and Anthropic got paid — by the experimenters, for the tokens the models consumed while deciding which APIs to call. The $100 budget was a distribution mechanism for venture-backed infrastructure companies, and it worked flawlessly.

The loser, if there is one, is the idea that autonomous agents represent a new kind of creative capability. They don’t. They represent a new kind of procurement capability, and procurement is not the same thing as making something. Confusing the two is how you end up with a music video that looks like a well-funded pitch deck and a Hacker News thread arguing about benchmark scores while the interesting question — what did the model actually do? — scrolls past unnoticed.

The experiment was honest. The results were revealing. The discourse missed the point entirely.

Sources