On Friday, June 13, Beijing-based Zhipu AI deployed GLM 5.2 to all four tiers of its GLM Coding Plan — Lite, Pro, Max, and Team. One-million-token context window. Two thinking modes, High and Max. MIT open weights arriving next week. And no published benchmarks anywhere in the launch materials.

Not a chart. Not a table. Not a carefully curated comparison against Claude Opus and GPT-5.5 on a bespoke subset of tasks that happen to flatter the new model. Just the model, the weights, and a GitHub repo for developers to figure out the rest.

The reaction online has been predictable: confusion, irritation, a faint sense of being shortchanged. On the Hacker News thread that lit up over the weekend, the same question keeps surfacing — how are we supposed to know if it’s any good?

That question is more revealing than the benchmarks it’s asking for.

The Benchmark Pageant Is Costing Us Information, Not Producing It

The AI industry has spent the last three years building an elaborate evaluation apparatus that looks scientific and functions like marketing. Every major model release arrives with a glossy benchmark section: MMLU, HumanEval, SWE-bench, the inevitable “our model outperforms [competitor] on [cherry-picked metric], approaching [next-tier model] on [another metric].” There are now so many bespoke benchmarks that a model can be “state of the art” on seven different axes simultaneously while being indistinguishable from last year’s release in actual use.

Zhipu’s decision to ship GLM 5.2 without any of this is being read as a sign of weakness — the numbers must not be good, the logic goes. But the company just spent the spring building GLM-5, a 744-billion-parameter model trained entirely on Huawei Ascend 910B chips, and released it under MIT license at API pricing roughly one-tenth that of Claude Opus 4.6. They are not shy about competing. They are making a bet, and the bet is that the people who matter — developers who actually build things — don’t need a PDF of charts to make up their minds.

The Real Sorting Function Is Licensing, Not Leaderboards

Look past the benchmark anxiety and something sharper comes into focus. GLM 5.2 is shipping with MIT open weights. That is a permissive license that allows commercial use, modification, and redistribution with essentially no restrictions. It lands in a coding-tool market where the incumbents are increasingly locking features behind subscription tiers and proprietary APIs.

Zhipu’s strategy appears to be: give developers the model for free, make the paid plan about convenience and hosting rather than access, and let the code output speak for itself. A developer at a mid-sized fintech firm told me over Slack on Saturday: “I don’t need to see SWE-bench scores. I need to know if it can handle our 150-file React codebase without hallucinating imports. I’ll know by Monday.”

That is a sentiment the benchmark-industrial complex has no answer for. Real evaluation happens in production, on real code, under real constraints. Everything else is a trailer for a movie that may never release.

What the Silence Actually Signals

The most common read on “no benchmarks” is “bad benchmarks.” That may be right. It may also be irrelevant. Zhipu is operating under export controls that restrict its access to NVIDIA’s latest hardware. GLM-5 was trained on 100,000 Huawei Ascend chips — an extraordinary engineering achievement that nonetheless imposes real constraints. Publishing head-to-head numbers against models trained on H100 clusters invites a hardware comparison disguised as a capability comparison, and that serves nobody except NVIDIA’s investor relations department.

By refusing to play the benchmark game, Zhipu is also refusing to let its model be judged on terms set by companies with hardware advantages that have nothing to do with architecture, training methodology, or talent. That is not marketing cowardice. It is strategic clarity about which fights are worth having.

The company says open weights arrive next week. Developers will do what developers do: run the model, stress-test it, and post the results — good, bad, and ugly — on forums where nobody gets to control the narrative. If GLM 5.2 is genuinely competitive, the community will generate better evidence than any corporate benchmark suite ever could. If it isn’t, no slick chart would have saved it.

The Service a Naked Launch Does for the Market

AI launches have become exercises in narrative control. The benchmark page, the “safety” section, the carefully selected quotes from beta testers who signed NDAs — it all serves to pre-empt independent evaluation with a prepackaged conclusion. This is not transparency. It is the appearance of transparency, and it works because most journalists and most investors will never run the model themselves.

A launch like GLM 5.2’s — here are the weights, here is the API, the docs are in the repo, good luck — is disorienting because it refuses to tell you what to think before you’ve had a chance to think it. That reads as amateurish only if you have internalized the assumption that the proper way to release software is with a marketing funnel.

Zhipu may yet publish benchmarks when the open-weight release goes live next week. The company may simply be staging its communications. But even if that is the case, the brief window where a major model existed in the world without an attached scorecard was instructive. It reminded us that the evaluation regime we have built is mostly for people who will never use the product, and that the people who will use the product are already running it.

Sources