On Wednesday, July 9, OpenAI released GPT-5.6 to the public. The launch came after months of what the company described as “government safety evaluations” — a process that, according to a Nextgov report, involved “working alongside government partners” to assess the model’s capabilities before anyone else could touch it.
The framing was careful. A June executive order had asked major AI developers to voluntarily submit their leading models for review. OpenAI complied. The public was told this was about safety. The public was told this was responsible.
The public was told wrong.
The Evaluation That Wasn’t
Let’s be precise about what happened here. The government did not audit GPT-5.6. It did not subject the model to adversarial red-teaming and then step back. It “worked alongside” OpenAI — the Nextgov article’s phrase — for months. That is not the language of a regulator. That is the language of a customer.
When a defense contractor “works alongside” the Pentagon on a new weapons system, nobody calls it a safety evaluation. They call it a procurement process. The only difference here is the label.
One former intelligence-community liaison, standing in the exhibitor hall at last month’s AI Expo in Washington, put it to me this way: “The safety review is the demo. By the time the public sees the model, the agencies have already figured out what they want to do with it. The rest is just the consumer version.”
This is not speculation. The pattern is already visible. Anthropic’s Mytho model — cited in the same executive order that framed GPT-5.6’s review — went through an identical process. Government partners got early access. Evaluations were conducted. And then, months later, a public release arrived, stripped of whatever capabilities the national security apparatus found most useful.
Two Models, Two Audiences
The practical result is a two-tier AI ecosystem. The most capable versions of these models — the ones with the fewest guardrails, the deepest reasoning, the most flexible tool-use — are not available to the public. They are available to the agencies that evaluated them.
This is not a conspiracy theory. It is the logical endpoint of a voluntary framework that asks companies to hand over their most powerful systems to the same institutions that have every incentive to keep the best tools for themselves. The executive order did not create a firewall between evaluation and adoption. It created a pipeline.
And the pipeline works. OpenAI gets a government seal of approval that insulates it from liability and regulatory uncertainty. The government gets first access to technology that would otherwise take years to develop through traditional contracting. The public gets a model with the sharp edges sanded off and a press release about safety.
The Uncomfortable Middle
Here is where the argument gets awkward for everyone.
For the AI safety crowd, the uncomfortable truth is that their preferred policy instrument — government evaluation of frontier models — is being used not to protect the public but to give the state a head start. Every call for mandatory testing is, in practice, a call for mandatory early access. The safety label is a fig leaf, and the fig leaf is working.
For the deregulation crowd, the uncomfortable truth is that the government is already inside the tent. The models are being shaped by national security priorities before they ever reach the market. There is no “free market” in frontier AI when the most powerful versions of the technology are routed through Langley and Fort Meade before they reach San Francisco.
And for the rest of us, the uncomfortable truth is that we are being asked to trust a process that was never designed for our benefit. The safety evaluation is not a consumer protection mechanism. It is a procurement mechanism with better PR.
What the Release Actually Means
GPT-5.6 is now publicly available. That is a real event, and it will have real consequences for the millions of people who use these tools. But the version the public received on Wednesday is not the version that matters most. The version that matters most is the one that spent months in government hands, being integrated, tested, and adapted for purposes that will never appear in a blog post.
The next time a frontier model is held back for “safety evaluation,” ask yourself: who is being kept safe from what? The answer, increasingly, is not the public. The answer is that the evaluation period is the procurement period, and the safety narrative is the price of admission.
OpenAI did not delay GPT-5.6 because it was dangerous. It delayed GPT-5.6 because the government asked to see it first. And the government asked to see it first because that is what the executive order was always designed to enable.
That is not safety. That is strategy. And it is working exactly as intended.