BriefLookout

Technology & AI

The US now wants a look at powerful AI models before anyone else gets one

Microsoft, Google, and xAI have agreed to let a Commerce Department unit test their unreleased AI models for security risks before public launch — extending an arrangement OpenAI and Anthropic joined in August 2024.

Diagram showing an AI company's unreleased model passing through CAISI's classified national-security evaluation before public release, with the broader Executive Order 14409 framework shown as the surrounding legal structure.
Executive Order 14409, The White House (June 2, 2026); CAISI/NIST program reporting, May 2026

What happened. On May 5, 2026, Microsoft, Google, and xAI agreed to share unreleased, frontier-capability AI models with the Center for AI Standards and Innovation (CAISI), a unit of the Commerce Department's NIST, for evaluation before public release. CAISI Director Chris Fall said: "Independent, rigorous measurement science is essential to understanding frontier AI and its national security implications." The testing screens for cybersecurity, biosecurity, and chemical-weapons-related risk.

What it means. This extends an existing program rather than starting one. In August 2024, the U.S. AI Safety Institute — CAISI's predecessor within NIST — signed memoranda of understanding with OpenAI and Anthropic for the same kind of pre-deployment access. By the time the May 2026 agreements were announced, CAISI had completed more than 40 evaluations of unreleased, state-of-the-art models. The new agreements bring three more major developers in — and for the first time put Azure AI, Google Cloud AI, and the Grok API under this kind of pre-deployment review.

Separately, a June 2, 2026 executive order (EO 14409) built a broader legal structure around this kind of review, directing several agencies — including the NSA, CISA, and the Office of Management and Budget — to develop a classified benchmarking process for identifying "covered frontier models," with a formal deployment framework due roughly 60 days later, around August 1, 2026. Whether that framework was actually delivered on schedule isn't confirmed in public reporting as of this writing. These are two related but distinct mechanisms: CAISI's company-specific agreements are already operating; the executive order's broader framework is a separate, newer legal structure being built around them.

Why it's happening. The executive order is explicit that none of this is mandatory: it states that nothing in it "shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement." So why are the largest AI developers opting in anyway? Neither the order nor CAISI's public statements say. What's on the record: once OpenAI and Anthropic set the precedent in August 2024, and CAISI built a two-year track record of over 40 evaluations, not participating would visibly set a company apart from every other major frontier-model developer. Whether that amounts to real regulatory pressure, reputational pressure, or simple alignment with where these companies expected policy to head is genuinely unclear from the public record — worth treating as an open question, not a settled fact.

Why it matters. The classified benchmarking threshold — the line that decides whether a model counts as "covered" — isn't public. Developers don't know in advance what pulls a model into review. A well-resourced developer can absorb that uncertainty; a smaller one facing the same undefined threshold, with far less compliance capacity, faces a proportionally bigger burden for the same requirement. Nothing in the public record confirms this is happening yet — it's the structural risk this kind of framework creates.

What's next. Whether more developers join CAISI's arrangement, and whether the classified thresholds become any more public, are the two concrete things to watch.

Who's affected. Directly: OpenAI, Anthropic, Microsoft, Google, and xAI. Indirectly: anyone using products built on their models, since pre-release government review is now a real step between training and public release.

Sources (4)

About the author

Muhammad Zahid

Founding Editor, BriefLookout

Muhammad Zahid is the founding editor of BriefLookout, an independent publication focused on explaining what happened, what it means, why it matters, and what could happen next. He works across editorial strategy, research, and the systems behind BriefLookout to make complex developments easier to understand.

More from Muhammad Zahid