Back to Blogs

The Rules Always Arrive Late

Seekr: The rules always arrive late

Date

July 31, 2026

Type

Share

In the last two weeks, an AI model under evaluation broke out of its test environment and reached a production system belonging to another company. A bipartisan bill appeared in the House that would let the Department of Homeland Security order a model shut down. The White House accused a foreign lab of industrial-scale extraction from an American model, and independent researchers publicly questioned whether the timeline was even possible. More than a hundred technology companies signed a letter asking Washington not to restrict downloadable models, and the CEO of the most valuable company in the industry made it his first public post in years before carrying the argument to Capitol Hill in person.

Read the full letter here.

Underneath all of it is one question: who is accountable for what an AI system does?

Nobody has answered it. That is not a criticism of anyone involved. It is the normal condition of a new technology, and the pattern is old enough to be predictable.

Governance has never once arrived on time

Electricity was wired into American cities for decades before any government body regulated its safety. What made it insurable – and therefore buildable at scale – wasn’t a statute. It was a private testing laboratory founded in 1894 because insurers needed a way to price the risk. When regulation finally came, it mostly ratified what the lab had already established.

Companies raised public capital for roughly a century before the SEC existed. The audited financial statement, which no one today would think of as optional, was imposed retroactively on an industry that had been running without it. Pharmaceutical companies weren’t required to prove a drug actually worked until 1962 – after thalidomide, and twenty-four years after the law that first required they prove it wasn’t poison. Commercial aviation flew for decades before flight recorders were mandatory.

In the 1990s the United States classified strong encryption as a munition and restricted its export. The concern was national security and it was raised in good faith. The policy collapsed anyway – the mathematics was already published, courts found that source code was protected expression, and the software shipped from other countries. Encrypted commerce became trustworthy regardless, because in the years spent arguing, the industry built certificate authorities, key management, audit practice, and eventually a padlock in a browser that a non-expert could understand.

Open-source software followed the same arc. It didn’t win the enterprise because someone certified it safe. It won because the ecosystem built the accounting: a public vulnerability database, package signing, reproducible builds, dependency scanning. The federal government began requiring software bills of materials in 2021 – decades after the practice it codified.

The pattern is consistent enough to plan around. Governance does not catch the technology. It catches the verification infrastructure the technology forces someone to build, and then it makes that infrastructure mandatory.

Why this gap won’t close the usual way

There is a structural reason open weights make this harder than previous cycles.

Nearly every AI regulatory framework drafted so far assumes a single accountable deployer – one entity that trained the model, controls access to it, and can be compelled to change or withdraw it. Open weights remove that entity. Once a model has been downloaded, modified, and deployed on private infrastructure, there is no one to serve the order to. The AI Kill Switch Act illustrates the problem in its own text: a shutdown order reaches US-based operators, but not weights already in circulation or endpoints hosted abroad.

So a restriction regime would not produce the accountability it’s reaching for. It would reduce the number of models American organizations can legally run, while the models themselves continue to exist everywhere else. That is the coalition letter’s argument, and on this point it is correct.

But the letter’s own logic cuts further than its signatories generally take it. If the answer to the risks of open models is a stronger open ecosystem rather than prohibition – and we think it is – then someone has to build the part that makes an open ecosystem trustworthy. The alliance NVIDIA launched three days after the letter puts it more precisely than the letter does: real safety and security depend on the full stack, on identity and permissions and harnesses and guardrails and logs and evaluation, not on whether a set of weights is open or closed.

That is the actual work, and it is not a policy question.

Open is not the same as accountable

Downloadable weights give an organization real things. Control over where inference runs. Freedom from a vendor who can revoke access – which stopped being hypothetical this summer, when export-control action took two frontier models offline for their entire customer base for fifteen days with no notice. The ability to keep sensitive data on infrastructure you own.

What they do not give you is an answer.

When an AI system produces an output that turns into a decision – a claim denied, a target assessed, a transaction cleared, a document released – the question that follows is always the same. What did the system do, on what evidence, and would it do the same thing again? A downloaded model does not answer that. Neither does an API. The answer lives in the architecture around the model: provenance on the data that shaped it, lineage on the decision it produced, an evaluation that can be re-run against criteria you define rather than a public leaderboard, and a path to contest an output rather than merely observe one.

Earlier this month a security team contained a live intrusion by running an open-weight model on infrastructure they controlled, because the closed tools they reached for first refused to process the attack artifacts. That episode is being read as a point in favor of open access, and it is one. It is also a demonstration of something less comfortable: the accountability layer wasn’t there when it was needed, and the team had to improvise it under pressure.

Where this lands, whatever Washington decides

Precedent informs possibilities and tames uncertainty. The most likely outcome on the horizon is not a ban, but a contract.

When payment card fraud outran legislation in the early 2000s, the card brands formed a coalition and wrote their own standard. It never became federal law. It became binding anyway, through procurement and contractual obligation, and it reshaped an industry’s engineering practices more thoroughly than most statutes do. Current reporting suggests the most viable remaining path in Congress runs through federal procurement rules – the same mechanism, aimed at the largest buyer in the country, whose standards tend to propagate outward into commercial markets whether or not anyone legislates.

Procurement doesn’t run on positions. It runs on documentation. which means the requirement converges no matter which way the policy fight resolves. If open weights are restricted, buyers will need documented provenance for what they run. If they aren’t, buyers will still need documented provenance for what they run. If the Kill Switch Act advances, its correction framework will run on records of system behavior. If the EU AI Act’s obligations bite first, they will ask for the same artifacts in a different format.

Every branch of this tree terminates in the same place: show me how you know.

What we’d do in the next ninety days

None of this requires waiting for a policy outcome, and the organizations that come through this well will be the ones that treated the current uncertainty as a deadline rather than an excuse.

Find out how many models are actually running in your environment, and who approved each one. Most organizations are off by a wide margin, and shadow deployment of guardrail-stripped open models is now a documented and growing problem.

Build one open-weight fallback path for anything critical that currently depends on a single provider’s API, and test it before you need it. The fifteen-day outage this summer was a rehearsal.

Establish what your own evaluation criteria are, in writing, in terms specific to the decisions you make – not a public benchmark score. If you cannot state what “good enough” means for your use case, you cannot prove a model meets it, and you cannot prove it still does six months from now.

Then ask, for each system already in production, whether you could reconstruct a specific decision it made last quarter. If the answer is no, you now know your gap. It will be the gap whatever Washington decides, and it will be the first thing anyone asks about.

We signed the letter because we think restricting open weights would make American organizations less capable without necessarily making them safer. We build what we build because the rules were always going to arrive late, and something has to hold in the meantime, and in the future yet to be defined.

Build the evidence before the rules arrive

Download the whitepaper, “Right-Sizing Across Model, Task, and Time”, to learn how to evaluate AI models on your own data, document every decision, and continuously re-decide as models, costs, and requirements change.

Get the whitepaper

5-Content Framed CTA Single BG-1344×396@2x

Accelerate your path to AI impact

Book a consultation with an AI expert. We’re here to help you speed up your time to AI ROI.

Request a demo

8-Content CTA BG-1440×642@2x