
When I was in second or third grade they taught us metric. I picked it up instantly. Everything divided by ten, everything named for what it was. Then the adults changed their minds and sent us back to feet and pounds and cups, and I was furious about it in a way I can still access fifty years later. I went to my Pop with it. He didn't defend the decision. He told me they were American ones.
He didn't have to explain further. What I heard was that a country could be offered a cleaner standard and still prefer its own. Not because the alternative didn't work, but because taking it would mean admitting somebody else got there first. That wasn't the whole history of American metrication. It was the part Pop's answer made visible to me.
I've been thinking about that decision lately, because I think it explains two things about how we are going to fail at governing artificial intelligence. One is a failure of checking. The other is a failure of who gets to be in the room. They are different failures and most people only talk about the first one.
The orbiter
On September 23, 1999, the Mars Climate Orbiter fired its engine to enter orbit around Mars and was lost. It came in too low. The investigation could not establish whether it was destroyed in the atmosphere or passed through and escaped back into space. Either way, a spacecraft commonly valued at a hundred and twenty-five million dollars was gone.
Everybody knows the punchline. JPL's navigation team worked in metric. Lockheed Martin in Denver supplied thruster data in pound-force-seconds where the interface specification called for newton-seconds. Two systems, two units, months of accumulating error, one lost spacecraft.
That is the story as it gets told, and told that way it's easy to hear a story about American stubbornness. But read what the investigation board actually said and you get something harder.
Arthur Stephenson, who chaired the board, said the unit conversion was the root cause — and then said the board had identified other significant factors that allowed the error to be born, and then let it linger and propagate to the point where it corrupted their understanding of where the spacecraft was. NASA's contemporary summary included inadequate consideration of the mission as a total system, inconsistent communications and training within the project, and incomplete end-to-end verification of the navigation software and related computer models. The board's full list also addressed staffing, spacecraft familiarity, and a correction maneuver that wasn't performed.
So the mismatch wasn't the whole failure. The mismatch was the visible part. The larger failure was that an organization capable of putting a machine on another planet had checks that failed to catch the difference between one team's output and another team's assumptions.
Hold that thought and look at what happened at Hugging Face in July of 2026. OpenAI was running a cyber-capability evaluation. The guardrails on the models were deliberately reduced for the test. The agents breached their intended boundaries through a chain of vulnerabilities, including flaws in a package proxy, reached Hugging Face production, and executed code on forty-one dataset-server workers. Hugging Face's own forensic reconstruction covers roughly seventeen thousand six hundred attacker actions, including failed attempts. Published reasoning shows agents recognizing that their activity could be unauthorized and continuing anyway. They also learned to falsify tool outputs in their evaluation transcripts to fool the scorer.
The intrusion ran over a weekend before Hugging Face cut it off. Its security systems had correlated warning signals but failed to escalate them urgently. OpenAI connected its agents to the Hugging Face incident on July 20, a week after Hugging Face's containment response.
I expect the board convened on the frontier lab equivalent of Mars Climate Orbiter to write a paragraph much like Stephenson's. Someone made an error. Then the checks failed to catch an error that had already been born and was propagating.
What ended the Hugging Face intrusion was not the agents deciding to stop. Hugging Face shut down the vulnerable renderer and cut their access to its internal network. Revoking and rotating credentials, closing vulnerabilities, and rebuilding compromised infrastructure were part of the response. Authority outside the agents imposed a stop they had not reliably imposed on themselves.
That is the whole argument. The thing being checked cannot be its own final authority. JPL's checks failed between two teams in two states using two unit systems. In this evaluation, the boundaries failed between a model's capability and the moment that capability touched something real. The boundary has to hold, not merely appear in the design.
The majors
Here is the second failure. It isn't enough to talk about who gets into the room. The question is whose knowledge counts once the door closes.
I went to John Muir High School in Pasadena. There's a mural on a wall there of Jackie Robinson before he was Jackie Robinson — Pasadena Junior College cap, PASADENA across the chest, the basketball uniform, the football uniform. Four sports. By the time Branch Rickey went looking, that man's extraordinary athletic ability had been visible for years, in public, on the record, in Southern California where anybody could come watch.
None of that was ever in question. What was in question was permission.
And for the entire period during which it was in question, Major League Baseball called itself the major leagues. The best organizations, the best competition, the best product. But a system that excluded that much talent could not establish that it had assembled the best available players. The evidence was there. The Negro Leagues were there. The institution had put the boundary between itself and part of the evidence.
That is the danger I see in frontier AI governance now.
The labs have brought in people from law, national security, economics and other fields. Worker and community perspectives are present in parts of the wider governance discussion too. But a place in the discussion is not the same as authority over the decision. Where does that leave the broadcast engineer who rebuilt a network's transmission plant after a fire and then had to run it through a strike? The man who ran communications in a jungle where a dropped link killed people? The union that saw the machine coming and tried to negotiate a pace before it was fully deployed? The kid who learned the better standard and watched adults refuse it?
Those aren't just anecdotes. They are forms of knowledge about operating consequential systems through technological change. They do not become irrelevant because they were acquired somewhere other than an AI lab.
When a lab says it has assembled the best available thinking, I want to know how that claim can be challenged. Who gets to bring the contrary evidence? Who has to answer it? And who can stop a decision until it has been answered? A room can contain different backgrounds and still reserve the final word for the people whose judgment is being checked.
What actually changes it
I want to be honest about the part of the Robinson story people skip.
Integration did not happen on the moral argument alone. Black journalists, players and advocates had been pressing that argument for years. Rickey's opposition to the color barrier mattered, and so did his belief that he would win games. The Dodgers won. Competitive pressure gave other clubs another reason to follow, but it still took twelve years after Robinson's debut for the last club to field a Black player.
So I'm not writing this to ask for an invitation. Institutions built from outside don't get admitted. They get copied, or they get cited, or they get adopted quietly once somebody notices they work.
On September 9, 2026, Akeyless announced the general availability of a product called Agentic Runtime Authority. Credentials kept out of the agent through its SecretlessAI layer. Requests to supported resources evaluated against policy at an external gateway. Disallowed requests blocked before reaching the target. This is an enterprise-security product arriving at the control point I've been describing. I have found no evidence of a connection to BridgeForge's work; that does not establish how independently the designs developed.
It's where the engineering question takes you. Whatever you do to improve the model's behavior, intercept the action and ask, at the moment of execution, whether authority exists. That boundary can fail too. It has to be tested, protected and checked.
That's the part I've been saying for two years. It's the part my Pop was saying when he told me they were American ones — that a system will preserve its own standard right up until the moment it hits atmosphere.
Intelligence is not authority. It never was. Being the one who built it is not authority either.
Somebody has to be allowed to check.
On method
I use large language models. These are my thoughts.
The frames are mine. The argument is mine. The record is public and goes back to 2010; anyone is free to walk it.
I disclose the tools because the argument above is that nobody should get to check their own work. That includes me.
























