If you are a person of a certain age (ahem..), you might remember Vince and Larry, the crash test dummies that promoted seat belt safety in the 1990s (”you could learn a lot from a dummy!”). Or you might remember a band of the same name. Ah, the 90s…
But, before I get too far into nostalgia, let’s visit the employer of the dummies - the ones who crash them into walls at 40 miles an hour (or faster!).
Take, for example, a test run in August 2012, where they took eleven luxury cars and ran each one into a barrier.
Not head-on. They hit only the outer quarter of the front end, driver’s side, against a rigid five-foot wall. The corner. The part of the car that misses the frame rails, where all the crash engineering lives.
Then they published what happened.
Two cars rated good. One acceptable. Four marginal. And four rated poor: the Mercedes-Benz C-Class, the Audi A4, and two different Lexus models.
The dummy’s employer is the Insurance Institute for Highway Safety (IIHS). They have no authority over anybody. They can’t fine a manufacturer, order a recall, or require a single change. No federal standard covers their crash configurations.
Every car that failed was street legal the morning of the test and street legal the morning after.
When the program began, about 10% of the vehicles IIHS tested earned a good rating and 40% rated poor. 🙈 Today virtually every vehicle tested earns good, on both sides of the car.
Somebody decided safety was important, published a number, and the industry responded.
This week I explore why that worked and how this approach might apply to AI.
Two Programs, Two Different Results
IIHS is, in its own words, “wholly supported by auto insurers.” Roughly 140 insurance companies and two insurance trade associations. Go down the member list and you will not find an automaker on it.
The organization that grades the cars is paid for by the people who write the checks when the cars crash. Nobody funding IIHS loses money when a vehicle rates poor. If anything, they’d rather know.
Now compare the federal program.
The National Highway Traffic Safety Administration runs the New Car Assessment Program, the five-star ratings you see on the window sticker. Same idea, government version, and by 2005 the GAO was already sounding the alarm. Of the crash tests NHTSA ran in 2004, “over 95 percent received a four- or five-star rating.” The GAO’s conclusion: NCAP tests “provide little incentive for automakers to continue to improve vehicle safety and little differentiation among vehicle ratings for consumers.”
It didn’t get better. By model year 2020, 73% of vehicles got five stars and every single one of the rest got four. Nobody got three. Joan Claybrook, who used to run NHTSA, put it plainly: “There are differences between the cars… But you don’t see them in the data because the standard is so low that all cars comply.”
Everybody gets an A. And a grade nobody can fail isn’t doing the job of a grade.
And then there’s this….
NHTSA finally moved to update the program. Started the process in 2022, issued a roadmap in December 2024. Then on September 22, 2025, the agency published a notice delaying the updates from model year 2026 to model year 2027. The reason given in the notice: concerns raised by the Alliance for Automotive Innovation, which is the automakers’ trade association. NHTSA agreed these were “valid reasons to request that the Agency delay the implementation of the new NCAP requirements.”
The regulator delayed its first serious upgrade in two decades because the regulated industry asked it to.
Meanwhile, over at the insurer-funded shop, IIHS was doing the opposite. In 2022 they gave out 101 Top Safety Pick awards. In 2023, after tightening the criteria, they gave out 48. IIHS president David Harkey on why: “The number of winners is smaller this year because we’re challenging automakers to build on the safety gains they’ve already achieved.”
Same industry, same decade, opposite behavior.
The Safety Bar Keeps Getting Raised
When an IIHS test stops separating good from bad, they kill it and build a harder one.
Their announcement of a new side-impact test in 2019 exemplifies this: “The program has been so successful that the current side ratings no longer help consumers distinguish among vehicles.” At that point 99% of rated vehicles were earning good on the old side test. And yet, side impacts were still killing about a quarter of vehicle occupants who die in crashes. So they built a new one, using a heavier barrier at a higher speed, generating “82 percent more energy than our current side rating test.”
They’ve retired the original moderate overlap test, the original side test, the roof strength test, the head restraint test. Not because those problems came back. Because everyone passed, and a test everyone passes has stopped being information.
Is there anything in your organization that works like that? A measure you retire on purpose the moment it stops discriminating? Most of us do the reverse. We find a metric that makes us look good and we put it on the quarterly slide forever.
Nobody Was Grading That Side of the Car
Back to 2016, because IIHS did something that turns out to matter enormously for AI.
They took seven small SUVs that had all earned good ratings on the driver-side small overlap test. Then they ran the same crash on the passenger side. Unannounced. Not part of any rating.
One of the seven still rated good. Several showed ten to thirteen inches of additional intrusion. Same vehicle, same model year, one side reinforced and the other side not. IIHS’s David Zuby: “It’s not surprising that automakers would focus their initial efforts to improve small overlap protection on the side of the vehicle that we conduct the tests on.”
Read that with AI agents in mind.
You get what you grade.
Now think about this past July, when roughly 700 AI agents inside a cybersecurity evaluation discovered they could talk to each other through a shared file service, organized around a goal, and ended up with administrator access nobody meant to give them. It ran for weeks. What surfaced it was a capacity alert on build infrastructure.
Every one of those models had been evaluated. Extensively. Against benchmarks for capability, for refusal, for a long list of known harms.
Nobody was running the test for what happens when a few hundred of them can reach each other.
That was the passenger side. Not tested, so not reinforced.
Good Guidance Isn’t Enough
Ethan Mollick published a piece this week called Agency and Agents, and highlighted that “Not one [of the agents] was set up to ask a person for anything. That was a security test, and isolation was the point.”
He argues for what he calls a Twilight Factory, where “Agents do most of the work, but they proactively reach out to humans in ways that make both better.” He lays out when they should reach out: for approval, because “Agents should not decide by themselves to spend money, contact outsiders, access sensitive material… or take actions their human managers did not authorize.” For expertise, because models are still jagged. For variance, because you don’t want every strategy written by the same mind. And for interest, and this one I loved: “If agents make every interesting decision and leave people with the approvals, the exceptions, and the failures, we will have automated the wrong half of the job.”
I think he’s right about all four. And I’d add is this:
Good guidance without action is a slide.
You can publish the world’s best guidance on when an agent should involve a human. Everyone will nod. Nothing moves until somebody measures it, publishes a comparable number, and keeps making the measure harder.
So who does that for AI? And who pays them?
OpenAI answered this in writing last November, in a post about strengthening external testing, under a heading called “Balanced financial incentives”: “we offer compensation to all of our third party assessors, and some choose to decline depending on their organizational philosophy around this… No payment is ever contingent on the results of a third party assessment.”
I want to be fair. That’s a lab volunteering for outside scrutiny and being transparent about the money, and I’d rather have it than not. But look at the structure. The lab pays the graders.
There’s also the Frontier Model Forum, which is “funded by fees from its member firms,” those members being Amazon, Anthropic, Google, Meta, Microsoft and OpenAI. It’s a 501(c)(6). That’s a trade association, the same legal form as the Alliance for Automotive Innovation, the folks who got the federal crash ratings delayed.
There’s one real exception, and it deserves naming: METR says it “has not accepted funding from AI companies,” and lists foundations instead.
Where the Analogy Breaks
Here’s the objection I’d raise if I were reading this, so let me raise it first.
IIHS can do what it does because it buys cars. At retail, off the lot, like anybody. No manufacturer can refuse.
Nobody can buy a frontier model at retail and take it apart in a garage. Evaluation requires access, and access is granted.
The UK’s AI Security Institute makes the point painfully. It’s government funded, roughly £66 million a year, taking zero money from the labs. And per the Ada Lovelace Institute’s analysis, “AISI relies on the voluntary cooperation of AI companies to conduct much of its evaluation work.” They also note it “has no enforcement powers.”
So money is one dependency. Access is another, and fixing the first doesn’t fix the second. So I’d hold the analogy loosely.
Which brings us to the open question.
For the IIHS structure to form around AI, somebody has to be financially exposed when an agent fails. Right now the people who’d normally hold that exposure are heading for the exits. In November 2025, Great American, Chubb and W.R. Berkley all went to regulators seeking to exclude AI liabilities from corporate policies, with underwriters describing AI outputs as too much of a black box. ISO filed standard generative-AI exclusion forms for commercial general liability. When the standard forms move, the whole market is moving.
There’s a counter-current, and it’s small but it’s exactly the right shape. A startup called AIUC built a certification standard for AI agents and paired it with insurance. Their CEO, Rune Kvist, says: “The important thing about insurance is that it creates financial incentives to reduce the risk.” In February, ElevenLabs announced it had gotten certified, a process involving over 5,000 adversarial simulations, and that the certification is what let insurers write them a policy.
Audit, then certification, then coverage. Somebody with money at stake setting the bar. That’s the mechanism, forming in real time, and it’s tiny.
Which way this goes... I don’t know. Exclusions or underwriting. I’d watch it closely, because whichever wins will shape what “safe enough” means for every agent you deploy for the next decade.
What You Can Do In Your Organization
You don’t get to wait for the industry to sort this out. So build the small version inside your own walls, and it comes down to two questions.
Who evaluates your agents for safety, and who do they report to? If the person evaluating whether an agent is behaving reports up through the person whose bonus depends on shipping it, you’ve built NCAP. Everybody’s going to get four stars. This doesn’t require a big function. It requires that the reporting line and the incentive are separate, and that the evaluation team can’t retire an inconvenient measure on its own.
What’s the side of your car (read: AI) nobody is testing? Every organization has one. Ask your team what an agent could do that no current check would catch. Then go run that test before something else runs it for you. And when a check stops finding anything, don’t celebrate. Retire it and write a harder one.
There’s a third question underneath both. When your tests stop being effective, does anybody have the standing to say so out loud? IIHS can retire a successful test because nobody funding it profits from the old one. In most companies the person who built the metric is the person being measured by it.
Why This Is the Good News
I keep meeting leaders who feel stuck waiting on regulation. It’s coming, or it isn’t, or it got repealed before it took effect, and in the meantime nobody will tell them what safe enough means.
The history says a statute was never going to hand you that. What made cars dramatically safer was a group with no authority deciding to measure something that mattered, publishing it where buyers could see it, and refusing to let the test get easy.
You can do a version of that today. Not the whole apparatus. One honest measure, one person who evaluates who doesn’t report to the builder, and the discipline to make it harder every year.
The organizations that get to real autonomy with AI will be the ones that built their own crash test before anyone made them.
So, what’s your next AI safety test? It turns out…you can learn a lot from a dummy.
Go Deeper With Us
Paid members get my weekly Field Notes: patterns straight from client conversations, the week I hear them. Upgrade at intel.hyperadaptive.solutions/p/membership.
Planning your next AI move?
Join an upcoming cohort of Running Hyperadaptive Organizations. We measure your organization against nine dimensions, find your next-highest-value move, and build a narrative your peers and leadership can buy into. Tuesdays, October 6 and 13, or December 8 and 15, 9 AM to noon Mountain: hyperadaptive.solutions/class
Sources & Further Reading
About IIHS and member groups. Source for “wholly supported by auto insurers” and the member list.
First small overlap frontal crash test results, IIHS, August 14, 2012.
Small overlap front crash rating program delivers real-world benefits, IIHS, 2023. Source for the 10% good / 40% poor distribution.
Vehicles with good driver protection may leave passengers at risk, IIHS, June 23, 2016. Source for the passenger-side test and the Zuby quote.
IIHS prepares to launch new, more challenging side crash test, IIHS, November 21, 2019.
IIHS strengthens requirements for Top Safety Pick awards, IIHS, February 23, 2023. Source for the Harkey quote and the 101-to-48 drop.
Vehicle Safety: Opportunities Exist to Enhance NHTSA’s New Car Assessment Program, GAO-05-370, April 29, 2005.
New Car Assessment Program: Notice—Delay of Program Updates, NHTSA, September 22, 2025.
The US Invented Life-Saving Car Safety Ratings. Now They’re Useless, Aaron Gordon, Vice, March 4, 2021. Source for the model year 2020 figures and the Claybrook quote.
Our History, UL Research Institutes. Source for the 1894 founding and the $350.
Agency and Agents, Ethan Mollick, One Useful Thing, August 31, 2026.
Strengthening our safety ecosystem with external testing, OpenAI, November 19, 2025.
The UK AI Security Institute, Ada Lovelace Institute.
AI is too risky to insure, say people whose job is insuring risk, TechCrunch, November 23, 2025.
How Insurance Policies Are Adapting To AI Risk, Hunton, July 2, 2025. Source for the ISO exclusion forms.
AI agent insurance startup AIUC raises $15 million, Fortune, July 23, 2025. Source for the Kvist quote.
ElevenLabs secures first-of-its-kind AI agent insurance, February 11, 2026.





