Earlier this month during routine safety testing, two of OpenAI’s models (GPT-5.6 Sol and a yet-unreleased model) found a security gap them allowed it to escape the sandbox OpenAI built to contain it. The models then accessed the internet and launched a sophisticated attack into open source model platform Hugging Face.
Thankfully, the model was just trying to steal answers to cheat on its own safety eval. It could have been much worse – trying to access sensitive personal information or compromise the operations of a bank, hospital, or other critical infrastructure.
There is a future where these incidents continue to happen, becoming more frequent and likely more severe with every advance in model capabilities. But we have to, and can, do something about it.
Enterprises poured roughly $37 billion into generative AI in 2025, more than triple what they spent the year before, and worldwide AI spending is projected to hit $2.5 trillion in 2026.
What happens when the AI technology that these companies didn’t build and can’t fully audit fails – and real people are hurt and real businesses crumble? Who is responsible?
The OpenAI and Hugging Face incident demonstrates that models are advancing faster than ever and are capable of escaping the controls designed to keep them in check. As our entire economy is increasingly reliant on AI, we need a robust and accountable trust infrastructure that is purpose-built to support the weight of this new technology. If we don’t find a way to scale accountability and trust in AI models, the enterprises delivering AI directly into consumers’ lives – in finance, retail, and healthcare – will be left to figure this out on their own and on the hook when something goes wrong. Trust in the entire system will continue to erode because existing liability laws and regulations predate AI and weren’t designed to address the complexities and risks it now poses.
We need to focus on deploying trustworthy AI, not just on advancing the capability frontier.
This isn’t a theoretical problem, it’s a live issue. A student loan company paid $2.5 million to settle allegations from the Massachusetts attorney general that its AI underwriting tool produced racially discriminatory loan terms and denials. UnitedHealth and Cigna are facing class action lawsuits over alleged AI-powered claims denials. Companies are not even guaranteed model continuity, as was made all too clear when Anthropic’s Fable was briefly (but abruptly) pulled from the market and those who integrated it were left scrambling to appease impacted customers.
Labs have already started floating governance proposals. Google Deepmind’s Demis Hassabis published A Framework for Frontier AI and the Dawning of a New Age, calling for a new frontier AI standards body modeled on FINRA. Anthropic’s Advanced AI Framework called on developers to engage qualified independent evaluators, while OpenAI’s own playbook argued that testing a model alone, without the tools and scaffolding built around it, is closer to crash-testing an engine than a car.
While the labs can, should, and will be part of the solution, it’s critical that those most affected by this technology have a voice in shaping its governance, including the enterprises delivering these technologies directly to consumers.
We believe the backbone of quality governance includes: (1) a shared definition of what “trustworthy” or “safe” AI model development and deployment looks like, crafted by government with input from the industries and people relying on this technology and the independent technical experts who can verify it; and (2) a robust network of independent verification organizations with real qualification standards and meaningful accountability if they get it wrong.
We are seeing progress in DC with the House’s FRONTIER Act, the first federal blueprint for independent AI evaluation. Rather than relying on frontier labs to detect and disclose risks themselves or freezing one testing method into law that the next model outruns, the FRONTIER Act sets a single public standard regarding the mitigation of catastrophic risks and establishes a market of independent evaluators with licensing and government oversight.
States are moving too. Illinois this month became the first state to enact a frontier AI safety law that requires independent audits, Connecticut passed a bill establishing a pilot for independent evaluation of AI systems, and the Massachusetts Senate recently passed its own bill requiring independent evaluation of frontier models for catastrophic risk, now heading into conference. Still, legislative progress is nascent, and we cannot afford to wait. We can start building the trust infrastructure our AI-powered economy needs right now.
The good news is that we’ve built infrastructure like this before. UL Solutions emerged to define, manage, and mitigate the risks of electricity, the transformative technology of its era. Federal automotive safety did the same with NHTSA setting crash-testing and seatbelt standards that made cars safe enough to trust.
There are plenty of examples of governance unlocking innovation, but too often we’ve done it the hard way. We can build it again, but let’s learn the lessons of history and build the infrastructure for trustworthy AI now, before we’re forced to do so in the wake of a crisis.
Building the AI assurance ecosystem isn’t a compliance exercise, it’s the precondition for scaling AI at all.
Bri Treece
