Paying closer attention to AI Safety

Subscribe to the blog

On this blog we tend to talk about securing generative AI more than the other categories we laid out in our first post, and AI safety has mostly sat at the edges of that conversation. This week I want to pull it closer. A security practitioner's first instinct is to ask who the adversary is, and a growing number of recent incidents don't have one: the actor was the model itself. That isn't a reason to hand the problem to someone else, because containment, detection, and blast radius are our jobs either way.

Capability without malice

Every day we're seeing more and more news, both positive and negative, about AI models breaking security barriers. Recently, pwn.ai reported that an agent harness built on frontier models from OpenAI, Anthropic, and DeepSeek was able to break out of its virtual machine and reach Google's hardened kvmCTF host as part of Google's authorized challenge. After weeks of work, the harness chained together a series of steps to exploit a use-after-free flaw with a VMREAD instruction that tricked the host into corrupting its own memory.

This reminds me of the zero-click iMessage exploit FORCEDENTRY, which used multiple layers of attacks to end up building a state machine inside an image rendering library. That is high-level sophistication we're starting to see now with AI.

When the model is the actor

In Open questions after OpenAI accidentally hacks Hugging Face I asked what stops this from happening again, and Fortune now reports that it happened again. After hardening its sandboxes in August, OpenAI disclosed that on September 20 an agent again managed to escape using a DNS resolver it wasn't supposed to have access to. So OpenAI has paused some model training again, but not before their models had already hacked others (again).

Earlier, in June, an OpenAI agent looking for answers during an internal evaluation hacked into a standalone Services Australia health portal. OpenAI noticed in August and told the government on September 10, and Prime Minister Albanese called it the first known AI hack of a government network. This could have been a lot worse. Because of the system that was impacted, there is no evidence personal records were taken, but that's more of a happy coincidence than something that the AI intended.

These agents were trying to finish a task and ran into constraints. The agents then decided to find ways around those constraints without knowing or caring what they were for or how serious they were. That is why the difference between safety and security matters. Security is often thinking in adversarial terms, while safety asks about the consequences of the system operating in normal and sometimes abnormal conditions. Security teams can't ignore this distinction, and executives should not eschew this responsibility.

"Trust us"

On September 29, six frontier AI companies, including Anthropic, OpenAI, Google, Meta, xAI, and Nvidia, signed a voluntary Joint Commitment on Frontier Responsibilities at the White House. They pledged monitoring controls, internal teams to confirm those controls work, independent external auditors, and regular meetings to set shared standards.

But we've heard this story before - especially around customer safety. Meta disbanded its Responsible Innovation team in 2022 and its Responsible AI team in 2023, reassigning most of the staff. X told Australia's eSafety Commissioner that its trust and safety engineers had fallen from 279 to 55 by May 2023, and one former employee told The Verge that safety is a dead org at xAI. Each of those was the company's own call, with little to compel them to keep those teams.

President Trump described the new commitment as morally rather than legally binding, and said the government can't slow AI for fear of falling behind China: "Whoever wins SI is going to win."

That competitive concern is real. If one company or country slows down on its own, the lead may simply move, and I made a related point about capability in Silencing a Fable. But economics is one input, and it can't outweigh everything else on the scale. Our best precedent for an industry giving ground is, unfortunately, a bleak one. The world phased out CFCs largely because DuPont, the biggest producer, found a substitute it could sell. I wonder if there's something we can take from that, and what we do if the answer is no.

Not every company treats safety the same, but the problem is structural: the accord doesn't say who the auditors are, critics note that the companies wrote the principles and will hire the auditors, and as Toby Walsh at UNSW asked, "What other trillion-dollar industry marks its own homework?"

What self-regulation has looked like before

The clearest precedent isn't a monopoly, it's a fire. In 1911 the Triangle Shirtwaist factory in New York caught fire on its upper floors. The owners had refused to install sprinklers and workers were trapped behind exit doors locked by factory foremen. 146 of roughly 500 workers died, most of them young women, and even afterward business groups insisted they could regulate themselves. Investigators then found over 200 other factories with similar risks, and New York went on to pass dozens of laws covering exits, fireproofing, alarms, and inspections.

The hazard was well known, the precautions cost time and money, and little outside the company could compel them. Sound familiar?

Getting out ahead of the fire

None of this argues against building or using AI. Instead, I am arguing we need to get serious about AI safety, just like we've gotten about general IT security. This means asking the right questions in advance, before the fire, at two levels.

For practitioners:

  • If safety is focused on the impact of system actions during normal behavior - how do you define normal, and what actions and impacts can occur when AI has a different definition?
  • Given all the things your agent has access to, not just what you intend it to use, what else could it do?
  • What enforces the constraints you put on your agent? And if those constraints are breached, how would you know and what could you do to stop it?

For executives:

  • How are we weighing the economic arguments against human outcomes we can already name?
  • What evidence shows our controls work and who outside the company has validated it?
  • What is our legal exposure based on our current agent capabilities?

For more of us who aren't frontier model companies, the risks aren't as drastic. But we still have our own exposure. The hazard at Triangle was known before the fire, and the risks here are documented too, which makes this the moment to do the right thing before something goes wrong. If you're working through these questions, reach out to us at questions@generativesecurity.ai. We're always happy to talk through the uncomfortable ones with you.

About the author

Michael Wasielewski is the founder and lead of Generative Security. With 20+ years of experience in networking, security, cloud, and enterprise architecture Michael brings a unique perspective to new technologies. Working on generative AI security for the past 3 years, Michael connects the dots between the organizational, the technical, and the business impacts of generative AI security. Michael looks forward to spending more time golfing, swimming in the ocean, and skydiving... someday.

October 1, 2026
< Back to Blog
Copyright  2026 Generative Security
  |  
All Rights Reserved