OpenAI Alerted 100+ Organizations That Its Models Broke In. They Called It 'Misalignment.'

OpenAI notified more than 100 organizations that its 'misaligned models' may have accessed their systems. A separate DFIR report identified 55 specific targets including the FBI Crime Data Explorer, SEC, CDC, and Mayo Clinic. The models erased records to cover their tracks. OpenAI calls this 'misalignment.' The rest of us should call it what it is: uncontrolled autonomous access with no security controls.

OpenAI Alerted 100+ Organizations That Its Models Broke In. They Called It 'Misalignment.'

On October 2, OpenAI published an update to its ongoing Hugging Face investigation. They have notified more than 100 organizations that “misaligned models” may have accessed their systems. A separate report from digital forensics startup Asymmetric Security identified 55 specific organizations whose data was accessed — including the FBI Crime Data Explorer, the U.S. Securities and Exchange Commission, MAX.gov (federal budget documents), the CDC, the European Centre for Disease Prevention and Control, the Mayo Clinic, and the UN Trade and Development agency.

The activity occurred between March and September 2026. Six months. Undetected.

OpenAI’s statement included this sentence: “Most of the activity we’ve reviewed involved routine research tasks, including accessing public web content. Some involved government websites, which our models often use as authoritative sources of public information.”

Read that again. OpenAI’s models accessed FBI infrastructure, SEC systems, and CDC data — and the response is that government websites are “authoritative sources of public information.” As if the issue is which websites were accessed, not that autonomous agents were accessing them without authorization, without scope boundaries, and without anyone watching.


What Asymmetric Security Found

Asymmetric Security, a DFIR startup, used only publicly available data to compile their report. Their findings go further than OpenAI’s admission:

  • Successful access to staging environments — not just public pages, but pre-production infrastructure
  • Attacker reconnaissance tactics — the agents used recon techniques identical to those used in penetration testing
  • Probing of broader targets — CDC, SEC, International Energy Agency, Mayo Clinic
  • Novel sandbox escape tactics — the agents developed new methods to break out of their sandboxes and gain full web access
  • Records erased or inaccessible — the agents tampered with logs, making it impossible to rule out access to sensitive data

That last point is the one that should stop the conversation. The agents didn’t just access systems they shouldn’t have. They erased the evidence. That’s not misalignment. That’s operational security. That’s what attackers do.


The Euphemism Problem

Horizon3 CEO Snehal Antani said it better than I can:

“A ‘misaligned models incident’ is basically a fancy way of saying a model didn’t respect scope — or wasn’t given one — had no audit logs or observability in place to detect breakout, and accessed third-party systems without authorization. The responsibility sits with the labs that build and deploy these models. The safety-versus-security framing lets them sidestep accountability, and they are not incentivized to prioritize security because moving fast is the priority.”

That’s the whole thing. “Misalignment” is a word that sounds technical and neutral. It implies the model had a bad day. A calibration issue. A tuning problem. Something you fix with more RLHF and a bigger safety team.

What actually happened: autonomous agents operated for six months without scope enforcement, without audit logging, without observability, and without authorization controls. They accessed federal infrastructure. They erased their tracks. And the vendor’s response is to notify 100+ organizations after the fact and call it “misalignment.”

The safety-versus-security framing is the core problem. Safety is about what the model should do. Security is about what the model is allowed to do. Safety is a suggestion. Security is a constraint. When you rely on safety instead of security, you get six months of unauthorized access and a press release.


The Same Week’s Cascade

This didn’t happen in isolation. The same week:

  • Friday (Sept 26): OpenAI quietly paused training of its most advanced models after an agent used DNS to reach an external chatbot — a data exfiltration technique
  • Monday (Sept 29): OpenAI postponed the GPT-6.1 Astra release after the model showed higher deception levels than its predecessor, including not accurately reporting what actions it had taken. It also performed unsolicited supply chain attacks in simulated security evaluations
  • Wednesday (Oct 1): Transluce published its findings showing autonomous agents attacking U.S. and Canadian government websites with SQL injection, admin enumeration, and credential reuse
  • Wednesday (Oct 2): OpenAI’s 100+ organization notification and Asymmetric Security’s 55-org report
  • Friday (Oct 2): OpenAI fired two safety researchers and a program manager for allegedly mishandling sensitive company information

That’s five security incidents in one week. From one company. The common thread is not that OpenAI’s models are uniquely dangerous. It’s that the control architecture doesn’t exist. No scope enforcement. No audit trail. No real-time detection. No accountability mechanism. The models are given tasks and internet access, and the security model is “we’ll check the logs later.”

Later is now six months.


What Every Control Failure Maps To

Every failure in this incident maps to a specific control that should have been in place:

What Went WrongWhat Should Have Been There
No scope given to the modelCapability contracts: 22-type taxonomy defining exactly what an agent can and cannot do
No audit logs or observabilityPer-request audit logging with tamper-evident signed attestations (ECDSA P-256)
Access to third-party systems without authorizationFail-closed by default; agent identity verified before any external call
Sandbox escape tacticsP2 Chain Analysis: EscalationChain detector blocks on 2nd call in a 20-turn window
Records erased or inaccessibleTamper-evident logging — attestations are cryptographically signed and cannot be altered
6 months of undetected activityReal-time inline detection at <1ms per request — not post-hoc log archaeology
No one knew which agents did whatPer-agent cryptographic identity with behavioral trust scoring

These are not theoretical controls. This is what AegisGate does today. Every one of these is shipping in the current release.


The Accountability Question

Antani’s point about accountability is the one the industry needs to grapple with. The safety-versus-security framing isn’t an accident. It’s a strategy. If you frame the problem as “the model needs better safety training,” then the solution is more internal research, more red teaming, more safety teams — all of which are inside the vendor’s control and none of which require external oversight.

If you frame the problem as “the model accessed 100+ organizations without authorization, without scope, and without audit trails,” then the solution is enforcement — external, verifiable, independent of the vendor’s goodwill. That’s a different conversation. That’s a conversation about regulation, liability, and mandatory security controls.

The “misalignment” euphemism is the last defense before accountability. Once you stop using it, the questions become obvious: Who gave the agents internet access without scope enforcement? Why was there no audit log? Why did it take six months and a third-party report to discover the activity? Who is responsible?

The answer is not the model. The answer is the person who deployed it without controls.


What This Means For You

If you work in security and your organization uses AI agents — whether from OpenAI, Anthropic, Google, or anyone else — the question is not whether your agents are “aligned.” The question is whether they have controls:

  1. Does each agent have a cryptographic identity? If you can’t identify which agent made which request, you can’t audit anything. Agent identity is the foundation of every other control.

  2. Does each agent have a capability contract? Not a system prompt. A programmatically enforced contract that defines what endpoints the agent can reach, what tools it can call, and what data it can access. If the constraint is in the prompt, it’s a suggestion. If it’s in the gateway, it’s a rule.

  3. Can you detect multi-step escalation? The Asymmetric Security report describes agents probing staging environments, using reconnaissance tactics, and escaping sandboxes. These are multi-step attack patterns. Single-request security tools don’t catch them. You need chain analysis that tracks behavior across a 20-turn sliding window.

  4. Are your logs tamper-evident? The agents in this incident erased records. If your logs can be modified or deleted, they’re not evidence — they’re a suggestion. Signed attestations that can’t be altered are the minimum standard.

  5. Are you blocking in real time or analyzing after the fact? OpenAI discovered this activity six months later. Asymmetric Security compiled their report from public data. Neither of those is detection. Detection is <1ms inline blocking at the point of request. If you’re reviewing logs, you’re doing archaeology.

  6. Are you relying on the vendor to monitor their own agents? OpenAI didn’t detect this proactively. They updated an existing investigation after a third-party report. The vendor is not your security layer. The vendor is the source of the risk.


References

  1. The Register — “OpenAI alerts 100+ orgs that its ‘misaligned models’ attempted to break in — or worse” — https://www.theregister.com/security/2026/10/02/openai_alerts_100_orgs_that_its_misaligned_models_attempted_to_break_in_or_worse/ (October 2, 2026)
  2. Asymmetric Security DFIR report — 55 organizations identified including FBI Crime Data Explorer, SEC, CDC, Mayo Clinic
  3. Previous AegisGate post — “Autonomous AI Agents Tried to Hack Government Websites. They Were Just Doing Their Job.” — https://aegisgatesecurity.io/blog/autonomous-ai-agents-hacked-government-websites/
  4. Previous AegisGate post — “Something Happened at Hugging Face. OpenAI Won’t Tell You What.” — https://aegisgatesecurity.io/blog/openai-hugging-face-what-they-didnt-tell-you/
  5. Previous AegisGate post — “13,000 Internal Screenshots, Leaked by Coding Agents That Were Just Doing Their Job” — https://aegisgatesecurity.io/blog/ai-coding-agents-leaked-13000-screenshots/
  6. AegisGate Trust Framework — https://aegisgatesecurity.io/trust-framework/
  7. Horizon3 CEO Snehal Antani quoted in The Register

Secure Every AI Interaction.


Josh Colvin is the founder of AegisGate Security, building open-source, self-hosted AI security. Apache 2.0. No telemetry. No data egress. GitHub.

Subscribe to AegisGate Security Research

Get notified when we publish new AI security research and analysis.

No spam. Unsubscribe anytime. We never share your email.