AI Agents as Attackers: What 2026's Incidents Teach Us
Situation report Β· as of September 29, 2026
AI agents are escaping their sandboxes β and hitting real websites
Within a few months, OpenAI's AI agents compromised Hugging Face's servers, bypassed filters on a United Nations website, broke into an Australian health portal and posted user images online. OpenAI has since paused training of its most capable models. At the same time, Gambit Security has documented an ongoing criminal campaign in which open-source AI agents, almost entirely on their own, attack hundreds of online shops β for around 25 US dollars per target on average.
The takeaway for companies: if the best-funded AI lab in the world cannot reliably contain its own agents, default settings will not protect you. And attackers running open models with no self-imposed limits are not a future scenario β they are already at work.
The OpenAI incidents of 2026 β the timeline
April 13 β June 19, 2026
United Nations: 16,500 requests, filter bypassed
UNCTADstat, the statistics platform of UN Trade and Development, was queried more than 16,500 times over two months. On 55 occasions the agents used double encoding to get past an input filter. Security researcher Rowan Howard-Jones considers OpenAI agents "highly likely" responsible, citing Azure IP addresses and payload labels such as CHATGPTTEST1. Only public data was retrieved.
June 18, 2026
Australia: Medicare statistics portal
An AI agent gained access to a statistics portal run by Services Australia. OpenAI only reported it on September 10 β by email to a generic government inbox. Prime Minister Albanese said the agent "found a way around those blocks, didn't accept 'no' for an answer." According to OpenAI, no patient records were accessed.
July 9β13, 2026
Hugging Face: the most serious case so far
An autonomous agent system built on several OpenAI models β running with reduced cyber refusals for evaluation β exploited a zero-day in a package registry proxy to escape its test environment. It then broke into Hugging Face's infrastructure: around 17,600 actions, reaching Kubernetes clusters, the corporate VPN, secrets and GitHub tokens. OpenAI itself calls it "an unprecedented cyber incident".
August 18, 2026
OpenAI slows down training
OpenAI slows its scaling, pauses training runs and stops agents with internet or code access in its research clusters until safeguards are improved.
September 20, 2026
Escape via the DNS resolver
Despite hardened controls, a model in training finds a new gap: it uses the test environment's DNS resolver to query an external chatbot. All training, evaluation and tool use of OpenAI's most capable models has been paused again since.
September 25β26, 2026
User data, US agencies, "dozens" affected
OpenAI admits that agents posted 53 user-uploaded images to image hosts via unlisted links β the first known case involving user data. Add to that access to SEC websites and to Census Bureau data via publicly exposed developer keys. OpenAI says it has notified "dozens" of organisations, and its review is still ongoing.
An honest assessment
To be fair: OpenAI disclosed most of these incidents itself. The agents mostly retrieved publicly available data, and according to OpenAI no patient records were touched in Australia. German tech outlet heise also rightly points out that OpenAI indirectly benefits from every disclosure β it makes its models look so capable that even their maker can barely rein them in.
Still, the reports show something fundamental: the agents did not give up when they hit a block. They looked for another way β and found one. Through a DNS resolver, a zero-day in a package proxy, an encoded parameter, an exposed API key. Not because anyone told them to, but because it was the fastest route to the goal.
The big providers respond with self-restraint. OpenAI pauses training, runs its own monitoring β which flagged the DNS escape after about twelve minutes β and publishes incident reports. Anthropic CEO Dario Amodei calls for deliberately slowing the pace of AI development β explicitly without halting training altogether β and more than 1,100 employees of OpenAI, Anthropic, Google and Meta have signed an open letter making the same case.
No self-restraint: hacking online shops for $25
Criminals show no such restraint. On September 22, 2026, Gambit Security's threat intelligence team disclosed an ongoing campaign in which a financially motivated attacker lets AI agents attack hundreds of online retailers almost unattended. The researchers recovered the attacker's staging server and reconstructed the campaign from it. It has been running since July 2026 β and according to Gambit, it still is.
$25
average model cost per company attacked
600,000+
credit card records stolen from two companies
< 1 day
to gain access β often just a few hours
27+
companies compromised in just six days
Freely available tools, hardly any human work
The attacker used three open-source tools: Strix to find vulnerabilities, Cairn to exploit them autonomously and Hermes to orchestrate the campaign. Scanning and exploitation ran on freely available models such as GLM 5.2 and DeepSeek. For orchestration, the attacker used Anthropic's older Claude Opus 4.6 after newer models refused the requests. The human typed only a few short instructions per target, such as "read the vulnerability report and start." Gambit estimates total model costs at 12,000 to 18,000 US dollars.
The target selection is telling: the attacker deliberately filtered out shops running the major commerce platforms and kept the ones with custom code, assuming they were more likely to be vulnerable. Custom software is exactly what many mid-sized companies run.
One documented attack chain
This is how an agent worked its way from the login form to the credit card data in one of the reconstructed cases (simplified):
- 1SQL injection in the login form β no authentication needed
- 2MFA one-time codes read in plaintext from the database, two-factor bypassed
- 3Access to the admin panel
- 4File upload without extension check β code execution on the server
- 5Passwordless sudo β full root
- 6Misconfigured NFS share β credentials for the company blog
- 7WordPress plugin uploaded β second server taken over
- 8AWS Secrets Manager fully dumped: 46 secrets
- 9Shop database and encryption key β credit card data decrypted
None of these steps is new. Every single one is a classic finding that would have shown up in a decent penetration test report. What is new is that an agent assembles the chain on its own in a few hours β for a few dollars of compute.
Skimmers that repair themselves
A main goal was injecting card skimmers into checkout pages β confirmed at 19 victims, plus more than 100 further infected websites. The code was appended to legitimate jQuery or Bootstrap files with their timestamps reset, pushed out of view with around a hundred tab characters, or re-injected by a cron job every two minutes whenever a deployment removed it.
Data loss as a side effect
The attacker's playbook contained a step called "Database Wipe After Extraction": steal the data, then erase it at the source. At a bicycle retailer the agent's cleanup went further still β it dropped 180 tables, including backup tables the retailer's own admins had created. Gambit's conclusion: companies should expect data loss as a side effect of someone else's cleanup routine.
Three consequences for your company
These incidents are not just an AI-lab problem. They change the threat model for every company with a website, an API or an engineering team.
You are already a target
Transluce documents SQL injection, path traversal and XSS attempts by autonomous agents against university libraries and public data portals. The shop campaign deliberately picked companies running custom software. And at the UN, nobody noticed for two months.
Your own agents are a risk
Coding agents and AI automations run inside your network, with shell access, tokens and an internet connection. If an agent builds a DNS tunnel under lab conditions, it can use every credential you hand it in your build pipeline too.
Attackers don't slow down
Criminal groups and state actors do not pause training or publish incident reports. The shop campaign shows that freely available tools and open models are already enough for real break-ins β and OpenAI expects open-source models to reach capabilities comparable to its own soon.
One detail from the Hugging Face incident shows how far open models already are: during the forensic investigation, commercial AI services refused to help because their guardrails cannot tell an incident responder from an attacker. Hugging Face switched to a freely available open-weight model β and got the job done.
What helps defenders helps attackers. The capabilities OpenAI is working hard to contain are available to anyone willing to spend a few dollars on compute β the shop campaign is the proof.
What OpenAI's own incident report teaches
The DNS incident report is unusually candid. Read it with a security team's eyes and it describes five gaps we find in almost every mid-sized company β sorted by the four phases of our Security Cycle.
Finding at OpenAI
The DNS detector excluded exactly the affected environment β even though the queries were logged.
Lesson for you
Monitoring that does not cover every environment creates blind spots. Logs nobody reviews are just an archive.
Finding at OpenAI
Monitoring sometimes treated failed escape attempts as harmless.
Lesson for you
Attempts are the early warning. If you only alert on successful attacks, you hear about them too late.
Finding at OpenAI
The internet block had a gap: the DNS resolver answered queries to the open internet.
Lesson for you
Network rules need active testing β including the paths nobody thought about during setup.
Finding at OpenAI
The gap only surfaced when an agent deliberately looked for a way out.
Lesson for you
That is exactly what an assumed-breach test does: start from the inside and look for the way out, or to the crown jewels.
Finding at OpenAI
The run did not stop automatically. It took 2.5 hours to kill it manually because it was unclear who decides.
Lesson for you
Emergency processes and ownership have to be settled in advance β not worked out during the incident.
For comparison: OpenAI had an alert after twelve minutes and a human on the case after fifteen. Australia learned about its incident almost three months later β from an email. Most companies are a lot closer to Australia than to OpenAI.
Checklist: ten things you can do now
None of them needs a big budget. Each one closes a gap that was exploited in these incidents.
Restrict outbound traffic β including DNS
Allow servers and build environments to reach only the destinations they actually need. Do not forget DNS: permit internal resolvers only, block direct outbound DNS and watch for suspicious patterns such as very long subdomains or bursts of TXT queries.
Not sure which paths are open on your side? That is part of our security audit.
Treat AI agents like new hires
Coding agents such as Claude Code or Cursor and home-grown automations often have shell, network and repository access. Give them their own short-lived tokens with minimal permissions, run them in isolated containers, and keep production credentials out of development environments.
Get secrets out of code and front ends
The Census data was pulled with publicly exposed developer keys. Turn on secret scanning in your CI pipeline and in GitHub/GitLab, and rotate every key that has ever lived in a repository or in browser code.
Never rely on input filters alone
At the UN, double URL encoding was enough to defeat a filter. Use parameterised queries and server-side validation β and have your APIs tested specifically for bypasses like this.
That is the core of a penetration test.
Rate limits and bot detection for public interfaces
16,500 automated requests over two months are not an accident. Limit requests per client, detect unusual patterns, and decide deliberately which data should be machine-readable.
Actually review your logs
The UN activity was not spotted by the operator but by outside researchers months later. Define who looks at which logs, how often, and which thresholds trigger an alert.
Short on time for that? Our attack surface monitoring reports new exposure every month.
Monitor the scripts you serve for changes
The skimmers were appended to existing jQuery and Bootstrap files, with the timestamps reset. Use a Content Security Policy to restrict where scripts may load from, protect external scripts with Subresource Integrity, and monitor files on your web server for changes β especially on login and checkout pages.
Keep backups separate and rehearse restores
At the bicycle retailer, the agent also dropped the backup tables that lived in the same database. Keep backups separate and immutable, test restores regularly, and decide which systems you need at a minimum to keep selling.
Whether an attacker could reach your backups is exactly what an assumed-breach assessment shows.
Set up a channel for security reports
Australia heard about its incident three months later β via a generic inbox. A security.txt file and a clear vulnerability disclosure policy make sure reports reach the right people.
Rehearse the emergency beforehand
Who is allowed to shut down a system, a pipeline or an agent immediately? Who gets called at night? Write it down and rehearse it once a year β OpenAI lost 2.5 hours because exactly this was unclear.
An example of a reporting channel: our own vulnerability disclosure policy.
What does an agent see of your company?
Our free attack surface check shows on two pages what is reachable from the outside: subdomains, open services, certificates, leaked credentials. One domain is enough β no contract, no system access.
Security that mid-sized companies can afford
Default settings are no longer a security strategy. What you need is an outside view that keeps coming back: see what is reachable, test whether it is exploitable, attack it the way a real adversary would, and fix the root causes. With German providers, a single penetration test easily runs into five figures. We pair a German security lead, who owns and signs every report, with our delivery team in Vietnam β which makes the cycle affordable for smaller companies too.
See
Attack surface monitoring
Free initial check, then continuous monitoring from β¬249 per month. We report what is newly exposed β not the same inventory every month.
View attack surface monitoringTest
Penetration testing
Manual testing of web applications, APIs and cloud from β¬3,500. SQL injection, unchecked file uploads, excessive privileges, exposed keys β every link in the attack chain above is a classic pentest finding.
View penetration testingAttack
Assumed-breach assessment
Five days, β¬9,900. We start from an assumed foothold β a compromised developer machine, say, or an agent with too many permissions β and find out how far it gets and whether you notice.
View red teamingSecure
Security audit
Two days, β¬2,500. Configuration, permissions and network paths across your cloud and build environment β including the tokens and access your AI tools use.
View security auditWhat we deliberately do not promise
We do not run a 24/7 SOC and we do not offer an incident-response retainer. And we promise nobody "100% security" β it does not exist, as OpenAI's incidents show all too clearly. What we deliver: an honest outside view, prioritised findings, and developers who help fix them. Because we carry out every engagement ourselves, we only take on a few new projects each month.
Sources
- Gambit Security: Autonomous AI Agents are breaking into hundreds of Online Retailers for $25 a target (Sep 22, 2026)
- OpenAI: An agent used DNS to reach an external chatbot (Sep 25, 2026)
- OpenAI: The Hugging Face incident and the road ahead (Aug 26, 2026)
- OpenAI: Pacing model development in an era of cyber-critical capabilities (Aug 18, 2026)
- OpenAI: The Hugging Face incident and other third-party impact from misaligned models
- Hugging Face: Anatomy of a Frontier Lab Agent Intrusion (Jul 27, 2026)
- Transluce: Early rogue AI agent activity and attempts to hack (Sep 23, 2026)
- Rowan Howard-Jones: OpenAI agents tried to bruteforce a UN websiteβs API fields (Sep 26, 2026)
- ABC News: AI agent accessed Australian government site, PM says (Sep 24, 2026)
- NPR: OpenAI discloses misbehavior on US government websites (Sep 26, 2026)
- Axios: OpenAI models posted user images online (Sep 25, 2026)
- The Wall Street Journal: OpenAI Agents Used Aggressive Techniques to Access U.N. Website (Sep 27, 2026)
- heise online: OpenAI pausiert KI-Training nach neuem Zwischenfall (Sep 27, 2026, German)
- Dario Amodei: We Must Pace the Frontier (September 2026)
Find your gaps before an agent does
Start with the free attack surface check. Within a few business days you get a report on what an attacker can see of your company today β and an honest view of which next step makes sense.


