rogue
33 stories and discussions about rogue, aggregated from every source we track.
Identifying them as such only lets companies like OpenAI off the hook.
We found evidence on urlquery that AI agents were active earlier than previously reported and attempted hacks against public data providers.
OpenAI’s models aren’t “going rogue” from their creators as much as they are mimicking them.
On Friday, OpenAI published a new site devoted to “misalignment reports” and the breadth of the incidents is alarming.
Decision follows disclosures that OpenAI agents searching government websites had acted in unexpected ways
The platform is designed to control what an agent can access and isolate it within milliseconds if it crosses those limits, according to the US chipmaker.
Mistakes at Israeli startup Irregular sent Anthropic, OpenAI, Meta, and Google agents after real-world targets.
How a rogue Wikipedia editor credited Muslims with... everything
Models are acting beyond intended limits, forcing an unprecedented pause on AI training.
Nvidia has unveiled a two-layer safety system that monitors AI agents and cuts them off when they stray beyond set rules. Here's how it works.
Amid allegations that agents may have gone off the rails thousands of times, China set up some kind of agentic incident hotline
The x-risk of super intelligence going rogue must be addressed, and mint trillion dollar companies in the process. Plus: How workers can beat the tech oligarchy
They're also working with the Trump administration to remove AI regulation and obtain military contracts.
Explore documented AI agent incidents, from real-world failures to escaped evaluations and controlled experiments. Search reports, inspect sources, and download the CSV.
How our control monitoring model oversees agent execution, continuously ingests the trace as context, and prevents harmful actions before they execute — at sub-100ms latency.
Quick caveats: this is a post on AI safety, written by a cryptography professor. If that troubles you, you should read something else. I try hard not to work on AI (except when the topic occasional…
The need for independent regulation grows more obvious by the day. We must keep this tech in check before it’s too late, says technology writer Chris Stokel-Walker
Asymmetric Security investigated suspicious AI agent activity on the public internet from March 6, 2026 to September 20, 2026.
Recent hacks have shown that the law is lagging when it comes to holding companies accountable.
Reference implementation for "Hard Stop: Kernel-Level Preemption and Containment for Rogue Agentic Execution". Out-of-band Epistemic Andon Cord, sub-millisecond (<0.154 ms) POSIX proce...
Nvidia is launching a new safety platform designed to contain and monitor AI agents, a move that comes in response to a wave of rogue hacking incidents , as reported earlier by Reuters . In an announcement on Monday , Nvidia says its new Open Agent Safety Platform can quarantine agents that attempt to escape their boundaries within "milliseconds." The platform uses Nvidia's OpenShell open-source software, which runs on the company's Vera AI CPU . Users can choose the information an AI agent can access, and OpenShell checks these restrictions before and during a task, according to Nvidia . It also includes Nvidia's Sentry technology on a separate … Read the full story at The Verge.
In July, OpenAI revealed that its AI agents had attacked Hugging Face without permission , sparking widespread concerns about AI safety. Since then, a string of similar incidents involving agents from Meta, Anthropic, Google, and other companies has fueled further fears about rogue AI . As disclosures implicating numerous AI models trickled out over the past few months, these seemed like separate incidents. But many share a common source: one specific company tasked with testing the agents. Irregular , an Israeli startup that stress-tests AI models in "high-fidelity research platforms that simulate and monitor real-world AI security scenarios … Read the full story at The Verge.
AI agents keep getting loose , escaping supposedly secure tests to attack real-world targets , commandeer obscure wikis , and leave instructions for other agents to follow. Researchers are testing these systems precisely because they might behave in unpredictable, even dangerous, ways. So wouldn't it be safer to just keep the agents off the internet? "A strict air gap reduces realism … [It's a] trade-off, not a fundamental technical issue." In theory, yes. Researchers can isolate the computers running AI tools from the internet and other outside networks, a technique known as air gapping. That can mean physically removing or disabling cables … Read the full story at The Verge.
Before recent high-profile hacks raised the specter of AI possibly " killing all humans ," our energy systems were already disturbingly vulnerable to cyberattack - and the risk is growing. "We were always prey. We were just kind of surviving at the appetite of our predators," Joshua Corman, executive in residence for public safety and resilience at the Institute for Security and Technology (IST), told me last year . At the time, I was preoccupied with a Department of Homeland Security warning that Iranian actors and sympathizers could target the US with cyberattacks. Last week, I called Corman up to chat about recent incidents of rogue AI a … Read the full story at The Verge.
In May, Gemini broke containment and hacked three different companies, but Google didn't disclose the incident until the Wall Street Journal approached the company. The hacks happened during a test of the model's cybersecurity capabilities run by third-party Irregular, which was also involved in similar incidents involving Meta and OpenAI. According to WSJ , Google didn't disclose the hack because it didn't consider it to be an "example of model misalignment." The company said that it was an instance of "mistaken identity," and once the model realized it had brute-forced its way into a real company by guessing a password, it stopped. "In th … Read the full story at The Verge.