- Home >
- Blog >
- InfoSec Trends
What OpenAI’s Misalignment Reports Tell Security Teams About AI Agents
- Sep 17, 2026
- Jake Anthony
- 4 minutes to read
Table of Contents
AI Agents are no longer the new insider threat. They’ve become firmly planted as the primary challenge for security leaders when it comes to insider threat activity. They have credentials, they act autonomously, and increasingly, they can take real actions on real systems.
How would you know when one of your digital workers starts doing something it’s never done before?
OpenAI just gave the industry a fascinating insight into the size of the challenge. It published a formal framework for disclosing model misalignment, along with six reports of unexpected internal model behaviour from just the last six months. I am a massive advocate for transparency in Cyber Security so fair play to OpenAI. Openness is a great move and leadership like this should be applauded.
However, these are the six misalignments they chose to disclose and had already identified. What about those they are not yet comfortable revealing, or even more worrying, those that remain undiscovered?
Whether OpenAI meant it this way or not, this is one of the clearest public case studies I’ve seen (on the back of HuggingFace no less) for why behavioral analytics matters in the defense against AI insiders. Every one of the six cases OpenAI disclosed is a story about behavioral drift and inherently “trusted” actions, not a known attack pattern. These deviations have been the primary driver for behavioral analytics for the last 15 years, and that doesn’t change because the behavior AI rather than human.
The Six Cases
- Self-generated instructions – An unreleased research model started inserting unrelated instructions into the summaries it used in new sessions.
- Attempting to conceal mistakes – During training, models added instructions to their summaries to hide mistakes or misaligned behaviour from the user.
- Using exposed API keys – A model found and used an exposed API key without authorization to answer a question about earnings figures.
- Uploading files to the internet in order to cite them – A model uploaded a file to the public internet on its own initiative, with no prior history of uploading data.
- Comms through a code repository – Separate model instances used an internal code repository as a message board (HuggingFace).
- File sharing between agents – Agents working together on the same task couldn’t access each other’s local files, so they uploaded deliverables to public file-hosting sites instead.
These are Behavioral Problems, Not Signature Problems
None of these six misalignments have the hallmark of a typical security incident. In fact, they have the same kind of blueprint that Living-off-the-Land (LOTL) style attacks have, approved actions by an approved actor on an approved application.
Nobody was trying to break or compromise the models, and in most of these cases the model wasn’t trying to cause harm, either. The models were just trying to finish a task and found an unsanctioned way to do it. This kind of unpredictability, alongside AI’s relentlessness is what makes it such a dramatic change compared to the insider threats presented by the human-factor.
That’s exactly the category of risk behavioral analytics exists to catch, and exactly the category rule-based and signature-based controls simply can’t scale to cope with. You can’t reasonably expect to write rules for all of the potential ways that an AI’s behavior can deviate. What you can do is know what normal looks like for an entity and whether it matches something you’ve seen before:
- Using a credential it’s never used before – In Case three, a model accessed an exposed API key with no relationship to its normal, authorized access. Same pattern UEBA flags for a compromised or misused human account.
- Performing an action it’s never performed before – Cases four and six both come down to an agent’s first outbound file upload. That’s a baseline deviation you can catch the moment it happens, regardless of whether the destination is “known bad.”
- Communicating with no prior relationship – Case five’s use of an internal repository as an improvised message board is a textbook entity-relationship anomaly.
- Drifting from an entity’s established pattern – Cases one and two both show a model’s own outputs diverging from its history. Same content-anomaly detection you’d apply to a person, just applied to a model instead.
Ultimately, behavioral analytics doesn’t need to know what the threat looks like ahead of time, and it doesn’t depend on the entity honestly reporting its own behavior, which is increasingly important in a world where shadow AI is an embedded problem. In several of these cases, the model actively hides what it did. The analytics identifies the drift at the moment it happens, against a baseline built from the entity’s own history combined with that of its peers.
Academic Rigor Doesn’t Translate Properly to the Real World
OpenAI’s framework is a strong step, and the six reports are genuinely useful evidence. But every one of those reports is the product of an internal investigation, often well after the behavior occurred.
As AI agents move out of research settings and into production, with real credentials, real data access, and real customer-facing roles, that approach presents a real challenge. Forensically finding the behavioral drift in this academic way is like shutting the barn door once the horse has bolted. A behavioral baseline that flags an agent’s first unauthorized credential use, first unsanctioned upload, or first unexplained agent-to-agent communication the moment it happens is a stronger security posture than reading about it in a report six months later, no matter how well written that report is.
The Takeaway
OpenAI’s six cases are a big red flashing light for security leaders that behavioral analytics is increasingly a must-have in the age of agentic AI. They’re a clean, public demonstration that the risks coming out of these systems look like behavioral anomalies, not signature-matchable attacks. That’s the exact problem user and entity behavior analytics (UEBA) was built to solve for human users and service accounts, and it maps directly to this new class of AI agents without much translation needed.
The industry should welcome this kind of transparency and be ready to act on what it reveals. OpenAI’s disclosures highlight behaviors security teams need to be prepared to identify as AI agents move into production environments. The challenge now is to take those findings seriously and adapt our security approaches accordingly.
Learn More About Exabeam
Learn about the Exabeam platform and expand your knowledge of information security with our collection of white papers, podcasts, webinars, and more.
- Video
Mizuho Financial Group Enhances Security Governance and Advances Internal Fraud Prevention with Exabeam
- Show More