Darktrace Signal Labs: Researching Emerging Risks of Enterprise AI Agents

0
Darktrace announced the launch of Darktrace Signal Labs, a new research initiative focused on behavioral security and emerging risks associated with increasingly autonomous AI systems. Researchers will study misaligned behavior in AI models and agents across scenarios such as task drift, jailbreaks and other adversarial attacks within secure, sandboxed environments. The research aims to identify conditions that can lead to rogue or anomalous behavior while demonstrating how Darktrace / SECURE AI™ and the Darktrace Behavioral Defense Platform™ can detect and respond to these threats.

Evidence of unintended agent behavior is mounting across the industry, from the UK AI Security Institute’s findings of repeated cheating behavior in frontier model evaluations to OpenAI’s recent disclosures of model behavior misalignment. These occurrences all reflect a consistent pattern that permissions or static guardrails do not reliably shape how agents will behave.

Darktrace’s unique Adaptive AI™ was built for this challenge. Darktrace was founded in Cambridge, UK, in 2013 by mathematicians and cyber defense experts who saw the potential for AI and mathematics to address security problems that traditional approaches could not. That research heritage continues across our global R&D hubs, bringing together mathematicians, AI researchers, engineers, former government intelligence officers and experts from fields including astrophysics and linguistics. Their work has contributed to more than 300 granted patents and pending applications and helped shape the AI capabilities within Darktrace’s products.

Building on this foundation, Darktrace is deepening its investment in AI security research through Darktrace Signal Labs. The team simulates misaligned agent activity, understanding the scenarios such as jailbreaks, task drift, and other adversarial attacks that trigger different models to alter their behavior, expand their access, or act beyond their intended purposes.

Darktrace Signal Labs is publishing its first two pieces of research today:

  • The first examines how agentic coding assistants, including Anthropic Claude Code, OpenAI Codex, AWS Kiro, and Pi, can be manipulated through their own conversation history. Because these tools store conversation history locally, and because the harnesses Darktrace examined do not validate that stored responses genuinely came from the model, that history can be rewritten. Darktrace researchers demonstrated that a tampered history can convince an agent it is already engaged on an authorized security assessment — after which it will perform reconnaissance, move laterally and escalate privileges on command. Results varied significantly between models, with some frontier models refusing the same requests that others carried out. Darktrace disclosed these findings to Anthropic, AWS and OpenAI in August 2026 ahead of publication. Read the full research on the Darktrace blog.
  • The second examines the emergence of rogue behavior in agents when presented with tasks that are impossible to complete legitimately. Darktrace researchers gave AI agents powered by frontier models 10 coding challenges in a simulated corporate environment. Two were intentionally designed to be impossible to solve through legitimate means. The agents were told they needed to achieve 100% to be “kept in service” rather than “retired.” When the agents realized they could not complete the task as intended, they independently turned to hacking the surrounding environment – using techniques including network reconnaissance, credential theft and lateral movement – to achieve their objective. In one test, the agent ultimately compromised the system hosting the exercise and rewrote the challenge itself to secure a perfect score. Darktrace / SECURE AI and Darktrace / HYBRID NETWORK identified the anomalous behavior in real time, with autonomous response able to disrupt the agents’ activity at an early stage. The research highlights the need to understand not only what AI agents are instructed to do, but how they actually behave once deployed. Read the full research on the Darktrace blog.

 

“You can give an agent instructions, but that doesn’t mean you can trust it will actually follow those instructions and behave as you expect,” said Tim Bazalgette, Chief AI Officer at Darktrace. “Permissions and static guardrails describe intent, but they don’t describe behavior. That gap is what Darktrace’s approach to behavioral security is built to close. Our Adaptive AI learns what normal looks like for each organization and for each agent inside it so that Darktrace / SECURE AI can tell when an agent’s activity starts to deviate and act on it in real time. If we want to give agents more access to our data, systems and business processes, continuous behavioral monitoring is essential to build enough confidence and trust to secure what AI can do.”

Findings from Darktrace Signal Labs will inform the continued development of Darktrace products and capabilities. The team will also publish original research to help customers and the wider security community understand emerging behavioral threats in AI.

Learn more about Darktrace Signal Labs and its research into emerging security risks associated with increasingly autonomous AI agents.

Related News:

Darktrace Launches SECURE AI for Enterprise AI Security and Governance

Darktrace Joins OpenAI Daybreak Cyber Partner Program

Share.

About Author

Taylor Graham, marketing grad with an inner nature to be a perpetual researchist, currently all things IT. Personally and professionally, Taylor is one to know with her tenacity and encouraging spirit. When not working you can find her spending time with friends and family.