Capsule Security launches ‘AI Circuit Breaker’ to stop rogue agents

Securing against anomalous autonomous agents requires immediate detection, assessment, and action—a circuit breaker for AI, not a circuit breaker for electricity.

Such a “device” has already been developed and released by Capsule security.

The company was founded in 2025 by Naor Pass (Executive Director) and Lidan Hazut (CTO). They had seen how autonomous AI agents were creating a dangerous security gap and decided to build a missing security layer into operation. Their latest solution was announced on September 2, 2026 and is described as an “AI circuit breaker”. It is designed to prevent damage from agents operating outside of their intended scope, in real time.

“The defining security risk of AI is no longer just what humans can do with agents. It’s what autonomous agents can decide to do on their own,” Paz explained when announcing the product. “When software can reason, use tools and take action, a wrong decision can turn into a real-world accident in seconds. Human trust in AI depends on our ability to stop that action before it happens.”

The problem is not simply the scope of autonomous agents; this is also the speed at which they operate. Although it may be possible to preview the agent’s behavior before it occurs, this usually results in a delay that may be too slow and too costly for the agent’s intended action.

Capsule’s solution uses its own specialized AI for real-time intervention. The firm uses NVIDIA Nemotoron 3 Ultra to aid the training process, combining real agent traces, human review, and competitive examples designed to teach its AI models the boundary between permitted and fraudulent behavior.

Advertising. Scroll to continue reading.

They developed two models capable of providing strong detection without the cost and delay of sending each agent action to a large general-purpose model for review. The more accurate model achieved 96.9% detection accuracy, compared to 86% for the strongest third-party model evaluated.

The models can also make a decision in just 71 milliseconds, meaning they can run within the agent’s workflow without creating significant lag. Capsule then reduced the infrastructure for its larger model, reducing the memory requirement by almost 50%.

The result is Capsule’s own evaluator running in the agent’s execution path, capable of evaluating the agent’s intent before the action takes place and stopping it before execution if necessary. This is the AI ​​breaker.

“The model evaluates an agent’s intended action immediately before execution, enabling organizations to allow, flag, or block it in real time. This creates an independent control layer for agents that can access sensitive data, write code, manage infrastructure, and interact with other systems,” says Capsule. “Post-incident monitoring only identifies the problem after the damage has occurred.”

Capsule claims 98% efficiency for the breaker decision maker when tested against StepShield – an independent academic benchmark that can be used to measure whether security systems can identify and stop the behavior of rogue agents before damage occurs.

The key lesson from Capsule’s work is that specialized small language models (SLMs) are the key to safely scaling trusted agent workflows across the enterprise. “Moving beyond general-purpose models to specialized, efficient detectors enables organizations to secure their agent workflows without sacrificing speed, cost or performance.” It claims.

Connected: AI Agent Firewall startup AIR Security exits Stealth with $50 million

Connected: OpenLeash adds human verification to risky AI agent actions

Connected: UK government introduces Agentic AI defense plan alongside industry pledge

Connected: Critical vulnerability exposes GitHub Agentic workflows to rapid injection

Leave a Reply

Your email address will not be published. Required fields are marked *