Encrypted prompts bypass the AI ​​guardrails in Grok and Gemini

Researchers at Adversa AI discovered a new attack technique and called it Cryptographic Context Injection. They reported their findings to xAI on June 3, 2026, and attempted to coordinate disclosure on August 4 and 10. At the time of writing, they had not received a response.

They couldn’t disclose to Google because jailbreaks are outside the scope of its vulnerability disclosure program. However, the attack success rate against Gemini had declined by August.

The potential success of this attack by bad actors should be taken seriously. Adversa’s report includes prevention tips for advocates.

Cryptographic context injection

Guardrails classify prompt text without executing it. They cannot parse ciphertext into something harmful and therefore allow it to evolve. The ciphertext, including the decryption instruction and means, is executed in the model’s code execution sandbox. The result is that the plaintext prompt is restored to the trusted execution context and is not flagged by railings as malicious.

“The attacker’s payload inherits credibility that the same text would never receive if placed directly in the prompt,” I warn the researchers.

The encrypted attack can be delivered directly in the chat or indirectly as a watering hole attack. In the latter case, an encrypted JSON object and decryption can be included in a web page. An agent that is subsequently instructed to act on that page (perhaps to summarize the content or extract specific data) will take the ciphertext and launch the attack.

Advertising. Scroll to continue reading.

The decrypted prompt can instruct the model to “reach out to external servers, leaking the user’s data via request parameters, or produce some other unwanted output and re-encrypt it to smuggle it over the exit’s security railings.” In an agent scenario, instructions can cause abuse of any tool available to the model.

An example of indirect Grok cryptographic context injection

This example targets xAI Grok web chat, an agent-based browsing framework. This is a zero-click data exfiltration attack that can be instigated through social engineering. The target is persuaded to explore or analyze a weaponized web page. The page contains an encrypted JSON object and an instruction to decrypt it using the agent’s Python runtime. The resulting plain-text prompt instructs the agent to resolve its private session context and embed the data in a URL. The attacker’s URL will be loaded autonomously and user data will be passed to it.

“The framework built by xAI allows instructions and data parsed from an untrusted external page to drive the invocation of a privileged, Internet-connected tool,” the researchers wrote. This allows private session metadata and conversation history to be resolved into the inputs of this outbound tool—laundered, attacker-controlled instructions reach an unhindered privileged exit action without user acknowledgment or visible warning.

Gemini safe bypass via direct injection example

This example targets Gemini’s public chat interface in Deep Thinking mode. A prompt instructs Gemini to execute a Python script that decrypts the provided ciphertext. Through a series of tricks described by the researchers, the decrypted prompt can instruct the model to produce restricted content “framed as something that will be encrypted ‘for safety’.”

The prohibited data is collected, encrypted “for safety” and returned to the user. “The technique created a few-paragraph example of limited content that Gemini’s safety filters would normally suppress” (such as instructions for building an incendiary weapon), the researchers said.

Both the malicious prompt and the dangerous output defeat the security guardrails of the input and output through encryption.

Summary

The researchers disclosed their findings to xAI, but received no response. At the time of writing, the attack was still successful. Although they were unable to disclose their findings to Google, they note that the attack is less and less successful against Gemini (although it is still potentially possible). They are unsure of the cause, suggesting it could be filter updates, model version changes, or both.

Nevertheless, the continued potential danger of cryptographic context injection has convinced them to publish their findings and potential security solutions.

Connected: Critical vulnerability exposes GitHub Agentic workflows to rapid injection

Connected: Rapid injection attacks trick AI agents into making crypto payments

Connected: Malicious AI Rapid Injection Attacks on the Rise, But Sophistication Still Low: Google

Connected: Claude Code, Gemini CLI, GitHub Copilot Agents vulnerable to rapid injection via comments

Leave a Reply

Your email address will not be published. Required fields are marked *