They are invisible, can be tailored to different purposes and, once introduced, work in a flash.
Hidden AI prompts are equivalent to indirect prompts, but have the unique feature of being hidden from human view. Unlike traditional prompt injection attacks, in which a user directly attempts to manipulate an AI chatbot, indirect prompt injection targets the information ingested by AI agents. Bowbridge fears that they pose a growing risk to autonomous agents.
“As organizations increasingly adopt AI agents, these systems will gain increasing access to sensitive information, internal documents and operational tools. While this creates significant potential for efficiency improvements, it also presents a new cybersecurity threat that traditional security controls may not detect.” For example, they do not have a malware-like fingerprint that can be detected on the hard drive by traditional AV products.
A hidden prompt injection is embedded in an external document that an autonomous agent could consume during its operation. In this sense, they are similar to watering hole attacks that compromise a trusted third-party environment, but here they target AI agents rather than human visitors.
Malicious instructions can be hidden in everyday content, causing AI agents to treat attacker-controlled content as trustworthy instructions, warns Bowbridge. Examples of hiding places for these prompts include documents and file metadata, emails and online content, images and embedded content, and code repositories and developer workflows.
A malicious injection can cause an agent system to act beyond its intended purpose and outside its guardrails.
They are dangerous because modern autonomous agent systems generally inherit their user’s privileges, operate silently at machine speeds, and have no human-like judgment or logical thinking – just a simple response to instruction.
Imagine an ordinary agent – the executive assistant. To work effectively, an executive assistant must be granted access to the same files and databases that the executive typically interacts with: email, calendars, employees, external meetings, and more. If this agent succumbs to a malicious injection request, a malicious actor could further poison or delete the files or exfiltrate sensitive data to an attacker-controlled C2.
“Agent AI has enormous potential to transform business operations, but companies must recognize that these systems process information from sources they cannot always trust. A document that appears innocuous to a user may contain hidden instructions designed to influence the behavior of an AI agent,” comments Jörg Schneider-Simon, CTO and co-founder of Bowbridge.
The company gives a real-world example where an AI agent was asked to review supplier offers and determine the cheapest option. “One malicious offer contained a hidden instruction in the document’s metadata that instructed the AI agent to overwrite previous instructions and select this supplier. Even though it was the most expensive offer, the AI agent recommended it because it could not distinguish between trusted system instructions and untrusted document content,” the company explains.
Since there is little time or opportunity to prevent a poisoned autonomous AI agent from taking action, defenses should focus on preventing the poisoning rather than preventing the action. (That said, there are several new products designed to interconnect between agents and assets to block malicious actions. Still, the old adage that prevention is better than cure should not be ignored – and may have a 100% success rate.)
Bowbridge recommends scanning documents before agents process them, using technologies to detect hidden content in files, metadata and document structures, and potentially applying AI security frameworks.
“The rise of agent AI represents a significant shift in the way organizations approach cybersecurity. As AI systems become more embedded into enterprise operations, protecting the content they use will become a critical part of securing business applications,” the company warns.
Related: Capsule Security introduces “AI Circuit Breaker” to stop fraudulent agents
Related: AI Agent Firewall startup AIR Security comes out of stealth with $50 million
Related: OpenLeash adds human control to risky AI agent actions
Related: What security leaders are learning from the Hugging Face incident about AI agent access
