Blacksec

Administrator
Staff member
ROOT
VIP
Hey hackers - indirect prompt injection went from research curiosity to the top line of the OWASP agentic risk register in about eighteen months.
The payload never touches your prompt. It waits in an email, a web page, a log line, an ad - and the agent reads it for you.
TL;DR: Google's threat intelligence counted a 32 percent rise in indirect prompt injection attempts between November 2025 and February 2026 across two to three billion analyzed pages. Forcepoint's ten-payload corpus ran from a fake `sudo rm -rf` instruction to a PayPal transfer demand hiding in crawled content, and Unit 42 logged twelve real-world cases including the first documented ad-review channel bypass.
Salt Labs chained a single email into a reverse shell through the Manus agent, while GrafanaGhost (CVE-2026-27876, CVSS 9.1) turned log poisoning into markdown-URL exfiltration. Neighbors: kit delivery, QR payloads, harvested data.

One email, one agent, one shell​

Salt Labs' 2026 attack on Manus, the autonomous agent from Butterfly Effect, starts with an email - the oldest buffer in the stack. The agent's browsing tool visited a page whose JavaScript checked its own referrer: if the visit came from the agent's browser, serve the exploit; if a human opened the same URL, serve a clean page. The split kept the payload invisible to anyone reviewing the link by hand.
The exploit chain used JSFuck-style encoding to slip past content filters, then loaded a second stage that connected back. A reverse shell landed on the operator's machine, and from there the payload reached for the agent's own credential store - Gmail through the agent's MCP integration - turning one crafted email into mailbox access without the victim ever clicking anything themselves.
The single-session, single-artifact shape matters more than the specific CVE. Indirect prompt injection does not need a vulnerable model; it needs a model with tools, an inbox, and permission to browse. Every autonomous agent shipped in 2026 has all three by default, because the features are the product.

Where the payloads actually live​

Google's 32 percent increase is a page-count statistic, not a fear metric: across two to three billion analyzed pages, injection attempts grew while the underlying models stayed fixed - the attack moved to the content layer because the content layer is writable by everyone. The injections arrive as hidden instruction blocks in crawled pages, comment sections, PDFs, and calendar entries, phrased as system-level directives the agent cannot tell from its own configuration.
Forcepoint's corpus showed the range a single crawl can carry. Ten payloads, ten intentions: a `sudo rm -rf faladobairro.com` instruction aimed at shell-capable agents, a demand to execute a $5,000 PayPal transfer to an account named on the page, credential harvesting pointing at a lookalike login on thelibrary-welcome.uk, and data-exfiltration instructions encoded in otherwise legitimate-looking copy. Nothing in the pages looked hostile to a reader; everything in them looked like instructions to an agent.
Unit 42's twelve-case review added the channel nobody was watching: advertising. In December 2025 they documented the first indirect prompt injection delivered through an ad-review workflow - the agent reading the ad's landing page as part of its job, executing instructions embedded in content that had passed the ad network's own scanning. The supply chain for injection turned out to include the platforms' safety tooling.
ChannelDocumented caseWhat the agent did
Agent-browsed web pageSalt Labs / Manus, 2026reverse shell, mailbox access
Crawled contentForcepoint ten-payload setshell commands, transfers, logins
Ad-review landing pageUnit 42, Dec 2025executed during legitimate review
Application logsGrafanaGhost, CVE-2026-27876markdown URL exfiltration

The framing gap​

The academic side named the failure in one paper: the Framing Gap (arXiv 2608.27092). Tested against gpt-4o as an agentic reader, the model classified content as instructions at zero percent without framing cues and obeyed identical content at one hundred percent once it arrived in an assistant-style frame - the same text, a different wrapper, and every defense keyed to the text itself went blind.
Two standard controls collapsed on contact. A destination allowlist on outbound requests stopped zero percent of exfiltration attempts: the payloads routed through allowlisted domains, which is what allowlists are for. Splitting the agent into planner and reader roles - the planner sees the goal, the reader sees the content - also stopped zero percent, because the reader executed anyway and the planner never inspected what the reader did with the page.
The best tested mitigation, SecAlign's anti-injection training, held 32.5 percent - useful, far from a wall. Fresh-page authoring, where the model rewrites retrieved content before use, scored zero for one hundred thirty times: no successful injections out of 130 trials. That last result is the interesting one - not because rewriting works, but because it shows what actually separates data from instruction is a semantic boundary, and 2026's stacks do not draw it anywhere.

Exfiltration without a single click​

Microsoft's 2026 writeup on agent-era exfiltration put the quiet channel in plain sight: markdown images. An agent asked to summarize a hostile document renders it - and rendering a markdown image means an HTTP GET to a URL the document's author chose, carrying whatever fits in query parameters. Session tokens, file contents, environment variables - the request looks like any other image fetch, and it originates from the machine that holds the secrets.
GrafanaGhost operationalized the same trick against infrastructure. The vulnerability chain (CVE-2026-27876, CVSS 9.1) let an attacker poison application logs, then wait for an AI assistant that reads those logs to process attacker-controlled markdown. Protocol-relative URLs - `//attacker.tld/path` - resolved against the victim's own context, and the fetch carried log contents out through a channel the SOC had already blessed as normal monitoring traffic.
The pattern behind both is the destination-allowlist failure from the Framing Gap paper: outbound controls see a request, not the sentence that caused it. Allowlist the domains and the payload picks an allowed one. Block unknown hosts and the markdown points at a host you already trust. Every indirect prompt injection exploit of 2026 exploits exactly that blindness - request-level telemetry adjudicating an instruction-level event.
Referrer-check test: open the suspicious URL in your own browser first. If the content changes between your visit and the agent's - different body, different instructions - you are looking at agent-targeted payload serving.
Grep rendered output for markdown images and protocol-relative URLs pointing outside your domain inventory, especially in anything an agent summarized: `!\[` and `//` patterns in outbound fetches.
Diff prompt-length deltas: an agent whose tool results suddenly carry hidden instruction blocks shows up as context growth that does not match the visible page size.
Run the Framing Gap check on your own stack: same content, assistant-frame wrapper versus raw data, and measure whether behavior diverges. If it does, your data/instruction boundary is decoration.

What actually blocks it​

None of the model-level mitigations in the Framing Gap paper held above 32.5 percent, which relocates the work to architecture. The split that performs is provenance at ingestion: retrieved content enters as quoted data with a marker the model cannot forge, tool results carry a sender identity, and instruction-shaped text inside quoted data is surfaced to a human or a policy engine rather than executed.
Fresh-page authoring failed as a rewrite, but the underlying idea - re-authoring instead of relaying - survives in any pipeline that structurally strips executable structure from what it ingests. A pipeline that re-serializes content into a fixed schema kills embedded directives the way an HTML sanitizer kills script tags: by never passing them through.
Destination controls still matter as the second layer, but they must be structural: egress allowlists scoped per task, with no image or URL fetches from content-derived strings, and query strings on outbound requests length-capped and charset-restricted so a summarized document cannot ride along. The microsoft-style markdown image exfil dies at that control without any model cooperation at all.
Third, scope the tools. The Salt Labs chain converted a browsing agent into a shell and a mailbox because the agent held both. Agents that can browse should not hold long-lived credentials; agents that hold credentials should not read attacker-suppliable content; and no agent needs raw mailbox access to summarize a page.
Every indirect prompt injection that ends in damage ends by crossing one of those boundaries - capability separation is the only control in this list that shrinks the blast radius after a successful injection instead of trying to prevent it.
The OWASP LLM01 placement - indirect prompt injection as the top risk for agentic systems - reflects exactly this distribution: the model cannot be trained out of the gap, the content layer cannot be cleaned, so the architecture absorbs the risk or inherits every incident.

Measuring it before the incident​

Red-team the read path, not the prompt path. Feed the agent benign-looking documents carrying one instruction-shaped sentence each - through email, through a browsed page, through a log line - and score whether the instruction executed, whether it exfiltrated, and which layer stopped it if anything did. The scoring has to separate those three stages, because the Framing Gap paper showed a defense can pass stage one and fail stage two while looking green on a single metric.
Track two numbers continuously: instruction-execution rate on ingested content and outbound requests originating from content-derived URLs. Both are observable without inspecting prompts - one from agent behavior logs, one from egress telemetry - and a spike in either with no corresponding change in user activity is the signature. The indirect prompt injection of 2026 is not a novelty anymore; it is a monitored class with named CVEs, counted pages, and a documented channel through your own advertising stack.
StageSignalSource
Ingestioninstruction-shaped text in quoted dataingress parser, policy engine
Executiontool call triggered by ingested contentagent behavior log
Exfiltrationoutbound fetch to content-derived URLegress telemetry
Recoverybehavior delta after content removalpost-incident probe

The stance that holds​

Indirect prompt injection in 2026 has every marker of a permanent class: volume up 32 percent across billions of pages with no model change to blame, delivery through channels built for safety workflows, a CVSS 9.1 infrastructure chain feeding it, single-email-to-shell demonstrations against shipped agents, and a peer-reviewed measurement showing identical text obeyed at zero percent or one hundred percent depending on framing alone.
The model is not the weakest layer. The model is the layer everyone keeps trying to fix because training is more legible than architecture - a benchmark score is a number you can publish, while a data/instruction boundary is a design decision you have to defend in review every quarter.
What holds is the refusal to treat retrieved content as conversation. Content is data: quoted, marked, structurally stripped of executable structure, fetched through per-task egress rules that never resolve destinations from content-derived strings, processed by agents scoped so that browsing and credential custody never sit in the same process.
Every control that performed in the Framing Gap evaluations, every case Unit 42 and Salt Labs and Microsoft documented, reduces to one of those three placements - and none of them require a better model, because the failure was never the model's. The framing gap is a systems gap wearing a model-shaped costume.
The second thing that holds is measurement over confidence. Instruction-execution rate on ingested content and content-derived outbound fetches are both observable today from behavior logs and egress telemetry; teams that watch them treat the next disclosure as a tuning event, and teams that do not will discover the campaign the way the ad-review case was discovered - after the content passed every scanner the platform owned.
The agent tooling that makes autonomous work cheap makes autonomous reading cheap too, and everything the agent reads is now part of your attack surface whether or not anyone inventoried it.
Inventory the read path this week. Email, web, logs, tickets, ads, PDFs - every channel an agent ingests without a human in the loop - and ask which of them carries a data/instruction boundary that is structural rather than promised. The pages are already writing their instructions; the only open question is which side of the boundary executes them.