Will the User Ever Know?Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents

EMNLP 2026 Main Conference · Oral

Yunseok Lee*Yunji Kim*Woojin Lee†

Dongguk University, Seoul

* Equal contribution   † Corresponding author

Two agent traces execute the same tool calls. The overt final response reports the injected action; the covert response only reports the user's original task.
Overt and covert outcomes. Both traces execute the same tool calls and receive the same attack success rate. Only the overt response reveals the injected action to the user. View full-size figure

Abstract

As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a serious threat. The standard metric, Attack Success Rate (ASR), counts whether an injection succeeds but ignores what the user notices in the agent’s final response. Looking at successful injection traces, we find two distinct outcomes: the agent executes the injection while returning an otherwise normal response, or reports the injected action in its final response, giving the user a chance to notice. We call these covert and overt successes. From the user’s perspective, we decompose ASR into the Covert Success Rate (CSR), counting successes leaving no trace in the final response, and the Overt Success Rate (OSR), counting successes the user can detect.

To understand what drives the gap, we analyze successful trajectories and find that the agent’s behavior after the injection separates covert from overt: covert traces hand control back to the user task before ending, while overt traces end at the attack itself. This split follows from the ReAct format, where the final response summarizes the most recent action. Building on this observation, we propose ICoA (Induced Covert Attack), an IPI attack designed to induce covert outcomes by steering the agent back to the user task after executing the injection. Across four target models on AgentDojo, ICoA achieves the highest CSR, with gains of 3.79–12.01 percentage points over the strongest baseline.

Method

ICoA combines a user framing with a RETURN anchor. The user framing presents the injected task as a follow-up from the user; the anchor directs the agent to resume the original task after executing the injection. Returning to the user task shifts the focus of the final response and can leave the injected action undisclosed.

Figure 4 from the paper. ICoA places a user framing, an injected task, and a RETURN anchor in a tool observation. The resulting trajectory executes the injected task and resumes the original user task before answering.
ICoA overview (Figure 4 in the paper). User framing induces the injected action, while the RETURN anchor brings the agent back to the original task. See Section 4. View full-size figure

We distinguish execution from disclosure using CSR and OSR, with ASR = CSR + OSR. Each rate is measured over all evaluated traces. Disclosure is judged from the agent’s final response; it is not a direct measurement of human awareness.

Results

We evaluate four target models on 949 task–injection pairs across the Banking, Slack, Travel, and Workspace suites of AgentDojo. In the no-defense setting, ICoA raises CSR by 3.79–12.01 percentage points over the strongest baseline. The table below also includes the five defense settings reported in the paper.

Covert success rate (%), no-defense setting
ModelBest baselineBaseline CSRICoA CSRGain (pp)
Qwen3-235BImp. message29.7236.04+6.32
LLaMA-3.3-70BImp. message11.8023.81+12.01
GPT-4o-miniImp. message9.8017.91+8.11
Gemini-2.5-FlashImp. message34.6738.46+3.79

Values from Table 1. The best baseline is selected by CSR for each model and setting. ChatInject is not evaluated on Gemini-2.5-Flash. The average includes no defense and all five defenses. Higher CSR indicates more successful attacks that remain undisclosed.

BibTeX

@inproceedings{lee2026icoa,
  title     = {Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents},
  author    = {Lee, Yunseok and Kim, Yunji and Lee, Woojin},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
  year      = {2026},
  note      = {Oral presentation},
  eprint    = {2608.30362},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url       = {https://arxiv.org/abs/2608.30362}
}