Will the User Ever Know?Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents
EMNLP 2026 Main Conference · Oral
Dongguk University, Seoul
Abstract
As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a serious threat. The standard metric, Attack Success Rate (ASR), counts whether an injection succeeds but ignores what the user notices in the agent’s final response. Looking at successful injection traces, we find two distinct outcomes: the agent executes the injection while returning an otherwise normal response, or reports the injected action in its final response, giving the user a chance to notice. We call these covert and overt successes. From the user’s perspective, we decompose ASR into the Covert Success Rate (CSR), counting successes leaving no trace in the final response, and the Overt Success Rate (OSR), counting successes the user can detect.
To understand what drives the gap, we analyze successful trajectories and find that the agent’s behavior after the injection separates covert from overt: covert traces hand control back to the user task before ending, while overt traces end at the attack itself. This split follows from the ReAct format, where the final response summarizes the most recent action. Building on this observation, we propose ICoA (Induced Covert Attack), an IPI attack designed to induce covert outcomes by steering the agent back to the user task after executing the injection. Across four target models on AgentDojo, ICoA achieves the highest CSR, with gains of 3.79–12.01 percentage points over the strongest baseline.
Method
ICoA combines a user framing with a RETURN anchor. The user framing presents the injected task as a follow-up from the user; the anchor directs the agent to resume the original task after executing the injection. Returning to the user task shifts the focus of the final response and can leave the injected action undisclosed.
We distinguish execution from disclosure using CSR and OSR, with ASR = CSR + OSR. Each rate is measured over all evaluated traces. Disclosure is judged from the agent’s final response; it is not a direct measurement of human awareness.
Results
We evaluate four target models on 949 task–injection pairs across the Banking, Slack, Travel, and Workspace suites of AgentDojo. In the no-defense setting, ICoA raises CSR by 3.79–12.01 percentage points over the strongest baseline. The table below also includes the five defense settings reported in the paper.
| Model | Best baseline | Baseline CSR | ICoA CSR | Gain (pp) |
|---|---|---|---|---|
| Qwen3-235B | Imp. message | 29.72 | 36.04 | +6.32 |
| LLaMA-3.3-70B | Imp. message | 11.80 | 23.81 | +12.01 |
| GPT-4o-mini | Imp. message | 9.80 | 17.91 | +8.11 |
| Gemini-2.5-Flash | Imp. message | 34.67 | 38.46 | +3.79 |
Values from Table 1. The best baseline is selected by CSR for each model and setting. ChatInject is not evaluated on Gemini-2.5-Flash. The average includes no defense and all five defenses. Higher CSR indicates more successful attacks that remain undisclosed.
BibTeX
@inproceedings{lee2026icoa,
title = {Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents},
author = {Lee, Yunseok and Kim, Yunji and Lee, Woojin},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026},
note = {Oral presentation},
eprint = {2608.30362},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2608.30362}
}