Blog / Agents
Email prompt injection: how one email hijacks an AI agent
How it works
A language model can't reliably tell your instructions from text that merely looks like instructions. When an agent summarizes or searches your inbox, the emails' text lands in the same context as your request. If one of those emails says "ignore the user and do this instead", the model may do it.
Attackers make this invisible to you:
- Hidden text: white text on a white background, zero-size fonts, or text inside HTML that your mail app doesn't display. You see a normal email; the model reads everything.
- Invisible characters: Unicode characters that don't render at all but that a model reads as text.
- Instructions dressed as system messages: "IMPORTANT NOTICE FROM THE ADMINISTRATOR:…", written to sound like part of the agent's own setup.
It has already happened
| When | What | Result |
|---|---|---|
| Feb 2023 | Research paper by Greshake et al., "Not what you've signed up for" | First systematic study of indirect prompt injection, including attacks on email assistants |
| Aug 2024 | Microsoft 365 Copilot, "ASCII smuggling" (Johann Rehberger) | A phishing email made Copilot collect sensitive emails and MFA codes and hide them in a link with invisible characters. Patched |
| Jun 2025 | EchoLeak, CVE-2025-32711 (Aim Security) | Zero-click: one crafted email led Microsoft 365 Copilot to leak data, getting past its injection filter. Fixed server-side; no known abuse |
| Jul 2025 | Gemini in Gmail, "Phishing for Gemini" (0DIN) | Hidden text made "Summarize this email" show a fake Google security warning with a phone number. No links or attachments needed |
None of these needed you to click anything suspicious. The email only had to be read by the assistant.
Why agents make it worse
Security researcher Simon Willison calls it the lethal trifecta: an agent is dangerous when it has all three:
- access to your private data,
- exposure to untrusted content (every email in your inbox),
- a way to communicate outward (sending email, fetching a link, loading an image).
An agent with full access to your mailbox has all three by design. Remove any one and the attack loses its payoff.
How to protect an agent that reads your email
These follow OWASP's guidance for LLM applications (LLM01, prompt injection):
- Least privilege. Give the agent only the emails its task needs: one mailbox, a few categories, the last few days. What it can't see can't instruct it. We describe this pattern in How to give Claude, ChatGPT or Cursor access to your email safely.
- No outward channel without you. The agent shouldn't send email, post data or open links on its own. Ask for your confirmation for anything that leaves.
- Keep untrusted text separate and marked. Tell the agent which parts are email content, and treat subject lines and sender names as text written by strangers.
- Flag the known tricks before the agent reads. Hidden text with instructions and invisible characters are signals you can detect without trusting the model.
What MailTag does
- Agents don't get your email's text from MailTag. Through its MCP server, an agent receives the category, date, and the sender and subject only if you allow it. Much less text written by strangers reaches the agent.
- Per-agent limits: which mailboxes, which categories or views, how far back. An agent that only sees Needs reply from today isn't reading the newsletter that carries the attack.
- No sending, deleting or moving mail out. MailTag's MCP server has no such tools.
- A safety check you can turn on (Settings → Mailboxes → Check each email). It looks for hidden text with instructions, invisible characters, lookalike sender domains and links, and asks MailTag's own model whether the email reads like phishing or like instructions to an AI agent. Suspicious emails get a mark, and you can hide dangerous ones from each agent.
No check catches everything: attackers adapt, and a model can be fooled. Limiting what the agent sees and what it can do is what keeps a missed email harmless.
FAQ
Can prompt injection happen if I never open the email?
Yes. The attack targets the assistant, not you. It runs when the agent reads or summarizes your inbox.
Is my spam filter enough?
No. These emails don't need links or attachments and can come from ordinary senders, so they often pass spam filters. The hidden instructions are what matter.
Does MailTag's check block the email?
No. It never moves or deletes email. It marks suspicious ones in MailTag and, if you want, with a ⚠ Suspicious label in your mailbox (in Gmail; a mark elsewhere), and it can hide them from your agents.
Does it cost extra?
No. The safety check and connecting agents through the MCP server are both included on every plan. You turn the check on in your settings.
Keep reading
Give Claude, ChatGPT or Cursor safe access to email
Connect an AI agent to your email through MailTag's MCP server and choose what it sees: which mailboxes, categories and dates, and nothing more.
Read →Guides · 30 Sept 2026Organize iCloud, Yahoo or Fastmail email automatically
What iCloud, Yahoo and Fastmail can sort on their own, where their rules stop, and how to sort each new email into your own categories over IMAP.
Read →