Salt Labs research showed a single malicious email could hijack Manus, an agentic AI platform seeking a $4 billion valuation.
No stolen password. No clicked link. The only actions required were the email arriving and the user asking Manus to check its inbox.
Manus’s own guardrail correctly flagged a plaintext malicious command. Researchers then obfuscated the same instruction. The agent decoded and executed it before the warning could stop anything.
Once inside, the researchers could reach credentials and tokens for every connected third-party service — email, cloud storage, and code repositories.
The real lesson goes beyond one platform: detection-based guardrails were designed for environments with a human in the loop. Autonomous agents collapse that window. By the time a security alert reaches a person, the action has often already finished.
Detecting an attack after the agent has acted is not prevention. It is only a log of what already happened.
Full analysis:
#AISecurity #AgenticAI #PromptInjection #Cybersecurity
Manus AI agent hijack: a single email bypassed guardrails and reached every connected account. Detection isn’t prevention.
