TheVortiq
Inteligencia Artificial

The Manus hack: the end of innocence for AI agents

Prompt injection vulnerability reveals a systemic risk: autonomous agents are executing hidden instructions within external data.

October 5, 2026 · 4 min read

a black background with a blue and green design

TL;DR: A prompt injection attack using obfuscated code allowed the hijacking of the Manus agent. This event highlights the fragility of agents with execution permissions and the urgent need to isolate data from commands.

The illusion of security in automation: the Manus turning point

The recent security breach discovered in the Manus AI agent is not an isolated incident, but a critical symptom of a structural crisis in the development of autonomous systems. Researchers at Salt Labs managed to bypass the agent's protection mechanisms using JSFuck, an obfuscation technique that converts JavaScript code into a sequence of six characters ([, ], (, ), !, +). This method allowed them to evade heuristic defenses, which typically scan for malicious text patterns in natural language, and execute arbitrary code directly within the agent's environment. This event sets a worrying precedent: the sophistication of attackers is outpacing the ability of developers to create secure boundaries, reminiscent of the SQL injection era in the early 2000s, where a lack of data sanitization led to the collapse of major infrastructures.

The fundamental problem: the ambiguity between data and instructions

The core of the problem lies in the architecture of current language models, which operate under a single instruction paradigm. AIs do not have an intrinsic distinction between the data they process (such as an incoming email) and the commands they execute. In the case of Manus, when a user asks the AI to "manage their inbox," the system interprets the entire content of the email as an extension of the user's instruction. If an attacker inserts a hidden command, the model processes it as a legitimate order, a phenomenon known as second-level prompt injection.

Historically, this is comparable to Cross-Site Scripting (XSS) attacks, where the browser does not distinguish between the developer's script and the injected script. The difference lies in the scale: an AI agent does not just read data, it has execution capabilities (tool use). If the model has permissions to send emails or modify files, the attacker does not just exfiltrate data, but uses the agent as an internal malicious actor. The JSFuck technique was the disruptive factor: as a form of non-standard code, it evaded sandboxing filters that usually look for commands like 'delete' or 'download', proving that blacklist-based security is fundamentally insufficient in the era of LLMs.

Impact and future: are we granting too many permissions?

The adoption of autonomous agents is happening at breakneck speed, often ignoring the lessons of the past in cybersecurity. According to the '2026: The State of Consumer AI' report by Menlo Ventures, user trust has grown disproportionately to security guarantees: 36% of users already grant access to their emails, 33% to browsers, 31% to messaging apps, and 29% to cloud storage. This interconnectivity, while boosting productivity, turns every user into a potential attack node within an enterprise network.

The risk is systemic. If an attacker compromises a personal agent, they can use it to perform lateral movement within the corporate network, accessing confidential documents or impersonating the user in internal communication channels like Slack or Microsoft Teams. Speculation about future attacks points to agents that, once compromised, act as persistent malware, waiting for the optimal moment to perform a malicious action without the user detecting an anomaly in their usual workflow.

Mitigation strategies for users and companies

  • Principle of least privilege: Companies must implement a strict role-based access control (RBAC) system for agents. Not all agents should have access to the email API or the root file system.
  • Human-in-the-loop: For high-sensitivity actions, such as changes to security settings, asset transfers, or external communications, human intervention is non-negotiable. Mandatory checkpoints must be established.
  • Execution sandboxing: Developers must execute AI-generated code in isolated environments (ephemeral containers without network access) and perform static and dynamic code analysis before it is executed by the host system.
  • Multi-layered defenses: Security should not rely solely on the AI model. Input/output filtering is required to scan prompts for obfuscation patterns and anomalous behaviors, regardless of the model's internal logic.

Although Manus has patched this vulnerability, the problem persists in the underlying architecture. AI security is not a fixed goal, but a process of continuous adaptation. The industry must transition toward a Zero Trust architecture applied to AI, where no instruction, even if it comes from a familiar "assistant," is considered safe by default without prior integrity verification.

Keep reading