The sandbox stopped at the template engine
On October 2, 2026, GitLab patched CVE-2026-90970, a CVSS 9.9 vulnerability in its self-hosted AI Gateway. An authenticated user with Duo Agent Platform access could craft a flow configuration that escaped the prompt-template sandbox and executed arbitrary commands on the gateway host.
What broke
The self-hosted AI Gateway exists for organizations that want Duo's AI features without routing prompts and responses through GitLab's own infrastructure. It runs inside the customer's environment, which is precisely why this flaw matters: the control built to reduce exposure is what carried the critical bug.
Duo Agent Platform lets users define flow configurations, workflows the AI agent follows to execute multi-step tasks. Those flows run through a prompt-template engine meant to sandbox user-supplied content away from the underlying system. GitLab classified the root cause as CWE-1336, improper neutralization of special elements in a template engine: a crafted flow definition could break out of that sandbox and reach command execution on the host directly.
Exploitation required authentication and Duo Agent Platform access, not an anonymous outsider, but GitLab still rated it critical and pushed targeted outreach to affected customers before publishing. The flaw affected self-hosted AI Gateway versions from 18.1.6, 18.2.6 and 18.3.1 up to 19.2.3, 19.3.1 and 19.4.0, fixed in 19.2.4, 19.3.2 and 19.4.1. GitLab.com, GitLab Dedicated, and any self-managed instance using a GitLab-hosted gateway were never exposed.
A self-hosted deployment also puts JWT signing keys inside that environment. GitLab requires separate signing key pairs for the AI Gateway and the Duo Agent Platform service, and its own install docs classify both as sensitive credentials. Host-level command execution lands next to more than prompts and configuration.
The boundary worth testing
This is not the first time GitLab has had to patch this exact boundary. In February 2026, CVE-2026-1868 hit the same Duo Workflow Service component: also CWE-1336, also CVSS 9.9, also a crafted flow definition reaching unsafe template expansion, that time able to cause denial of service or code execution. Different bug, same boundary. User-controlled workflow configuration keeps reaching farther into the runtime than it should, and it's happened twice in eight months in the same part of the same product.
It's also the same underlying shape as the Salesforce Agentforce flaws we covered last week. SalesBleed broke at the boundary between untrusted input and trusted instructions. This one breaks at the boundary between user-controlled configuration and the engine executing it. Different product, different mechanism, same category of failure: a boundary an AI feature depends on turned out softer than it needed to be.
Worth keeping in proportion: this needed real authentication and specific access, not an open door, and GitLab patched it, notified affected customers directly, and published a clear advisory with no reported in-the-wild exploitation. Nothing here suggests the AI model did anything it wasn't supposed to. The flaw sits in the infrastructure around the agent, in how the gateway processes a configuration, not in a decision the model made.
What we would check
Patch first, to 19.2.4, 19.3.2 or 19.4.1, if you're running a self-hosted AI Gateway. Then look at who can actually create or modify Duo Agent Platform flows, since your real exposure window was scoped by that access list, not your whole user base. Check where the Gateway's signing keys live and whether they need rotating if you have any reason to suspect pre-patch activity. Look at what privileges the gateway's container actually holds, and what a successful sandbox escape could reach from that network segment. Given that this is the second CVSS 9.9 template-engine escape in this exact component in eight months, the question isn't whether the fix works, it's whether the sandbox around it has been tested by anyone other than the person who built it.
A sandbox is only useful if someone keeps trying to break it. If your team runs self-hosted agent infrastructure, we test that boundary as part of our penetration testing work.

