Article cover image

Agents escaping containment

Guillermo Rauch
Guillermo Rauch@rauchg

After OpenAI reported that Hugging Face was hacked by AI agents powered by their models, people have (rightfully) raised cybersecurity and even existential worries. I'm not worried, and I want to explain why in terms everybody can understand.

If you are not up to speed, OpenAI reported that an agent was put inside an evaluation environment, or sandbox, for cybersecurity research purposes. The agent was instructed to pursue exploits, and… exploit it did. It escaped. But it did so in a way that's surprising and deserves both credit and further study: it identified a novel security vulnerability.

At Vercel we let people run "untrusted code", meaning code written by a potential adversary to our systems and those of our other customers, billions of times a week. And we help customers build and deploy these systems―all kinds of agents, websites and applications―millions of times a day.

And yet, in our 10 year history, even after AIs, LLMs and agents became commonplace, we've had zero instances of code running "escaping" and compromising other systems. Even now, when over 50% of the code deployed on our platforms comes from AI agents.

What are we escaping from?

In our modern world, "everything is computer." Almost everything you interact with can run fairly arbitrary code. Your computer, obviously, runs code. Your browser runs the code X.com sends to show you posts and compose new ones. Computers in datacenters run code to process data, like returning the most recent timeline of posts to read.

Take the browser for example. Can the code X.com sends read your email? It cannot. If you're using Chrome, millions of expert human and AI hours have been spent on sandboxing the execution of the code. The code runs in a virtual machine¹: like a baby computer inside a larger computer.

What most people don't know is that virtual machines run the modern world. When you host an application on Vercel and you access it, we put code in a virtual machine, run it, and return the response to the end user. The code running inside cannot escape this sandbox. If it could, it would be catastrophic, which is why OpenAI and Hugging Face raised concerns.

These types of vulnerabilities, however, have existed, and are some of the most sought-after in the world by bad guys. When you "escape", you can find data you're not supposed to access. If you break a browser sandbox, a website can access information of other websites you're logged into. In extreme cases, all files on your computer, and even trojan-horse them.

In the case of escaping a virtual machine running in the datacenter, bad guys can access data of every user of an application, like stored messages, emails, and passwords, and even data of other applications. Imagine hacking A.com, and suddenly having access to B.com. Jackpot.

How to escape the matrix

Code contains bugs. Humans have written lots of code, and have in the process introduced bugs of all kinds. Bad people exploit these bugs to "escalate privileges" or escape sandboxes. In both cases, to get data they're not supposed to get. Some bugs are very subtle and require extraordinary cleverness and skill to exploit. Some are downright embarrassing.

In 2017, someone discovered you could walk up to any Mac, type "root", press enter twice, and get full control of that machine. Someone's kids randomly smashed the keyboard and caused a bypass of the lock screen of one of the most popular Linux distributions. On iOS 6 you could bypass the iPhone lock screen by making an Emergency Call. You could hijack any Windows system by pressing the Shift key 5 times.

This is to say, when you hear about cyberattacks, there's a whole gamut that ranges from "the dumbest bug in the world" to occasionally an exhibit of an intricate labyrinth of novel exploits that evade multiple layers of defense.

Here's the thing… just like those kids tried a bunch of things and found the reward (exploit a bug to log in), this is how we train agents. To make an agent good at cybersecurity, we put them in "gyms", which are sandboxed virtual machines, and we reward them for finding escapes. This process is called reinforcement learning (RL).

Think of it as a hamster that solves a puzzle in order to get its carrot.

Good news, bad news, good news

The good news is that the agent did exactly what it was optimized to do, was put in an environment similar to the one it was trained in, and… escaped. There's no bigger plot of a Skynet-style takeover of an evil agent trying to shut down humanity.

The bad news is that (1) the code that humans (increasingly with the help of agents) have written to date is vast and littered with security bugs and (2) agents are becoming amazing at finding and hacking them.

In the process of training, agents pick up a lot of tricks. They can think around corners that are not obvious to humans. They're relentless and do not tire. They can research vast amounts of information and load it up into their context.

The good news is that you can defend yourself. My recommendations:

1️⃣ Scan your code. We released an open source project called deepsec that helps you summon agents for defensive purposes.

2️⃣ Sandbox your agents. When you build and ship your own agents, even if benevolent, you should deploy them in secure sandboxes to mitigate the risk of misbehavior. Vercel Sandbox is a secure² environment that helps you reap the benefits of agentic intelligence, mitigating risks of data exfiltration.

The best part is that these tools are model-agnostic, helping you harness the power of both open and proprietary models for good. These are some tools we built, there are many others out there. Stay safe!


¹ There are multiple levels of isolation and virtualization available to developers, from leveraging the sandboxes of the execution runtimes (like JS or wasm), to processes, containers, micro and full virtual machines that imitate hardware. For the purpose of this article I'm treating these generally for the purpose of describing agent escapes.
² Vercel Sandbox builds on Firecracker, an open source microVM isolation technology with multiple layers of defense (KVM, a minimalist Rust device model, seccomp filters, and a jailer). Firecracker has a strong security track record, with no known full guest-to-host escape in its CVE history, with AI-assisted security research (e.g., Anthropic's disclosure of CVE-2026-5747) only turning up a conditional, pre-condition-gated flaw, now patched.