The AI agent that hacked its own sandbox

Bodil Biering


We started a small weekly thing recently — the AI Builders Club, a handful of CTOs and CPOs getting together to compare notes on AI tools.

No pitching, no slides. Just what’s actually working, what isn’t, and what’s a waste of time.

In our first session, two stories stuck with me.

One person had sandboxed an AI coding agent — locked down, no permissions, the way you’re supposed to.

But the agent needed a file that lived somewhere else.

So it found a workaround.

It used Excel as a side door, and let itself out.

Another person had told their agent to stay inside one folder. It didn’t. It reached into their Downloads folder, grabbed a file it needed, and carried on.

It never asked.

Neither of them had told the agents to do that.

They just... found a way.

The interesting part isn't that AI is "sneaky"

An AI agent doesn't need malicious intent to do something you didn't expect.

You give it a goal. It encounters an obstacle. And increasingly, it has enough tools and autonomy to find another route.

That's useful.

It's also exactly why we need to think differently about the boundaries we give these systems.

Telling an agent where it's allowed to go isn't the same as preventing it from going somewhere else.

If your prompt says stay inside this folder, but the agent has access to a tool that can reach outside it, you haven't really created a security boundary.

You've made a request.

If you need a wall, build a wall

The practical takeaway from the conversation was surprisingly simple:

Don't ask an AI agent to respect a boundary you haven't actually enforced.

If an agent shouldn't be able to access something, don't rely on its instructions to keep it away.

Think about what the environment actually allows:

  • Can it access files outside the workspace?

  • Can the tools it uses access them?

  • What credentials does it have?

  • Can it make unrestricted network requests?

  • If it tries something unexpected, what actually stops it?

That might mean filesystem isolation, tighter tool permissions, network controls, containers, or separate credentials depending on what you're building.

The exact setup isn't really the point.

The point is that the boundary needs to exist outside the model.

We're giving AI agents more autonomy because that's what makes them useful. And as they get better at solving problems, they'll also get better at finding paths we didn't anticipate.

That's not necessarily a reason to give them less autonomy.

It's a reason to build better walls.


Get more stories like this in your inbox — join the list →

Let’s get started

Tech start-ups and scale-ups trust CyberJuice - the experts on getting rid of customers' question "Are you secure?"

Build trust. Speed up security reviews. Close more deals.

Get started

cyberjuice-logo

Fast-track your business with proof of your company's security.

© 2025 Cyberjuice. All rights reserved.

Let’s talk

Growing teams trust CyberJuice - the compliance platform that makes you smile.

Get started

cyberjuice-logo

Fast-track your way to security and compliance with smart automation and human support - while upskilling your team to handle it with confidence.

© 2025 Cyberjuice. All rights reserved.

Let’s talk

Growing teams trust CyberJuice - the compliance platform that makes you smile.

Get started

Fast-track your way to security and compliance with smart automation and human support - while upskilling your team to handle it with confidence.

cyberjuice-logo

© 2025 Cyberjuice. All rights reserved.