How GitHub's sandboxing could help us keep AI Coding Agents under control

AI coding agents can now inspect repositories, run commands and make changes across a project. GitHub's new sandboxing approach raises an important question about how much access we should actually give them.

By Lorenzo Mugnai · · Custom Software & Integrations · 7 min read

I use AI coding tools regularly on my own software projects, and they've become useful surprisingly quickly.

What I've also noticed is how easily the level of access we give them can creep up.

It starts fairly innocently. You ask an AI tool to explain a piece of code or suggest a change. Then you let it look through the repository because that's much more useful than copying individual files into a chat. Before long it's running tests, changing several files at once, installing dependencies and executing commands in the terminal.

And that's the point where I think we need to understand a little more about what we're actually allowing these tools to do.

I'm not particularly worried about an AI agent suddenly deciding to do something malicious. I'm more concerned about it doing something perfectly logical based on the information it has, without understanding something about the system or environment that a developer knows.

A coding agent needs more than access to the code

Part of what makes the latest coding agents so useful is that they can actually work with a project.

They can search through a repository, inspect build files, run tests, look at the output and use what they find to decide what to try next. Some can run shell commands, install packages and interact with external services.

That's much closer to how I would investigate a problem myself.

The difference is that my development environment contains far more than the code I'm currently working on.

Like most developers, I may have access to private repositories, SSH keys, cloud services, local configuration, databases and different development environments. Some of that access is there simply because I've accumulated it while working on a project.

If I ask an agent to fix a failing test, how much of that does it actually need?

Probably the repository. Probably the build tools. It may need to run the tests.

Does it need access to my entire home directory? My SSH keys? Every service my machine can connect to? Production credentials?

Probably not.

And yet unless we understand how the tool is configured, we may be giving it exactly that.

GitHub is starting to put boundaries around this

This is why a recent change from GitHub caught my attention.

GitHub has introduced local sandboxing for its Copilot coding agent. The idea is that the agent can still work on a project, but developers can restrict what it has access to.

That can include which files it can read and change, whether it can access the network and whether Git and GitHub credentials are available to it.

GitHub has also designed the feature so that a session fails if the operating system cannot enforce the requested restrictions, rather than simply continuing without them.

What I like about the idea isn't really the sandboxing technology itself. It's the principle behind it.

Instead of asking what can we allow this agent to do?, we start by asking what does it actually need to do this piece of work?

That's already a familiar principle in software engineering. We generally don't give an application unrestricted access to every database and service available just because it's convenient. We give it the permissions it needs.

I'm beginning to think we should look at coding agents in much the same way.

What happens when an agent does something we didn't expect?

Imagine asking an agent to investigate a failing test.

It looks through the code, runs the test suite and decides that the problem appears to involve a database connection. It examines the configuration and starts trying different things to work out what's happening.

Nothing particularly unusual there. That's probably similar to what a developer would do.

But suppose the repository contains configuration for several environments.

A developer who knows the system may recognise one of those connections and think, I don't want to run that against there.

Would the agent?

Maybe. But I don't think that's something I'd want to rely on.

And this is where I think there's a danger of treating approval prompts as the complete solution. If an agent has successfully run ten commands while investigating something, it's quite easy for the eleventh approval to become another button we click without giving it much thought.

I'd rather have a development environment where the agent couldn't connect to something sensitive in the first place unless we'd deliberately decided that it should.

This gets more complicated when you're working in a team

This is where I think the conversation becomes particularly interesting for development teams.

When I'm using an AI coding agent on one of my own projects, I can decide how much access I'm comfortable giving it. But once several developers are working on the same application, those individual decisions could start to affect the whole team.

One developer might use an agent inside a restricted environment. Another might let it run commands directly on their machine. Someone else may have credentials or network access that the rest of the team doesn't have.

Suddenly the way AI agents are being used becomes part of the team's engineering practices, even if nobody has actually discussed it.

I don't think that means teams need a huge AI policy before anybody can use these tools. That would probably become outdated almost as quickly as it was written.

But I do think teams should understand what the tools they're using are capable of doing.

If we're happy for agents to write tests, refactor code and investigate failures, what access do those tasks require? Are there environments they shouldn't be able to reach? Should they be allowed to install new dependencies? When should a developer need to approve something?

Those seem like reasonable engineering conversations to have.

We still need to understand the changes they're making

Access is only one side of this.

An agent can stay entirely within the repository and still make a change that nobody on the team properly understands.

It might update a build configuration, introduce another dependency or remove some code that looks redundant.

Sometimes that code really is redundant. Sometimes it exists because of a problem somebody discovered three years ago and nobody ever wrote down why.

I've worked on enough existing systems to know how much knowledge sits outside the code itself.

There are business rules, integrations, strange edge cases and decisions that made perfect sense when they were made but aren't obvious when you look at an individual file.

AI doesn't make that context disappear.

In fact, as agents become capable of making larger changes more quickly, I think understanding that context becomes more important.

If an agent changes twenty files in a few minutes, somebody still needs to understand what changed and why. A green build helps. Passing tests help. Neither automatically tells us that we've made the right change to the system.

That doesn't mean manually checking every line an agent produces forever. Our development practices will change as the tools improve.

But I wouldn't be comfortable getting to the point where the answer to "why does the system work this way?" becomes "the agent changed it and the tests passed."

How should a team approach this?

I don't think the answer needs to be particularly complicated.

A sensible starting point would be to give an agent access to what it needs to do the work, rather than everything the developer happens to have access to.

For a lot of development work, that's probably the repository, the build tools and whatever is required to run the application and its tests locally. If it needs something else, that can be a deliberate decision rather than something the agent inherits automatically.

I'd also want developers to be able to see what an agent has done: the commands it ran, the files it changed, any dependencies it introduced and any configuration it touched.

Most importantly, I'd still expect the team to understand the change being delivered.

That last part matters to me more than any particular sandboxing feature.

AI coding agents are going to get better. I'm already finding them useful on my own projects, and I can see them taking on considerably more development work over the next few years.

But increased capability is probably a reason to understand their boundaries better, not a reason to stop thinking about them.

We shouldn't need to be frightened of letting an agent run a command. We should know what it can run, what it can reach and what happens if it gets something wrong.

That's the point where I'd be much more comfortable letting it get on with the work.