AI's impact on rewriting old software
AI coding agents are making large software rewrites much faster. But the difficult part of replacing an old system may never have been writing the code.
By Lorenzo Mugnai · · Custom Software & Integrations · 12 min read
I've always been fairly cautious about rewriting existing software. It's not that a rewrite is never the right thing to do. Sometimes a system really has reached the point where continuing to patch it up doesn't make much sense. But there's usually a lot more hiding inside an old application than you realise when you first look at the code.
That's why I found GitHub's recent rewrite of the Copilot runtime interesting. They moved it from TypeScript running on Node.js to Rust, and this wasn't a small application. Around 430,000 lines of production TypeScript passed through the migration and the resulting runtime ended up at more than 800,000 lines of production Rust. Most of that new code was written using AI coding agents.
The numbers are impressive, but what caught my attention was GitHub's estimate of what the same project would have looked like without AI. They believe it could have needed a development team for a year or two. Instead, most of the migration was led by one developer over a few months, with support from the wider team, while development of Copilot continued around it.
That did make me reconsider one of the reasons I've traditionally been wary of rewrites. A big part of the problem has always been the cost. You could have a development team spend a year rebuilding something the business already has, and during that time they're not building all the other things the business needs. Even if everyone agrees the old system isn't particularly pleasant to work with, it can be a difficult business case to make.
If AI can take a substantial amount of that implementation work away, perhaps the calculation isn't quite the same as it was a few years ago. There are probably systems sitting in businesses today where modernisation has been discussed several times and put off because nobody can justify the amount of development effort involved. AI could make some of those conversations worth having again.
The further I got into GitHub's write-up, though, the less interested I became in the amount of code the agents had produced and the more interested I became in what happened during the migration.
GitHub found dozens of regressions. Some functionality wasn't completely carried across, there were differences in behaviour between the old and new implementations, and there were problems involving state, object lifetimes and the boundaries between the runtime and the applications using it. One thing that particularly caught my attention was what GitHub describes as incorrect test oracles. In some cases, the thing being used to decide whether the new implementation was behaving correctly wasn't actually checking the right behaviour.
Anyone who's spent much time working on existing systems will probably recognise the wider problem. It's tempting to think that if we've got the old code and a decent set of tests then we've got a specification for the new system, but I've rarely found existing software to be quite that tidy. Tests can be incomplete and documentation gets out of date. A business rule might have been added years ago to deal with an edge case and never written down anywhere else. Sometimes there aren't many tests or much documentation in the first place.
I had a look at some of the discussions developers were having around legacy rewrites as well. One that caught my attention involved a Rails application dating back to 2014 that processed real money. The team had spent two years trying to get agreement to rewrite it and had finally got the go-ahead, only to become nervous about actually starting. There were no tests and parts of the codebase that people were reluctant to touch.
I can understand that. Getting permission to rewrite a system doesn't suddenly tell you everything the old system does.
The advice from other developers was largely about capturing that behaviour before replacing it: putting integration and end-to-end tests around the important flows and migrating the application piece by piece. I found similar discussions from developers actually using Copilot for this sort of work. One team migrating a PHP/Symfony application to Laravel found that Copilot became inconsistent when it didn't have enough architectural context, while another developer described a TypeScript-to-Rust migration where the agent kept reinventing architectural decisions between sessions.
These are individual experiences rather than evidence that every AI-assisted rewrite will behave the same way, but I think they're useful because the problems sound very familiar. AI might make producing the new code considerably quicker, but it still needs a reliable way of understanding what the old code was supposed to do.
You can easily come across a strange condition in an old piece of code and wonder why it's there. It might be redundant or something that hasn't been needed for five years. On the other hand, it could handle an obscure business rule that affects a particular type of transaction a couple of times a year. Perhaps there's a test explaining it, or an old Jira ticket, or somebody on the team remembers why it was added. Sometimes the code itself is the only record you've got.
This is one area where I think AI coding tools can be genuinely useful. On my own projects I've found their ability to explore a codebase, follow references and look at related code particularly helpful. An agent can inspect considerably more code than I can in the same amount of time and help piece together relationships that might otherwise take quite a while to find.
There's still a difference between finding a piece of code and understanding the business reason behind it. If that knowledge was never recorded in the repository, the agent is in much the same position as a developer joining the project for the first time. It can make a sensible interpretation from the information available, but that interpretation can still be wrong.
The way GitHub carried out its rewrite becomes quite important when you look at it from that perspective. They didn't point an agent at the TypeScript repository and ask it to come back when everything was Rust. The migration happened incrementally, with relatively small pieces of the TypeScript implementation being replaced while the main branch remained shippable.
Over roughly fourteen and a half weeks, GitHub shipped 135 releases: 100 prereleases and 35 stable releases. The new implementation was being exercised throughout the migration rather than sitting separately for months waiting for one big switch at the end. GitHub also used the prereleases to limit initial exposure, watch for problems and get fixes out quickly.
There's nothing particularly new about that approach. We've been breaking risky changes into smaller pieces, comparing old and new behaviour and gradually replacing parts of systems for years. What I find interesting is that AI doesn't seem to make those engineering practices less important. If anything, being able to produce code much more quickly might make them more important.
If an agent can generate in a day the amount of code that might previously have taken a developer considerably longer, we still need to understand and verify that change. At some point the bottleneck starts to move. Writing the code becomes quicker while reviewing it, testing it and understanding whether it does the right thing still takes time.
There was another detail in GitHub's account that made me think about what we actually mean by modernising software. Their aim during the migration was to preserve behaviour rather than redesign everything at the same time. That makes sense to me. Changing the programming language, architecture and behaviour together would make it considerably harder to work out what had caused a problem.
The consequence is that although the new implementation is written in Rust, GitHub says much of it still consists of TypeScript-shaped algorithms rendered in Rust. They haven't yet done the broader redesign work that Rust's ownership model, concurrency model and in-process architecture could allow.
It made me wonder what we actually mean when we say we've modernised a system. If we use AI to translate an old application very efficiently into newer technology, have we actually modernised it or have we created a newer version of the same thing?
That might be perfectly fine if the technology itself is the problem. GitHub had specific reasons for wanting to move away from Node.js and V8. The existing approach meant SDK consumers had to run the Copilot runtime in a separate process, bringing startup and memory costs, cross-process communication and another process to monitor and debug. GitHub wanted lower overheads, more predictable resource use and a runtime that could be embedded directly into applications rather than always running out of process. Those requirements eventually led them to Rust.
The resulting performance improvements were substantial. In one workload GitHub measured the time for creating a client and session, completing a turn and tearing everything down falling from 5.25 seconds to around 55 milliseconds when running the Rust runtime in-process. In another test, memory use for ten clients fell from 1,383 MB above baseline to 126 MB. GitHub is careful not to claim that this makes the new runtime universally a particular number of times faster, because the results depend on the workload, but it does show there was a real engineering reason for doing the work beyond simply replacing an old technology with a newer one.
With an older business application, I'd want to understand what we're trying to fix before deciding that rewriting it is the answer. An old framework may genuinely be becoming difficult to support, or perhaps security and dependency problems are starting to accumulate. In other cases the pain comes from the way the application has grown over the years: tightly coupled components, difficult deployments, poor test coverage or important business knowledge that only exists in the heads of a couple of people. Rewriting the code in a newer language doesn't automatically make those problems disappear.
GitHub isn't the only organisation experimenting with AI-assisted migration at this sort of scale either. Google has been using specialised coding and reviewing agents for C++ to Rust migration work, including work around Fuchsia and other substantial codebases. Again, what interests me isn't simply that an AI agent can generate a lot of Rust. The resulting code still has to be tested, reviewed and understood before anyone can be confident putting it into production.
There does seem to be a pattern emerging. AI can do an extraordinary amount of the implementation work, but people are still spending a lot of effort proving that the result is correct.
I think that's where the GitHub story has changed my view a little. I still wouldn't look at an old application and assume that because we now have capable coding agents we should rewrite it. The risks that made rewrites difficult haven't disappeared. What may have changed quite significantly is the cost of one part of the job.
GitHub says the AI token bill for its migration came to around $120,000, covering roughly 136 billion tokens. There was obviously human engineering effort around that figure as well — guiding the work, reviewing changes, fixing problems and supporting the migration — so it would be misleading to compare $120,000 directly with the cost of employing a development team for a year or two. Even so, the difference in what a relatively small number of people can now attempt is difficult to ignore.
For a business with an ageing system, that could make some options worth reconsidering. A modernisation project that didn't make financial sense three years ago might look different if a significant amount of the repetitive implementation work can now be handled by agents.
I'd still want to know a lot about the existing system before deciding to do it: what behaviour has to be preserved, how good the tests are, where the business knowledge lives, whether the system can be changed gradually rather than replaced in one enormous step, and what we're actually trying to improve in the first place.
AI can help investigate the existing code, create tests, trace dependencies and produce replacement code at a scale that would have seemed fairly unrealistic not long ago. What it can't do is magically recover business knowledge that was never recorded or decide which of fifteen years' worth of accumulated behaviour still matters.
When I think about the difficult legacy systems I've worked with, physically writing the replacement code was rarely the thing that concerned me most. It was understanding what I could safely change without breaking something that somebody, somewhere, still relied on.
AI may be starting to make the first part considerably cheaper. I'm not sure it's made the second part any easier.
References and further reading
GitHub — Migrating the GitHub Copilot runtime to Rust, using Copilot
The main source for the scale of the migration, how GitHub approached it, regressions, performance results and AI costs.
Read GitHub's engineering write-up
Reddit / r/softwarearchitecture — Finally convinced leadership to let us rewrite the legacy app. Now everyone is terrified to start
A developer discussion around a Rails application handling real money, no existing tests and the risks of beginning a rewrite.
Reddit / r/GithubCopilot — Anyone using Copilot effectively for refactoring a large legacy codebase?
Experiences of using Copilot during larger migrations, including test generation, architectural context and keeping agents consistent across the work.
Reddit / r/GithubCopilot — Migrating codebases between proprietary frameworks
Discussion around using coding agents with more than 20 years of legacy code and little or no existing test coverage.
Google Fuchsia — C++ to Rust migration tooling
Google's work on agent-assisted C++ to Rust migration, including separate coding and reviewing roles.