Key takeaways
  • A rewrite has to reproduce everything the old system does, and the only complete record of that is the code.
  • Before deciding anything, map the architecture, business rules, integrations, real data, security gaps and dead code.
  • The decision usually lands in one of three places: fix it in place, replace it piece by piece, or start over.
  • An architecture and code review gives you that picture in 30 days at a fixed price, and your own team can act on the report.
01

Why rewrites go wrong

In April 2000, Joel Spolsky wrote about Netscape's decision to rewrite its browser from scratch and called it "the single worst strategic mistake that any software company can make." His reasoning has aged well. Old code looks ugly because it is full of fixes, and each fix is something somebody learned the hard way: a bug a customer hit, an edge case in a date calculation, or a workaround for a partner's broken file format. In his words, "When you throw away code and start from scratch, you are throwing away all that knowledge."


A rewrite team usually starts from what the business says the system does. That description is always shorter than the truth. The gaps show up late, during parallel running or after go-live, when the new system produces a different number from the old one and nobody can explain why.

There's a second problem, and it involves everyone else. Software engineer Hyrum Wright put it this way "With a sufficient number of users of an API, it does not matter what you promise in the contract: all observable behaviors of your system will be depended on by somebody." A system that has been running for a decade has had a lot of users. Some of them rely on behavior that was never meant to be a feature, and they won't tell you until it's gone.

02

The code is the specification

Old systems do have documentation, sometimes a lot of it. The trouble is that it describes the system as someone intended it on the day they wrote it. The code describes what the system does today, including every change made since by people who didn't update the docs.

So when you ask what this system does, the honest answer is in the source. That's uncomfortable, because the source might be RPG, PL/SQL, PowerBuilder, VB6, or a .NET Framework app with half its business logic inside stored procedures. It's also the only place where the answer is complete.

03

What to find out before you decide?

Nobody needs to understand every line before making this call. Answering these six questions is enough.

04

How is it actually put together?

Draw the real architecture: the applications, the databases, the scheduled jobs, the file drops, and what calls what. Then compare it with the diagram on the wiki. The differences are usually where the risk sits, because they're the parts nobody planned for.

05

Which business rules live in the code?

Pricing exceptions, eligibility checks, the rounding rule finance agreed to years ago, and the special case for one large customer that someone added on a Friday and never removed. These rules are the requirements for whatever comes next. If you can't list them, you can't test a replacement against them, and you'll find them one production incident at a time.

06

What depends on it?

List every integration, including the unofficial ones: the nightly export another team picks up, the report someone in finance runs straight against a production table, and the spreadsheet with a live database connection. None of these will be in an architecture document, and they'll be the first things to break.

07

What does the data really look like?

Over the years, columns get reused. A field called `notes` ends up holding account codes. A status of 9 means something only one person remembers. A null date means "never" in one table and "unknown" in another. A migration plan built from the schema instead of the data will run into every one of these.

08

Where are the security and compliance gaps?

Look for unsupported runtimes and operating systems, libraries with published vulnerabilities, credentials hard-coded in config files, and audit trails that record the change but not who made it. These matter whether you rewrite or not, and some of them can't wait for a two-year program to finish.

09

What can you throw away?

Every old system carries code that never runs and features nobody uses. Finding them is the cheapest win in the whole exercise, because every module you delete is one you don't have to rewrite, migrate or test.

10

Fix it, replace it in pieces, or start over

Once you can answer those questions, the decision is usually less dramatic than the original argument made it sound. It tends to land in one of three places.

If the architecture is sound and the pain is concentrated in a few modules, fix those modules. Upgrade the runtime, put tests around the parts you're about to change, and refactor where the changes keep landing. It's the least exciting option and often the right one.

If the system is too tangled to fix but too important to switch off, replace it a piece at a time. New components sit beside the old system and take over one function after another while the rest keep running. Martin Fowler named this approach the strangler fig, after the vine that grows around a tree until it can stand by itself. Each step is small enough to roll back.

And sometimes a full rewrite really is the answer. The platform has no upgrade path, or the system is small and well understood, or the business has changed so much that the old rules no longer apply. That's a legitimate conclusion. It's one you should reach after reading the system, though, and not before.

11

What an architecture and code review gives you

Reading a legacy system properly takes time your team probably doesn't have, because they're the people keeping it running. An architecture and code review does that reading for you.

The review runs 30 days at a fixed price agreed before we start. We go through your codebase and architecture, map how the system fits together, find the hidden architectural debt, and flag the security and compliance gaps. Then we write down, in plain language, what is broken and what it would take to fix it.

What happens next is up to you. Your own engineers can take the report and do the work. That's a complete outcome, and the report is written for it. If you'd rather we do the upgrade, that's a separate project with its own scope and price, and it carries less risk because nobody is guessing what the system does anymore.

Either way, you know what you're dealing with before anyone signs up for a multi-year rewrite.

12

Common questions

How long does an architecture and code review take?

Thirty days, at a fixed price we agree before we start.

Do we have to hire you to fix what the review finds?

No. The report stands on its own and is written for your engineers to act on. If you want us to do the upgrade, we scope and price it as a separate project.

Nobody on our team understands the system anymore. Is a review still worth it?

That's one of the better reasons to do one. When the people who built a system have moved on, the code is the only complete record of what it does, and somebody needs to read it before anything changes.

Is a full rewrite ever the right answer?

Yes. If the platform has no upgrade path, or the system is small and well understood, starting over can be the sensible call. A review tells you which situation you're in before the budget is committed.

How is this different from running a code scanning tool?

Scanners are good at what they do. They'll list warnings, outdated dependencies, and style problems. They won't tell you why a module exists, what another team depends on, or whether the system is worth keeping. A review ends in a recommendation you can act on and defend to the people paying for it.


How do we start?

Book a call. We'll talk through the system and tell you plainly whether a review is worth it. If it isn't, we'll say so, and the call costs you nothing.