Engineer sketching a system diagram on a whiteboard
The diagram is where the review starts, not where it ends.

Search for a software architecture review checklist, and most of what comes back is written for a design that doesn't exist yet: a proposal on a whiteboard, reviewed before anyone writes code. That's useful work. It's not much help when your problem is a system that has been running for twelve years and nobody is quite sure how it's built anymore.

Reviewing an existing system is a different job. You're not asking whether a design will work. You're asking what was actually built, how it behaves, and what it's costing you. The answers sit in code, configuration, and logs. The architecture document may be the least reliable source in the building, because it describes what someone intended on the day they wrote it.

This checklist is for that second job. Work through it in order, and write down where each answer came from.

Key takeaways
  • Review the system that's running, not the diagram of it. The differences are where the risk is.
  • Check boundaries, data, integrations, failure handling, security, deployment, operations and change history.
  • Every answer should come with evidence: a file, a log, a config setting, or a person who confirmed it.
  • Finish with a short list of decisions the business can make, not a long list of findings.
01

Before you start, collect evidence instead of slides

Get read access to the source repositories, the deployment scripts and server configuration, the database schemas, the list of scheduled jobs, and a week or two of logs. Then book an hour each with the people who run the system and the people who use it.

Someone will offer you the architecture deck first. Take it, but read it last. If you read it first, you'll spend the review confirming the diagram instead of checking it.

1. Boundaries: What is actually in the system?

List every piece that runs: applications, services, databases, batch jobs, file drops, and the scripts living on somebody's server. Mark which ones are in source control. Anything that runs in production without being in version control is a finding on its own.

Then put your list next to the official diagram. Every box on the diagram that you couldn't find, and every running piece that isn't on the diagram, goes in the report.

A good answer is one page that names every running part and where its code lives.

2. Data: Where does it live, and who writes to it?

Find out which databases and schemas exist and which applications write to each important table. Pay attention to tables that more than one application writes to. Shared writes are coupling, and they're what makes "small" changes turn into cross-team projects.

Look for data stored in two places. Customer details in the billing system and again in the CRM is normal. Not knowing which copy wins when they disagree is the problem.

A good answer names a system of record for each important thing the business tracks: customers, accounts, orders, and invoices.

3. Integrations: What talks to what?

List every connection in and out: APIs, file transfers, message queues, database links, and emails that trigger work somewhere else. For each one, note who owns the other end and what happens when it's unavailable.

Don't take the integration list from the wiki. Check firewall rules, scheduled jobs, and outbound connections in the logs. The integrations nobody documented are the ones that break during a migration, which is why [knowing what depends on a system](/insights/before-you-rewrite-a-legacy-system) matters before anyone changes it.

4. Failure handling: what happens when something breaks?

Pick the three flows the business cares about most, such as taking a payment, issuing an invoice, or closing the month, and walk each one step by step. At every step, ask what happens if it fails halfway.

Are there retries? Can a retry create a duplicate? Does anyone get told, and how? If the honest answer is "a customer calls us," write that down. It's a valid finding and an important one.

5. Security: Where are the obvious gaps?

You're looking for the gaps that don't need a penetration test to find:

  • Runtimes, operating systems,/contact and libraries that are past their vendor's support date
  • Credentials stored in source code or plain-text config files
  • Accounts shared between people, or between systems
  • Sensitive data that isn't encrypted or masked where it should be
  • Audit logs that record a change but not who made it

Each of these is cheap to find and expensive to discover the hard way.

6. Deployment: How does a change reach production?

Trace one recent change from merged code to production. Count the manual steps. Note who did them and whether anyone else could have.

Then ask whether you can roll back, and whether anyone has actually tried. A rollback plan that has never been run is a hope, and the review should say so plainly.

7. Operations: How do you know it's healthy?

What's monitored, and what isn't? Where do the logs go, and how long are they kept? Is there a runbook, and was it updated after the last incident?

The test here is simple. If the system misbehaved at 2 a.m., would someone find out before the users did?

8. Change history: Where does the work keep landing?

Version control history is the most honest document you have. Find the files and modules that change most often and the ones that collect the most bug fixes. Those hot spots are usually where the design is fighting what the business needs now.

Also note the modules nobody has touched in years and nobody can explain. Some of them are stable and fine. Some of them are dead code that still gets patched, tested, and migrated for no reason.

02

Turning findings into decisions

A review that ends in a list of two hundred findings rarely gets acted on. People read the first page and file the rest.

Group the findings into decisions the business can actually make. Fix now covers security gaps and single points of failure. Fix while you're there covers the hot spots, which get cleaned up as part of work that's already planned. Leave alone is for code that's stable and understood. Retire is for dead code and features nobody uses. Give each decision its evidence and a rough size, and put the full finding list in an appendix.

03

Doing it yourself, or bringing someone in

Your own team can run this checklist. The catch is time and distance. The people who know the system best are busy keeping it running, and they've stopped noticing the parts that look normal to them. A reviewer from outside the team reads the system cold, which is exactly what this kind of review needs.

An architecture and code review is that outside reading, with the code examined as well as the architecture. It runs 30 days at a fixed price agreed before we start. We map how the system fits together, find the hidden architectural debt, and flag the security and compliance gaps. The report says, in plain language, what is broken and what it would take to fix it. Your team can act on it, or we can do the upgrade as a separate project with its own scope and price.