Key takeaways
  • If only one person understands a system, losing them is a bigger risk than any bug in it.
  • The code shows what the system does. Often only that person knows why.
  • Before they go, capture the access, the map, the reasons behind the strange parts, the runbook and who depends on it.
  • Read the code first, so their remaining time goes on the questions only they can answer.
01

Why one person is a bigger risk than it looks

The code tells you what a system does. It rarely tells you why. Why does one customer get a different rounding rule? Why does the export run at 2:40 a.m. instead of 2:00? Why is there a table called `temp_fix_2014` that three reports still read from?

Somebody knew the answer when that code went in, and often the only person who still knows is the one everyone asks. When they leave, you keep the code and lose the reasons. The next change to that part of the system becomes a guess, and before long the team stops touching it at all. Plenty of systems end up frozen this way. Nobody decided it. Changing them just got too risky.

That person also carries knowledge that isn't in the code anywhere: which vendor contact actually answers the phone, what to do when the nightly job fails, which server has to come back up first.

We've argued before that before you rewrite a legacy system, you should read the code, because it's the only complete record of what the system does. This is the other half of that problem. The code records what. The why lives in people, and people leave.

02

Signs you have a bus factor of one

You probably know already. These are the usual tells:

- One name shows up on almost every incident ticket for the system.

- Changes sit in a queue until that person has time.

- Nobody else has ever deployed it, restarted it or recovered it after a failure.

- Releases get scheduled around their holidays.

- When you ask what a particular job or table is for, the answer is "check with them."

- They hold logins, keys or vendor contacts that nobody else has.

If several of those sound familiar, the risk is real, and it grows every month that nothing changes.

03

What to capture before they go

Start with whatever would hurt most if it disappeared tomorrow.

04

The access

Accounts, passwords, license keys, certificates and their renewal dates, vendor support contracts, and the contacts who actually respond. It's the easiest thing on this list to collect and the most embarrassing one to lose, so move it into a shared vault first.

05

The map

How the system is put together: the applications, the databases, the scheduled jobs, the file transfers, and what calls what. It doesn't need to be pretty. A rough diagram that's correct is worth more than a polished one from 2016.

06

The reasons behind the strange parts

Every old system has code that looks wrong but is there on purpose. Go through the special cases, hard-coded values and odd schedules, and write the reason for each one down next to it. This is the hardest knowledge to recover later, because nothing else records it.

07

The runbook

How to deploy a change. How to restart things, and in what order. What to do when the nightly job fails, and how to tell whether the rerun worked. Have someone else do each of these while the expert watches, rather than the other way round. If the new person can't do it unaided, it hasn't been captured yet.

08

Who depends on it

The reports, exports and spreadsheets that other teams pull from the system, especially the unofficial ones. The expert usually knows who will call when something breaks. Get those names, and find out what each of them uses.

09

Why an exit interview won't cover it

The usual plan is to book a few long sessions and ask the expert to explain the system. That helps, but it misses a lot. People describe a system the way they think about it, which is a summary, and the parts they stopped noticing years ago are exactly the parts nobody else knows. They can also only answer the questions you think to ask.

It works better the other way round. Read the code and the architecture first, make a list of everything unclear or surprising, and take those specific questions to the expert. An hour on "why does this job delete rows older than 91 days?" will teach you more than a day of "tell us about the system."

Timing matters too. A retirement usually comes with months of warning. A resignation might come with a couple of weeks. Plan for the short version, because that's the one that leaves you exposed.

10

Where an architecture and code review fits

An architecture and code review is the reading half of that plan, done for you. The review runs 30 days at a fixed price agreed before we start. We go through your codebase and architecture, map how the system fits together, find the hidden architectural debt, and flag the security and compliance gaps. The report says, in plain language, what is broken and what it would take to fix it.

Done while your expert is still around, it means their remaining time goes on the questions the code can't answer instead of on explaining the basics.

The report is yours to keep. Your team can use it to take over the system, and a new hire can start from it instead of from nothing. If you later want us to do the upgrade work, that's a separate project with its own scope and price.

11

Common questions

What does bus factor mean?

It's the number of people who would have to leave suddenly before nobody could keep a system or project running. A bus factor of one means a single departure stops the work. Some teams call it the truck factor.

Can't we just ask them to write documentation?

You should, but don't rely on it alone. People write down what they think is important, and the gaps are the things they don't think to mention. Documentation written in someone's last two weeks also tends to be rushed. Pair it with somebody reading the code and asking questions.

How long does knowledge transfer take?

For any system that matters, longer than a notice period. Start with the access and the runbook, which are quick to collect, then work through the map and the reasons behind the strange parts.

Should we hire a replacement first?

If you can find one, yes. But a new hire can't absorb years of history in a few weeks of overlap. Give them a map and a list of known problems, and they'll be useful much sooner.

How long does an architecture and code review take?

Thirty days, at a fixed price we agree before we start.

How do we start?

Book a call. Tell us about the system and who knows it, and we'll tell you plainly whether a review is the right next step. If it isn't, we'll say so, and the call costs you nothing.