Skip to main content

Key-person risk after an engineer leaves


When the one engineer who understood a platform leaves, the risk is rarely the code. It is the operational knowledge and access that lived with that person: which keys unseal the secrets store, where the infrastructure state is, and which credential quietly expires next month.

We worked through this with an AI software company whose lead engineer had been the main author of 15 of its 22 repositories. The engagement produced a 28-item continuity and risk register, and its most valuable finding was almost anticlimactic: a finished product had not launched because one deploy credential had expired. The case study has the outcome. This post is the checklist we now use, written so you can run it on your own estate before you need it.

1. Who holds the secrets, and the keys to the secrets

Start with the secrets store itself. A self-hosted vault that must be unsealed after a restart is only as recoverable as its unseal keys and root token. If those were held by one person, the platform is one reboot away from an outage nobody can fix.

  • Find the unseal key shares and recovery keys, and confirm at least two people can reach them.
  • If they cannot be found while the vault is still running and unsealed, export what you need and stand up a re-keyed instance now, not after an unplanned restart.
  • Check cloud key vaults for purge protection and for single-copy keys, such as a licence-signing key whose loss would invalidate everything it signed.

2. Infrastructure state that lives on one laptop

Terraform and OpenTofu need their state file to manage existing resources safely. If no repository configures a remote backend, the state was almost certainly local, on the departed engineer’s machine.

Without it, an apply may try to recreate resources that already exist, which can take down DNS, mail or the site. Until state is recovered or rebuilt with imports, treat the live infrastructure as manually managed and do not apply. Then move state to a remote backend with locking, and run plan and apply from CI rather than a workstation.

3. Deploy credentials about to expire

Service-principal secrets and other pipeline credentials have expiry dates, and nobody notices until a deploy fails. That was the single credential in our case: the pipeline could no longer authenticate to the cloud, so a product that had been built was never shipped, and it looked like a product problem rather than an access problem.

  • List every service principal and app registration your CI uses, with its owner and expiry date.
  • Give each one a named owner. An ownerless service principal is one nobody will rotate.
  • Where the platform supports it, replace stored secrets with workload identity federation (OIDC), so there is no secret to expire or leak.

4. Committed secrets in history

Search full git history, not just the current tree, across every repository the person touched. Look for API keys, webhook and token-signing secrets, default client secrets in bootstrap code, committed SSH public keys, and build artifacts that were tracked by accident because a .gitignore was broken.

Rotate first, then purge. A committed SSH public key also needs a sweep of every authorized_keys file and cloud-init or user-data template, because removing it from the repository does nothing to the hosts that already trust it.

5. Access revocation and break-glass

Disabling the directory account is the start, not the end. Access that the identity provider does not control survives it:

  • source-control organization membership and personal access tokens
  • VPN client certificates and SSH keys
  • release-signing keys, which may need re-keying of a release chain
  • vendor dashboards for payments, email, DNS and registrars
  • hard-coded admin or owner identities in bootstrap code, which will provision the departed person as owner on the next fresh deployment
  • groups and resources whose only owner was the disabled account

Then check the other direction: make sure no single remaining person holds every admin role, add a second owner to each subscription, and confirm break-glass access works without any one individual.

6. Separate core IP from reproducible output and demos

An estate with dozens of repositories and thin documentation is hard to value and hard to protect. Classify each repository by what it is: defensible core IP, supporting plumbing, output that can be rebuilt from something else, stubs, or demos. Rank by value, maturity and authorship.

This tells you where to spend scarce attention. Core IP gets cross-training, consolidation of divergent copies, and backups of the inputs it depends on. Demos and stubs get archived so they stop cluttering the picture.

7. Rebuild documentation from the code

When the author is gone, the code is the only reliable record. Write the architecture and a short deep-dive per repository from what the code actually does, not from old design documents, which describe intent. Where they disagree, note it.

Then capture the tribal steps. Have a new engineer perform a witnessed end-to-end deploy of a non-production environment and write down every manual step. That one exercise surfaces most of what lived only in someone’s head.

8. Verify ground truth in the live cloud

Repository contents tell you what was intended. The live cloud tells you what exists. Run a read-only audit with a temporary, least-privilege identity, and remove that identity when the assessment closes.

It is also where we made our own mistake, and it is worth telling. Early in that engagement we reported that there was no production environment: the subscriptions named for production appeared empty. That was wrong. The audit identity had been granted access to the real production subscription after its last CLI sign-in, and the CLI was showing a stale, cached subscription list that contained only older, empty ones. Refreshing the list showed production running.

The lesson is broader than one CLI flag. Your tools show you a view, and the view can be stale, filtered or scoped to the wrong identity. Before you report that something does not exist:

# refresh the subscription list rather than trusting the cache
az account list --refresh --output table
  • Confirm which identity you are signed in as and what it can see.
  • Cross-check with a second source: billing, DNS, the CI pipeline’s target, or the people who run it.
  • Treat a finding of absence with more suspicion than a finding of presence.

We corrected the report and the risk register once we found it. Measured, not assumed, applies to our own findings too.

How we can help

We run continuity assessments that map the estate, build the risk register, and check every finding against the live cloud before it reaches a report. Read the case study, or see our services for how we scope this work.

Back to blog

Not sure where to start? Start with an assessment.

A senior review of your app, cloud estate, or AI platform, scoped and quoted before work starts, that ends in a prioritized plan, so you decide what to fix and when.

Talk to an engineer