DR workarounds that never merge are defects with a timer
By Matthew Gray on Oct 6, 2026
A failover exercise we ran for a client hit three defects, and any one of them would have stopped a real recovery. All three had already been found in the previous exercise fifteen months earlier, worked around on a tag that sat on no branch, and never fixed. A workaround that does not merge is not a fix: it is a known defect waiting for the next disaster.
What the exercise was
The estate runs in East US 2, with Central US as the failover region. The failover plan is infrastructure-as-code: the same Terraform that describes production is pointed at the second region and applied, workspace by workspace, to rebuild the landing zone, the network and five application workloads at production sizes. Data comes back from Azure Backup through cross-region restore.
We repeated the previous exercise against the current code. The networking and governance layers applied cleanly. The application layer failed four times on the first day before it went through.
Why imported resources hide failover bugs
The root cause of all three failures was the same, and it applies to most estates that moved into Terraform after they were built.
Production resources had been imported into Terraform, not created by it. An imported resource only ever goes through the read and update paths. Terraform compares the configuration with what exists, finds nothing to change, and moves on. The create path, meaning the exact API calls Azure would receive when building the resource from nothing, had never run in that region.
A failover build is nothing but create calls. So the first time that code really ran was during the exercise. If it had been a real incident, the first time would have been the disaster.
A clean terraform plan against production tells you the configuration matches what exists. It tells you nothing about whether the same configuration can build that estate from nothing.
The three defects
| # | Defect | Azure’s response | The old workaround |
|---|---|---|---|
| 1 | UltraSSD-only performance attributes set on Premium_LRS managed disks | Rejected at create time: those attributes apply only to UltraSSD | Comment out the whole disk resource |
| 2 | Boot-diagnostics storage URIs hard-coded to accounts in the primary region | StorageAccountLocationMismatch: diagnostics storage must be in the VM’s region | Set every URI to null |
| 3 | A standalone managed disk with the same name as the disk the VM’s os_disk block creates | ConflictingUserInput: the disk already exists and can only be attached | Comment out the whole disk resource |
Each one was invisible in the primary region for the same reason. The disks with Ultra-only attributes had been imported, so Azure never validated those attributes. The diagnostics accounts happened to sit in the same region as the VMs. And the duplicate disk name never clashed, because both resources already existed when they were imported.
Why the workarounds never merged
The workarounds from fifteen months earlier were unconditional, and that is probably why they stayed on a tag. Setting the boot-diagnostics URIs to null works in the failover region. Merged to the main branch, it would have removed production’s dedicated diagnostics accounts on the next apply. Commenting out the disk resource would have had the same kind of effect. Nobody merges a change that breaks production in order to fix a path that is only used in a disaster, so the fixes stayed on a tag, and the defects stayed in the main configuration.
The fix is to make each change conditional, so that production plans “No changes” and the failover path still works:
boot_diagnostics {
# Primary region keeps its dedicated account; anywhere else uses managed boot diagnostics
storage_account_uri = var.location == var.primary_location ? var.diag_storage_uri : null
}
resource "azurerm_managed_disk" "vm_disk" {
count = var.manage_adopted ? 1 : 0
# ...
}
The Ultra-only attributes were simply removed, because they never applied to Premium disks anyway. The standalone disk was gated behind an “adopted resources” flag that is true only for the imported production estate. The VM’s os_disk now builds its name from the same expression, so it no longer refers to the gated resource. Before each fix merged, we confirmed that every production workspace still planned “No changes” and kept its resource count.
Proving the fixes were fixes
A defect you got past on the day is still a workaround. To show these were fixed, we tore the failover environment down and ran a second, clean failover the next day from the merged code. All three workspaces applied on the first attempt, including the application layer that had failed four times the day before. Both failover environments were then destroyed with no errors, and Central US went back to its baseline from before the exercise.
To be fair about the record: one failure on the first day was not a real defect. It was a race between deleting orphaned disks and creating VMs that needed those names. It happened only because we were applying fixes over and over to a half-built environment, and it cannot happen in a clean-slate failover. We logged it and did not count it as a finding.
The backup half worked, including for a VM that no longer existed
The second half of the recovery plan is Azure Backup with cross-region restore. We restored a VM that had been decommissioned from production six months earlier. It no longer existed in the primary region, and its backup had been replicated to the secondary region. The restore took 21 minutes from an application-consistent recovery point, and the VM came up running in Central US.
Restoring a system that is completely gone is a stronger test than restoring one that still exists, because nothing is left to fall back on.
What else the exercise surfaced
- Privileged access expired mid-recovery, twice. Just-in-time elevation lapsed between the Terraform build and the backup restore, and the cross-region restore failed on a permission check. Elevation also had to be at subscription scope. A tenant-level Global Administrator role was not enough. The previous exercise ran in one sitting, so it never ran into this. We recommended a standing break-glass DR role, or at least a step in the runbook to re-activate elevation.
- Private IPs are not preserved. Allocation was dynamic, so failover addresses differed from production. Anything with a hard-coded peer address, such as firewall route tables, needs checking during a real cutover.
- The failover estate drifts from Terraform immediately. Within minutes, Microsoft Defender auto-provisioned an extension onto all five VMs. The resource count went from 33 to 38, and those resources sit outside Terraform’s state.
Checks for your own estate
- Find out which production resources were imported rather than created. Their create path is untested in every region.
- Search the repository for DR, failover or exercise tags and branches that never merged. Each one is a list of known defects.
- Run the failover build into an empty region at least once, then run it a second time from merged code.
- Make region-specific fixes conditional, and require production to plan “No changes” before they merge.
- Restore something that has been decommissioned, not just something that is still running.
- Time how long your just-in-time elevation lasts against how long a recovery actually takes.
How we can help
We run DR exercises against the real estate and the real code, and we write up what failed as well as what worked. If your failover plan has only ever been tested on paper or against imported resources, get in touch and we will help you find out whether it builds.