A failover exercise we ran for a client hit three defects, and any one of them would have stopped a real recovery. All three had already been found in the previous exercise fifteen months earlier, worked around on a tag that sat on no branch, and never fixed. A workaround that does not merge is not a fix: it is a known defect waiting for the next disaster.
Read Moreazure-backup
Twice on the same SQL Server estate on Azure VMs, the database engine kept running while the service around it failed. In the first incident, customers lost access for 71 minutes because a backup job was retrying against a credential that did not exist. In the second, seven databases went without full backups, and nothing alerted until their transaction-log backups stopped and point-in-time recovery was already lost. In both cases “is the service running?” was the wrong check.
Read More