Skip to main content

SQL Server was down, but it never crashed


Twice on the same SQL Server estate on Azure VMs, the database engine kept running while the service around it failed. In the first incident, customers lost access for 71 minutes because a backup job was retrying against a credential that did not exist. In the second, seven databases went without full backups, and nothing alerted until their transaction-log backups stopped and point-in-time recovery was already lost. In both cases “is the service running?” was the wrong check.

The client is a regulated financial-services SaaS platform. Its production SQL Server runs on Azure VMs in Always On availability groups, protected by Azure Backup. Both incidents are summarised in our case study. This post covers the mechanisms.

Incident 1: a 71-minute outage with no crash

Customers reported that the database was unreachable for 71 minutes early one morning. The obvious explanations were all ruled out:

  • No service stop or start events, and no SQL Server startup messages.
  • No availability-group failover. The primary node owned every listener and IP resource throughout.
  • The private endpoint to the backup storage account resolved correctly, and port 443 was reachable.

The Application event log had the real story. Every few seconds the same three errors repeated: 18210 (BackupIoRequest write failure), 3041 (BACKUP LOG failed), and 3634 (the operating system returned error 13, “The data is invalid”, on DeleteFile against a blob URL). The SQL VM IaaS extension was issuing BACKUP LOG ... TO URL against an Azure storage account in a tight retry loop.

Then we listed the credentials. sys.credentials and the database-scoped credential views in master and msdb returned zero rows. SQL Server had no SAS or managed-identity credential for any URL. Every log backup failed at authentication, and the extension tried again at once. Our working explanation is that each attempt briefly tied up the backup machinery, and at that retry rate client connections appeared to hang. The engine stayed up.

Both nodes showed the same pattern. That pointed to configuration, not a sick host. Meanwhile the system-database full backups to a local target kept succeeding every four hours. The VM had two backup pipelines, one working and one broken, and only the broken one never reached msdb history.

One limit on our evidence: the error pattern was still in the log after customers’ connections recovered. The data does not explain why the outage ended when it did. We reported that openly and did not tidy it up.

The fix was to stop the retry loop, then reissue the credential. Managed identity was the preferred route, because it leaves no SAS to expire or get dropped. We also recommended an alert on those three event IDs against the backup URL. That night, customers noticed the problem first.

Incident 2: log backups that stopped without a clear failure

Months later, transaction-log backups had stopped for 7 of 30 protected databases across two availability groups. Point-in-time recovery for those seven ended at their last log backup. Four of them could no longer truncate their logs. The largest log file had reached 141.2 GB, on volumes with 12 to 14 percent free.

Three behaviours combined to cause this.

1. Azure Backup will not extend an old log chain. If the latest full backup behind a log chain is more than 15 days old, Azure Backup refuses the next log backup. It first tries a “remedial” full backup to re-base the chain.

2. The remedial full can only run on the primary. For databases in an availability group, Azure Backup takes full backups only on the primary replica. All the groups were set to AUTOMATED_BACKUP_PREFERENCE = SECONDARY, so log backups ran on a secondary, where the remedial full is not allowed. The error says so directly: UserErrorLatestFullPitIsOldRemedialFullNotPossibleOnNode.

3. Why the fulls were missing: concurrent configuration jobs. When protection was first configured in bulk, 15 ConfigureBackup jobs failed with “another configure protection operation is in progress for this item”. It was a race within the batch, and nobody retried them. Those items still showed as protected under the hourly policy and still received log backups, but no daily full was ever scheduled for them.

The signature was in the job counts: the vault took exactly 23 full backups a day against 30 protected items. On the primary, msdb showed hundreds of fulls for one instance and none at all for the two affected ones.

An unplanned failover hid the problem for a while. It put the node that took log backups into the primary role, the remedial fulls succeeded, and the clock reset. Fifteen days after that, the failures came back. A second, unrelated change then blocked the repair. A dormant default SQL instance had been disabled, and the Azure Backup plugin’s configure path always connects to that instance before doing anything. We found this by reproducing the failure on a single item before touching the other six.

The repair was to rediscover the workloads, re-apply the existing policy one item at a time (each waiting for its configure job to finish, because concurrency caused the original failure), and take an on-demand full backup of each database. The largest was about 1.7 TB and took 2 hours 31 minutes. No failover, no instance restart, no policy change. Log truncation resumed, and a follow-up shrink of three log files moved the log volumes from 11–13 percent free to 27–28 percent.

What the repair could not do: about three days of point-in-time recovery for those seven databases is permanently gone. We documented that as a data-loss statement and did not leave it out.

Why the alerting missed it

The existing alert fired on failed backup jobs. For 15 days this defect produced no failed jobs. It produced no full-backup jobs at all. Once the chain broke, failures arrived quickly: about 168 failed log jobs a day and 50 Sev1 alerts in a week. By then the recovery gap had already started.

The lesson is to alert on absence. “A protected item with no successful full backup in more than two days” would have caught this the day after the batch failed.

Checks you can run today

On each SQL Server that backs up to URL, confirm a credential exists:

SELECT name, credential_identity FROM sys.credentials;

Check the latest full and log backup for every database. A database whose newest full is approaching 15 days old needs attention:

SELECT d.name,
       MAX(CASE WHEN b.type = 'D' THEN b.backup_finish_date END) AS last_full,
       MAX(CASE WHEN b.type = 'L' THEN b.backup_finish_date END) AS last_log
FROM sys.databases d
LEFT JOIN msdb.dbo.backupset b ON b.database_name = d.name
GROUP BY d.name;

Run that on every replica. In an availability group the history is split across nodes. Also find logs that are waiting on a backup, and check where backups are preferred:

SELECT name, log_reuse_wait_desc FROM sys.databases WHERE log_reuse_wait_desc = 'LOG_BACKUP';
SELECT name, automated_backup_preference_desc FROM sys.availability_groups;

Then, outside SQL Server:

  • In the vault, compare the number of full backups taken per day with the number of protected items. They should match.
  • Look for ConfigureBackup jobs that failed and were never retried.
  • Alert on the event IDs 18210, 3041 and 3634 together.
  • If you disable a default instance that looks unused, find out first whether a backup agent depends on it.

How we can help

Both incidents were solved from evidence the estate already held: event logs, msdb history and vault job counts. Collecting that evidence is most of the work. If your SQL estate has backups you assume are working, talk to us and we will measure them.

Back to blog

Not sure where to start? Start with an assessment.

A senior review of your app, cloud estate, or AI platform, scoped and quoted before work starts, that ends in a prioritized plan, so you decide what to fix and when.

Talk to an engineer