Your CI watchdog can't live on the pool it watches
By Matthew Gray on Oct 6, 2026
A self-hosted CI runner pool on a logistics platform we operate had zero runners online for 6 hours 38 minutes, and nothing reported it. The watchdog built to catch exactly that ran on the same pool, so it waited in the queue behind the outage it was supposed to detect. A health check that shares a failure domain with the thing it checks only works when you don’t need it.
What happened
Every runner in the pool went offline: 0 of 5. At the peak, 28 workflow runs were stuck in the queue, including an Android release and an iOS release. All 57 self-hosted jobs in the repository were pinned to that one Azure virtual machine scale set. That included the deploys, the infrastructure apply and the restore drill.
The pool had two automated controls: a watchdog workflow that checks runner health, and a reimage workflow that repairs bad instances. Both declared the same runs-on labels as every other job:
runs-on: [self-hosted, linux, azure]
With one sick instance, that works. When the pool is empty, the detector and the fix both sit in the queue next to the deploys they exist to protect.
We found this during a scheduled audit pass. The census that measures the estate also runs in CI, so the pass started with the estate unmeasurable. That is a second version of the same mistake.
How the evidence erased itself
The watchdog’s scheduled runs did more than queue. They cancelled each other.
The watchdog workflow set cancel-in-progress: true on its concurrency group. Each hourly scheduled run cancelled the previous one, which was still waiting for a runner. The reimage workflow correctly set cancel-in-progress: false, because two reimages at once would be dangerous. But GitHub Actions allows only one pending run per concurrency group. When a third run arrives, the pending one in the middle is cancelled anyway.
In the 12 hours covered by the last 100 runs, the record read:
| Result | Runs |
|---|---|
| Success | 37 |
| Cancelled | 30 |
| Queued | 28 |
| Failure | 3 |
Thirty cancellations look like someone tidying up. A person reading the run history the next morning would see no alerts, very few failures and some cancelled runs, which is exactly what a human pressing “cancel” looks like. The record of the outage was being deleted while the outage was still running.
When the pool healed, the new runners weren’t equivalent
The scale set recovered on its own. Four new instances registered, and at least one came up without the Azure CLI. It reported online and took work. Everything that did not touch Azure ran fine on it. Everything that did failed with:
Unable to locate executable file: az
Across the last 60 runs, that one instance had six failures. Every other instance combined had one. The failed jobs included a backend deploy, the estate census, and two that matter most here: the pool’s own watchdog and its reimage workflow. Hours after the pool caused an outage, an unhealthy member of the pool knocked out the controls that keep it healthy.
The cause was in the provisioning template. The cloud-init file installed the CLI before registering the runner, so the order was right. But the install was best effort:
- [ bash, -lc, "command -v az >/dev/null || curl -sSL <installer-url> | bash" ]
A failed download did not stop cloud-init, and the registration script ran next anyway. Two lines above, a comment in the same file called the Azure CLI “a hard requirement, not a convenience”. The comment was right, but the code treated the CLI as optional.
The error also blames the wrong thing. “Unable to locate executable file” looks like a PATH problem in the workflow, so the person who wrote the workflow ends up debugging a broken machine.
What to change
Take the health check out of the failure domain
At minimum, whatever acts on total pool loss must run somewhere that does not depend on the pool. The first change we proposed was to move the reimage workflow onto a hosted runner. Options, roughly in order of independence:
- A GitHub-hosted runner for the watchdog and the repair job. They need little compute and should almost never run on your own pool.
- A separate schedule outside CI, such as an Azure Function, Logic App or cron job on unrelated infrastructure, that calls the GitHub API for runner status.
- An external probe that alerts through a channel that does not depend on CI at all.
Make concurrency fit the job
For a watchdog, cancelling a pending run in favour of a newer one sounds harmless, but it destroys the evidence. Prefer a concurrency group with cancel-in-progress: false and a short timeout-minutes, so stale runs fail visibly and are not silently cancelled. Remember the one-pending-run rule. Even with false, a backlog will be cancelled. That is one more reason the watchdog should not be queuing at all.
Bake and test runner images
Install tools into a versioned image, not at boot from the internet. If you must install at boot, fail closed:
set -euo pipefail
command -v az >/dev/null || install_azure_cli
az version >/dev/null # refuse to register if this fails
./register-runner.sh
Then run a smoke job against each new image before it joins the pool, with a check for every tool your workflows assume. A runner that cannot authenticate is worse than no runner. It looks like capacity, and its failures land on the workflow.
Alert on queue age, not just runner count
Runner count tells you whether machines exist. Queue age tells you whether work is moving, and that is what users actually feel. A job whose labels no online runner carries does not fail. It waits. On the same platform, a later release build sat in the queue for 58 minutes for that reason before someone cancelled it and dispatched it again.
Useful signals, all available from the GitHub API:
- The oldest queued run, alerted at a threshold that fits your pipeline.
- Online runners per label set, not just in total.
- The rate of cancelled runs for workflows that people rarely cancel by hand.
- The failure rate per runner. One instance with six failures against one everywhere else is a broken machine, not flaky tests.
Write down the path for when the pool is what broke
The restore drill on this platform had run and passed. What had never been tested was running it with no runners. Where a manual recovery document exists, exercise it at least once with CI deliberately unavailable. If your recovery objective assumes CI, CI is part of your disaster-recovery scope.
The general rule
A watchdog has to survive the failure it watches for. That applies to CI runners, to log pipelines that alert on missing logs, and to monitoring hosted on the cluster it monitors. Ask of each control: “If the thing it watches is completely gone, does this still run, and does anyone hear about it?”
How we can help
We run CI and platform operations for clients, and we audit them by measuring the live estate, not by reading the configuration. If your pipelines depend on a single runner pool, talk to us about where your watchdogs run and what they would miss.