Skip to main content

14 articles

Platform Reliability & FinOps

Your CI watchdog can't live on the pool it watches

A self-hosted runner pool was down for 6 hours 38 minutes. Its watchdog ran on the same pool. What went wrong, and how to monitor CI from outside.

App Stabilization & Rescue

What we find in apps built fast

The failure classes a stabilization assessment keeps finding in apps built quickly, including with AI builders, and how to check your own.

Platform Reliability & FinOps

What breaks when you move a GitHub repo

Moving repositories between GitHub organizations without breaking CI/CD: the hidden dependencies, and the order to change them in.

Platform Reliability & FinOps

Two SQL Server incidents where nothing crashed

Two SQL Server on Azure VM incidents where the engine never stopped: a backup retry storm, and log backups that quietly ended for seven databases.

Cloud Security & Compliance

Key-person risk after an engineer leaves

A continuity checklist for when the one engineer who understood your platform leaves: secrets, state, credentials, access and ground truth.

Applied AI Engineering

Fail-safe tools for a voice AI receptionist

Lessons from building our own AI phone receptionist: tools that degrade safely, caller ID matching, conversation tuning, and live-call verification.

Applied AI Engineering

Cutting an agent platform's LLM spend by 79%

How a multi-agent platform's metered model spend fell from about $107 to $22 a day: attribution, three cache layers, model routing and loop caps.

Subscribe by RSS

Not sure where to start? Start with an assessment.

A senior review of your app, cloud estate, or AI platform, scoped and quoted before work starts, that ends in a prioritized plan, so you decide what to fix and when.

Talk to an engineer