Platform Reliability & FinOpsYour load balancer health probe may be lying
Three Azure health probe failures on one production estate, and what a probe should check so that it can tell a healthy backend from a broken one.
Notes from the field on cloud, security, and AI engineering.
29 articles
Platform Reliability & FinOpsThree Azure health probe failures on one production estate, and what a probe should check so that it can tell a healthy backend from a broken one.
App Stabilization & RescueHugo's default code theme failed WCAG AA contrast on 10 of our 12 posts with code. How we tested nine themes with axe-core, and the parser trap we found.
Applied AI EngineeringAzure Communication Services retires as a standalone offering in 2028, and new phone numbers are already restricted. What it means for voice AI agents.
Cloud Security & ComplianceHSM-backed describes who holds the key, not how strong the encryption is. How to state Azure encryption-at-rest coverage per resource type, with evidence.
App Stabilization & RescueHow to keep pre-silicon firmware honest: fail-closed stubs, a sanitized host test harness, marked bench seams and a protocol simulator.
Applied AI EngineeringHow Baseline caps LLM spend per collection, tenant and org, makes BigQuery partition pruning mandatory, and why projected savings need measuring.
Cloud Security & ComplianceHow to make an Azure Custom Script Extension safe to run against VMs that already carry an EDR agent, and keep the site token out of the repository.
App Stabilization & RescueShovel Solutions' strategy for retiring an insecure legacy app: a live switch, delta loads, paced invites, a read-only legacy database and rehearsed restores.
App Stabilization & RescueA post-mortem: a template-encoding bug sent every website inquiry to a dead URL for six days. How we found it, sized the damage and stopped it recurring.
Platform Reliability & FinOpsWhy remote students can't hear the teacher, why most studio platforms run on Zoom, and why a small studio should rent before it builds.
Platform Reliability & FinOpsSince July 15, 2026, renaming or transferring a GitHub repo changes its OIDC subject claim. How to update Azure, AWS and Google Cloud trust before the move.
Cloud Security & ComplianceHow Omnisnia enforces tenant isolation in GORM callbacks that refuse unscoped queries, with an explicit system scope, tests, and where RLS fits.
Platform Reliability & FinOpsHow we measure traffic, Core Web Vitals, conversions and failures on our own site with Application Insights and Front Door logs, without cookies.
Platform Reliability & FinOpsWhat we learned moving our site's redirects and cache headers into Azure Front Door Standard rules: status codes, rule order, path matching and limits.
Applied AI EngineeringHow we generate llms.txt and llms-full.txt from Hugo content with custom output formats, the gotchas we hit, and why we did it while it is only a proposal.
A self-hosted runner pool was down for 6 hours 38 minutes. Its watchdog ran on the same pool. What went wrong, and how to monitor CI from outside.
A private endpoint existing does not mean traffic uses it. Two DNS problems that kept Azure OpenAI calls on the public endpoint, and how to check yours.
The failure classes a stabilization assessment keeps finding in apps built quickly, including with AI builders, and how to check your own.
Moving repositories between GitHub organizations without breaking CI/CD: the hidden dependencies, and the order to change them in.
Two SQL Server on Azure VM incidents where the engine never stopped: a backup retry storm, and log backups that quietly ended for seven databases.
How S3 access logging, CloudTrail data events and GuardDuty malware scanning fed each other until one AWS account's bill went from about $1.7K to $7.8K.
Applied AI EngineeringHow a customer-facing app, a backend job and an AI receptionist open support cases in Omnisnia CRM with scoped, revocable support tokens.
Applied AI EngineeringHow we send website enquiries and AI receptionist calls into Omnisnia CRM through one public web-to-lead endpoint, and why a public token is safe there.
A continuity checklist for when the one engineer who understood your platform leaves: secrets, state, credentials, access and ground truth.
Lessons from building our own AI phone receptionist: tools that degrade safely, caller ID matching, conversation tuning, and live-call verification.
A failover exercise hit three defects that had been found and worked around 15 months earlier. Why imported Terraform hides them, and how to fix them.
Four lessons from reducing Sentinel and Log Analytics cost for a healthcare client, including the cheaper table plan that silently emptied 33 detections.
How a multi-agent platform's metered model spend fell from about $107 to $22 a day: attribution, three cache layers, model routing and loop caps.
The control pattern we design to before an autonomous AI agent touches a Microsoft 365 tenant: identity, network, policy, logging and a kill switch.
No articles match that search.
A senior review of your app, cloud estate, or AI platform, scoped and quoted before work starts, that ends in a prioritized plan, so you decide what to fix and when.
Talk to an engineer