Skip to main content
52
findings — 9 critical, 11 high
25
resolved in about two weeks
1 → 27
VMs actually sending monitoring data
$470K+
three-year savings case identified

Client: A regulated data-exchange SaaS platform serving large financial institutions, running on Azure across six subscriptions in a hub-and-spoke network. About 1,187 resources and roughly $45K a month in spend — an estate that had grown without ever being fully assessed.

The challenge

Nothing about the estate’s security, resilience, or cost was documented. A read-only baseline audit surfaced 52 findings — 9 critical, 11 high, 16 medium, 16 low — including:

  • Observability gaps hiding in plain sight. No activity-log forwarding, monitoring agents installed on 25 VMs but only one actually sending data, and network flow logging off everywhere.
  • Identity risk. Privileged roles permanently active, too many global administrators, no true break-glass account, and a privileged automation credential that had already expired.
  • Resilience gaps. Production SQL on a single node with no cross-region disaster recovery, and about 17 TiB of live customer files with only 3 of 24 shares backed up.
  • Cost waste. A $3,365-a-month charge for a resource that had been deleted, and a SQL fleet averaging about 5% CPU on 48 provisioned cores.

What we did

We delivered the audit, then the remediation, then the forward architecture — all through a peer-reviewed, Infrastructure-as-Code-gated workflow.

  • Restored observability: activity-log forwarding on every subscription, monitoring data from 1 VM to 27, and flow logs on everywhere.
  • Hardened identity: real break-glass accounts, and every automation identity moved to keyless OIDC federation — no more long-lived or expired secrets.
  • Rebuilt resilience: VM backup coverage from 25/27 to 31/31, protected file shares from 3 to 36, and production SQL extended to a two-node cluster with distributed availability groups and a cross-region disaster-recovery replica in a new second region.
  • Encrypted to standard: 100% of disks on HSM-backed RSA-4096 customer-managed keys with automatic rotation.
  • Attacked the cost: about $17K a year of idle infrastructure removed, the phantom $3,365-a-month charge surfaced, and a database-platform migration proposed with projected savings of $470K–$540K over three years.

Along the way, a 71-minute customer-facing SQL “outage” — in which the database engine never actually stopped — was traced forensically to a single missing storage credential.

Results

OutcomeResult
Security findings25 of 52 resolved in about two weeks
IdentityKeyless OIDC federation replaced all long-lived and expired privileged secrets
ResilienceSecond Azure region with distributed availability groups; backups 25/27 → 31/31; file shares 3 → 36
Cost~$17K/year removed; $3,365/month phantom charge surfaced; $470K–$540K three-year case proposed
Platform28-module Infrastructure-as-Code registry with an explicit-confirm change workflow

The client's identity is anonymized. The environment, findings, and results are drawn from a real Hat Boy Software engagement; figures are representative and rounded.

More case studies

Not sure where to start? Start with an assessment.

A senior review of your app, cloud estate, or AI platform, scoped and quoted before work starts, that ends in a prioritized plan, so you decide what to fix and when.

Talk to an engineer