<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Blog on Hat Boy Software, Inc.</title><link>https://hatboysoftware.com/blog/</link><description>Recent content in Blog on Hat Boy Software, Inc.</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><lastBuildDate>Tue, 06 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://hatboysoftware.com/blog/index.xml" rel="self" type="application/rss+xml"/><item><title>Caging an autonomous AI agent inside an enterprise tenant</title><link>https://hatboysoftware.com/blog/caging-autonomous-ai-agents/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/caging-autonomous-ai-agents/</guid><description>&lt;p>Autonomous agents that read mail, post in Teams and edit SharePoint are arriving in enterprise tenants faster than the controls around them, and each one is a new non-human user that takes instructions from any text it reads. Treat it like an untrusted contractor with admin tools: contain it, log it and make it stoppable, and design those controls in before it is switched on.&lt;/p></description></item><item><title>Cutting an agent platform's LLM spend by 79%</title><link>https://hatboysoftware.com/blog/cutting-llm-spend-79-percent/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/cutting-llm-spend-79-percent/</guid><description>&lt;p>A multi-agent software development platform was spending about $107 a day on model calls before it had any product traffic, and no one could say which agent was spending it. Multiplied across the tenants it planned to run, that was not a viable cost base. The fix was attribution first, then caching, routing and hard caps, each verified against a full day of real traffic rather than a projection.&lt;/p></description></item><item><title>Cutting Microsoft Sentinel cost without blinding your detections</title><link>https://hatboysoftware.com/blog/sentinel-cost-without-blinding-detections/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/sentinel-cost-without-blinding-detections/</guid><description>&lt;p>The usual levers for a Microsoft Sentinel bill (cheaper table plans, ingest-time filtering, moving old data out of the workspace) can each do exactly what the change record says while quietly breaking something else. For a regional healthcare group, a cheaper table plan had left 33 of the workspace&amp;rsquo;s 83 enabled detections running against no data, and nothing reported an error. Measure detection coverage and the whole bill after every change, not only the setting you changed.&lt;/p></description></item><item><title>DR workarounds that never merge are defects with a timer</title><link>https://hatboysoftware.com/blog/dr-workarounds-must-merge/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/dr-workarounds-must-merge/</guid><description>&lt;p>A failover exercise we ran for a client hit three defects, and any one of them would have stopped a real recovery. All three had been found in the previous exercise fifteen months earlier and worked around on a tag that sat on no branch, so the main configuration never changed. A workaround that does not merge is not a fix. It is a known defect waiting for the next disaster.&lt;/p></description></item><item><title>Fail-safe tools for a voice AI receptionist</title><link>https://hatboysoftware.com/blog/fail-safe-tools-for-voice-ai-agents/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/fail-safe-tools-for-voice-ai-agents/</guid><description>&lt;p>A phone receptionist that drops a caller loses a customer, and a voice agent&amp;rsquo;s tools fail in ways a chat agent&amp;rsquo;s don&amp;rsquo;t: the caller hangs up mid-sentence, a setting is missing, or the phone network refuses a transfer while someone is listening. Building our own AI receptionist taught us that every tool needs a safe fallback that still captures the caller, and that telephony features need proof on live calls, not just passing unit tests.&lt;/p></description></item><item><title>Key-person risk after an engineer leaves</title><link>https://hatboysoftware.com/blog/key-person-risk-after-an-engineer-leaves/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/key-person-risk-after-an-engineer-leaves/</guid><description>&lt;p>When the one engineer who understood a platform leaves, the risk is rarely the code. It is the operational knowledge and access that lived with that person, and it shows up as outages nobody can fix, launches that stall, and access nobody can account for in an audit. The lesson: inventory secrets, state, credentials and access before you need them, and check every finding against the live cloud.&lt;/p></description></item><item><title>One intake pipeline for our web form and AI receptionist</title><link>https://hatboysoftware.com/blog/web-form-and-ai-receptionist-into-omnisnia/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/web-form-and-ai-receptionist-into-omnisnia/</guid><description>&lt;p>People reach Hat Boy Software in two ways: the contact form on this site, and a phone line answered by our AI receptionist. Both now post to the same public web-to-lead endpoint in Omnisnia, the CRM built by Nandeshou, so every enquiry lands in one queue, tagged with the channel it came from.&lt;/p></description></item><item><title>Support tickets from your apps, straight into the CRM</title><link>https://hatboysoftware.com/blog/support-tickets-from-your-apps-into-omnisnia/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/support-tickets-from-your-apps-into-omnisnia/</guid><description>&lt;p>When a customer reports a problem from inside your app, the ticket should land where your team already works, next to the customer&amp;rsquo;s record, not in a shared inbox someone checks twice a day. Omnisnia, the CRM built by Nandeshou, does this with support tokens: one plain HTTPS POST from any app, website, backend job or phone agent opens a support case.&lt;/p></description></item><item><title>The logging feedback loop that more than quadrupled an AWS bill</title><link>https://hatboysoftware.com/blog/aws-logging-feedback-loop/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/aws-logging-feedback-loop/</guid><description>&lt;p>A client&amp;rsquo;s monthly AWS bill went from roughly $1.7K to $7.8K in a single month, with no meaningful change in application usage. Three security and logging features, each reasonable on its own, had been connected so that each one&amp;rsquo;s output became input for the others. Logging has to be designed as a system, because a single console change can turn it into a loop that pays for its own growth.&lt;/p></description></item><item><title>Two SQL Server incidents where nothing crashed</title><link>https://hatboysoftware.com/blog/sql-down-but-never-crashed/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/sql-down-but-never-crashed/</guid><description>&lt;p>Twice on the same SQL Server estate on Azure VMs, the database engine kept running while something around it failed. In the first incident, customers lost access for 71 minutes, most likely because a backup job kept retrying against a storage credential that did not exist. In the second, seven databases went weeks without a full backup, and nothing alerted until their log backups started failing, by which point point-in-time recovery was already being lost. In both cases &amp;ldquo;is SQL Server running?&amp;rdquo; was the wrong question.&lt;/p></description></item><item><title>What breaks when you move a GitHub repo</title><link>https://hatboysoftware.com/blog/what-breaks-when-you-move-a-github-repo/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/what-breaks-when-you-move-a-github-repo/</guid><description>&lt;p>Transferring a repository to another GitHub organization looks like a single button, and for history, issues and pull requests it is. What breaks is everything outside the repository that refers to it by owner and name: cloud logins, tokens, infrastructure code and deployment controllers, and any one of them can stop deploys on transfer day. The fix is to find those references first and change them in the right order.&lt;/p></description></item><item><title>What we find in apps built fast</title><link>https://hatboysoftware.com/blog/what-we-find-in-apps-built-fast/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/what-we-find-in-apps-built-fast/</guid><description>&lt;p>Apps built quickly, whether by a small team or with an AI builder such as Bolt, Lovable or Cursor, tend to fail in the same few places once real users, data and money arrive: one customer seeing another&amp;rsquo;s data, duplicate payments, and edits that silently vanish. Most of these can be found in a short, focused review and fixed at the root, which is why the usual answer is a stabilization pass, not a rebuild.&lt;/p></description></item><item><title>Your Azure OpenAI private endpoint may not be carrying your traffic</title><link>https://hatboysoftware.com/blog/azure-openai-private-endpoint-public-path/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/azure-openai-private-endpoint-public-path/</guid><description>&lt;p>A regional healthcare group&amp;rsquo;s architecture said its Azure OpenAI traffic used private endpoints, and its audit evidence relied on that. In practice every call from its application networks resolved to a public IP and went through the public endpoint. A private endpoint only carries traffic if DNS sends clients to it, and two separate DNS problems meant it did not.&lt;/p></description></item><item><title>Your CI watchdog can't live on the pool it watches</title><link>https://hatboysoftware.com/blog/ci-watchdog-on-the-pool-it-watches/</link><pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate><guid>https://hatboysoftware.com/blog/ci-watchdog-on-the-pool-it-watches/</guid><description>&lt;p>A self-hosted CI runner pool on a client&amp;rsquo;s logistics platform had zero runners online for 6 hours 38 minutes, and nothing reported it. Every deploy, the infrastructure apply and the restore drill depended on that pool, and so did the watchdog built to catch exactly this, so it waited in the queue behind the outage it was meant to detect. A health check that shares a failure domain with the thing it checks only works when you don&amp;rsquo;t need it.&lt;/p></description></item></channel></rss>