Skip to main content

Our contact form silently dropped every inquiry for six days

A post-mortem: a template-encoding bug sent every website inquiry to a dead URL for six days. How we found it, sized the damage and stopped it recurring.

From the day our new contact form shipped until we noticed six days later, it did not deliver a single inquiry. Every visitor who used it saw an error asking them to email us instead, and nothing anywhere told us it was happening. This post covers the bug, how we sized the damage without request logs, and the checks that now stop it shipping or flag it within minutes.

What broke

The contact form on hatboysoftware.com shipped on September 30, 2026. It posts by fetch() straight to the public web-to-lead endpoint of Omnisnia, our CRM, so each submission becomes a lead. We described that design in one intake pipeline for our web form and AI receptionist.

The site is built with Hugo. The endpoint URL and the form’s status messages live in site configuration, and the template writes them into an inline <script>. We piped each value through Hugo’s jsonify function to turn it into a JavaScript string:

<script>
  var endpoint = {{ .formEndpoint | jsonify }};   // before: double-encoded
  var endpoint = {{ .formEndpoint }};             // after: one quoted literal
</script>

That was one encoding too many. Hugo templates use Go’s html/template, which is context-aware: a string inserted into a script is already escaped and wrapped in quotes as a JavaScript string literal. The Go documentation shows a plain string becoming a quoted literal in a JavaScript context. Adding jsonify on top produced a string whose value included the quote marks.

So the endpoint was no longer https://… but "https://…". A string that starts with a quote mark is not an absolute URL, so fetch() treated it as a relative path and resolved it against the page:

intended:   https://<crm-host>/api/v1/public/lead/<token>
requested:  https://hatboysoftware.com/contact/%22https://<crm-host>/...%22

Before the fix, the rendered endpoint string contained literal quote marks around the URL, so fetch() posted to a relative path under /contact/ and got a 405. After the fix, it is a plain absolute URL and the CRM answers 202 Accepted. The endpoint line as the browser received it, before and after the fix. The host path and token are elided.

Our site is static files behind Azure Front Door. A POST to a static path gets a 405 (Method Not Allowed). The form’s code did the right thing with that: it showed the error message, which asks the visitor to email us. The status messages had the same flaw, so they appeared on screen wrapped in literal quote marks.

The contact form filled in with sample details for Jordan Example, with a red error below the Send button that reads, inside literal quote marks, “Sorry — that did not send. Please email info@hatboysoftware.com and we will pick it up there.” The form as every visitor saw it, rebuilt from the code before the fix. Note the quote marks around the message.

How we found it

Not through monitoring, because there was none for this path. On October 6 we were taking screenshots of the form’s success message for a blog post and noticed the stray quote marks. Quote marks in a status message meant a double-encoded value, and the endpoint came from the same template, so we traced the request in the browser and found the relative URL.

The fix was the two-line change above, applied to every value in the script. It went out as release 1.4.0 the same day. A live test submission from the production site then returned 202 Accepted from the CRM.

The contact form after a successful submission: the fields are cleared and a green message below the Send button reads “Thanks — we have your note and will be in touch shortly.” with no quote marks. The same submission with the fix in place. The CRM response is simulated locally with a 202.

Sizing the damage without logs

The harder question was how many people we had lost. We had no request logs. Front Door diagnostic logs are not enabled by default, and we had not turned them on. The storage account behind the site had no logging either.

What we did have was Front Door’s built-in platform metrics, which Azure Monitor keeps for 93 days with no setup. The request-count metric can be split by HTTP status and by client country. That gave us a daily count of 405 responses going back well before the form existed.

The picture was noisy, but readable:

  • Before the form existed, the site still returned anywhere from single digits to about 400 405s a day. They came in tight batches of about 35 to 42, the pattern of automated scanners trying POSTs against everything.
  • After launch, we looked for 405s that did not fit that pattern: isolated single requests from the US, outside any batch.
  • Our estimate is at most about four real attempts over the six days, possibly fewer. It is an estimate: metrics carry no path, user agent or body, so we cannot prove any one of them was a person using the form.

Bar chart of daily HTTP 405 responses at Front Door from September 25 to October 6, 2026. Before the form launched on September 30 the site already returned between 2 and 410 a day; after launch the daily counts, 45 to 125, look no different, which is why real failed submissions could only be estimated. Daily 405s from Front Door platform metrics. Scanner traffic produced them before the form existed, so the form’s failures hide inside the noise.

We could not recover who those visitors were, because the submissions never reached any system we control. If you used our contact form between September 30 and October 6, 2026, your message did not reach us, and we’re sorry. Please email info@hatboysoftware.com or send the form again; it works now.

What we changed

We wanted every layer to catch this failure on its own: the build, the edge, and the browser.

A build check that tests the rendered page

A check script now runs on every pull request. It builds the site minified, as production does, opens the rendered contact page, and fails if the endpoint or any status message appears double-quoted, or does not appear as a plain string literal at all. We confirmed it fails against the old template before relying on it.

- name: Build
  run: hugo --minify -d build
- name: Verify contact form submits to the CRM
  run: python3 scripts/check_contact_form.py build config.toml

The point is that it inspects the build output, not the template. The template looked correct to everyone who read it.

Edge logs and an alert on failed form posts

Front Door now sends its access logs to a Log Analytics workspace. An alert fires on any failed POST to a path under /contact/ within a 15-minute window. A form that posts to our own site instead of the CRM is the signature of this exact bug, and it is a request the site should never receive from a real visitor.

Browser telemetry on every form outcome

The edge can’t see a failure between the visitor’s browser and the CRM, because that request never touches our site. So the form now reports its own outcome through Azure Application Insights: success, validation error, throttled, error, network failure, or the spam honeypot. It also records the CRM call’s HTTP status. A second alert fires on any error or network outcome, or a failed call to the CRM. The setup is cookie-free and loads after the page.

Lessons

Test the deployed artifact, not the template. The bug was invisible in source. It existed only in the rendered page, and only a check against built output could see it.

A UI that handles errors gracefully also hides them. The form did exactly what we designed it to do when a submission failed: it apologized and offered an alternative. That was right for the visitor and silent for us. Any path that ends in “we’ll be in touch” needs an end-to-end check that the message actually arrived.

Logging is cheap insurance you only miss when you need it. Access logs for a small site cost little, and their absence turned a simple question into a statistical estimate.

How we can help

We find and fix the failures in production systems that look like success: forms, integrations and background jobs that fail quietly. We add the build checks, logs and alerts that make them visible. If something in your stack matters this much and nobody is watching it, see our app stabilization service.

Related articles

Platform Reliability & FinOps

Your CI watchdog can't live on the pool it watches

A self-hosted runner pool was down for 6 hours 38 minutes. Its watchdog ran on the same pool. What went wrong, and how to monitor CI from outside.

Applied AI Engineering

Adding llms.txt and llms-full.txt to a Hugo site

How we generate llms.txt and llms-full.txt from Hugo content with custom output formats, the gotchas we hit, and why we did it while it is only a proposal.

← All articles

Not sure where to start? Start with an assessment.

A senior review of your app, cloud estate, or AI platform, scoped and quoted before work starts, that ends in a prioritized plan, so you decide what to fix and when.

Talk to an engineer