Firmware written before the boards exist tends to fill up with stubs that quietly report success, so the demo works and nobody can say how much of it is real. The cure is an honest inventory, stubs that fail closed, a host test harness that runs the real code under sanitizers, and a clear line between what is tested and what only the hardware can prove.
We recently went through this on a pre-production, battery-powered IoT device: an nRF52840 running Zephyr, an NFC reader, a small display, a Thread mesh back to an on-site edge server, and a cloud tier above that. The boards weren’t fabricated yet, nothing had shipped, and no Zephyr toolchain was available in the working environment. The question was the one any reviewer, buyer or new engineer eventually asks: how much of this firmware is real?
Stubs that report success are the real risk
A stub that crashes gets noticed. A stub that returns a plausible value doesn’t. Our inventory found several of the second kind:
- An NFC read that always returned “no tag.” From the server’s side, that looks the same as a working device with nothing in range.
- A battery gauge that always read 100%. The edge server could never flag a device about to go flat.
- A shared development authentication key compiled into every image. Left in place, anyone holding one firmware file could have signed messages as any device.
- A content download that returned success without downloading. It bumped the reported version, so the server believed new content was in place while the screen showed nothing.
None of these would fail a demo. Each would fail in the field, two of them silently.
Replace each stub with logic that fails closed
The rule we applied: every function either does the real work, or returns an answer that makes the caller take the safe path. “Not yet implemented” must never look like “worked.”
- NFC. The tag-selection protocol is now real ISO 14443A: wake-up, the anticollision cascade, select, assembly of 4-, 7- or 10-byte UIDs, and the check-byte validation. Only the reader chip’s transceive call sits behind a seam.
- Battery. The gauge converts the fuel-gauge chip’s state-of-charge register into a clamped percentage. On a bus error it holds the last good value, so a single bad read never reports a false empty, and it shows 100 only until the first real read.
- Authentication. No usable key ships in the image. Each unit’s key is to be read at runtime from provisioned storage on the chip (a bench seam, below), and signing and verification refuse to operate until a key is present.
- Content download. The version only advances once a flash step confirms the content is in place, and that step defaults to false.
The authentication change is the pattern in miniature (simplified):
int auth_sign(const char *msg, char *out, size_t out_len) {
if (!key_provisioned) {
return -1; /* refuse; never sign with a default key */
}
return hmac_sign(unit_key, msg, out, out_len);
}
Signing stayed cross-compatible with the server: a fixed test vector produces the same signature in the C firmware and the Node.js edge code.
A host harness: real modules, mocked kernel
Without a target toolchain, we compiled the real firmware modules with the host’s gcc against a thin mock of the Zephyr APIs they call. Each suite builds with AddressSanitizer and UndefinedBehaviorSanitizer, with recovery disabled so the first error fails the run, and each module also goes through GCC’s static analyzer:
gcc -std=c11 -g -O1 -Wall -Wextra -I<mocks> -I<src> \
-fsanitize=address,undefined -fno-sanitize-recover=all \
test_<module>.c -o test_<module> && ./test_<module>
gcc -std=c11 -fanalyzer -I<mocks> -I<src> -c <src>/<module>.c -o /dev/null
The GCC manual (current online edition, checked October 2026) describes AddressSanitizer as instrumenting memory accesses “to detect out-of-bounds and use-after-free bugs,” and UndefinedBehaviorSanitizer as instrumenting computations “to detect undefined behavior at runtime.” With -fno-sanitize-recover, “only the first detected error is reported and program then exits with a non-zero exit code,” which is what you want in CI. AddressSanitizer halts on the first error by default; the flag matters most for UndefinedBehaviorSanitizer, which otherwise keeps going. The static analyzer (-fanalyzer) looks for problems along paths that cross function boundaries. The manual is candid that it is “neither sound nor complete” and “only suitable for use on C code in this release,” so treat it as a bug finder, not a proof.
The harness now runs 156 checks with no failures, across the state machine, the uplink and downlink message handling, the display logic, the NFC protocol, the battery gauge and the NFC driver glue. It also pushes 400,000 fuzzed inputs through the message parser and formatter under the sanitizers. The server-side services carry their own Node.js suites, 283 tests, which stayed green throughout.
The harness paid for itself before any stub was replaced. Fuzzing and the sanitizers found three bugs of familiar kinds: a sequence-number lockup in duplicate detection, a missing escape in a message formatter, and an integer overflow in a time calculation. All three were fixed in pre-production code, before any hardware existed, and each now has a regression test so it can’t quietly return.
What the host harness proves, and what only the boards can.
Mark silicon-only behavior as a bench seam
Some behavior can only be proven on real hardware: SPI and I²C timing, the RF field, flash writes, and above all sleep current. We wrapped each piece behind a named function, a bench seam, wrote the code to the datasheet, and recorded it as ready for bring-up rather than done.
| Area | Done and host-tested | Waiting on hardware |
|---|---|---|
| NFC reader | Tag-selection protocol; vendor-driver glue tested against a fake | Building against the vendor library, antenna tuning, presence detection |
| Battery gauge | Register conversion, last-good caching | I²C binding and real register reads |
| Wake on motion | Motion latch, wake-time policy | Interrupt configuration on the accelerometer |
| Deep sleep | Timeout and wake policy | Low-power state entry; measured sleep current |
| Per-unit keys | Fail-closed signing | Reading keys from chip storage; provisioning flow |
The gap inventory says it plainly: a battery-life figure is a bench measurement, not code. A reviewer can find every unproven line by searching for the seam functions, and the bring-up runbook maps each one to a hardware step.
A protocol simulator so the platform runs without hardware
The other half of testing without hardware is the rest of the system. The devices send one JSON object per UDP datagram, carrying a device ID, a sequence number, battery level, firmware version and an event such as “tag seen” or “heartbeat,” with a message authentication code. A small simulator sends exactly those datagrams, signed the same way as real firmware, and walks a pool of virtual devices through a realistic usage cycle, from assignment through use to return.
Because the edge server can’t tell the simulator from a real device, the whole stack, from the operator dashboards to cloud sync and update rollouts, can be built, demonstrated and tested end to end with no boards. The simulator uses a seeded random generator, so a run can be replayed exactly, which turns “it glitched during the demo” into a reproducible bug report.
The simulator replaces the devices, not the platform: everything from the edge server up runs for real.
What comes next
Once the toolchain is in place, Zephyr’s own tooling is the natural next layer: the ztest framework for unit tests, and the Twister test runner to build and run them on native_sim, emulators such as QEMU, and real boards. The host harness doesn’t replace that layer. It means that by the time the boards arrive, the logic has already been tested, and the only open questions are the ones that need hardware to answer.
How we can help
We review and stabilize systems that were built quickly or ahead of their hardware. That means inventorying what is real, making stubs fail closed, adding test harnesses that run in CI, and producing a bring-up plan that keeps finished work separate from unfinished. See our services for how we engage.