Draft for review. Not yet approved for publishing.
Case study · August to September 2026
Green is a claim, not a fact
A certification harness that could declare success without evidence. A watchdog accusing a healthy store. 136 information sources shown as “active” that nothing ever ran. This is how Falkor learned to prove its own status.
At a glance
- Problem
- Status surfaces reported green for dead links, idle sources and unproven certifications, and red for a healthy store.
- Fix
- Every status is derived from evidence about the thing it describes, and every gate is proven to fail before it is trusted to pass.
- Proof
- 19 negative controls, priority-1 alerts from 1 to 0 with no false greens, and 7,215 of 7,218 browser tests passing with retries still at zero.
The problem
A home system with dozens of services lives or dies by its status lights. When the dashboard says everything is green, you stop looking. That’s only safe if green means green.
Falkor’s status layer had grown fast, and by late summer it could no longer be taken at its word. It showed some things as healthy that weren’t, and raised alarms about things that were fine. Both errors are corrosive: a false green hides a real problem, and a false red teaches you to ignore alarms.
The evidence
Measured one at a time, the failures looked like this.
- A dead button that looked alive. Falkor offered an Open button for two side apps while both were down. The button depended on a registry setting that says whether an app can be launched, not on whether anything was listening. Pressing it could only ever produce a connection error.
- 136 sources that never ran. Of 459 enabled information sources, 136 had no execution path at all, yet every surface showed them as enabled, active and visible. Nothing was broken. Nothing ran them, either, and nothing said so.
- A watchdog accusing a healthy store. A priority-1 alert claimed a data store kept failing. All 32 of its timestamps came from a single 30-minute outage. The watchdog counted raw timestamps, while the reconciler correctly merged them into one episode: two components, two definitions of “the same incident”.
- A report that counted “didn’t run” as “skipped”. The test framework’s own summary folded tests that never executed into its skipped count. In a deliberately aborted run, 14 “skipped” tests had never run at all, so an aborted run could look like a clean one.
- A certification harness that could pass without evidence. It declared the system certifiable when a list of commands exited with code 0. That meant it could report green in six specific ways:
- It didn’t require a post-reboot check, and would accept one left over from an earlier boot.
- Nothing tied the result to the build actually deployed.
- Nothing compared state before and after, so the run could change the very thing it was certifying.
- Required user journeys had no stable identity, so a renamed test left them green.
- A “reasoned skip” made the skip gate greener.
- Nothing checked for leftover test artifacts.
- A matcher that matched everything. One required journey was matched by the pattern
/(status|watchdog|truth)/. Against the real test suite, it matched 57 of 692 files by file path alone. Any of those 57 unrelated tests passing would have “proved” that no surface reports a false green. - A gate nobody could pass. The boot pre-check and the certification step that read its receipt disagreed on the file location, the schema and every field name. They even formatted boot times differently. I could reboot, watch the pre-check pass, and certification would still answer “no receipt”, forever.
The investigation
These weren’t unrelated bugs. Each one was a status claimed by something that hadn’t checked the thing it described:
- a registry flag standing in for health;
- an intention reported as activity;
- an exit code standing in for a certification;
- a common word standing in for a behaviour.
The pattern was the defect.
The decision
I set one rule for everything that followed: a status must be derived from evidence about the thing it describes, and every gate must be shown to fail before it is trusted to pass. That meant no raised timeouts to make a red go away, no added retries and no weakened assertions.
The fix
- Launch buttons follow health. A single rule decides whether an app can be opened now, started and then opened, opened manually, or not launched at all. The link exists only while health is green at that moment. A counter of dead Open links must stay at zero, and a contract test checks the rule across its entire input space, so a dead button is structurally impossible rather than merely absent today.
- Every source gets exactly one disposition. Each enabled source either runs in an automatic lane or carries a named, reversible reason and a clear action. “Uncovered” is zero by construction: 333 automatic, 126 explained, 0 uncovered.
- One definition of an incident. Recurrence is classified by one function that the reconciler and the watchdog both call.
- Certification requires evidence, not exit codes.
- Every gate must be green and every acceptance clause satisfied.
- Required journeys carry stable IDs and are matched by an exact tag or file name; a broad matcher is rejected outright.
- Results are parsed from the real test report, and a test that didn’t run can never count as a skip.
- Promises that only a person can verify are listed separately, and can never satisfy an automated requirement.
- One owner for the receipt. The pre-check and the certification step now share a single contract from one library, and only a genuine PASS reads as a pass. “Not finished yet” can never read as “passed”.
The verification
- Red before green. Nineteen focused negative controls drive each certification failure directly. They cover a missing pre-check, the wrong boot, the wrong build, a mutated store, a missing, skipped or mocked journey, and leftover artifacts. The harness stayed red, correctly, until real evidence existed.
- Live, not hypothetical. After the incident fix, priority-1 alerts went from 1 to 0, with zero false greens and zero false reds. No threshold was lowered and no evidence deleted.
- A full certification run (22 Sep 2026). 7,215 of 7,218 browser tests passed, and 52 of 56 gates were green. Every red was root-caused rather than retried. One was a real product defect. The rest were test-side defects that surfaced only under the full battery’s load or in one particular launch environment, such as a click landing before the page was ready to respond, or a test mock registered after the page had started. All were fixed at the cause, and retries stayed at zero.
Lessons
- A green exit code is not a certification.
- A broad matcher proves only that somebody used a common word.
- A gate nobody can satisfy is as false as one anybody can.
- Controls that must fire matter more than controls that must not. The tests that count are the ones proven to fail.
- Raising a timeout turns a real observation into a decorative gate.