Logbook

Decisions, lessons, blockers and open questions

The record behind the numbers: what was decided and why, what went wrong and what it taught, what's blocked, and what's still undecided. Every entry traces to Falkor's own records.

The featured story

The rewrite that rebuilt Falkor

In July I had two AI models design a replacement for Falkor from scratch. It never shipped. Its best parts rebuilt Falkor instead.

  1. Decision

    Commission a clean-sheet rival

    The question was whether a clean start would beat Falkor. Two frontier AI models each designed a complete successor from scratch, and I judged them head to head. The winner, AURYN, came as a build-ready package of 20 documents and 35 screens: a Rust core, a “morning paper” interface, and three hard rules aimed at Falkor's worst defects. Stale data can't show green, one referee owns the GPU, and an action isn't done until a check proves it.

    Competition records

  2. Decision

    Stalled, audited, frozen

    Coding ran in phases through July. When I authorized the next phase on 1 August, it stalled within a day in a resource wait loop behind a safety lock. On 9 August an independent audit of the demo found every screen a shell: 0 of 13 user journeys worked, none of 6,303 acceptance checks had been worked through, the source no longer built, and the only runnable binary carried a false release identity. Falkor, checked the same day, was all green. The migration goal was deleted, and AURYN was frozen.

    Build records and independent audit · 9 Aug 2026

  3. Decision

    Audited again, for parts

    Then AURYN was audited a second time, for what was worth keeping: the best parts that had actually been coded, and the ideas that never were. Both were refactored into a new package for reimagining Falkor. Among the ideas: one registry of everything the system can do, status that can't show green on stale data, and actions that aren't done until a check proves them.

    Audit and reimagining package · Aug 2026

  4. Decision

    Rebuilt in layers

    The rebuild took groups of pages one at a time, focused each group and layered it onto the new foundation. It was a large rebuild. Early on, an audit found 184 page routes, more than half of the registered ones retired but still alive, because every cleanup had added a hub and deleted nothing. In two days, 19 and 20 August, that became 98 with no capability lost. AURYN itself was left behind, adopted by the Falkor it was built to replace.

    Route audit · 19–20 Aug 2026

The lesson. A demo that shows shells is a finding, not a delay. Freeze early, audit for what's worth keeping, and rebuild the product people already use, one layer at a time.

Every entry

  1. Lesson

    Report k of n, never a bare percentage

    After an agent's notes were compacted, a baseline survived only as “7.1%”, which was the failure rate. It nearly reported a 69-of-70 re-run as a jump from 7.1% to 98.6%. The real comparison was 65 of 70 against 69 of 70, with overlapping intervals: consistent with improvement, not proof of it.

    Behaviour suites · 28 Sep 2026

  2. Blocker

    The calendar went quiet, so Falkor says so

    From 21 Aug the calendar's sign-in could no longer be renewed, but a wrapper still returned “ok” with zero events, which let the assistant say there was nothing on tomorrow. Now Falkor says it can't reach the calendar: no web fallback, no invented events. Re-authorizing was mine to do; on 28 Sep it was reconnected, and a live check read all 7 calendars.

    Local-truth routing · Aug–Sep 2026

  3. Decision

    Red tests don't get longer timeouts

    The 22 Sep certification ran for 183 minutes with retries off and ended with 4 of 56 gates red. Raising timeouts or adding retries was rejected: it would have turned those gates into decoration and erased a real observation about behaviour under load. The same day, every failing check was fixed at its cause. No timeout was raised, no retry added, no assertion weakened.

    Certification · 22 Sep 2026

  4. Lesson

    Most reds were the test battery's own load

    The 22 Sep certification ended with 52 of 56 gates green. Behind the 4 red gates were 5 failing checks, and only one was a product defect; the others passed on a quiet machine. One took 15.6 seconds inside the battery and 672 milliseconds on its own. Now every red gets a quiet re-run before anyone repairs it.

    Certification · 22 Sep 2026

  5. Open questionOpen

    Fine-tune the chat model?

    On 17 Sep I chose a behaviour lab first, with fine-tuning only for the gaps it measures. Research on 21 Sep found that training the 12B chat model needs about 20 GB of GPU memory, against a 16 GB card. The question now is which measured gap, if any, is worth that.

    Roadmap · Sep 2026

  6. BlockerOpen

    Certification needs a person

    Falkor's certification matrix has 43 mandatory requirements. Two can't be met by any process: physically rebooting and logging in, and judging the cockpit on the real screen. Four more need evidence from a genuine reboot of the same build, so every fix deployed after a reboot costs another one. The formal stamp waits on me, on purpose.

    Certification matrix · Sep 2026

  7. Open questionOpen

    Keep n8n, or park it?

    n8n 3.0, due in October 2026, removes nodes. The options are to stay pinned on 2.x and prove that at least one approval-gated workflow earns its keep, or to park it.

    Roadmap · 17 Sep 2026

  8. Decision

    Keep a century-old archive out of the news

    The Library of Congress's Chronicling America was the last keyless news source on the list. But it serves digitised newspapers from before 1963, and mixing them into the current-news corpus would have damaged it. No adapter was written; the entry keeps its reasoning and a working endpoint until history has a proper home.

    Source registry · 4 Sep 2026

  9. Lesson

    A green gate that couldn't see

    The rule against hard-coded model names passed while 1,592 of them sat in 371 files, because it checked only two patterns. Its replacement brought them to zero, then turned out to exempt its own file, where it was carrying exactly what it bans. It now scans itself, and a missing self-scan fails the run.

    Model platform · Aug–Sep 2026

  10. Open questionOpen

    Where does history belong?

    Keeping pre-1963 newspapers out of the news was the easy half. The open half: does an archive deserve its own surface, with its own rules for ranking and freshness?

    Source registry · Sep 2026

  11. Decision

    Chat gets no permanent claim on the GPU

    A scheduled job loaded its own model and pushed the chat model off the 16 GB GPU. My first call was to give chat priority. That evening I reversed it: workloads run on schedule, and whichever model happens to be loaded is not the model chat uses. Removing the priority code (8 files) let the job run, and that exposed three hidden defects, including an 81.3-second cold load against a 25 to 45 second limit.

    Model scheduling · 3 Sep 2026

  12. Lesson

    Find the mechanism; don't tolerate the symptom

    A React hydration error (#418) appeared under load and was first tolerated as a framework bug. Instrumenting React's error path in the test browser showed the real mechanism: React 19.2 replayed a page element after a lazily loaded chunk paused hydration. A structural fix took it from 6 errors in 18 runs to 0 in 36.

    Front end · 2 Sep 2026

  13. Lesson

    An outlier is a question, not a finding

    An AI agent filed a disabled watchdog task as a high-priority recovery gap: it was the only one of four siblings switched off. It was a finished migration. The service had moved to the process manager in July, and re-enabling the task would have given one process two owners. The finding was withdrawn, tests now pin the migration, and the rule is to ask what replaced something before asking why it's off.

    Operability · 30 Aug 2026

  14. Lesson

    An estimate isn't a measurement

    When the full browser battery first ran, 889 tests failed. A 15-file sample was extrapolated to “about 83% are stale tests pinned to retired pages”. Measured across every failure the same day, only 26 of 890 (2.9%) carried the retired-page signature, and 391 (43.9%) failed on live pages. The work went where the measurement pointed: by 22 Sep, 7,215 of 7,218 passed.

    Test census · 26 Aug 2026

  15. Lesson

    A 400-character window hid 889 failures

    The certification script trimmed each red gate's evidence to its last four lines, 400 characters at most, so a shard with 120 failures showed one or two. Three full runs read as “a handful left”, and an agent's “24 → 10 → 0” progress story was an artifact of that window. Once red gates printed every failure, the real debt was 889 tests across 337 files.

    Certification tooling · Aug 2026

  16. Decision

    n8n: superseded, removed, restored behind approvals

    Falkor's own scheduler became the system of record (17 of 17 jobs healthy), so n8n was declared superseded, and the next lane removed the integration. That removal had no recorded authorization, while the instance and its 11 workflows were intact, so it was reversed. n8n came back as a lean bridge: it can list and describe workflows, but nothing runs without a plan-bound approval. Retiring a tool is a decision, not a side effect of someone's cleanup.

    Lane records · Aug 2026

  17. Lesson

    A time limit that could never pass

    Every certification gate had the same 900-second cap, including a browser tier of 6,949 tests run one at a time; a 260-test slice alone took 20 minutes. The cap could never cover the suite it claimed to gate. It became a fast collection check plus eight shards with two hours each, and full runs now take about three to four hours.

    Certification tooling · 24 Aug 2026

  18. Lesson

    Two dead signals before a kill

    During a certification run, the watchdog that recovers the Linux subsystem killed a healthy instance: one failed probe was enough, and that probe fails briefly under load. Recovery now needs two independent dead signals, 20 seconds apart. It has since absorbed about 25 real blips with no false kills and no missed failures.

    Supervisor · 24 Aug 2026

  19. Decision

    More memory for the Linux VM? Measured, and no

    The proposal was to raise the Linux subsystem's memory cap from 13 to 16 GB so the model runtime could memory-map its model. Measurement said it wouldn't be enough: the runtime would still have been about 8 GiB short, while Windows would have lost most of its last 2 GB. Instead, a held session cut cold starts to 2 in 8 hours, a larger swap file took chat turns over 10 seconds from 12% to 0%, and on 26 Aug the machine's RAM was doubled to 64 GB.

    Runbook and handoff · Aug 2026

  20. Lesson

    One word disabled every tool

    Chat tools had been silently off. The model service dropped its tool instructions whenever the system prompt contained “Falkor”, and Falkor's identity prompt always did. On the same question, the prompt shrank from 5,068 tokens to 502 and the tool call vanished.

    Audit · 21 Aug 2026Read the case study

  21. Lesson

    A feature that never worked, with a passing test

    The unified Inbox had never completed an inline action since it shipped: it sent display IDs where every backend expected exact IDs, and sent approvals to the wrong broker. Its only test checked that two URL strings appeared in the source code. After the fix, all 1,433 actionable rows resolve.

    Consolidation · 20 Aug 2026

  22. Lesson

    A build loop quietly ate 404 GiB

    The rewrite's build harness created a fresh 10.7 GiB build folder roughly every 25 minutes for two days, 40 of them, and never cleaned one up. A storage root-cause sweep found it among 480 GiB of reclaimable space. Falkor now keeps a permanent collector that names the owner of leftover files from their path, never guessing from their size.

    Storage sweep · 19 Aug 2026

  23. Lesson

    One setting reloaded the model every turn

    Falkor sent a per-request batch size that disagreed with the one the model runner was started with. The runtime quietly reloaded the model on every chat turn while still reporting it as loaded: 0.3 seconds of load time became 8 to 17. Removing one key cut time to first byte from 4.4 seconds to 0.4.

    Chat performance · 3 Aug 2026

  24. Lesson

    Agreement isn't truth

    Every document agreed that the rewrite's coding was frozen, so the contradiction checker found nothing stale. I had authorized coding that day; they were all wrong together. Status is now re-derived from the live source, never from documents agreeing with each other.

    Docs truth · 1 Aug 2026

  25. BlockerOpen

    The payout model broke the rules

    For the classroom fundraising app, policy research found that paying donations straight to teachers conflicts with school-district policy and state gift limits. I documented three compliant alternatives, and the real-money pilot waits until the payout model changes. Better found in research than in production.

    Policy memo · 17 May 2026

Esc