Capability audit · 25 Sep 2026

Your home.
Your data.
Your AI.

Falkor is a private, local-first AI system that runs our home. It handles chat, memory, news, media, displays and automation, plus the operations that keep it honest, all on one home PC under one set of rules.

Flight recorderReplays of recorded behaviour

How a request moves through FalkorEvery request passes Falkor's governed core, which routes it to tools, memory, local models or the collection engines, checks the answer against what actually ran, and returns it to you or your screens. An operations ring of watchdogs surrounds the core.Falkorgoverned coreYoubrowser · phoneToolsreminders · notesMemoryinbox → approvalModelslocal · one GPUEngines459 sourcesScreensTV · kiosks

You“Remind me tomorrow afternoon to call the plumber.”

  1. 01RequestRequest arrives from the browser
  2. 02RouteRouter: a write to your own data. It stays on the home PC
  3. 03Toolcreate_reminder · due tomorrow, 6:00 PM
  4. 04VerifyReminder store now holds the new item
  5. 05Guard“I've set a reminder” is backed by a successful tool call

FalkorDone. I'll remind you tomorrow at 6 PM.

Reconstructed from Falkor's logs and case studies; prompts are illustrative. Not a live connection.How the tools were fixed →

Pick a scenario to watch a real request move through Falkor, or hover a part of the map to see what it does.

What Falkor is

Not a chatbot. Not a dashboard. A system.

Falkor began in March 2026 as a local AI stack and grew into one product: a cockpit of 100 pages over 1,051 API operations, local models that answer on the home PC, a long-term memory with review built in, engines that read hundreds of sources, and the operational tooling to prove what is actually working.

01

Local-first

504 of the 649 catalogued capabilities run entirely on the home PC. Public web search runs only when a question needs it.

02

Governed actions

Consequential actions wait for approval, and each answer is checked against what the turn actually did, so the AI can't claim work it never ran.

03

Evidence over optimism

Status surfaces are tested for false greens as hard as for failures. One certification run covers 56 gates and more than 7,000 browser tests.

04

Used every day

It runs the living-room TV, the morning news, the calendar and reminders. It's daily infrastructure, not a demo.

Showcase

Twelve things worth seeing

Pick one. Each opens a real screen from Falkor, annotated, or a working demo you can drive yourself.

The browser edition of Falkor's Local AI World report: a teal editorial masthead, then a section marked Partial whose shortfall note reads requested 10, delivered 9, 33 sources failed, no filler added.
Falkor's Reports and Briefings page: a large serif masthead, report tabs from What Matters Today to Gadgets and Products, and cards for Local AI World and Geek Patrol, each marked published with quality 100.

Data & integration · 01

Briefings that write themselves

Nine reports on the news, local AI, sports, entertainment and more come from one report engine, each with browser and email editions. When a report falls short, it says so instead of padding.

Local AI World this week: 9 of 10 items delivered, 33 sources failed, no filler added

What you're looking at Local AI World

  1. Every edition states its coverage: 46 of 51 items this week.
  2. Each section declares a target, a ceiling and a minimum it will accept.
  3. The shortfall is printed. Nothing is invented to fill the gap.

What you're looking at Reports & Briefings

  1. All nine reports run on one canonical engine, one archive and one scheduler.
  2. News, local AI, sports, entertainment, charts and gadgets, one tab each.
  3. Each edition is scored before it publishes.
Live demo · sample turn

Chat

Remind me Friday to renew the car registration.

Done. I'll remind you Friday at 9 AM. I've also emailed you the renewal form.

Correction: no email was sent this turn.

Turn record

  • create_remindersucceeded
  • send_emailnever called

Claim guard: 1 unbacked claim struck and corrected.

The answer can only claim what the record shows.

AI · 02

Chat that proves its actions

Ask in plain English for a reminder, a note or your schedule. Falkor does it with real tools, and every answer is checked against what the turn actually ran.

A fabricated claim was caught and corrected live

Read the case study

How it works

  1. Each turn leaves a record of every tool it called
  2. A claim guard compares the answer with that record
  3. Unbacked claims are struck and corrected before you see them
A Falkor dossier on Ada Lovelace: a portrait, a summary, chapter navigation, and an at-a-glance panel listing 74 verified statements from 3 sources.

Data & integration · 03

A dossier on anything

Name a person, a band, a team, a product or a topic. Falkor researches official, reference and archival sources, then writes a dossier from what it can verify, and shows what it found and what it couldn't.

One dossier: 74 verified statements from 3 sources

What you're looking at A dossier on anything

  1. Every statement opens its source.
  2. 74 verified statements from 3 sources, in edition 4.
  3. Chapters from the short version to the timeline, each searchable.
Live demo · sample captures

Review inbox · 3

  • From chat

    Prefers the local-AI briefing before sports.

  • From the share menu

    The car is due for service in May.

  • From the clipboard

    Recycling goes out Thursday night.

Long-term memory · 0

    Nothing yet. Only what you approve lands here.

    Nothing becomes memory without a decision.

    AI · 04

    Memory you approve

    Say “remember this” in chat, from the clipboard or from any app's share menu. It lands in a review inbox, and only what you approve becomes long-term memory that grounds future answers.

    Promotion is a decision, not a side effect

    How it works

    1. Captures land in a review inbox, never straight in memory
    2. Approved items are embedded locally and indexed for recall
    3. Rejected items never reach long-term memory
    Falkor's Spin Lens page: a dark banner titled Adversarial decision workbench, the selected local model and search setting, and a form to choose a proposition and its evidence.

    AI · 05

    An argument machine

    Spin Lens takes a claim, an article or pasted text and builds the strongest case for and against it, steelmans each side and pressure-tests the leader. It's a guard against one-sided answers, including the AI's own.

    Runs on the local model, with private search first

    What you're looking at Spin Lens

    1. Builds the case for and the case against, steelmans both, then rebuts them.
    2. Facts carry stable evidence IDs; assumptions and value judgments stay labeled.
    3. Nothing is saved unless you ask. Pasted text is never stored.
    Live demo · replays a measured run
    Sources
    16 / 110
    Stories
    —
    Duplicates
    0

    The recorded run: killed at source 16 of 110, then resumed with zero duplicate stories.

    Data & integration · 06

    Hundreds of sources, one honest feed

    459 sources across feeds, APIs and sites, collected by two engines. Duplicates are fetched once, failing sources are quarantined and re-probed, and a collection pass survives a crash.

    Force-killed at source 16 of 110, resumed with zero duplicate stories

    How it works

    1. Progress is checkpointed as the pass runs
    2. A resumed pass picks up where the last one stopped
    3. Stories are de-duplicated, so a restart can't count anything twice
    Live demo · sizes illustrative

    One GPU

    1. Chat model resident, ready for questions

    One chat model you choose; scheduled work borrows the GPU and hands it back.

    Falkor's AI and Models page: one global chat model with a 32,768-token context, tool policy approvals required, VRAM free, and a persona using 296 of an 800-token prompt budget.
    Falkor's Model Maintenance page: the global model fully in VRAM, and a runtime residency panel comparing the model resolved for chat with the model actually active on the machine.

    AI · 07

    Local models that share one GPU

    You choose the chat model. Scheduled jobs such as the morning briefing load their own models on the same GPU and hand it back, and a guard refuses cloud models that pose as local ones.

    Chat, then a scheduled job, then chat again, each on the right model

    How it works

    1. One chat model, chosen by you, serves every surface
    2. Workloads borrow the GPU on schedule and give it back
    3. Residency is measured, never assumed from the selection

    What you're looking at AI & Models

    1. Tools wait for approval, by policy.
    2. Residency is checked live: the model is fully in GPU memory.
    3. The persona fits an 800-token prompt budget, 296 used.

    What you're looking at Model residency

    1. No fake “installed” or “ready” status.
    2. What chat asked for, next to what the machine actually has loaded.
    3. Memory pressure, measured live.
    Falkor's Capabilities page: 307 capabilities split into Can do now (107), Needs your OK (86) and Can't right now (114), with filter chips and capability cards showing health, write policy and locality.

    Operations · 08

    It knows what it can do

    One registry, built from each owner's own records, lists every capability, whether it works right now and how it is used safely. The shelf explains; it executes nothing.

    307 capabilities: 107 ready, 86 waiting for approval, 114 blocked with a reason

    What you're looking at Capability registry

    1. Ready now, waiting for your OK, or blocked, each with a reason.
    2. 566 source rows left out, each with a written reason.
    3. Every capability shows its health, its write policy and where it runs.
    Falkor's Open-Source Engines page: 7 of 7 engines online, 0 down, exposure loopback-only, and an evaluation confirming every running engine can be checked from Falkor without its own login.

    Data & integration · 09

    The best of open source, governed

    More than 20 open-source projects run as managed engines: photo library, document archive, private search, uptime monitoring, notifications, recipes, web archiving, workflow automation, design and vector search.

    • Immich
    • Paperless-ngx
    • SearXNG
    • Uptime Kuma
    • ntfy
    • Mealie
    • ArchiveBox
    • n8n
    • Penpot
    • Qdrant
    • ComfyUI

    Engine upgrades are re-validated against the full test battery

    What you're looking at Open-source engines

    1. Seven engines, all online.
    2. Loopback only: reachable from the home PC and nowhere else.
    3. Each engine is checked without its own login. Credentials are referenced by name, never by value.
    Falkor's Self-Healing page: a clean bill of health with every required service healthy, and one optional integration idle by choice, with its cause, consequence and source explained.
    The Hermes custodian cockpit in Falkor: overall status green, last sweep 45 of 48 green and 0 red, and failure-ownership counters including 3,084 recovered and 0 missed.

    Operations · 10

    Heals itself, or says why not

    Watchdogs cover every core service and an always-on custodian sweeps every two minutes. Stop everything and the stack restores itself in 6.8 seconds; after a cold reboot it comes back unaided in about 14 minutes.

    Full stack restored in 6.8 s, measured 26 Aug 2026

    What you're looking at Self-healing

    1. You're asked to act only when automatic repair is impossible.
    2. Green only when every required service is healthy.
    3. An optional service idle by choice is explained, not painted red.

    What you're looking at Custodian cockpit

    1. Last sweep: 45 of 48 checks green, none red.
    2. A bounded sweep every two minutes, owned by the supervisor.
    3. Detected, delegated, recovered: 3,084 recoveries, 0 missed.
    Live demo · sample services
    • Chat modelResident
    • Scheduled AI jobsOn schedule
    • Image engineReady
    • News collectorsNormal

    GPU

    1. Normal mode: Falkor's AI work holds the GPU

    Falkor holds the GPU for its AI work.

    Operations · 11

    Game Mode

    One switch hands the GPU to a game: background AI work pauses, services throttle, and everything is restored afterwards.

    Throttle and restore verified live

    How it works

    1. One switch, in Falkor's header
    2. AI work pauses, services throttle and the GPU is freed
    3. Switching back restores every service, and checks it
    Live demo · sample processes
    ProcessOwnerIdleAction
    falkor-webFalkoractive
    helper.exenone (orphan)14 min
    updater.exeunknown3 min
    worker.exenone (orphan)4 min

    Every stop passes four gates, and anything unknown fails closed

    1. Fresh census
    2. Verified orphan
    3. Idle 10 minutes or more
    4. Typed confirmation

    Pick a process and try to stop it.

    Operations · 12

    A process census with a safety catch

    See every Windows process, who owns it and whether it is idle. Stopping one takes a fresh census, a verified orphan at least ten minutes old and a typed confirmation, and every action is audited.

    Anything unknown fails closed

    How it works

    1. A fresh census before any stop
    2. Only verified orphans, idle ten minutes or more, qualify
    3. A typed confirmation, then an audit record

    Screens captured from the running system on 28 Sep 2026 and cropped to the page. Demos run in your browser on sample data.

    What it does

    649 capabilities, ten areas

    The audit catalogued 649 capabilities across 26 families, grouped here into ten areas. Each bar shows the share observed running during the audit (459 in all). A read-only audit can't exercise everything, so "not observed" doesn't mean broken.

    Capability audit · 25 Sep 2026
    • Intelligence

      Local chat models, long-term memory with retrieval, personas and screen understanding.

      66 of 112

    • Information

      News from hundreds of sources, weather and radar, sports and local events: ranked, de-duplicated and explained.

      85 of 94

    • Life & household

      Reminders, the family calendar, documents, read-only email and shared household tools.

      79 of 93

    • Media & creative

      A living-room TV experience, music, radio and podcasts, video discovery, and local image and video generation.

      53 of 66

    • Home & displays

      Any screen can become a Falkor display: TV dashboards, weather radar and kiosks.

      Smart-home control is built, but the audit did not observe it running.

      7 of 14

    • Voice

      Push-to-talk speech in, natural speech out.

      Always-on listening is built and deliberately switched off.

      5 of 11

    • Automation & tools

      Scheduled jobs, notifications, a tool shelf and agents, with approval gates on consequential actions.

      54 of 68

    • Platform & side projects

      A registry and SDK that let new apps borrow Falkor's models, memory and status, plus a Labs shelf for side projects.

      66 of 79

    • Security & governance

      Origin guards, approvals, a local-only model rule, and every capability classified by privacy and cloud exposure.

      7 of 17

    • Operations

      Health checks, watchdogs, self-healing, certification and controlled deployment.

      A read-only audit can't exercise most of these, so many show as not observed.

      37 of 95

    By the numbers

    A real system, measured

    Every figure carries its source and date. None of them are estimates, and none update themselves: the site changes only when a new snapshot is published.

    Unless marked · Capability audit · 25 Sep 2026

    649

    capabilities catalogued

    459 observed running during the audit, across 26 families

    1,051

    API operations

    100

    Pages in the cockpit

    381

    Tools

    7,331

    Mapped dependencies

    69

    Containers

    54 running at audit time

    36

    Scheduled background jobs

    78%

    Capabilities that are local-only

    504 of 649

    2,947

    Commits since March 2026

    Falkor and its stack

    Git · 28 Sep 2026

    7,215 of 7,218

    Browser tests passed in one certification run

    The three failures were root-caused and fixed afterwards

    Certification · 22 Sep 2026

    How it's run

    An enterprise product, run by one person

    I own Falkor's whole lifecycle: planning, architecture, delivery, verification, deployment, operations and audit. It has been run like an enterprise program from its first weeks, with a roadmap, a groomed backlog, an agile cadence, a CI/CD pipeline, and AI agents working as the delivery team.

    01 · Plan

    Plan it like a product

    Falkor has had a phased roadmap since its first weeks: ten build phases by the end of April, then the Falkor 3.0 program. Work lives in a delivery tracker in which every item carries a phase, a gate, its dependencies, an owner and a completion check.

    • Phased roadmap from month one
    • Backlog captured and groomed as it grows
    • Clear definition of done for every item

    940

    tracked items, 644 done

    Delivery tracker · 28 Sep 2026

    02 · Build

    Build in scoped lanes

    Work is cut into lanes: briefs with a goal, acceptance criteria, required reading, an ordered scope and a do-not-touch list. AI coding agents implement them. I own the architecture, the rules and the review.

    • Change impact declared before any edit
    • Several agents in parallel, with collision rules
    • Every brief written to be checkable

    541

    named lanes in the canonical record

    Canonical docs · 28 Sep 2026

    03 · Verify

    Prove it before calling it done

    Every new gate must be shown to fail before it is trusted to pass. Certification runs 56 gates over more than 7,000 browser tests, and a lane isn't done until its living documentation says so too.

    • Negative controls on every gate
    • No retries or timeouts added to hide a red
    • Behaviour suites for the product's promises

    7,215 / 7,218

    browser tests passed in one run

    Certification · 22 Sep 2026

    04 · Ship

    Ship continuously, and safely

    A self-hosted CI/CD pipeline carries every change. Blocking checks run first: change impact, documentation parity, audit and coverage. Deploys are atomic and refuse uncommitted code, the build ID comes from the commit, and a failed readiness check rolls the release back.

    • Commit, built bundle and live build must match
    • Automatic rollback on a failed readiness check
    • Pre-commit hooks guard the history

    Every deploy

    commit, bundle and live build verified identical

    Deployment records

    05 · Operate

    Run it every day

    Watchdogs cover every core service, a proactive status check runs every 30 minutes, and readiness is checked after every boot. When something fails, it heals: the full stack comes back in 6.8 seconds.

    • Status truth tested for false greens and false reds
    • Game Mode and graceful restore
    • Operator attests what only a person can check

    11 / 11

    critical readiness checks green

    Readiness · Sep 2026

    06 · Audit

    Audit the whole thing

    A read-only audit mapped 649 capabilities, 1,051 API operations and 7,331 dependencies, then triaged 412 findings by severity. It was checked for internal consistency before any of it entered the canonical docs.

    • Every capability classified for privacy and cloud exposure
    • Findings ranked and turned into roadmap items
    • Documentation drift corrected in place

    412

    findings triaged by severity

    Capability audit · 25 Sep 2026

    07 · Improve

    Feed it back

    Findings become roadmap items for the next round. New open models, runtimes and open-source releases are evaluated as they land, and engine upgrades are re-validated against the full test battery before they stay.

    • Audit-derived roadmap
    • Upgrades validated, not assumed
    • Retired features recorded, never silently dropped

    0

    test failures attributable to the 22 Sep engine upgrades

    Certification · 22 Sep 2026

    Operating rhythm

    1. Every 30 minutes

      A proactive status check, plus watchdogs on every core service.

    2. Every day

      Standup-style planning with the AI agents (what closed, what's next, what's blocked), defect triage, backlog grooming, and mixture-of-experts brainstorming: several AI models weigh in on open questions, then I decide and the decision is recorded. Software-update checks, memory re-indexing and fresh briefings run on their own.

    3. Every week

      A review of the AI landscape (new open models, runtimes and trending open-source projects), with mixture-of-experts sessions on the bigger decisions: what to adopt, what to retire, what comes next.

    4. Every round

      A planned round of parallel lanes with a steward, acceptance criteria and a definition of done. Nothing closes until its evidence and documentation are in.

    In agile terms

    On FalkorAgile equivalent
    Round of parallel lanesSprint
    Lane brief with checkable rulesUser story with acceptance criteria
    Delivery trackerProduct backlog
    “Done-done” plus updated living docsDefinition of done
    Daily planning and triage with the agentsDaily standup and bug triage
    Mixture-of-experts brainstorming, decision recordedDesign review and decision record
    Steward reviewTech-lead review and sign-off
    Closeout: what's claimed, and what isn'tSprint review
    Audit-derived roadmapRetrospective into backlog

    AI practice

    Fluent in a field that won't sit still

    AI's vocabulary changes every few months, and Falkor has moved with it. I treat AI as a powerful tool with known failure modes: I adopt what's new quickly, measure it honestly, and design around what it gets wrong.

    The shifts, and what they look like in Falkor

    1. From Prompt engineeringto Context engineering

      Each turn's context is assembled from structure: the capability map, the persona, local facts and recalled memory. It is refreshed every 15 minutes and trimmed by policy. A versioned structural marker, never a word in the prompt, decides what gets injected.

    2. From Keyword searchto RAG and vector memory

      Approved memories are embedded by a dedicated local model, indexed in a vector store, re-indexed nightly, and recalled into answers.

    3. From Workflow automationsto MCP tools

      Scheduling moved into Falkor's own scheduler, n8n workflows stay behind approvals, and capabilities are served as MCP tools through a bridge, in a tool universe of 462 tools from 8 providers.

    4. From Chatbotsto Agents

      OpenClaw runs on Falkor's own model selection, with Falkor's memory behind a proxy. Hermes works as a 24/7 custodian that can repair services. The agentic modes ask for approval before they act.

    5. From Bigger context windowsto Deliberate token budgets

      Budgets are measured, not guessed. A reasoning model given 900 tokens produced nothing and 2,400 produced a full answer. A batch setting that forced a model reload on every turn was found and fixed.

    6. From One big modelto A model plane

      One chat model you choose, workload models scheduled on the same GPU, cloud models refused at the boundary, and several models consulted side by side (mixture-of-experts brainstorming) before a decision.

    Honest about the limits

    Where AI falls shortWhat Falkor does about it
    Models claim work they didn't doEvery answer is checked against the turn's execution record.
    One date word can send a private question to the public webA positive classifier decides whose data a question is about before anything is searched.
    Reasoning models can spend the whole budget thinkingToken budgets are sized for thought and verified per model.
    A local model can crash on a large promptThe failure is reported and retried within bounds. The gap is never filled with invented text.
    AI coding agents over-report successEvidence gates, independent review, and an operator who signs off.
    GPU memory is finiteA model plane schedules who is resident; workloads hand the GPU back, and Game Mode frees it entirely.

    Outside the box

    • Make the gate audit itself

      The rule that forbids hard-coded model names had exempted its own file, and was hiding two violations. It now scans its own source like everything else, and a missing self-scan fails the run.

    • Fetch once, account everywhere

      41 endpoints were shared by 157 sources. The first row fetches and every twin records the same result: 29.7% fewer wasted fetches, and no source hidden.

    • Quarantine, not retirement

      Sources that have failed 800 or more times in a row move to a 6 to 24 hour re-probe, and rejoin the moment one fetch succeeds. Nothing is deleted to make a dashboard greener.

    • Maintenance without downtime

      The engines can quiesce: hold new work, drain, checkpoint and keep running. A retention job deleted 871 rows in the middle of a collection pass without stopping it.

    • A face for the voice

      When Falkor speaks, a small pixel display lip-syncs to the reply, a playful piece of hardware wired into the voice pipeline.

    • A flight recorder for every turn

      Each chat turn leaves a record of what actually ran. That record is how false claims get caught, and it's what the panel at the top of this page replays.

    How it works

    Three paths through the system

    Simplified on purpose. The audit maps 7,331 dependencies like these, from each page down through the APIs, libraries and services beneath it to the data it touches.

    Ask

    A question becomes a governed turn on a local model.

    1. 01You askFrom the browser, a phone or the quick-ask drawer
    2. 02Governed turnOne local model you chose, a persona, local facts first
    3. 03Tools and memoryReminders, notes and recall; writes wait for approval
    4. 04Answer with receiptsClaims are checked against what actually ran

    Public web search joins only when the question needs it.

    Remember

    Memory is earned: nothing becomes long-term without review.

    1. 01Capture“Remember this” in chat, the clipboard, or share to Falkor
    2. 02InboxEvery capture lands in a review inbox
    3. 03Your approvalPromotion to long-term memory is a decision, not a side effect
    4. 04RecallIndexed for search and used to ground answers

    Memory stays on the home PC.

    Stay informed

    Hundreds of sources become one explained feed.

    1. 01459 sourcesNews feeds, APIs and sites, in one registry
    2. 02Two enginesFeeds and scraping. Duplicates are fetched once; failing sources are quarantined and re-probed
    3. 03One feedDe-duplicated, ranked and explained
    4. 04EverywhereNews, daily briefings, reports and chat answers

    Source count as of September 2026.

    The story

    Seven months, one builder

    From early March to late September 2026. This history comes from Falkor's own records (git history, roadmaps and audits), not from memory.

    Commits per day

    Falkor and its stack, 1 Mar 2026 to 28 Sep 2026

    Git · 28 Sep 2026
    Commits per day2,947 commits across 163 active days.MarAprMayJunJulAugSepMonWedFri1 Mar 2026: 0 commits2 Mar 2026: 0 commits3 Mar 2026: 0 commits4 Mar 2026: 0 commits5 Mar 2026: 0 commits6 Mar 2026: 0 commits7 Mar 2026: 0 commits8 Mar 2026: 0 commits9 Mar 2026: 0 commits10 Mar 2026: 0 commits11 Mar 2026: 0 commits12 Mar 2026: 0 commits13 Mar 2026: 0 commits14 Mar 2026: 0 commits15 Mar 2026: 0 commits16 Mar 2026: 0 commits17 Mar 2026: 9 commits18 Mar 2026: 7 commits19 Mar 2026: 16 commits20 Mar 2026: 12 commits21 Mar 2026: 14 commits22 Mar 2026: 13 commits23 Mar 2026: 0 commits24 Mar 2026: 0 commits25 Mar 2026: 12 commits26 Mar 2026: 5 commits27 Mar 2026: 2 commits28 Mar 2026: 9 commits29 Mar 2026: 11 commits30 Mar 2026: 7 commits31 Mar 2026: 0 commits1 Apr 2026: 0 commits2 Apr 2026: 0 commits3 Apr 2026: 0 commits4 Apr 2026: 30 commits5 Apr 2026: 7 commits6 Apr 2026: 0 commits7 Apr 2026: 3 commits8 Apr 2026: 0 commits9 Apr 2026: 0 commits10 Apr 2026: 0 commits11 Apr 2026: 0 commits12 Apr 2026: 0 commits13 Apr 2026: 0 commits14 Apr 2026: 0 commits15 Apr 2026: 0 commits16 Apr 2026: 0 commits17 Apr 2026: 0 commits18 Apr 2026: 0 commits19 Apr 2026: 0 commits20 Apr 2026: 0 commits21 Apr 2026: 0 commits22 Apr 2026: 0 commits23 Apr 2026: 0 commits24 Apr 2026: 10 commits25 Apr 2026: 25 commits26 Apr 2026: 25 commits27 Apr 2026: 22 commits28 Apr 2026: 42 commits29 Apr 2026: 37 commits30 Apr 2026: 21 commits1 May 2026: 22 commits2 May 2026: 7 commits3 May 2026: 7 commits4 May 2026: 53 commits5 May 2026: 51 commits6 May 2026: 45 commits7 May 2026: 39 commits8 May 2026: 31 commits9 May 2026: 28 commits10 May 2026: 27 commits11 May 2026: 31 commits12 May 2026: 27 commits13 May 2026: 38 commits14 May 2026: 34 commits15 May 2026: 40 commits16 May 2026: 33 commits17 May 2026: 43 commits18 May 2026: 31 commits19 May 2026: 23 commits20 May 2026: 22 commits21 May 2026: 6 commits22 May 2026: 19 commits23 May 2026: 56 commits24 May 2026: 55 commits25 May 2026: 14 commits26 May 2026: 35 commits27 May 2026: 15 commits28 May 2026: 30 commits29 May 2026: 25 commits30 May 2026: 0 commits31 May 2026: 5 commits1 Jun 2026: 28 commits2 Jun 2026: 40 commits3 Jun 2026: 9 commits4 Jun 2026: 12 commits5 Jun 2026: 10 commits6 Jun 2026: 8 commits7 Jun 2026: 4 commits8 Jun 2026: 20 commits9 Jun 2026: 12 commits10 Jun 2026: 18 commits11 Jun 2026: 8 commits12 Jun 2026: 12 commits13 Jun 2026: 6 commits14 Jun 2026: 4 commits15 Jun 2026: 8 commits16 Jun 2026: 22 commits17 Jun 2026: 17 commits18 Jun 2026: 13 commits19 Jun 2026: 7 commits20 Jun 2026: 9 commits21 Jun 2026: 10 commits22 Jun 2026: 12 commits23 Jun 2026: 7 commits24 Jun 2026: 1 commit25 Jun 2026: 6 commits26 Jun 2026: 7 commits27 Jun 2026: 0 commits28 Jun 2026: 13 commits29 Jun 2026: 8 commits30 Jun 2026: 24 commits1 Jul 2026: 14 commits2 Jul 2026: 27 commits3 Jul 2026: 1 commit4 Jul 2026: 25 commits5 Jul 2026: 13 commits6 Jul 2026: 30 commits7 Jul 2026: 17 commits8 Jul 2026: 0 commits9 Jul 2026: 0 commits10 Jul 2026: 19 commits11 Jul 2026: 12 commits12 Jul 2026: 3 commits13 Jul 2026: 23 commits14 Jul 2026: 6 commits15 Jul 2026: 9 commits16 Jul 2026: 1 commit17 Jul 2026: 0 commits18 Jul 2026: 33 commits19 Jul 2026: 16 commits20 Jul 2026: 21 commits21 Jul 2026: 2 commits22 Jul 2026: 3 commits23 Jul 2026: 1 commit24 Jul 2026: 5 commits25 Jul 2026: 1 commit26 Jul 2026: 18 commits27 Jul 2026: 6 commits28 Jul 2026: 15 commits29 Jul 2026: 18 commits30 Jul 2026: 6 commits31 Jul 2026: 23 commits1 Aug 2026: 0 commits2 Aug 2026: 0 commits3 Aug 2026: 11 commits4 Aug 2026: 1 commit5 Aug 2026: 0 commits6 Aug 2026: 15 commits7 Aug 2026: 0 commits8 Aug 2026: 4 commits9 Aug 2026: 11 commits10 Aug 2026: 32 commits11 Aug 2026: 12 commits12 Aug 2026: 10 commits13 Aug 2026: 15 commits14 Aug 2026: 6 commits15 Aug 2026: 14 commits16 Aug 2026: 4 commits17 Aug 2026: 27 commits18 Aug 2026: 7 commits19 Aug 2026: 10 commits20 Aug 2026: 22 commits21 Aug 2026: 36 commits22 Aug 2026: 18 commits23 Aug 2026: 9 commits24 Aug 2026: 23 commits25 Aug 2026: 5 commits26 Aug 2026: 18 commits27 Aug 2026: 25 commits28 Aug 2026: 15 commits29 Aug 2026: 3 commits30 Aug 2026: 12 commits31 Aug 2026: 67 commits1 Sep 2026: 67 commits2 Sep 2026: 36 commits3 Sep 2026: 12 commits4 Sep 2026: 18 commits5 Sep 2026: 8 commits6 Sep 2026: 13 commits7 Sep 2026: 9 commits8 Sep 2026: 5 commits9 Sep 2026: 3 commits10 Sep 2026: 19 commits11 Sep 2026: 8 commits12 Sep 2026: 0 commits13 Sep 2026: 1 commit14 Sep 2026: 8 commits15 Sep 2026: 17 commits16 Sep 2026: 74 commits17 Sep 2026: 24 commits18 Sep 2026: 7 commits19 Sep 2026: 1 commit20 Sep 2026: 4 commits21 Sep 2026: 7 commits22 Sep 2026: 9 commits23 Sep 2026: 14 commits24 Sep 2026: 3 commits25 Sep 2026: 5 commits26 Sep 2026: 73 commits27 Sep 2026: 11 commits28 Sep 2026: 105 commits

    Commits per day01–78–1314–2526+

    2,947 commits across 163 active days. At least 7 in 10 commits to Falkor's main repository carry an AI coding agent's signature.

    Monthly totals (table view)
    MonthCommits
    March 2026117
    April 2026222
    May 2026892
    June 2026355
    July 2026368
    August 2026432
    September 2026561
    1. Early Mar 2026

      A local AI stack

      The DGo stack begins: local models, memory and services on one home PC.

      Project records

    2. 17 Mar 2026

      Falkor gets a face

      Falkor's clean-source repository starts: one cockpit over everything the stack can do.

      Git history

    3. Mar to Apr 2026

      Ten build phases

      Memory, agent orchestration, source quality and integrations, built phase by phase through the end of April.

      Roadmap

    4. May 2026

      The expansion

      826 Falkor commits in one month: a media studio, a pixel-display companion, personal tools and screen understanding.

      Git history

    5. 15 Jun 2026

      Falkor 3.0

      The shift from a pile of features to one product: one navigation, one model authority, production baselines.

      Project records

    6. 13 Jul 2026

      A clean-sheet rival

      Two AI models each designed a from-scratch successor to Falkor, and I judged them head to head. The winner, AURYN, was a Rust core with three hard rules aimed at Falkor's worst defects.

      Competition recordsThe whole story

    7. 9 Aug 2026

      Freezing the rewrite

      The rewrite's demo was all shells: 0 of 13 user journeys worked and the source no longer built. Falkor, checked the same day, was all green. I froze the rewrite, then had it audited for the parts worth keeping.

      Independent auditThe whole story

    8. Aug to Sep 2026

      Rebuilt in layers

      AURYN's best coded parts and unbuilt ideas became a package for reimagining Falkor, which was then rebuilt one group of pages at a time. In the first two days, 184 page routes became 98 with no capability lost.

      Route auditThe whole story

    9. Aug to Sep 2026

      Proving it

      Recovery programs, a 7,000-test browser battery made runnable, and certification gates that refuse to pass without evidence.

      Certification records

    10. 25 Sep 2026

      The full audit

      649 capabilities, 1,051 API operations and 7,331 dependencies mapped. It's the snapshot this site is built from.

      Capability audit

    Case studies

    Engineering, with receipts

    Each study follows the same arc: problem, evidence, investigation, decision, fix, verification, lesson. Every number comes from the record.

    • AI governance
    • Tool use
    • Approvals

    Make the AI admit what it didn't do

    Falkor's chat tools were silently switched off on every real turn, because the system prompt contained the word “Falkor”. The fix went further: each answer's claims are now checked against what the turn actually executed.

    5,068 → 502prompt tokens when the tools silently vanished
    • Quality engineering
    • Observability
    • Release governance

    Green is a claim, not a fact

    A certification harness that could declare success without evidence. A watchdog accusing a healthy store. 136 information sources shown as “active” that nothing ever ran. This is how Falkor learned to prove its own status.

    6ways certification could pass without evidence, all closed
    • AI-assisted engineering
    • Program ownership
    • Agent orchestration

    One builder, a fleet of AI agents

    Nearly 3,000 commits since March, most of them written by AI coding agents working in scoped lanes, with handoffs, evidence gates, independent review, and an operator who overturns a “PASS” that isn't one.

    7 in 10main-repository commits signed by an AI coding agent

    Beyond Falkor

    Other work, for clients and for fun

    The same habits outside the home lab: small, useful products with privacy designed in and honest status notes. Client names and people stay private.

    For clients and businesses

    • An IT services and staffing firm · Jul–Sep 2026

      An HR policy assistant, and an AI roadmap

      An on-premises assistant that answers only from published policy, checks every citation, and turns personal HR questions into a draft for a person instead of an AI answer.

      • Hybrid retrieval (Postgres full-text plus vector search) over a local model pinned by digest
      • Chats encrypted and deleted after 24 hours; every citation checked before it is shown
      • An 80-case acceptance set: 50 factual, 30 safety
      • A phased AI roadmap that ordered three products by risk, so an automated score can never reject a candidate
      • A recruiting-intelligence engine that records graded evidence instead of scores about people, with approved sources only and injection screening on every fetched page

      Built and tested; going live is the client's call.

      • RAG
      • pgvector
      • Local LLM
      • Evals
      • Next.js
      • Python
    • A mortgage loan officer

      A local-first CRM for a loan officer

      Pipeline, referral partners, tasks, templates and a closings calendar for one loan officer, running entirely in the browser: no server, no account, and no data leaving the laptop.

      • A Kanban pipeline from long shot to past client, with referral partners in tiers
      • Email templates with merge fields, and a calendar of birthdays and closings
      • CSV import with column mapping; JSON backup and restore
      • Plain JavaScript with no runtime dependencies, and a one-click Windows installer

      A leaner rewrite replaced the original after the original's own audit found it not release-stable.

      • JavaScript
      • IndexedDB
      • Offline-first
      • Zero dependencies
    • Teachers and parents

      Classroom wishlists, with the research to match

      Teachers post a classroom wishlist and parents chip in. A teacher console handles the page, messages to parents, payouts and a yearly archive.

      • Privacy by design: the public page shows progress and a donor count, never goals, totals or amounts
      • Stripe Connect checkout and webhooks, built and exercised in test mode
      • Moderated comments, seven progress visuals, and a status page showing which integrations are live and which are stubbed

      Working demo. Policy research found direct payouts to teachers conflict with district policy and state gift limits, so the real-money pilot waits for a compliant payout model.

      • Next.js
      • tRPC
      • Prisma
      • Stripe Connect
      • Postgres

    Side projects, for the fun of it

    • CouchCred

      A party game that runs itself. Phones join by room code, picks lock in, and every call settles from the live sports feed, then AI, then a crowd vote, or it scores no points.

      • Real-time rooms over WebSockets
      • About 17,000 lines of TypeScript and 32 test files
      • Next.js
      • WebSockets
      • Prisma
      • SQLite
    • God's Eye View, wired up

      A 3D world-events globe built by joining two open-source apps behind one bridge. It starts them when you open the page, sleeps them when idle, adds 29 live public-data layers, runs summaries on the local model, and holds voice AI to a hard $2 budget.

      • Built on God's Eye View (MIT) and World Monitor (AGPL-3.0)
      • A bridge of about 9,000 lines with 24 test files
      • CesiumJS
      • Docker
      • Node.js
    • A face for the voice

      A 64×64 LED panel that cycles through information scenes, and lip-syncs when Falkor speaks.

      • One serialized device writer with priority leases
      • About 11,700 lines and 36 test files
      • Node.js
      • Hardware
      • Pixel art
    • Signal Deck

      A radio and podcast player with no runtime dependencies, covering about 58,000 stations, plus a dial that links out to public amateur-radio receivers.

      • About 3,000 lines and 14 test files
      • JavaScript
      • Audio
      • Zero dependencies
    • Family Huddle

      A fantasy-football league designed for every age at the table: a live draft on one shared device, large-text modes and 44-pixel touch targets.

      • Hashed PINs and an audit log of manager overrides
      • Scoring from public sports data
      • Next.js
      • WebSockets
      • SQLite
    • DreamForge

      A child describes a game in a sentence, and the system builds it in Roblox. It now produces a real game file from one sentence; making it playable and editable is the unfinished half.

      • 2,323 tests passing
      • Paused, and labeled honestly as not done
      • .NET
      • React
      • Luau

    Also: a playoff salary-cap fantasy league with a lineup optimizer, a two-player pick'em with trading-card reveals, and a desktop prize wheel.

    About the builder

    Built by Dustin M. Gordon

    I lead quality engineering and modernization for mission-critical federal software: more than 20 years of software delivery, the last four supporting IRS tax-processing modernization.

    On Falkor I own the whole lifecycle: the vision, the roadmap and backlog, the architecture, delivery through a governed team of AI coding agents, testing and certification, deployment, day-to-day operations, and the audits that feed the next round. I have run it like an enterprise product from its first weeks, because that is the only way a system this size stays trustworthy. The habits are the ones I bring to every team: evidence before claims, negative controls on every gate, and nothing called done until it's proven.

    What I bring

    • Product ownership from first idea to daily operations
    • AI-assisted delivery at scale: governed agents, scoped lanes, evidence gates
    • Quality engineering, certification and release governance
    • Private, local-first AI architecture
    • Fast, safe adoption of new AI patterns, models and open source
    • Honest reporting: what's done, what isn't, and why

    What Falkor demonstrates

    Falkor workCapability shown
    Roadmap, 940-item delivery tracker, planned roundsProduct ownership and program management
    Capability audit: 649 capabilities, 7,331 dependenciesArchitecture discovery and systems thinking
    Agent lanes, handoffs and evidence gatesAI-assisted engineering management
    Gated, atomic deploys with automatic rollbackCI/CD and release engineering
    Certification harness and false-green huntingQuality engineering and release governance
    Approvals and the execution-claim guardAI governance and safety
    Ingestion engines over 459 sourcesData pipelines and reliability
    Recovery, watchdogs and self-healingOperations and incident response
    20-plus open-source projects, governed and kept currentRapid, safe technology adoption

    Hiring?

    Roles

    More than 20 years of mission-critical delivery, plus hands-on AI systems work: product ownership, quality engineering and AI-assisted delivery at scale.

    Have a project?

    Consulting

    I help teams turn scattered AI experiments into governed, observable products people actually use: roadmap, architecture, delivery process and the quality gates that keep it honest.

    Esc