Local-first
504 of the 649 catalogued capabilities run entirely on the home PC. Public web search runs only when a question needs it.
Falkor is a private, local-first AI system that runs our home. It handles chat, memory, news, media, displays and automation, plus the operations that keep it honest, all on one home PC under one set of rules.
Flight recorderReplays of recorded behaviour
You“Remind me tomorrow afternoon to call the plumber.”
FalkorDone. I'll remind you tomorrow at 6 PM.
Reconstructed from Falkor's logs and case studies; prompts are illustrative. Not a live connection.How the tools were fixed →
Pick a scenario to watch a real request move through Falkor, or hover a part of the map to see what it does.
What Falkor is
Falkor began in March 2026 as a local AI stack and grew into one product: a cockpit of 100 pages over 1,051 API operations, local models that answer on the home PC, a long-term memory with review built in, engines that read hundreds of sources, and the operational tooling to prove what is actually working.
504 of the 649 catalogued capabilities run entirely on the home PC. Public web search runs only when a question needs it.
Consequential actions wait for approval, and each answer is checked against what the turn actually did, so the AI can't claim work it never ran.
Status surfaces are tested for false greens as hard as for failures. One certification run covers 56 gates and more than 7,000 browser tests.
It runs the living-room TV, the morning news, the calendar and reminders. It's daily infrastructure, not a demo.
Inside Falkor
Captured from the running system on 28 Sep 2026 and cropped to the page itself. Select any screen to see it full size; arrow keys page through them.

Local AI World
The weekly local-AI briefing, and it admits when it falls short: 9 of 10 items, 33 sources failed, no filler added.

Spin Lens
An argument machine: the strongest case for and against, with facts tied to evidence and assumptions labeled.

A dossier on anything
Research first, then writing from what could be verified, with the source one click from every statement.

Reports & Briefings
Nine source-backed reports from one engine, each with a browser and an email edition.

Capability registry
What Falkor can do right now: 107 ready, 86 waiting for approval, 114 that say why not.

Open-source engines
Self-hosted engines that only the home PC can reach, each checked without its own login.

Self-healing
Repairs what it safely can, and asks for a person only when a repair can't be automatic.

Custodian cockpit
An always-on agent sweeps every two minutes. As of 28 Sep: 3,084 recoveries, none missed.

AI & Models
One chat model for every surface, tools behind approvals, and a persona held to a prompt budget.
Showcase
Pick one. Each opens a real screen from Falkor, annotated, or a working demo you can drive yourself.
Data & integration · 01
Nine reports on the news, local AI, sports, entertainment and more come from one report engine, each with browser and email editions. When a report falls short, it says so instead of padding.
Local AI World this week: 9 of 10 items delivered, 33 sources failed, no filler added
What you're looking at Local AI World
What you're looking at Reports & Briefings
Chat
Remind me Friday to renew the car registration.
Done. I'll remind you Friday at 9 AM. I've also emailed you the renewal form.
Correction: no email was sent this turn.
Turn record
create_remindersucceededsend_emailnever calledClaim guard: 1 unbacked claim struck and corrected.
The answer can only claim what the record shows.
AI · 02
Ask in plain English for a reminder, a note or your schedule. Falkor does it with real tools, and every answer is checked against what the turn actually ran.
A fabricated claim was caught and corrected live
Read the case studyHow it works
Data & integration · 03
Name a person, a band, a team, a product or a topic. Falkor researches official, reference and archival sources, then writes a dossier from what it can verify, and shows what it found and what it couldn't.
One dossier: 74 verified statements from 3 sources
What you're looking at A dossier on anything
Review inbox · 3
Prefers the local-AI briefing before sports.
The car is due for service in May.
Recycling goes out Thursday night.
Inbox clear.
Long-term memory · 0
Nothing yet. Only what you approve lands here.
Nothing becomes memory without a decision.
AI · 04
Say “remember this” in chat, from the clipboard or from any app's share menu. It lands in a review inbox, and only what you approve becomes long-term memory that grounds future answers.
Promotion is a decision, not a side effect
How it works
AI · 05
Spin Lens takes a claim, an article or pasted text and builds the strongest case for and against it, steelmans each side and pressure-tests the leader. It's a guard against one-sided answers, including the AI's own.
Runs on the local model, with private search first
What you're looking at Spin Lens
The recorded run: killed at source 16 of 110, then resumed with zero duplicate stories.
Data & integration · 06
459 sources across feeds, APIs and sites, collected by two engines. Duplicates are fetched once, failing sources are quarantined and re-probed, and a collection pass survives a crash.
Force-killed at source 16 of 110, resumed with zero duplicate stories
How it works
One GPU
One chat model you choose; scheduled work borrows the GPU and hands it back.
AI · 07
You choose the chat model. Scheduled jobs such as the morning briefing load their own models on the same GPU and hand it back, and a guard refuses cloud models that pose as local ones.
Chat, then a scheduled job, then chat again, each on the right model
How it works
What you're looking at AI & Models
What you're looking at Model residency
Operations · 08
One registry, built from each owner's own records, lists every capability, whether it works right now and how it is used safely. The shelf explains; it executes nothing.
307 capabilities: 107 ready, 86 waiting for approval, 114 blocked with a reason
What you're looking at Capability registry
Data & integration · 09
More than 20 open-source projects run as managed engines: photo library, document archive, private search, uptime monitoring, notifications, recipes, web archiving, workflow automation, design and vector search.
Engine upgrades are re-validated against the full test battery
What you're looking at Open-source engines
Operations · 10
Watchdogs cover every core service and an always-on custodian sweeps every two minutes. Stop everything and the stack restores itself in 6.8 seconds; after a cold reboot it comes back unaided in about 14 minutes.
Full stack restored in 6.8 s, measured 26 Aug 2026
What you're looking at Self-healing
What you're looking at Custodian cockpit
GPU
Falkor holds the GPU for its AI work.
Operations · 11
One switch hands the GPU to a game: background AI work pauses, services throttle, and everything is restored afterwards.
Throttle and restore verified live
How it works
| Process | Owner | Idle | Action |
|---|---|---|---|
falkor-web | Falkor | active | |
helper.exe | none (orphan) | 14 min | |
updater.exe | unknown | 3 min | |
worker.exe | none (orphan) | 4 min |
Every stop passes four gates, and anything unknown fails closed
Pick a process and try to stop it.
Operations · 12
See every Windows process, who owns it and whether it is idle. Stopping one takes a fresh census, a verified orphan at least ten minutes old and a typed confirmation, and every action is audited.
Anything unknown fails closed
How it works
Screens captured from the running system on 28 Sep 2026 and cropped to the page. Demos run in your browser on sample data.
What it does
The audit catalogued 649 capabilities across 26 families, grouped here into ten areas. Each bar shows the share observed running during the audit (459 in all). A read-only audit can't exercise everything, so "not observed" doesn't mean broken.
Capability audit · 25 Sep 2026Local chat models, long-term memory with retrieval, personas and screen understanding.
66 of 112
News from hundreds of sources, weather and radar, sports and local events: ranked, de-duplicated and explained.
85 of 94
Reminders, the family calendar, documents, read-only email and shared household tools.
79 of 93
A living-room TV experience, music, radio and podcasts, video discovery, and local image and video generation.
53 of 66
Any screen can become a Falkor display: TV dashboards, weather radar and kiosks.
Smart-home control is built, but the audit did not observe it running.
7 of 14
Push-to-talk speech in, natural speech out.
Always-on listening is built and deliberately switched off.
5 of 11
Scheduled jobs, notifications, a tool shelf and agents, with approval gates on consequential actions.
54 of 68
A registry and SDK that let new apps borrow Falkor's models, memory and status, plus a Labs shelf for side projects.
66 of 79
Origin guards, approvals, a local-only model rule, and every capability classified by privacy and cloud exposure.
7 of 17
Health checks, watchdogs, self-healing, certification and controlled deployment.
A read-only audit can't exercise most of these, so many show as not observed.
37 of 95
By the numbers
Every figure carries its source and date. None of them are estimates, and none update themselves: the site changes only when a new snapshot is published.
Unless marked · Capability audit · 25 Sep 2026649
capabilities catalogued
459 observed running during the audit, across 26 families
1,051
API operations
100
Pages in the cockpit
381
Tools
7,331
Mapped dependencies
69
Containers
54 running at audit time
36
Scheduled background jobs
78%
Capabilities that are local-only
504 of 649
2,947
Commits since March 2026
Falkor and its stack
7,215 of 7,218
Browser tests passed in one certification run
The three failures were root-caused and fixed afterwards
How it's run
I own Falkor's whole lifecycle: planning, architecture, delivery, verification, deployment, operations and audit. It has been run like an enterprise program from its first weeks, with a roadmap, a groomed backlog, an agile cadence, a CI/CD pipeline, and AI agents working as the delivery team.
01 · Plan
Falkor has had a phased roadmap since its first weeks: ten build phases by the end of April, then the Falkor 3.0 program. Work lives in a delivery tracker in which every item carries a phase, a gate, its dependencies, an owner and a completion check.
940
tracked items, 644 done
02 · Build
Work is cut into lanes: briefs with a goal, acceptance criteria, required reading, an ordered scope and a do-not-touch list. AI coding agents implement them. I own the architecture, the rules and the review.
541
named lanes in the canonical record
03 · Verify
Every new gate must be shown to fail before it is trusted to pass. Certification runs 56 gates over more than 7,000 browser tests, and a lane isn't done until its living documentation says so too.
7,215 / 7,218
browser tests passed in one run
04 · Ship
A self-hosted CI/CD pipeline carries every change. Blocking checks run first: change impact, documentation parity, audit and coverage. Deploys are atomic and refuse uncommitted code, the build ID comes from the commit, and a failed readiness check rolls the release back.
Every deploy
commit, bundle and live build verified identical
05 · Operate
Watchdogs cover every core service, a proactive status check runs every 30 minutes, and readiness is checked after every boot. When something fails, it heals: the full stack comes back in 6.8 seconds.
11 / 11
critical readiness checks green
06 · Audit
A read-only audit mapped 649 capabilities, 1,051 API operations and 7,331 dependencies, then triaged 412 findings by severity. It was checked for internal consistency before any of it entered the canonical docs.
412
findings triaged by severity
07 · Improve
Findings become roadmap items for the next round. New open models, runtimes and open-source releases are evaluated as they land, and engine upgrades are re-validated against the full test battery before they stay.
0
test failures attributable to the 22 Sep engine upgrades
Every 30 minutes
A proactive status check, plus watchdogs on every core service.
Every day
Standup-style planning with the AI agents (what closed, what's next, what's blocked), defect triage, backlog grooming, and mixture-of-experts brainstorming: several AI models weigh in on open questions, then I decide and the decision is recorded. Software-update checks, memory re-indexing and fresh briefings run on their own.
Every week
A review of the AI landscape (new open models, runtimes and trending open-source projects), with mixture-of-experts sessions on the bigger decisions: what to adopt, what to retire, what comes next.
Every round
A planned round of parallel lanes with a steward, acceptance criteria and a definition of done. Nothing closes until its evidence and documentation are in.
| On Falkor | Agile equivalent |
|---|---|
| Round of parallel lanes | Sprint |
| Lane brief with checkable rules | User story with acceptance criteria |
| Delivery tracker | Product backlog |
| “Done-done” plus updated living docs | Definition of done |
| Daily planning and triage with the agents | Daily standup and bug triage |
| Mixture-of-experts brainstorming, decision recorded | Design review and decision record |
| Steward review | Tech-lead review and sign-off |
| Closeout: what's claimed, and what isn't | Sprint review |
| Audit-derived roadmap | Retrospective into backlog |
AI practice
AI's vocabulary changes every few months, and Falkor has moved with it. I treat AI as a powerful tool with known failure modes: I adopt what's new quickly, measure it honestly, and design around what it gets wrong.
From Prompt engineeringto Context engineering
Each turn's context is assembled from structure: the capability map, the persona, local facts and recalled memory. It is refreshed every 15 minutes and trimmed by policy. A versioned structural marker, never a word in the prompt, decides what gets injected.
From Keyword searchto RAG and vector memory
Approved memories are embedded by a dedicated local model, indexed in a vector store, re-indexed nightly, and recalled into answers.
From Workflow automationsto MCP tools
Scheduling moved into Falkor's own scheduler, n8n workflows stay behind approvals, and capabilities are served as MCP tools through a bridge, in a tool universe of 462 tools from 8 providers.
From Chatbotsto Agents
OpenClaw runs on Falkor's own model selection, with Falkor's memory behind a proxy. Hermes works as a 24/7 custodian that can repair services. The agentic modes ask for approval before they act.
From Bigger context windowsto Deliberate token budgets
Budgets are measured, not guessed. A reasoning model given 900 tokens produced nothing and 2,400 produced a full answer. A batch setting that forced a model reload on every turn was found and fixed.
From One big modelto A model plane
One chat model you choose, workload models scheduled on the same GPU, cloud models refused at the boundary, and several models consulted side by side (mixture-of-experts brainstorming) before a decision.
| Where AI falls short | What Falkor does about it |
|---|---|
| Models claim work they didn't do | Every answer is checked against the turn's execution record. |
| One date word can send a private question to the public web | A positive classifier decides whose data a question is about before anything is searched. |
| Reasoning models can spend the whole budget thinking | Token budgets are sized for thought and verified per model. |
| A local model can crash on a large prompt | The failure is reported and retried within bounds. The gap is never filled with invented text. |
| AI coding agents over-report success | Evidence gates, independent review, and an operator who signs off. |
| GPU memory is finite | A model plane schedules who is resident; workloads hand the GPU back, and Game Mode frees it entirely. |
The rule that forbids hard-coded model names had exempted its own file, and was hiding two violations. It now scans its own source like everything else, and a missing self-scan fails the run.
41 endpoints were shared by 157 sources. The first row fetches and every twin records the same result: 29.7% fewer wasted fetches, and no source hidden.
Sources that have failed 800 or more times in a row move to a 6 to 24 hour re-probe, and rejoin the moment one fetch succeeds. Nothing is deleted to make a dashboard greener.
The engines can quiesce: hold new work, drain, checkpoint and keep running. A retention job deleted 871 rows in the middle of a collection pass without stopping it.
When Falkor speaks, a small pixel display lip-syncs to the reply, a playful piece of hardware wired into the voice pipeline.
Each chat turn leaves a record of what actually ran. That record is how false claims get caught, and it's what the panel at the top of this page replays.
How it works
Simplified on purpose. The audit maps 7,331 dependencies like these, from each page down through the APIs, libraries and services beneath it to the data it touches.
A question becomes a governed turn on a local model.
Public web search joins only when the question needs it.
Memory is earned: nothing becomes long-term without review.
Memory stays on the home PC.
Hundreds of sources become one explained feed.
Source count as of September 2026.
The story
From early March to late September 2026. This history comes from Falkor's own records (git history, roadmaps and audits), not from memory.
Falkor and its stack, 1 Mar 2026 to 28 Sep 2026
2,947 commits across 163 active days. At least 7 in 10 commits to Falkor's main repository carry an AI coding agent's signature.
| Month | Commits |
|---|---|
| March 2026 | 117 |
| April 2026 | 222 |
| May 2026 | 892 |
| June 2026 | 355 |
| July 2026 | 368 |
| August 2026 | 432 |
| September 2026 | 561 |
Early Mar 2026
The DGo stack begins: local models, memory and services on one home PC.
Project records
17 Mar 2026
Falkor's clean-source repository starts: one cockpit over everything the stack can do.
Git history
Mar to Apr 2026
Memory, agent orchestration, source quality and integrations, built phase by phase through the end of April.
Roadmap
May 2026
826 Falkor commits in one month: a media studio, a pixel-display companion, personal tools and screen understanding.
Git history
15 Jun 2026
The shift from a pile of features to one product: one navigation, one model authority, production baselines.
Project records
13 Jul 2026
Two AI models each designed a from-scratch successor to Falkor, and I judged them head to head. The winner, AURYN, was a Rust core with three hard rules aimed at Falkor's worst defects.
Competition recordsThe whole story
9 Aug 2026
The rewrite's demo was all shells: 0 of 13 user journeys worked and the source no longer built. Falkor, checked the same day, was all green. I froze the rewrite, then had it audited for the parts worth keeping.
Independent auditThe whole story
Aug to Sep 2026
AURYN's best coded parts and unbuilt ideas became a package for reimagining Falkor, which was then rebuilt one group of pages at a time. In the first two days, 184 page routes became 98 with no capability lost.
Route auditThe whole story
Aug to Sep 2026
Recovery programs, a 7,000-test browser battery made runnable, and certification gates that refuse to pass without evidence.
Certification records
25 Sep 2026
649 capabilities, 1,051 API operations and 7,331 dependencies mapped. It's the snapshot this site is built from.
Capability audit
Logbook
Numbers show what happened. The logbook shows why: decisions and their reasons, mistakes and what they taught, blockers, and the questions still open.
The featured story
Case studies
Each study follows the same arc: problem, evidence, investigation, decision, fix, verification, lesson. Every number comes from the record.
Falkor's chat tools were silently switched off on every real turn, because the system prompt contained the word “Falkor”. The fix went further: each answer's claims are now checked against what the turn actually executed.
A certification harness that could declare success without evidence. A watchdog accusing a healthy store. 136 information sources shown as “active” that nothing ever ran. This is how Falkor learned to prove its own status.
Nearly 3,000 commits since March, most of them written by AI coding agents working in scoped lanes, with handoffs, evidence gates, independent review, and an operator who overturns a “PASS” that isn't one.
Beyond Falkor
The same habits outside the home lab: small, useful products with privacy designed in and honest status notes. Client names and people stay private.
An IT services and staffing firm · Jul–Sep 2026
An on-premises assistant that answers only from published policy, checks every citation, and turns personal HR questions into a draft for a person instead of an AI answer.
Built and tested; going live is the client's call.
A mortgage loan officer
Pipeline, referral partners, tasks, templates and a closings calendar for one loan officer, running entirely in the browser: no server, no account, and no data leaving the laptop.
A leaner rewrite replaced the original after the original's own audit found it not release-stable.
Teachers and parents
Teachers post a classroom wishlist and parents chip in. A teacher console handles the page, messages to parents, payouts and a yearly archive.
Working demo. Policy research found direct payouts to teachers conflict with district policy and state gift limits, so the real-money pilot waits for a compliant payout model.
A party game that runs itself. Phones join by room code, picks lock in, and every call settles from the live sports feed, then AI, then a crowd vote, or it scores no points.
A 3D world-events globe built by joining two open-source apps behind one bridge. It starts them when you open the page, sleeps them when idle, adds 29 live public-data layers, runs summaries on the local model, and holds voice AI to a hard $2 budget.
A 64×64 LED panel that cycles through information scenes, and lip-syncs when Falkor speaks.
A radio and podcast player with no runtime dependencies, covering about 58,000 stations, plus a dial that links out to public amateur-radio receivers.
A fantasy-football league designed for every age at the table: a live draft on one shared device, large-text modes and 44-pixel touch targets.
A child describes a game in a sentence, and the system builds it in Roblox. It now produces a real game file from one sentence; making it playable and editable is the unfinished half.
Also: a playoff salary-cap fantasy league with a lineup optimizer, a two-player pick'em with trading-card reveals, and a desktop prize wheel.
About the builder
I lead quality engineering and modernization for mission-critical federal software: more than 20 years of software delivery, the last four supporting IRS tax-processing modernization.
On Falkor I own the whole lifecycle: the vision, the roadmap and backlog, the architecture, delivery through a governed team of AI coding agents, testing and certification, deployment, day-to-day operations, and the audits that feed the next round. I have run it like an enterprise product from its first weeks, because that is the only way a system this size stays trustworthy. The habits are the ones I bring to every team: evidence before claims, negative controls on every gate, and nothing called done until it's proven.
| Falkor work | Capability shown |
|---|---|
| Roadmap, 940-item delivery tracker, planned rounds | Product ownership and program management |
| Capability audit: 649 capabilities, 7,331 dependencies | Architecture discovery and systems thinking |
| Agent lanes, handoffs and evidence gates | AI-assisted engineering management |
| Gated, atomic deploys with automatic rollback | CI/CD and release engineering |
| Certification harness and false-green hunting | Quality engineering and release governance |
| Approvals and the execution-claim guard | AI governance and safety |
| Ingestion engines over 459 sources | Data pipelines and reliability |
| Recovery, watchdogs and self-healing | Operations and incident response |
| 20-plus open-source projects, governed and kept current | Rapid, safe technology adoption |
Hiring?
More than 20 years of mission-critical delivery, plus hands-on AI systems work: product ownership, quality engineering and AI-assisted delivery at scale.
Have a project?
I help teams turn scattered AI experiments into governed, observable products people actually use: roadmap, architecture, delivery process and the quality gates that keep it honest.