The Kernel (lib/kernel/) could already answer four questions, and every surface on the platform read the same answer to each:
| Question | Module |
|---|---|
| What can this platform do? | capabilities.ts |
| Where does this request belong? | intent.ts |
| Where does a document live? | channels.ts |
| Is the produced work broken? | verify.ts |
It could not answer the one a person actually holds in their head: where am I in this, and what happens next? Nothing in the tree modelled the work — only the tools.
That gap did not look like one gap. It looked like nine. The Studio's status bar counted warnings. The Library card showed a thumbnail and a date. The Flow console ranked a next move without asking whether the current one was finished. An agent finished the moment its own source parsed, because "done" had no definition beyond "not syntactically broken". Each surface had invented a private idea of progress, and no two of them agreed — so a person moving between them could not tell whether "3 warnings" here was better or worse than "checked" there.
Four modules close it, and one vocabulary comes out of them. A fifth, memory.ts, closes the question the router itself could never answer — how does this particular person work? — and it is here rather than in a studio because a habit learned in ⌘K should not be a fact ⌘K alone knows.
1. phases.ts — the stages
Seven, in the order work actually happens:
intake → frame → build → verify → critique → refine → publish
Each stage names the question it answers in the user's voice ("What have I got?", "What am I actually making?", "Is it sound?", "Is it any good?"), what material it consumes and produces, and what evidence closes it. Exactly one is required — verify, and only because it is mechanical: a document whose source does not parse is not a matter of taste. Everything else is skippable, which is the honest model. Plenty of good work goes straight from build to publish, and a lifecycle that scolds you for that is one people learn to ignore.
The module owns the stages and nothing else. Which agent serves a stage is a fact about the agent, so it lives on the preset (AGENTS[].phase); which action serves one is a fact about the action, so it lives on the action seed. Both are projected. Tagging a new specialist phase: "critique" makes it appear in every critique affordance on the platform with no second edit — the same property that makes adding an engine enough to make it routable.
Every stage has at least one way to perform it, and tests/kernel-phases.test.ts fails if that stops being true. A stage the platform names in its vocabulary and cannot deliver is worse than one it does not name.
2. compose.ts — the whole programme, derived
intent.ts answers "where does this belong?" in at most three steps: a converter, a destination, a specialist. That shape is exactly right for a command palette and exactly wrong for what people actually ask for, which is a job:
"photo of the process on that whiteboard, as a BPMN model I can hand to the team, checked, with a figure for the deck"
Multi-step sequences did exist — eight hand-written pipelines in lib/flow/pipeline.ts — so a job nobody had thought to curate had no route at all, and adding an engine added no sequences.
The registry already declares what every move can be handed and what it leaves behind (accepts / produces). That is a graph: material kinds are the nodes, capabilities are the edges. Finding a programme is a shortest path over nine nodes — cheap, and complete: every route the registry admits is found, including the ones nobody wrote down.
Two properties are load-bearing:
- It cannot propose a step that cannot be taken. Coherence is the construction rule, not a check applied afterwards. The class of bug where a suggested workflow hands a photo to the BPMN modeller is unreachable rather than merely tested for.
- It is deterministic and it shows its working. Same inputs, same programme, same order, including every tie-break. Every step carries a
why, because a sequence whose reasoning is invisible is one people stop following the first time it is wrong.
composeLifecycle() is the other half: not "where does this request belong" but "I have drawn the thing — what is left?".
This only works because the material graph is honest. It used to sayprosewas produced by zero capabilities,databy zero,inkby zero. Three whole classes of route were missing and no surface could tell. Every studio now declares what it actually takes and leaves, andmaterialGraphGaps()fails the test suite if a kind ever becomes unreachable again.
3. critique.ts — the taste, held in one place
verify.ts refuses to have opinions, and is right to: a lint that can hold a correct diagram hostage is worse than no lint, because the author's only escape is to damage the work until the checker is happy.
So the taste lives here instead. Seven dimensions, each scored 0..1 with the findings that cost it, and each finding carrying the concrete edit that would raise it:
fidelity · completeness · structure · legibility · rigour · accessibility · portability
Three rules keep it honest:
- Nothing here can block anything. A
Critiquehas nook. The gate isverify.tsand it stays there. - A dimension it cannot assess is not scored. A Mermaid flowchart has no numbers to be rigorous about; scoring it 100% would inflate every flowchart above every dataset figure. Unassessable dimensions are reported as such and left out of the average.
- Every finding names its fix. A criticism you cannot act on is an insult with a percentage attached.
Two things it can see that nothing else could:
- Fidelity, because it holds the request. "You asked for a PRISMA flow and this is drawn in Mermaid." "The request names 1,284 records and the figure never states it." Only the Kernel is positioned to make that observation, and until now nothing did.
- Accessibility, computed rather than guessed. Every colour pair the source declares is simulated under protanopia, deuteranopia and tritanopia and scored by CIEDE2000, plus a greyscale luminance check — reusing the colour science
lib/science/accessibility.tsalready had and which nothing outside the Science studio could reach.
4. readiness.ts — one word, every room
Five levels, in order, each earned by evidence rather than asserted:
| Level | Means |
|---|---|
empty | nothing a person would miss |
draft | it exists, and something in it is broken |
sound | every mechanical check passes |
reviewed | a quality read was taken and it held up |
publishable | it would survive somebody else opening it |
The ordering is unforgiving in one direction: you cannot reach sound with a blocking defect however beautiful the figure is, and you cannot reach publishable with a colour pair a twelfth of readers cannot separate however correct it is. Everything else is advice.
Two rules stop the average flattering:
- **No single dimension may sit below a floor and still count as reviewed.** A weighted average is the right way to score and the wrong way to gate — a two-node stub scored 83% because six dimensions had nothing to complain about and the one that did was outvoted.
- An omission is never a pass. The number of unassessed dimensions is stated on the badge, in the API response, and in the text handed to a model.
readiness.ts composes the three modules above rather than forming a fourth opinion, which is what stops the fifth vocabulary becoming the fifth disagreement.
5. memory.ts — the one signal the router used to throw away
The intent router is a frozen substring scan over a frozen registry. That is the right shape for the platform's front door — microseconds, offline, signed out, the same answer in a test as in a browser, and a rationale that prints why it chose what it chose.
It also means the front door was the one part of the platform that could never improve. Somebody who has opened PlantUML forty times and Mermaid never still got Mermaid first for the word "UML", every time, forever. The evidence to know better was already on disk, in three ledgers kept for other reasons — the Pulse, the Lineage, and nothing at all for the one that mattered: a person looking at what the router put first, ignoring it, and choosing something else. That signal was discarded at the exact moment it appeared.
recall(ledgers, now) turns those ledgers into a small set of multipliers the router applies last, after the words and after the material. Four rules make that safe, and each one exists because the obvious version of this feature is worse than no feature at all.
It reorders near-ties; it does not overturn clear matches. The affinity multiplier is clamped to [0.85, 1.35]. The ceiling is the number that carries the argument: at 1.35 a habit can lift a capability past a rival it was within about a quarter of, and cannot lift it past one that beat it outright. It was set by taking the real score gaps in tests/kernel-intent and choosing the largest value that leaves every one of them intact. Someone who types "gantt" gets Gantt whatever their history says, because a router that can be argued out of an exact match by last week's habits is a router people stop typing into.
It needs a quorum. One correction is a mis-click as often as a signal. Two of the same correction is a preference, and only then does an exact repeated query pin. The quorum counts occurrences, not decayed weight — two deliberate corrections made in March are two corrections, and a quorum that summed decayed weight would score them as a fraction of a preference and never fire.
It forgets. Every signal decays on a fortnight half-life; the ledger has a ninety-day horizon. A vocabulary shifts — somebody who meant PlantUML by "UML" in March may mean Mermaid by it in September — and a memory with no horizon holds the March answer forever, with no way for the person to disagree except to stop using the product. Below three decayed events it emits nothing at all, so a first session routes exactly as it did before the layer existed.
It explains itself. Every bias carries a why string and the plan's rationale says it out loud: "PlantUML (engine) — matched "uml" — you have picked this over mermaid before". A ranking that moved for reasons the user cannot see is indistinguishable from a ranking that is broken.
The Kernel never reads storage. Ledgers are passed in, the same way lib/flow/state.ts and lib/flow/pipeline.ts take theirs — which is what lets the same function run on the server, in a test, and against a hypothetical history a settings screen wants to preview. The browser side is lib/hooks/useRecall.ts (one hook, subscribed to all three ledgers, so every omnibox biases identically) and lib/platform/corrections.ts (a bounded local ledger, best-effort writes, nothing that leaves the browser).
Settings shows what it has learned in plain words and erases it on its own; "Erase personalization" takes it too, because that button says it removes usage history and this is the ledger that decides what you see first.
Where it shows up
| Surface | What it shows |
|---|---|
| Studio status bar | the level, the reason, the seven dimensions, the stages left as buttons |
| Each studio's document bar | the same chip in Lab, BPMN, Weave, Science, Sketch and Figure |
/library | a badge on every card, and a Needs work sort |
/flow | "Where the work stands" — every document, weakest named |
| ⌘K | the router's alternatives when it is unsure, the derived programme when the job is more than one move, and a bias learned from the times it was wrong |
| Settings | what ⌘K has learned, in one sentence, with its own erase button |
| The agent loop | critique_diagram, plan_work, preflight_figure, and a finish gate that verifies what the source actually DRAWS |
POST /api/v1/critique | the whole verdict, with a gate parameter and a pass boolean for CI |
Each of them reads the Kernel. None of them computes a second reading of its own — which is the entire point.
The stage specialists
The agent roster was organised on two axes — which studio's canvas you can write, and which domain you know about — and both answer who. None answered when. A platform with ten ways to build something, no one whose job is to decide what is being built, and no one whose job is to say whether the result is any good, has a hole in the middle of it that no amount of building talent fills.
Five specialists fill it, and each is defined as much by what it refuses to do as by what it does:
frame— the framing analyst. Never builds. Produces a brief: the deliverable, its reader, the decision it supports, the notation and why it beats the runner-up, the altitude, what is in and out of scope, what would make the figure wrong, and how correctness will be established. Deep tier — a wrong brief is the most expensive artefact this platform can produce and the cheapest one to get right.verifier— adjudicates somebody else's project across every figure and every checker. Never edits: if it fixes something it has become the author, and nobody is left to check it.critic— reads it the way the person who will reject it reads it. Never edits. Starts from the deterministic critique, then adds what no checker can see: is the abstraction right, is anything asserted that is not true, what is missing that a reader will assume is absent on purpose.refine— acts on the critique and **must measure the delta it claims**. Never redesigns; never fixes by deletion.publish— the preflight. Column width, type-size floors at final size, contrast, colour-vision safety, greyscale, embedded assets, alt text, and an export recipe.
That discipline is the whole value. An agent that quietly slides into building is just the builder again, and the stage goes unowned a second time — which is exactly what happened to the two old flows, both of which ended in a review preset that finds problems and repairs them in one pass, so nothing was ever reported that was not also silently changed.
