TL;DR: Timothy’s chat turns suit questions. They are the wrong shape for “refactor this module” or “research this and write it up”, work that takes dozens of model turns, must survive restarts, and needs checking. That work runs as a mission: a goal driven through fixed phases by a pure state machine, with every transition written to Postgres. Three rules shape it.
First, state lives in rows, not in a transcript. Every worker turn is seeded from a packet (goal, plan, progress notes, git log), not from the last turn’s transcript. Resuming after a crash is a new session plus that packet, never a replay.
Second, the harness decides when work is done, never the model. A worker’s DONE is a request for verification. A unit’s passes flag flips only on a check the harness ran itself.
Third, every loop has a brake, and most brakes park rather than throw work away. A stopped mission pauses with a typed reason, keeps its branch, and tells me.
The problem
What does long work need that a chat turn does not have?
A chat turn lives in one request and one growing context. Long work breaks both: the context fills with tool output, the process restarts, and a model that says “all tests pass” may never have run them.
| Requirement | Detail |
|---|---|
| Durable | A brain restart mid-mission loses nothing; the driver picks it up from the row. |
| Honest | “Done” means a check passed, not that a model said so. |
| Bounded | No mission can loop forever or spend without limit. |
| Unattended | Missions run on a schedule and reach me when they need me. |
Challenge 1: how does work outlive a conversation?
Why not just give a chat turn more tools and more time?
| Option | How it works | Why not (or why) |
|---|---|---|
| A. A long chat turn | One agent loop runs until the model stops | Unbounded context, a restart kills it, and the model decides when it is finished. |
| B. An agent framework’s loop | The framework holds plan and memory in process | The state I most need to inspect and resume lives in someone else’s objects. |
| C. A state machine over rows | A pure Step(state, input) decides every transition; workers are short, fresh sessions | Restartable, testable without a model, and every decision is an event row. |
I chose C. Step has no I/O: it takes phase, status and counters, and returns the next state plus the events to append. Store.ApplyTransition is the only writer of mission state, in one transaction under a row lock. Phase and status are separate axes, so “in build, paused for budget” is one readable state.
- Fixed phases.
discoverexplores,plansubmits units (artifacts, acheck_cmd, acceptance criteria),buildworks one unit per turn,provereviews, andresultdelivers (push, pull request, destinations) with zero model calls. Lighter flows skip phases; the flow is fixed at creation and no tool call can change it. - Fresh sessions. A native worker’s whole context is the rendered packet as one user message. Durability lives in the plan, the progress log and git history.
- Sentinel tools. Each phase reports through a tool call (
submit_plan,mission_status,review_verdict), so results are structured, not prose.
Challenge 2: who decides that the work is done?
Can the worker that wrote the code also grade it?
| Option | How it works | Why not (or why) |
|---|---|---|
| A. Trust the worker’s DONE | The worker reports success | Models claim success on unfinished work, and on goals that cannot be done. |
| B. An LLM reviewer alone | A second model reads the result | Still a claim about a claim. It cannot see what never ran. |
| C. Harness evidence, then a reviewer | Go checks artifacts and runs check_cmd; a fresh reviewer judges what checks cannot | Evidence the model cannot fake, plus judgement where existence checks fall short. |
I chose C. After every DONE the harness re-checks every unit, so an earlier unit’s regression is caught too. A failing check costs a worker turn, with its output in the next packet. The reviewer is a fresh session on its own review_route, which exists so review can run on a different provider than the worker: a model is a weak judge of its own mistakes.
- Gates must fail first. At plan time a sandbox probe runs each
check_cmdagainst the untouched tree. A gate that already passes proves nothing, so the plan is rejected. - Artifacts before commands.
CheckArtifactsconfirms declared files exist, are non-empty and sit inside the workspace before any model-written command runs. - Findings need evidence. A blocking finding must name a changed file and quote it, or Go demotes it to minor. A criterion marked
not_metforces rework even if the reviewer approved.
Challenge 3: who does the work?
Should Timothy write code itself, or hand it to a coding CLI?
| Option | How it works | Why not (or why) |
|---|---|---|
| A. Native only | Every unit runs on Timothy’s agent loop | Fine for research and writing; dedicated coding CLIs have far more tuning for code. |
| B. Delegate everything | Every mission goes to a CLI | Loses Timothy’s skills, memory, knowledge base and permission chain. |
| C. Both behind one contract | Runner.RunWorker takes a packet and returns a WorkerVerdict either way | The state machine and verification never know which kind ran. |
I chose C. A coding mission may name a harness (claude, codex, cursor, opencode or pi). The CLI runs detached in the mission’s sandbox, the brain polls its output, and it reports through a DONE/RETRY/BLOCKED schema. It may resume its own CLI session between units, but the packet is always complete. Everything else runs native.
- An explicit harness is a contract. If no route entry can serve it, the mission pauses as infra and resumes when the cooldown ends. A silent native fallback would change tools and prompt mid-mission.
- One sandbox per mission.
sandboxd, the only Docker-socket holder, runs it as an unprivileged user with a read-only root filesystem, no capabilities, resource caps, and onlyPATHandHOMEin its environment. - A clone per mission. Each coding mission gets its own clone and branch, one commit per finished unit.
- Delegated review is optional. A CLI reviewer runs read-only; any failure falls back to native review.
Challenge 4: how does it run when I am not watching?
What starts a mission at 06:00, and what stops one that never ends?
| Option | How it works | Why not (or why) |
|---|---|---|
| A. Cron creates missions directly | A ticker inserts mission rows | Notifications and run tracking become side effects a crash can lose. |
| B. A message broker | Triggers and outcomes go through a queue | Another service to run, for one user’s volume. |
| C. An events inbox in Postgres | Side effects are events rows written in the same transaction as the change; a drainer delivers them after commit | Nothing is lost between commit and send, in the database I already run. |
I chose C. Schedules are automations: an agent, a mission template, triggers (cron, webhooks, channel messages) and destinations. Each run is an ordinary mission created through the API’s path. When a mission reaches done, failed, paused or waiting for input, ApplyTransition writes an events row; consumers notify me, reply in the originating channel, and close the automation run.
The brakes (mission ceilings are operator settings, read per turn):
| Brake | Stops |
|---|---|
| Budget | Ledger spend reaching the mission’s cap: pause |
| Backoff | Three worker failures in a row: pause |
| Stall | The same failure twice: one automatic replan, then pause |
| Rework | Review rounds reaching max_iterations: pause, branch kept |
| Iterations | max_iterations (default 8) retries in a phase: fail |
| Automations | Runs per hour, concurrency mode, a dedup key per boundary, a breaker after three failed runs |
Every harness change must pass make canary: a real mission against the live stack, with a randomly drawn goal so caches cannot flatter it. It fails on any permission park, no harness-verified unit, or too many turns. A sibling canary inverts the test: its goal names a missing file, and it passes only if the mission fails honestly without fabricating it.
What broke
The rollback that deleted good work. Retries used to roll the working tree back. A delegated CLI wrote a module and its test, reported RETRY to fix one function, and the rollback deleted both. RETRY means unfinished, not wrong, so no retry rolls back now; only a review rework does.
Lessons
- Put long-running state in rows a pure function moves; transcripts are for reading, not resuming.
- A model’s “done” is a request for verification, and only harness evidence answers it.
- A check that passes before the work starts proves nothing, so test the gate too.
- Give review its own route so it can run on another provider; let the harness have the final say.
- A brake should park with a reason whenever the work might still be worth keeping.
- Side effects that must survive a crash belong in the same transaction as the change that caused them.