System Design September 15, 2026 ~9 min read

Designing a Self-Hosted AI Assistant (Part 7): Missions

TL;DR: Timothy’s chat turns suit questions. They are the wrong shape for “refactor this module” or “research this and write it up”, work that takes dozens of model turns, must survive restarts, and needs checking. That work runs as a mission: a goal driven through fixed phases by a pure state machine, with every transition written to Postgres. Three rules shape it.

First, state lives in rows, not in a transcript. Every worker turn is seeded from a packet (goal, plan, progress notes, git log), not from the last turn’s transcript. Resuming after a crash is a new session plus that packet, never a replay.

Second, the harness decides when work is done, never the model. A worker’s DONE is a request for verification. A unit’s passes flag flips only on a check the harness ran itself.

Third, every loop has a brake, and most brakes park rather than throw work away. A stopped mission pauses with a typed reason, keeps its branch, and tells me.

The problem

What does long work need that a chat turn does not have?

A chat turn lives in one request and one growing context. Long work breaks both: the context fills with tool output, the process restarts, and a model that says “all tests pass” may never have run them.

RequirementDetail
DurableA brain restart mid-mission loses nothing; the driver picks it up from the row.
Honest“Done” means a check passed, not that a model said so.
BoundedNo mission can loop forever or spend without limit.
UnattendedMissions run on a schedule and reach me when they need me.

Challenge 1: how does work outlive a conversation?

Why not just give a chat turn more tools and more time?

OptionHow it worksWhy not (or why)
A. A long chat turnOne agent loop runs until the model stopsUnbounded context, a restart kills it, and the model decides when it is finished.
B. An agent framework’s loopThe framework holds plan and memory in processThe state I most need to inspect and resume lives in someone else’s objects.
C. A state machine over rowsA pure Step(state, input) decides every transition; workers are short, fresh sessionsRestartable, testable without a model, and every decision is an event row.

Mission phases, loops and pause reasons
Figure 1: The mission pipeline, with rework, the one automatic replan, the flows that skip prove, and the typed reasons a mission can park on.

I chose C. Step has no I/O: it takes phase, status and counters, and returns the next state plus the events to append. Store.ApplyTransition is the only writer of mission state, in one transaction under a row lock. Phase and status are separate axes, so “in build, paused for budget” is one readable state.

  • Fixed phases. discover explores, plan submits units (artifacts, a check_cmd, acceptance criteria), build works one unit per turn, prove reviews, and result delivers (push, pull request, destinations) with zero model calls. Lighter flows skip phases; the flow is fixed at creation and no tool call can change it.
  • Fresh sessions. A native worker’s whole context is the rendered packet as one user message. Durability lives in the plan, the progress log and git history.
  • Sentinel tools. Each phase reports through a tool call (submit_plan, mission_status, review_verdict), so results are structured, not prose.

Challenge 2: who decides that the work is done?

Can the worker that wrote the code also grade it?

OptionHow it worksWhy not (or why)
A. Trust the worker’s DONEThe worker reports successModels claim success on unfinished work, and on goals that cannot be done.
B. An LLM reviewer aloneA second model reads the resultStill a claim about a claim. It cannot see what never ran.
C. Harness evidence, then a reviewerGo checks artifacts and runs check_cmd; a fresh reviewer judges what checks cannotEvidence the model cannot fake, plus judgement where existence checks fall short.

Harness verification and the review round
Figure 2: A worker's DONE triggers a commit and a harness check of every unit; only harness-passed units reach the reviewer, and only they can flip passes.

I chose C. After every DONE the harness re-checks every unit, so an earlier unit’s regression is caught too. A failing check costs a worker turn, with its output in the next packet. The reviewer is a fresh session on its own review_route, which exists so review can run on a different provider than the worker: a model is a weak judge of its own mistakes.

  • Gates must fail first. At plan time a sandbox probe runs each check_cmd against the untouched tree. A gate that already passes proves nothing, so the plan is rejected.
  • Artifacts before commands. CheckArtifacts confirms declared files exist, are non-empty and sit inside the workspace before any model-written command runs.
  • Findings need evidence. A blocking finding must name a changed file and quote it, or Go demotes it to minor. A criterion marked not_met forces rework even if the reviewer approved.

Challenge 3: who does the work?

Should Timothy write code itself, or hand it to a coding CLI?

OptionHow it worksWhy not (or why)
A. Native onlyEvery unit runs on Timothy’s agent loopFine for research and writing; dedicated coding CLIs have far more tuning for code.
B. Delegate everythingEvery mission goes to a CLILoses Timothy’s skills, memory, knowledge base and permission chain.
C. Both behind one contractRunner.RunWorker takes a packet and returns a WorkerVerdict either wayThe state machine and verification never know which kind ran.

Native and delegated workers
Figure 3: Native and delegated workers return the same verdict to the same harness checks.

I chose C. A coding mission may name a harness (claude, codex, cursor, opencode or pi). The CLI runs detached in the mission’s sandbox, the brain polls its output, and it reports through a DONE/RETRY/BLOCKED schema. It may resume its own CLI session between units, but the packet is always complete. Everything else runs native.

  • An explicit harness is a contract. If no route entry can serve it, the mission pauses as infra and resumes when the cooldown ends. A silent native fallback would change tools and prompt mid-mission.
  • One sandbox per mission. sandboxd, the only Docker-socket holder, runs it as an unprivileged user with a read-only root filesystem, no capabilities, resource caps, and only PATH and HOME in its environment.
  • A clone per mission. Each coding mission gets its own clone and branch, one commit per finished unit.
  • Delegated review is optional. A CLI reviewer runs read-only; any failure falls back to native review.

Challenge 4: how does it run when I am not watching?

What starts a mission at 06:00, and what stops one that never ends?

OptionHow it worksWhy not (or why)
A. Cron creates missions directlyA ticker inserts mission rowsNotifications and run tracking become side effects a crash can lose.
B. A message brokerTriggers and outcomes go through a queueAnother service to run, for one user’s volume.
C. An events inbox in PostgresSide effects are events rows written in the same transaction as the change; a drainer delivers them after commitNothing is lost between commit and send, in the database I already run.

From cron boundary to notification
Figure 4: Starting a scheduled mission and reporting its outcome both travel as events rows, so neither is lost in a crash.

I chose C. Schedules are automations: an agent, a mission template, triggers (cron, webhooks, channel messages) and destinations. Each run is an ordinary mission created through the API’s path. When a mission reaches done, failed, paused or waiting for input, ApplyTransition writes an events row; consumers notify me, reply in the originating channel, and close the automation run.

The brakes (mission ceilings are operator settings, read per turn):

BrakeStops
BudgetLedger spend reaching the mission’s cap: pause
BackoffThree worker failures in a row: pause
StallThe same failure twice: one automatic replan, then pause
ReworkReview rounds reaching max_iterations: pause, branch kept
Iterationsmax_iterations (default 8) retries in a phase: fail
AutomationsRuns per hour, concurrency mode, a dedup key per boundary, a breaker after three failed runs

Every harness change must pass make canary: a real mission against the live stack, with a randomly drawn goal so caches cannot flatter it. It fails on any permission park, no harness-verified unit, or too many turns. A sibling canary inverts the test: its goal names a missing file, and it passes only if the mission fails honestly without fabricating it.

What broke

The rollback that deleted good work. Retries used to roll the working tree back. A delegated CLI wrote a module and its test, reported RETRY to fix one function, and the rollback deleted both. RETRY means unfinished, not wrong, so no retry rolls back now; only a review rework does.

Lessons

  1. Put long-running state in rows a pure function moves; transcripts are for reading, not resuming.
  2. A model’s “done” is a request for verification, and only harness evidence answers it.
  3. A check that passes before the work starts proves nothing, so test the gate too.
  4. Give review its own route so it can run on another provider; let the harness have the final say.
  5. A brake should park with a reason whenever the work might still be worth keeping.
  6. Side effects that must survive a crash belong in the same transaction as the change that caused them.

// end of article — process exited with code 0

// share:Twitter / XLinkedIn
// llm:View as Markdown

// comments

loading...