System Design September 8, 2026 ~9 min read

Designing a Self-Hosted AI Assistant (Part 4): Never Lose a Turn

TL;DR: Timothy is the self-hosted AI assistant I built in Go. A chat turn can call tools, wait for my approval, and stream from a slow local model, so it often outlives the browser tab that started it. This part is about making sure a turn that ran always leaves a trace. Three rules carry it.

First, a turn belongs to the session, not to the request. Closing the tab ends the HTTP request, not the turn. Only the stop button or a 30-minute ceiling can.

Second, a returning tab reattaches; it does not guess. A running turn buffers its events in memory. A new tab replays that buffer and follows the turn live, through the same code path as the tab that started it.

Third, every stream ends with exactly one terminal event. The rule holds at every hop from the provider to the browser, and a backstop in the persistence path covers any hop that still fails.

The problem

Why can’t a chat turn just be an HTTP request?

The obvious Go handler runs the whole turn on the request’s context, so everything stops when the client goes away. But I reload pages and my laptop sleeps. Each of those used to kill an answer seconds from finishing, and the session then showed a failure that was not one.

The opposite mistake is worse: a turn that ends with nothing recorded at all. The requirements:

RequirementWhat it means
Survives the clientA closed tab, reload or dropped connection never kills a turn
StoppableAn explicit stop works from any tab
WatchableAny tab can watch a running turn live
ExclusiveTwo sends on one session never interleave in the event log
Always evidencedA turn that ran leaves an answer, a partial, or a named failure

Challenge 1: who owns a turn?

What should cancel a turn, if not the browser?

Detached turn lifecycle
Figure 1: Closing the tab ends only the request, while the turn runs on its own context until it persists, and stop takes the same exit as the deadline.

OptionHow it worksWhy not (or why)
A. Request-scoped contextThe turn runs on the request’s contextA reload or a dropped connection kills the turn
B. Job queue and workerPersist a job, a worker runs it, clients pollA broker and a second copy of state the event log already holds, for one user
C. Detached turn with its own context (chosen)The turn runs in brain on a context the request cannot cancel, registered per sessionNo new infrastructure; the event log stays the only durable state

I chose C. Brain builds the turn’s context with context.WithTimeout(context.WithoutCancel(reqCtx), 30*time.Minute). It keeps the request’s values (auth, trace) and drops its cancellation. The relay watches the request only to know whether this tab is still listening. When the tab goes, the relay switches to drainAndPersist: it consumes the turn to the end and persists it as if the client had stayed.

  • Stop is explicit. Closing a tab used to double as a stop button, so a real one had to exist: POST /v1/sessions/{id}/stop calls the stored cancel function. The turn then exits through the same path as the deadline: partial text becomes a pending_state, otherwise a turn_failed event.
  • The ceiling must outlive the permission park. A turn can wait 10 minutes for my approval of a tool call. A TURN_TIMEOUT override exists for slow CPU-only backends, and values of 10 minutes or less are rejected.
  • Partials are checkpointed. The relay writes a pending_state every 2 seconds while text grows, so even a killed brain loses little.

Challenge 2: how does a returning tab catch up?

Is the persisted transcript enough for a tab that opens mid-turn?

Re-fetch versus reattach
Figure 2: A re-fetch sees only what is durable, while /live replays the turn so far and then follows it, falling back to a re-fetch when it cannot finish.

OptionHow it worksWhy not (or why)
A. Re-fetch the transcriptPoll GET /v1/sessions/{id} until the turn completesOnly as fresh as the last checkpoint, with no live tokens or tool calls
B. Persist every token, tail the logWrite each chunk to session_events, stream from the databaseFloods the append-only log with deltas both projections must ignore
C. In-memory replay buffer (chosen)GET /v1/sessions/{id}/live replays the turn’s events so far, then follows itNo extra writes, same wire format as the original stream

I chose C. Each running turn has a turn broadcaster: a buffer of every event so far plus a set of live subscriber channels. subscribe() takes the snapshot and registers the channel under one lock, so no event can fall into the gap between replay and live. The stream is identical to POST .../messages and ends with the same meta event, so the web client feeds both through one reducer.

  • The registry is the truth. turn_active on the session endpoint is read from the broadcaster map, not a separate flag. The entry is removed only after the terminal persist is durable, so turn_active: false guarantees a fetch sees the finished turn.
  • Slow watchers are cut loose. A subscriber more than 64 events behind is closed rather than allowed to stall the turn. It sees no meta, so the client re-fetches once.
  • Re-fetch remains the fallback. A 404 no_active_turn means the turn just finished, so the client re-fetches. A transport error means it is probably still running, so the client reattaches, up to five times.

Challenge 3: what stops two turns on one session?

What happens when I double-click send, or two tabs send at once?

One turn per session
Figure 3: The slot is claimed before the first append, so the losing request writes nothing and becomes a viewer of the winning turn.

OptionHow it worksWhy not (or why)
A. Database lockA row or advisory lock per sessionA second source of truth for “a turn is running” that can drift
B. Queue the second messageRun it after the current turnThe second send is usually an accident (a double-click or a racing retry), and a queue is more state to own
C. Exclusive claim, then 409 (chosen)turnBegin registers the broadcaster or failsThe same map answers “is a turn running?” and “may I start one?”

I chose C. turnBegin checks and inserts under one mutex. A loser gets ErrTurnInFlight, which the API maps to 409 turn_in_flight.

  • Claim before the first append. The check runs before the user_message is written, so a loser never leaves a message with no turn behind it.
  • The loser becomes a viewer. The web client answers a 409 with a toast and an attach to /live, so a double-click is harmless.

Challenge 4: how does every stream end exactly once?

Can the end of a turn be best-effort?

Every provider’s wire format is translated into one event vocabulary with a one-line contract: a stream ends with exactly one terminal event, done (possibly preceded by incomplete) or error, and the channel closes immediately after.

OptionHow it worksWhy not (or why)
A. Trust each producerEach layer sends its terminal if it canAny one lost send erases the turn
B. Backstop onlyPersistence synthesizes a failure when nothing arrivesA generic reason that hides which layer broke
C. Defend every hop, plus a backstop (chosen)Each layer guarantees its terminal; persistence refuses to record nothingReal reasons reach the user

I chose C. The details that make each hop honest:

  • Headers before the provider dial. Brain’s gateway client allows 30 seconds for response headers, and a cold-loading local model can be slower. The gateway writes 200 and flushes before its first provider attempt; later failures, including chain_exhausted, arrive as SSE error events.
  • Silence has a name. A provider that finishes with zero content is a failed attempt (empty_output) that fails over to the next provider, and the ledger books it as incomplete, never ok. A completed turn with no text, reasoning or tool calls becomes turn_failed with empty_response. The agent loop keeps an incomplete sticky, so a later done cannot erase it.
  • Heartbeats. The chat stream and /live write a : ping comment every 20 seconds, so no proxy sees an idle connection.

What broke

Where the terminal event was lost
Figure 4: The terminal event was lost at five places on its way to the browser, and persistTurn is the backstop beneath all of them.

A turn of about thirty minutes against a slow model on a CPU-only box ended with nothing in session_events. Detached turns made long turns possible, and long turns exposed every weak hop. A proxy killed the connection after about 100 idle seconds. The provider driver’s final emit lost a race against ctx.Done() and closed its channel empty. The gateway sent no terminal after content had streamed. Brain’s gateway client skipped its synthetic gateway_stream_cut because the context was already done, and the agent loop’s final send raced the same deadline.

Three of the five came down to one pattern: select { case out <- ev: case <-ctx.Done(): } for an end-of-stream event. When the deadline fires as the terminal is sent, both cases are ready and Go picks one at random. The gateway now names a terminal-less close stream_cut, and brain sends terminals with a 5-second bound instead (emitFinal in the loop). Underneath, persistTurn turns “no text and no failure” into turn_failed with code turn_aborted. The lesson: racing cancellation is fine for progress events and wrong for the event that proves the work happened.

What’s next

Part 5 turns to what Timothy remembers afterwards: typed, staged memory and a knowledge base the model searches through a tool.

// end of article — process exited with code 0

// share:Twitter / XLinkedIn
// llm:View as Markdown

// comments

loading...