TL;DR: Timothy is the self-hosted AI assistant I built in Go. A conversation in Timothy is not a list of messages. It is an append-only log of typed events in one Postgres table, session_events, and everything else is computed from it. Three rules shape it.
First, the log is the only truth, and it only grows. No event is ever edited. Every change, including a summary that replaces old turns, is a new row with the next sequence number.
Second, what the model sees and what I see are different documents. Two pure functions read the same rows: one builds the model’s context, the other builds the rich transcript the web UI replays. Neither is stored.
Third, the projected prefix never changes between turns. New content is appended at the end, so provider prompt caches keep hitting. The prefix moves only at a compaction event, which is rare and deliberate.
The problem
What does a chat session need to survive?
A chat session looks simple until turns run tools and stream for minutes. Then it has to handle all of these:
| Requirement | Detail |
|---|---|
| Crash safety | Killing the brain mid-answer loses at most a couple of seconds of text. |
| Full replay | Reloading the page shows everything: reasoning, tool calls, cost, failures. |
| Bounded context | The model’s context stays under a token budget without forgetting names, dates or commitments. |
| Cache friendliness | Turn N+1 resends turn N’s messages byte for byte, so the provider can reuse its cache. |
| Fast follow-ups | A question sent seconds after an answer must see that answer as complete. |
The UI wants everything, the model wants as little as possible, and the cache wants nothing to change.
Challenge 1: how do you store a conversation?
Why not a messages table?
| Option | How it works | Why not (or why) |
|---|---|---|
| A. Messages table, edited in place | One row per message; summaries and partial answers overwrite rows. | Every summary destroys history, and the UI and the model are forced to share one shape. |
| B. One JSON document per session | Load, modify, save the whole conversation. | Concurrent writers (a streaming turn, a permission answer, a late summary) lose each other’s writes. |
| C. Append-only event log with projections (chosen) | Typed events with a gapless per-session seq; views are computed on read. | Writes are tiny and never touch history. Any view can be rebuilt. |
I chose the event log. Every requirement above turned into “append one more event”: a checkpoint, a summary, a failure, a late distillation. Nothing reaches back to edit.
- Gapless sequence numbers.
Appendlocks thesessionsrowFOR UPDATEand insertsMAX(seq) + 1in the same transaction. Writers to one session queue behind each other; different sessions never contend. - Kinds are added, never changed. There are ten:
session_started,user_message,assistant_turn,tool_execution,turn_memory,compaction_applied,pending_state,permission_request,permission_resolvedandturn_failed. New fields are additive, so old rows still decode. - Tool traffic stays out of the model’s history. Within a turn the model sees tool results directly. Across turns,
tool_executionholds only a digest for the UI, and a separateturn_memoryevent carries the distilled residue the model needs: files changed, failures and why, key findings. - Attachments are converted once. A PDF’s markdown is stored inside the
user_messageat send time. Images are stored as references and turned into bytes only just before the gateway call.
Challenge 2: how do you keep the prompt cache warm?
Why not trim the context to fit?
Provider prompt caches match on a byte-identical prefix, so rewriting any earlier message means a cache miss and a higher bill.
| Option | How it works | Why not (or why) |
|---|---|---|
| A. Sliding window | Drop the oldest messages until the context fits the budget. | Once the budget is reached the prefix changes every turn, and old facts disappear silently. |
| B. Per-turn rewriting | Middleware compresses or reformats older messages on each request. | An “optimisation” that busts the cache on every turn. |
| C. Append-only projection (chosen) | LLMContext never trims; new events become new messages at the tail. | The prefix stays identical between turns and changes only when a compaction event lands. |
I chose the append-only projection. LLMContext takes a budget argument and deliberately ignores it: the compactor enforces the budget, in rare and deliberate steps.
- Late residue is its own message. Distilling a turn into
turn_memoryis an LLM call that finishes seconds after the answer. Theassistant_turnis appended the moment the stream ends; the residue lands later as a new event rather than an edit. - Per-turn extras ride the system prompt’s tail. Recalled memories and a hinted skill can change every turn, so they go after the stable part of the system prompt.
- A property test enforces it. It grows a log one event at a time and fails if any earlier projected message changes.
- One sanctioned exception. A spliced partial answer (Challenge 4) vanishes when the turn completes. That is one cache miss per recovered interruption, cheaper than duplicating content or adding an LLM call to the failure path.
Challenge 3: how do you compact without forgetting?
What do you throw away when the context is full?
| Option | How it works | Why not (or why) |
|---|---|---|
| A. Truncate | Drop old turns. | Loses exactly the commitments and decisions a long conversation exists to keep. |
| B. One-shot rewrite | Summarize the whole session and replace the stored history. | Irreversible: one bad summary is permanent. |
| C. Incremental compaction event (chosen) | Summarize the oldest half of the live messages into a compaction_applied event. | The log keeps every original row; only the model’s view changes. |
I chose incremental compaction. A compaction_applied event says “in the model’s context, replace everything up to replaces_through_seq with this summary”. The UI still shows every original message, plus a divider noting which ones the model now sees only as a summary.
- Trigger.
MaybeCompactruns after every completed turn and again before each send. The budget is 60% of the context window of the model that served the last turn, falling back to a setting (60,000 tokens by default). Tokens are counted with tiktoken’so200k_baseencoding, not a bytes-divided-by-three guess. - Boundary. The cut is half the live messages, extended so it never ends on an unanswered question, with at least one turn left live. It maps back to an event
seq; if the mapping does not line up, nothing is compacted. - Facts first. Before summarizing, the turns about to be replaced go to memoryd for extraction, and the memory ids are recorded in
facts_extracted. Whatever the summary drops still survives as memory. - Summary rules. The summarizer must keep every name, date, number, commitment, decision and open question. A truncated or empty summary is rejected, and the next turn tries again.
Challenge 4: what survives a crash mid-answer?
What if the brain dies while the model is still talking?
| Option | How it works | Why not (or why) |
|---|---|---|
| A. Write at the end of the turn | Persist only when the stream completes. | A crash or deploy during a three-minute answer loses all of it. |
| B. Write every chunk | One row per streamed token. | Thousands of rows per turn for no extra safety. |
| C. Periodic checkpoints (chosen) | Append a pending_state with the full partial text every two seconds, plus one on any abnormal end. | Bounded loss, a handful of rows, and the same projection code handles recovery. |
I chose periodic checkpoints. Each pending_state holds the whole partial text, so only the newest matters, and a SIGKILL loses at most one flush interval.
- Live means newest and not superseded. A later
assistant_turnsupersedes it, and so does a compaction whose boundary covers it. A user message does not: the question after an interruption must still see the partial. - The model is told. The partial projects as an assistant message ending in
[this response was interrupted mid-stream; continue from it]. The UI shows it as an interrupted turn. - Writes outlive the request. Checkpoints and final writes use a context detached from the HTTP request, so a closed tab or a shutdown still leaves a durable row.
- A turn never leaves zero events. If there is no partial text, a
turn_failedevent records the reason instead, and the model sees[previous turn failed: ...]next time.
What broke
Turn memory as an assistant message. The residue first rode in the assistant role. Weaker models imitate what they appear to have written, so one began echoing [turn memory] finding: ... into live answers, and the distiller then fed on its own output. Residue now rides as a user-role note, which models are far less inclined to imitate.
Empty summaries. Reasoning-heavy models spent the whole output cap thinking and returned an empty summary. The summarize call now asks for low reasoning effort and allows 3,000 output tokens, and an empty result is refused rather than stored.
What’s next
Part 4 follows a single turn from the browser to the provider and back, and shows why it must always end in exactly one terminal event.