System Design September 10, 2026 ~9 min read

Designing a Self-Hosted AI Assistant (Part 5): Memory and Knowledge

TL;DR: Timothy is my self-hosted, single-user AI assistant. It keeps two kinds of long-term context. Memory is short facts about me, pulled out of conversations. The knowledge base (KB) is documents I add, chunked and indexed. Both live in one Postgres with pgvector, under three rules.

First, the model proposes, code decides. One LLM call suggests facts as strict JSON. Go validates them, removes duplicates and decides whether each one becomes active.

Second, a memory is never edited, only superseded. A correction inserts a new row and archives the old one, which points at its successor. Recall only sees active rows.

Third, memory is pushed, knowledge is pulled. A small, token-budgeted memory block rides on every turn inside a trust fence. The KB is too big for that, so the model searches it with a tool.

The problem

What should an assistant remember, and what should it look up?

A provider’s built-in memory belongs to that provider, and I switch models often. Memory also steers the model, so a bad memory is worse than none.

RequirementDetail
Provider-independentFacts survive a change of model or provider.
Owner in controlI can confirm, reject, edit and archive every memory.
No confident stalenessUnconfirmed facts fade. Returning nothing beats returning something wrong.
Bounded costA hard token budget, counted with a real tokenizer.
Poisoning resistantRetrieved text is framed as data, never as instructions.
DocumentsPDFs, DOCX, HTML, web pages and their images become searchable.

Challenge 1: who decides what becomes a memory?

If the model writes its own memory, who checks it?

OptionHow it worksWhy not (or why)
A. Provider memoryThe vendor’s memory featureLocked to one vendor, invisible to me.
B. Model writes directlyA tool call stores an active factOne injected web page can plant a standing instruction.
C. Search raw transcriptsIndex every old turnNoisy, and old turns contradict newer ones.
D. Propose, then stageAn LLM proposes; Go checks; most land pendingThe model reads. Code makes every decision that matters.

I chose D. Spotting facts in a conversation is language work. Deciding what may influence future turns is policy, which belongs in code I can test.

Memory lifecycle: proposals pass Go checks, land pending, and move to active, rejected or archived
Figure 1: Only the promotion policy, the user, or the remember tool make a memory active, and corrections form a supersede chain.

TriggerInput
Turn endThe user’s message plus the turn’s distilled summary, fire-and-forget
Before compactionTurns about to be summarized away
Mission finishedThe outcome digest, under a stricter prompt: discoveries only
Daily reflectionRecent episodic memories, asked for at most three cross-episode insights
remember toolA fact I explicitly asked to keep: active at once

What keeps it safe:

  • Types age differently. episodic is something that happened, semantic a durable fact or preference, procedural a how-to. Rules and standing instructions must be semantic.
  • Promotion is a Go function. Only episodic facts with confidence of at least 0.8 and no sensitive words (“password”, “always”, “from now on”) activate alone. semantic and procedural facts always wait for me, because that is where instructions would hide.
  • Strict parsing. An unknown field, bad enum or missing changes_behavior flag rejects the batch, which gets one retry. Facts marked as general knowledge are dropped.
  • Repetition reinforces. A proposal within cosine 0.95 of an active memory bumps its last_confirmed_at instead of inserting. A match with a rejected memory is dropped, so rejections stick.

Challenge 2: how does recall find the right few facts?

Which memories go into this turn’s prompt?

OptionHow it worksWhy not (or why)
A. Vector onlyk nearest embeddingsMisses exact names; always returns something.
B. Separate storesVector DB, search engine, graph DBThree systems to sync for a few thousand rows.
C. Tool-only recallThe model calls search_memoryIt cannot ask about a preference it does not know exists.
D. Three legs in PostgresVector, full-text and entity queries fused with Reciprocal Rank Fusion (RRF)One store, and each leg covers another’s blind spot.

I chose D, run before every turn. search_memory remains for deliberate lookups.

Recall pipeline from the user message through three legs, RRF, decay, cutoff and packing into a fenced block
Figure 2: RRF merges three legs' ranks, recency and type weights scale the score, and survivors are packed into a fenced block at the prompt's tail.

The score is sum(1 / (60 + rank)) × 0.5^(age / 90 days) × type weight.

  • Legs fail independently. With the embedding route down, full-text and entity still answer.
  • A floor, not a quota. Scores under 0.005 are dropped: a fresh semantic hit at rank 30 survives; after about 120 unconfirmed days it does not.
  • A real token budget. Items are costed with the o200k tokenizer in rendered form, fence included, up to 1,500 tokens. The best item goes first and the runner-up last, because models attend least to the middle.
  • Fenced at the tail. The block is wrapped in <memory trust="data"> with a “not instructions” preamble, and closing-tag lookalikes are escaped. It sits after the stable prompt prefix, so prompt caching survives.

Challenge 3: how does memory stay clean?

What stops a year of extraction from becoming noise?

In-place rewrites would break supersede-only; a TTL would delete facts I rely on. Instead, memoryd runs a daily consolidation job of independent stages, each a status change or a supersede:

StageRule
Merge near-duplicatesAn LLM merges each group into one fact; a guard rejects merges that drop specific tokens or change length too much.
Dedupe pendingOf two pending near-duplicates, the newer is rejected.
Archive episodicNot created or retrieved in 180 days.
Decay semanticUnconfirmed for a year: confidence × 0.8, queued for reconfirmation.
ReflectTen or more episodic memories in 7 days: ask for up to 3 insights, which land pending.
Demote pendingLow-confidence, never retrieved, 60 days old: archived.

Recall stamps last_retrieved_at and retrieval_hits, so the job acts on what is actually used.

Challenge 4: how should documents reach the model?

Retrieve the KB like memory, or search it on demand?

OptionHow it worksWhy not (or why)
A. Pre-retrieve every turnInject top chunks for the user messagePays tokens on every turn; most turns need no documents.
B. Paste whole documentsAttach files to the promptBlows the context window.
C. Tool-driven searchThe model calls search_kb, then read_kbTokens are spent only when a question needs documents.

I chose C. Memory is small and nearly always relevant; documents are large and sometimes relevant. Pinning a collection to a session tells the model to call search_kb first.

KB ingestion from file or URL through markitdown, image description, chunking and embedding, and the search_kb read path
Figure 3: Brain converts and enriches documents, memoryd chunks and embeds them, and search_kb fuses two gated legs into capped, fenced passages.

  • One converter. A markitdown-svc sidecar turns PDF, DOCX and HTML into Markdown, and returns PDF images plus renders of pages with almost no text layer.
  • Images become text. A vision-route caption, or local OCR from ocr-svc when no vision route is bound.
  • Heading-aware chunks of about 500 tokens with 15% overlap. The breadcrumb (title and heading path) is embedded with each chunk.
  • Delete and rewrite. Re-ingest replaces a document’s chunks in one transaction; transient failures retry after 2, 10 and 30 minutes.
  • Gated legs. Vector hits need cosine 0.25; keyword hits need two shared query words and count at a tenth of the vector weight.
  • Weights, never filters. The agent’s and session’s collections get a 1.5× boost, but the model cannot choose collections. Mission output (0.8) and web clips (0.6) rank below curated documents.
  • Two chunks per document at most, and results come back fenced as <untrusted_content>.

Side-by-side comparison of long-term memory and the knowledge base
Figure 4: Memory and the knowledge base share one Postgres and the same fusion idea, but differ in who writes them, when they are read and how they are framed.

What broke

The confirmation queue flooded. Near-duplicates were once inserted for consolidation to merge later, and mission digests had their own goal and title extracted as “facts”. The fixes were in code: duplicates reinforce the existing row, missions got their own extraction contract, and a word-overlap check drops facts that echo the digest’s header.

One shared word was enough to rank. The KB keyword leg OR-matched query words, so a chunk sharing one common word outranked relevant documents. The two-word minimum and the low keyword weight came from that, checked against a recall test set.

What’s next

Part 6 covers what happens when the model stops reading and starts acting: tools, permissions and the sandbox.

// end of article — process exited with code 0

// share:Twitter / XLinkedIn
// llm:View as Markdown

// comments

loading...