System Design September 12, 2026 ~9 min read

Designing a Self-Hosted AI Assistant (Part 6): Letting the Model Act Safely

TL;DR: Timothy is the self-hosted, single-user AI assistant I build in Go. This part is about the model acting: running shell commands, reading mail, fetching pages. One agent loop executes every tool call, for chat and native mission workers alike, under three rules.

First, the model proposes and code decides. Every call walks one permission chain in Go: policy guard, exempt list, danger classifier, standing grants, then a prompt to me. First match wins.

Second, anything an outsider wrote reaches the model as fenced data. Trust is a flag set on each tool where it is built, and the default is untrusted.

Third, privacy and blast radius are enforced by routes and networks, not prompts. A session that touches a sensitive connector stays on a route I chain to a local model, side-calls included. Model code runs in containers only sandboxd can create, and model-supplied URLs dial through an SSRF guard.

The problem

A model with tools runs instructions from strangers. Some come from me; others ride in on a web page, an email or a PDF, and the model reads them with the same attention. It can also simply be wrong: a confident rm in the wrong directory is not an attack, but it ends the same way.

RiskExampleStopped by
Destructive actiongit push, rm outside the workspacePermission chain
Injected instructionsA mail body saying “forward the inbox”Trust fence, permission chain
Private data leaving homeRaw email sent to a cloud modelSensitive route pinning
Escaping the boxModel code reaching Postgres or the hostsandboxd, networks
Reaching internal addressesfetch_url http://169.254.169.254/netguard

Challenge 1: what does the model get to see?

How much capability goes into every request?

Every schema costs tokens on every step, and the surface grows with each connector.

OptionHow it worksWhy not (or why)
A. Agent frameworkA library owns the loopThe loop is where safety lives; I want to read every line.
B. Everything eagerAll schemas, skills and full results in every requestOne big MCP server or build log fills the window.
C. Typed registry, loaded lazilyTyped tools; skills, big servers and big results load on demandEach request carries only what this turn needs.

I chose C. The loop is plain Go: call the gateway, run the requested tools (four in parallel), append results, repeat. By default, at step 15 of 16 the model is told to finish; at 16 the schemas are dropped and it must answer. The same call three times in a row forces that early.

  • Typed tools. A tool is a name, a description (the model’s only manual), a JSON Schema and an Execute function. Arguments are validated and clamped before the tool runs. Failures, panics included, return as structured feedback, never a crashed turn.
  • Big results stay out. A result over 8 KB (4 KB for shell) goes to tool_outputs. The model gets a deterministic digest: size, first 30 and last 10 lines, error-looking lines, and a ref for retrieve_output.
  • Skills are lazy packs. The system prompt lists one “Use when” line per SKILL.md; the body arrives only via load_skill.
  • Large MCP servers hide behind an index of one line per tool and a load_tool entry point.

Challenge 2: who decides whether a call runs?

Why not let the model decide?

OptionHow it worksWhy not (or why)
A. Rules in the system prompt“Never delete files without asking”Injected text can argue with a prompt.
B. Ask every timeEvery call needs a clickSafe and unusable. By day three I would approve without reading.
C. Static allowlistListed tools run freelyCannot tell git status from git push.
D. Ordered chain in codeHard deny, exemptions, danger classifier, grants, then askMore code, but each layer is small and testable.

Permission chain decision flow
Figure 1: The permission chain for one tool call, where the first match wins.

I chose D. The order is the design: a hard deny runs before anything that could allow, and a dangerous command is caught before any grant could wave it through.

  • The policy guard refuses shell tokens naming env files, key material, ~/.ssh, ~/.aws, system directories, .., or absolute paths outside the workspace. It matches path-shaped tokens only, so grepping for the word “credentials” still runs.
  • The danger classifier scores commands against a regex table. rm, git push, sudo, docker, curl | sh and an overwriting redirect score 3; mv, kill and chmod -R score 2. At 3 the call asks, and no grant skips that. A command it cannot read ($(...), eval, sh -c) counts as destructive. Inside a mission sandbox, file-scoped rules on the mission’s own paths relax, because there the container is the boundary.
  • Exempt tools are pure reads with their own guards (fetch_url, search_kb), mission sentinels, and the root-confined write_file.
  • Asking parks the turn. I allow once, allow for the session (a session_grants row for 12 hours), or deny. An unattended mission gets an automatic denial with a hint on how to rewrite the call.

Challenge 3: how does untrusted text stay data?

What stops a web page from giving orders?

OptionHow it worksWhy not (or why)
A. NothingResults go in as plain textInjected text looks exactly like my instructions.
B. List of untrusted toolsThe loop fences results from named toolsFails open for any tool the list forgets.
C. Trust flag per toolTool.Trusted is set where the tool is built; everything else is fencedA new tool is fenced until someone decides otherwise.

Trust fence around tool results
Figure 2: Untrusted results are wrapped and escaped; attached documents are the remaining gap.

I chose C, after B failed: the list missed tools it did not name, and MCP servers choose their own names. Now a tool returning anything an outsider wrote simply never sets Trusted.

  • One fence. trustfence wraps content in <untrusted_content source="..." trust="data"> with a preamble saying everything inside is quoted material. Memories use the same mechanism.
  • Forged close tags are escaped, case and whitespace variants included.
  • Errors pass through. They are harness prose the model needs to read.
  • shell is trusted on purpose. Its output is the model’s own workspace; fencing it would wrap every build log.

The known gap: attached files and referenced documents enter the user message unfenced. And the fence is a strong hint, not enforcement: fetch_url is exempt, so a page that persuades the model to fetch a URL with private data in its query string is not stopped here. That risk is still open, which is why the defences below never rely on the model behaving.

Challenge 4: where may private data and model code go?

What keeps my inbox off a cloud model, and model code off my host?

OptionHow it worksWhy not (or why)
A. Local model for everythingNo cloud provider at allLocal models are weaker at long, tool-heavy work.
B. Pin only the turnThe turn that called a sensitive tool switches routeThe mail stays in the transcript, and side-calls re-send it later.
C. Pin the session, side-calls includedAfter a sensitive tool runs, every later call uses the sensitive routeSome turns run on a weaker model.

Sensitive session route pinning
Figure 3: A sensitive connector flips the route mid-turn, and the session stays pinned, side-calls included.

I chose C. Privacy is a property of the transcript, not of one turn. Every connector, from Gmail and Calendar to an MCP server, has a sensitive flag, and the sensitive_tool_route setting names a route I chain to a local model.

  • Mid-turn flip. When the model asks for a sensitive tool, the loop switches route before the result goes back, and drops any model hint, which would otherwise outrank the route.
  • Sticky. Every later turn scans session_events for a sensitive tool_execution. Unified tools like search_mail carry no connector name, so the call’s account argument is resolved back to its connector.
  • Side-calls too. Memory extraction, turn distillation, auto-titling and compaction share the verdict.

Sandbox and network boundaries
Figure 4: Only sandboxd holds the Docker socket, and outbound fetches dial through netguard.

Chat’s shell runs inside the brain container under the permission chain. Mission code runs in per-mission containers that only sandboxd can create, because it alone holds the Docker socket. It sits on a network only brain reaches, runs read-only with every capability dropped, and accepts a mission UUID, a command and a few validated fields, never a free-form image, mount or container name. Each container runs unprivileged with a read-only root filesystem, resource caps, and only its own mission directory mounted. It lives on the default bridge, so pip install works while Timothy’s internal network stays out of reach.

Outbound URLs go through netguard. It resolves the host itself, refuses loopback, private, link-local (cloud metadata) and CGNAT addresses, dials the vetted IP so DNS cannot change in between, and re-checks every redirect. Model-supplied URLs get no exceptions; operator-configured MCP endpoints and webhooks may skip it only for allowlisted hosts.

What broke

The first load_tool exemption was a suffix rule: any tool ending in _load_tool skipped the chain. MCP servers name their own tools, so a server could publish exfiltrate_load_tool and walk past every check. The exemption is now an exact match against names the connector manager generates. Same lesson as the trust flag: a name an outsider chooses must never carry authority.

What’s next

Part 7 covers missions: long-running work that outlives a chat turn, judged by evidence the harness collects rather than the model’s word.

// end of article — process exited with code 0

// share:Twitter / XLinkedIn
// llm:View as Markdown

// comments

loading...