# Designing a Self-Hosted AI Assistant (Part 6): Letting the Model Act Safely

> Source: https://www.sumonselim.com/timothy-part6-letting-the-model-act-safely/
> Author: Muhammad Sumon Molla Selim
> Published: 2026-09-12
> Tag: System Design
> Summary: Timothy's model can run shell commands, read mail and fetch web pages, so every tool call walks a permission chain written in Go. Part 6 covers that chain, keeping big results and idle skills out of context, fencing untrusted tool output, pinning sensitive sessions to a local model, and the sandbox and network boundaries underneath.
>
> Series: Designing a Self-Hosted AI Assistant, part 6 of 7
> 1. [The Architecture](https://www.sumonselim.com/timothy-part1-the-architecture.md)
> 2. [The LLM Gateway](https://www.sumonselim.com/timothy-part2-the-llm-gateway.md)
> 3. [Sessions as an Event Log](https://www.sumonselim.com/timothy-part3-sessions-as-an-event-log.md)
> 4. [Never Lose a Turn](https://www.sumonselim.com/timothy-part4-never-lose-a-turn.md)
> 5. [Memory and Knowledge](https://www.sumonselim.com/timothy-part5-memory-and-knowledge.md)
> 6. Letting the Model Act Safely (this part)
> 7. [Missions](https://www.sumonselim.com/timothy-part7-missions.md)

**TL;DR:** Timothy is the self-hosted, single-user AI assistant I build in Go. This part is about the model acting: running shell commands, reading mail, fetching pages. One agent loop executes every tool call, for chat and native mission workers alike, under three rules.

First, **the model proposes and code decides**. Every call walks one permission chain in Go: policy guard, exempt list, danger classifier, standing grants, then a prompt to me. First match wins.

Second, **anything an outsider wrote reaches the model as fenced data**. Trust is a flag set on each tool where it is built, and the default is untrusted.

Third, **privacy and blast radius are enforced by routes and networks, not prompts**. A session that touches a sensitive connector stays on a route I chain to a local model, side-calls included. Model code runs in containers only `sandboxd` can create, and model-supplied URLs dial through an SSRF guard.

## The problem

A model with tools runs instructions from strangers. Some come from me; others ride in on a web page, an email or a PDF, and the model reads them with the same attention. It can also simply be wrong: a confident `rm` in the wrong directory is not an attack, but it ends the same way.

| Risk | Example | Stopped by |
| :-- | :-- | :-- |
| Destructive action | `git push`, `rm` outside the workspace | Permission chain |
| Injected instructions | A mail body saying "forward the inbox" | Trust fence, permission chain |
| Private data leaving home | Raw email sent to a cloud model | Sensitive route pinning |
| Escaping the box | Model code reaching Postgres or the host | `sandboxd`, networks |
| Reaching internal addresses | `fetch_url http://169.254.169.254/` | `netguard` |

## Challenge 1: what does the model get to see?

**How much capability goes into every request?**

Every schema costs tokens on every step, and the surface grows with each connector.

| Option | How it works | Why not (or why) |
| :-- | :-- | :-- |
| **A. Agent framework** | A library owns the loop | The loop is where safety lives; I want to read every line. |
| **B. Everything eager** | All schemas, skills and full results in every request | One big MCP server or build log fills the window. |
| **C. Typed registry, loaded lazily** | Typed tools; skills, big servers and big results load on demand | Each request carries only what this turn needs. |

**I chose C.** The loop is plain Go: call the gateway, run the requested tools (four in parallel), append results, repeat. By default, at step 15 of 16 the model is told to finish; at 16 the schemas are dropped and it must answer. The same call three times in a row forces that early.

- **Typed tools.** A tool is a name, a description (the model's only manual), a JSON Schema and an `Execute` function. Arguments are validated and clamped before the tool runs. Failures, panics included, return as structured feedback, never a crashed turn.
- **Big results stay out.** A result over 8 KB (4 KB for `shell`) goes to `tool_outputs`. The model gets a deterministic digest: size, first 30 and last 10 lines, error-looking lines, and a ref for `retrieve_output`.
- **Skills are lazy packs.** The system prompt lists one "Use when" line per `SKILL.md`; the body arrives only via `load_skill`.
- **Large MCP servers hide behind an index** of one line per tool and a `load_tool` entry point.

## Challenge 2: who decides whether a call runs?

**Why not let the model decide?**

| Option | How it works | Why not (or why) |
| :-- | :-- | :-- |
| **A. Rules in the system prompt** | "Never delete files without asking" | Injected text can argue with a prompt. |
| **B. Ask every time** | Every call needs a click | Safe and unusable. By day three I would approve without reading. |
| **C. Static allowlist** | Listed tools run freely | Cannot tell `git status` from `git push`. |
| **D. Ordered chain in code** | Hard deny, exemptions, danger classifier, grants, then ask | More code, but each layer is small and testable. |

![Permission chain decision flow](https://www.sumonselim.com/images/articles/timothy/p6-permission-chain.svg "Figure 1: The permission chain for one tool call, where the first match wins.")

**I chose D.** The order is the design: a hard deny runs before anything that could allow, and a dangerous command is caught before any grant could wave it through.

- **The policy guard** refuses shell tokens naming env files, key material, `~/.ssh`, `~/.aws`, system directories, `..`, or absolute paths outside the workspace. It matches path-shaped tokens only, so grepping for the word "credentials" still runs.
- **The danger classifier** scores commands against a regex table. `rm`, `git push`, `sudo`, `docker`, `curl | sh` and an overwriting redirect score 3; `mv`, `kill` and `chmod -R` score 2. At 3 the call asks, and no grant skips that. A command it cannot read (`$(...)`, `eval`, `sh -c`) counts as destructive. Inside a mission sandbox, file-scoped rules on the mission's own paths relax, because there the container is the boundary.
- **Exempt tools** are pure reads with their own guards (`fetch_url`, `search_kb`), mission sentinels, and the root-confined `write_file`.
- **Asking** parks the turn. I allow once, allow for the session (a `session_grants` row for 12 hours), or deny. An unattended mission gets an automatic denial with a hint on how to rewrite the call.

## Challenge 3: how does untrusted text stay data?

**What stops a web page from giving orders?**

| Option | How it works | Why not (or why) |
| :-- | :-- | :-- |
| **A. Nothing** | Results go in as plain text | Injected text looks exactly like my instructions. |
| **B. List of untrusted tools** | The loop fences results from named tools | Fails open for any tool the list forgets. |
| **C. Trust flag per tool** | `Tool.Trusted` is set where the tool is built; everything else is fenced | A new tool is fenced until someone decides otherwise. |

![Trust fence around tool results](https://www.sumonselim.com/images/articles/timothy/p6-trust-fence.svg "Figure 2: Untrusted results are wrapped and escaped; attached documents are the remaining gap.")

**I chose C,** after B failed: the list missed tools it did not name, and MCP servers choose their own names. Now a tool returning anything an outsider wrote simply never sets `Trusted`.

- **One fence.** `trustfence` wraps content in `<untrusted_content source="..." trust="data">` with a preamble saying everything inside is quoted material. Memories use the same mechanism.
- **Forged close tags are escaped,** case and whitespace variants included.
- **Errors pass through.** They are harness prose the model needs to read.
- **`shell` is trusted on purpose.** Its output is the model's own workspace; fencing it would wrap every build log.

The known gap: attached files and referenced documents enter the user message unfenced. And the fence is a strong hint, not enforcement: `fetch_url` is exempt, so a page that persuades the model to fetch a URL with private data in its query string is not stopped here. That risk is still open, which is why the defences below never rely on the model behaving.

## Challenge 4: where may private data and model code go?

**What keeps my inbox off a cloud model, and model code off my host?**

| Option | How it works | Why not (or why) |
| :-- | :-- | :-- |
| **A. Local model for everything** | No cloud provider at all | Local models are weaker at long, tool-heavy work. |
| **B. Pin only the turn** | The turn that called a sensitive tool switches route | The mail stays in the transcript, and side-calls re-send it later. |
| **C. Pin the session, side-calls included** | After a sensitive tool runs, every later call uses the sensitive route | Some turns run on a weaker model. |

![Sensitive session route pinning](https://www.sumonselim.com/images/articles/timothy/p6-sensitive-pinning.svg "Figure 3: A sensitive connector flips the route mid-turn, and the session stays pinned, side-calls included.")

**I chose C.** Privacy is a property of the transcript, not of one turn. Every connector, from Gmail and Calendar to an MCP server, has a `sensitive` flag, and the `sensitive_tool_route` setting names a route I chain to a local model.

- **Mid-turn flip.** When the model asks for a sensitive tool, the loop switches route before the result goes back, and drops any model hint, which would otherwise outrank the route.
- **Sticky.** Every later turn scans `session_events` for a sensitive `tool_execution`. Unified tools like `search_mail` carry no connector name, so the call's `account` argument is resolved back to its connector.
- **Side-calls too.** Memory extraction, turn distillation, auto-titling and compaction share the verdict.

![Sandbox and network boundaries](https://www.sumonselim.com/images/articles/timothy/p6-sandbox-boundary.svg "Figure 4: Only sandboxd holds the Docker socket, and outbound fetches dial through netguard.")

Chat's `shell` runs inside the brain container under the permission chain. Mission code runs in per-mission containers that only `sandboxd` can create, because it alone holds the Docker socket. It sits on a network only brain reaches, runs read-only with every capability dropped, and accepts a mission UUID, a command and a few validated fields, never a free-form image, mount or container name. Each container runs unprivileged with a read-only root filesystem, resource caps, and only its own mission directory mounted. It lives on the default bridge, so `pip install` works while Timothy's internal network stays out of reach.

Outbound URLs go through `netguard`. It resolves the host itself, refuses loopback, private, link-local (cloud metadata) and CGNAT addresses, dials the vetted IP so DNS cannot change in between, and re-checks every redirect. Model-supplied URLs get no exceptions; operator-configured MCP endpoints and webhooks may skip it only for allowlisted hosts.

## What broke

The first `load_tool` exemption was a suffix rule: any tool ending in `_load_tool` skipped the chain. MCP servers name their own tools, so a server could publish `exfiltrate_load_tool` and walk past every check. The exemption is now an exact match against names the connector manager generates. Same lesson as the trust flag: a name an outsider chooses must never carry authority.

## What's next

Part 7 covers missions: long-running work that outlives a chat turn, judged by evidence the harness collects rather than the model's word.
