Skip to content

The Drift Log · Governing agents that write code

Policy‑Driven Preview Sandboxes Prevent Unauthorized Agent Writes

2 October 2026 · 3 min read · 716 words · established

A crystal shard held at a red-lit gateway in a black wall

A preview sandbox isolates agent write proposals, ensuring only policy‑approved changes reach the real repository.

Unchecked agent writes break the repository contract

When an autonomous agent receives a prompt to modify code, many teams let the agent call the filesystem directly. The agent’s output is written to the repository without any intermediate review. The failure mode is clear: a mis‑prompt, a buggy prompt‑template, or a compromised token can cause the agent to introduce insecure imports, delete files, or alter licences. The repository ends up with changes that were never vetted by human policy, and the damage may only be discovered after a CI failure or a security scan. In practice the breach is hard to trace because the write happened as a normal git commit, not as a logged request.

A preview sandbox isolates the write path

The remedy is to interpose a preview sandbox between the agent and the repository. The agent’s proposal – a diff, a set of file creations, or a deletion list – is sent to an isolated environment that mirrors the target repository’s layout but has no write‑through to the real storage. Inside the sandbox the proposal is applied to a copy of the current tree. Policy checks then run against the resulting state. Only when every check passes does the system forward the proposal to the Build‑Intent gate where licensing, invariants and entitlement are verified.

The sandbox is deliberately short‑lived. It is created per‑proposal, populated with the current HEAD, and destroyed after the policy verdict. Because the sandbox never touches the production filesystem, any malicious or malformed proposal is confined. The result is a clean separation: the agent can experiment freely, but no side‑effect reaches the repository until the policy gate says “allowed”.

Policy gate implementation on the MCP server

The MCP server is the single entry point for agents that use the SHPBL library. By configuring the server to accept only preview requests, you force every agent write to follow the sandbox route. The flow is:

  1. Agent sends a POST /preview with its intended file operations.
  2. The server spawns a sandbox, applies the operations, and runs the policy suite.
  3. If the suite returns approved, the server translates the proposal into a Build‑Intent request.
  4. The Build‑Intent gate performs the final entitlement check and, on success, writes the changes to the real repository.

A minimal TypeScript client call illustrates the pattern:

import { ShpblClient } from '@shpbl/sdk';

const client = new ShpblClient({ endpoint: 'https://mcp.example.com' });

async function proposeChange(diff: string) {
  const preview = await client.previewSandbox(diff);
  if (!preview.allowed) throw new Error('Policy rejected');
  const result = await client.buildIntent(preview.proposalId);
  return result;
}

The client never receives a direct write token; it only ever receives a proposal identifier that the server validates. This design satisfies agent governance by making policy enforcement unavoidable, and it keeps the policy gate logic centralised on the MCP server.

Aligning with the broader governance model

The preview sandbox is one layer of the governance stack described in the pillar post on governing agents at the Build‑Intent gate. It complements the later metering and audit stages. By rejecting unauthorised proposals early, you reduce the load on downstream verification and keep the audit log focused on truly intent‑driven actions. The approach also respects the principle that the gate – not the model – is the critical control point.

If you already have a policy suite (e.g., licence compliance, static‑analysis thresholds, or naming conventions), you can plug it into the sandbox without changing the agent code. The sandbox simply reports a boolean allowed flag and an optional diagnostics payload. Because the sandbox runs on a copy of the repository, the diagnostics are reproducible and can be stored alongside the proposal for later review.

What to do on Monday

  1. Enable preview mode on your MCP server. The configuration flag is sandbox.enabled = true.
  2. Redirect all agent write endpoints to the /preview path. Existing agents that call the plain HTTP API will need only a URL change.
  3. Load your policy scripts into the sandbox container. A typical policy file lives in /policies/agent-governance.js.
  4. Run a dry‑run against a non‑critical repository. Verify that a harmless proposal is accepted and that a deliberately malformed diff is rejected.

Once those steps are verified, you can roll the change out to production repositories. The result is a repository that only ever mutates under the watch of a policy gate, with the preview sandbox guaranteeing that no unauthorised agent write ever reaches the real filesystem. This simple change removes a whole class of supply‑chain risk without altering the agent’s core logic.

Keep reading

Next in the log

The Strategic Master Library · written and reviewed under the house's own epistemic rules: nothing claimed that we cannot show.