Case study · Tool · Aug 2026 – present

Agent Platform

A control plane for safely delegating software work to autonomous coding agents.

Role
Solo: spec, design and CLI (in progress)
When
Aug 2026 – present
Repo
TypeScript · ★ 0 · last push Sep 2026

01 The problem

Coding agents can already write plenty of code. The hard part is trusting it: an agent reports “done”, a locally correct shortcut gets merged, later agents copy it, and the architecture slowly degrades.

The goal, from the spec: maximize architecture-approved, verified engineering output per dollar and per minute of human attention, for one developer delegating work without watching every run.

02 What I built

Solo, in progress. I started with the specification, then built the MVP it scopes: declarative artifacts (JSON Schemas, workflow templates, role contracts, policies) plus a small file-based control-plane CLI.

agent init installs a .agent/ scaffold and default policies into a target repo; from there the CLI creates missions, compiles workflow templates into task graphs, and runs the claim → start → submit → gate loop, with commands for evidence bundles, architecture checks, retrospectives, replayable eval cases and a tier-gated memory.

03 Highlights

  • Wrote the v3 spec: missions, workflow templates, task and role contracts, review gates, model routing, memory and fleet security.
  • Built the “agent” control-plane CLI in TypeScript: it scaffolds a target repo, validates artifacts against JSON Schemas and runs the claim → start → submit → gate loop.
  • Workers never certify their own work: a task passes only with an evidence bundle, independent verification, and spec, quality and architecture review.
  • Routes each task by risk to a capability tier (frontier, strong, mid, cheap), mapped to models by policy instead of hard-coded vendors.

04 Key decisions

  1. Workers never certify their own work

    An agent’s own “done” is the least trustworthy signal in the loop, so completion is decided elsewhere. The evidence contract’s rule: a PASS without evidence is a FAIL.

  2. Tiers, not vendors

    Model catalogs change faster than anything else in the repo, so routing only ever names a tier. The active harness profile in policies/models.yaml is the one place a tier becomes a model.

  3. Schemas as the source of truth

    Every artifact shape (mission, task, result, review verdict, policies…) is a JSON Schema, validated with Ajv, so agents and the CLI agree on the contract.

  4. Roles as contracts

    Each pipeline stage is a role contract (inputs, outputs, acceptance bar, failure behaviour) that can be bound to any harness: the platform’s own skills, pinned MIT skill packs, or an operator’s native install.

05 Architecture

  1. 01MissionA human writes the goal and budget
  2. 02CompileWorkflow template → task DAG
  3. 03RouteRisk → capability tier
  4. 04WorkBounded worker in a worktree
  5. 05VerifyIndependent checks + evidence
  6. 06Review & mergeSpec, quality, architecture gates

From a human mission to merged, verified work. Every arrow is a file on disk, not a hidden call.

06 Stack

TypeScript
The CLI and its tests (Node’s built-in test runner).
Node.js
JSON Schema (Ajv)
15 schemas, one per artifact shape.
Commander
The agent command surface.
YAML
Missions, workflow templates and policies humans can read and diff.

07 Links & sources

The project’s links and the public sources this write-up draws on.

08 Also on GitHub

All repositories on GitHub →