AI · Project
Claudia Control Plane, my own Claude Code agent harness
A Next.js dashboard that drives a fleet of coding agents on a Mac mini at home, billed as a flat-rate subscription instead of per-token API credits. I create a job, approve it, and a router on the Mac mini claims it, runs it in an isolated workspace, and reports back with a summary and a ledger of everything it decided while I was asleep.
The single design rule is that autonomy lives inside one bounded job. Everything around the job — scheduling, claiming, model selection, workspace isolation, outcome classification, recovery — stays deterministic and auditable. Here is how it fits together.

The architecture
Five layers, left to right: what I see in the browser, the control plane on Vercel, the database, the agent runtime on the Mac mini, and the outside services it talks to. Click any box to see what it does.
Interfaces · browser
- Dashboard
- Job console
- Chat
- Career console
- Memory browser
Control plane · Vercel
- Next.js app + API
- Vercel crons
Data · Neon
- Neon Postgres + pgvector
- Voyage AI embeddings
Agent runtime · pallas
- colima · Docker engine
- task-router
- Claude Code CLI
- email-poller
- vault-indexer
- Doctor + remediation
External services
- 1Password
- Tailscale
- Claude Max
- Kimi coding endpoint
- GitHub
- Gmail
Click any box to see what it does. Drag the canvas to pan; the page scrolls normally. Rust arrows trace the main flow. Thin grey lines are everything around it, and dashed grey lines are supporting config and platform.
How a job runs
Every unit of work is a “job” with a lifecycle: draft, pending, approved, running, succeeded. The router polls for approved jobs, cuts a fresh git worktree so two agents can never share a checkout, runs headless Claude Code inside it, and classifies the outcome from a structured summary block the agent has to produce. A self-report is a claim, so the router cross-checks it rather than taking “done” at face value.
Rendering diagram…
Why it is built this way
The diagrams show what it is. These are the calls that shaped it — the part that actually matters when someone asks how I think about building systems.
Flat rate, not metered
Drive the heavy agents through a coding CLI billed as a subscription, and never through a per-token API.
WhyThe API-metered equivalent was heading toward roughly $1,500 a month for the same volume of work.
Trade-offI run and maintain a machine at home instead of going fully serverless, and a quota ceiling can pause work.
The harness is the product, the model is a setting
A job carries a model field. Setting it re-points the same CLI at a different vendor and pins every internal model slot to that vendor, keeping the skills, the worktree, the turn budget, and the summary contract intact.
WhyThe engineering value is in the harness around the model, so the backend behind it should be a configuration value rather than a rewrite.
Trade-offEach backend brings its own quota shape and failure modes, and I have to keep two subscriptions honest.
No model on the serverless side
The Vercel app does not call a model API. Anything that needs one enqueues a job for the runtime instead.
WhyTwo billing paths for the same work is how a subscription-first system quietly becomes a metered one.
Trade-offA handful of legacy call sites are still migrating. They are tracked as named violations rather than quietly grandfathered.
Containers, because macOS is a desktop OS
The agents run as containers on a headless Docker engine, not as raw launchd jobs.
WhyThree macOS behaviours kept breaking headless work: permission dialogs that block a background process with nobody to click Allow, a credentials file that other system processes invalidated about weekly, and restart semantics that cost tens of hours to debug. None of them have a fix at the macOS layer.
Trade-offOne more layer to understand, and the container host itself now needs its own watchdog.
One job, one worktree
Every repo-backed code job gets its own git worktree on a per-job branch, cut fresh from the base branch.
WhyTwo agents editing one checkout is a corruption bug waiting to happen, and the repo a job edits can be this repo.
Trade-offA rebase step at the end of every job. A conflict lands as a review branch rather than a merge, which is slower and correct.
Humans approve before agents act
Every job waits in a pending state; nothing executes until it is approved.
WhyAutonomy is a setting, not a default. I decide how much rope each kind of job gets.
Trade-offA beat of latency per job, traded for control.
Fail loud, no fallbacks
A missing secret throws. An unset auth token returns a 503 instead of skipping the check. A quota block pauses the queue rather than rerouting onto the metered path.
WhyA silent fallback once meant data could have been unrecoverable with no warning during development. A silent reroute would have quietly rebuilt the bill I built this to avoid.
Trade-offMore upfront wiring, and errors surface immediately instead of in production.
The healer is watched, and nothing heals the healer
The Doctor can fix what it detects, but only by naming an action from a hardcoded whitelist that a separate process on the Mac mini executes. Read-only diagnosis jobs propose everything else. There is no tier that edits code, touches secrets, or deletes anything.
WhyAgent failures are silent and cryptic, so something has to act. An automated fixer with an open command channel is a far worse failure than the one it fixes: even a compromised control plane can at most restart a container.
Trade-offThe whitelist has to be extended by hand for each new fault, and some faults have no safe automatic fix at all.
Memory is a database, not a prompt
Long-term memory lives in Postgres with pgvector and is retrieved by relevance.
WhyThe agent stays consistent across sessions without stuffing everything into the context window.
Trade-offAn embedding and retrieval step on every relevant call.
Stack
Boring, well-supported pieces wired together carefully.
Next.js 16·React 19·Tailwind 4·Neon Postgres·pgvector·Claude Code CLI·Kimi k3·Voyage AI·Docker on colima·1Password service accounts·Tailscale·launchd·UptimeRobot·Vercel