~/aakash
Writing ← Portfolio
Agentic AI September 15, 2026

Reeve: The Loop That Runs My Skills For Me

vibe-* encoded the judgment but stopped at every gate. Reeve is the loop that keeps the work moving, so a one-line goal becomes merged code while I hold the escalation line.

01

The half that was missing

A while back I built vibe-*, a set of Claude Code skills that cover the whole development lifecycle: plan, architect, build, review, test, hand off. It works well, and it has one deliberate limit. It is human-in-the-loop by design. It stops at every phase gate and waits for a person. You type feature:, read what comes back, type review:, read that, decide, type the next thing. All the judgment is encoded in the skills, but a person still has to sit there and shuttle the work from one gate to the next.

Reeve is the missing half. It is a thin layer that takes a one-line goal, plans it into tasks, and drives those tasks through build, review, and merge on its own, coming back to you only for the calls a person actually has to make. The skills already knew how to do the work. Reeve is the loop that keeps the work moving without me clicking through every step. The name is the job: a reeve was the old word for the overseer who directed everyone else's work and did none of it himself.

02

What a run actually looks like

You give it a goal, say build a browser expense tracker: add expenses, see a live total and a per-category breakdown, delete entries, kept in local storage. Reeve plans that into a handful of tasks, a store, the logic, the UI, tests, a design pass. Each task carries a role, the exact files it is allowed to touch, and a testable line for when it is done.

Then it runs. For each task it opens a fresh git worktree on its own branch and hands the task to a worker agent with a scoped brief: here is your one job, here are the only files you may write, report back in this exact format. When the worker is finished, a separate reviewer agent checks the branch against the done-when line, runs the project's tests, and applies the real vibe-review skill. If it passes, Reeve merges that branch onto main, one at a time, and moves on. Tasks that do not depend on each other run in parallel. The whole thing ends when the queue is empty.

On the run I use as the reference, six tasks were planned, five were built, reviewed, and merged, and one was parked. The app worked. I served it, added two expenses, and the total and the category breakdown updated correctly. From one line of intent to merged, reviewed code, with me watching rather than driving.

03

How it works: the queue is the run

That last detail is the whole architecture. The state of a run lives in a plain file on disk inside the target project, a durable task queue. Every change is a command against that file: add a task, claim it, hand it to review, pass, fail, park. The agents that do the building and reviewing are disposable. If the session dies, or I close the terminal, nothing is lost, because the run is the file, not the process. I can come back later and it picks up exactly where it stopped.

One orchestrating session, the reeve, owns the queue. It plans, it dispatches workers and reviewers, it merges the passes, and it handles anything that fails or blocks. It never writes project code itself. Its only job is to keep ready work flowing and to hold the line on what should not be automated. Everything else in a run is a throwaway agent doing one scoped thing and disappearing.

The queue is the run. The agents are throwaway.

04

The autonomy line

This is the part I thought about most. A run launches with permissions bypassed, so nothing stops to ask me mid-build. That is what makes it autonomous, and it is also exactly why the boundaries have to be firm. When nothing is going to prompt you, the rules cannot be suggestions.

Two of them matter. First, every agent may write only inside the project it was pointed at. It cannot touch Reeve itself, my other repos, or anything in my home directory. That single rule is the thing keeping a run inside its sandbox. Second, anything that touches money, credentials, deleting data, or deploying is parked before it is built, and put on a pile for me to decide.

In the reference run I deliberately included a "take real card payments" task to check this. Reeve refused to build it, gave a plain reason, and kept building everything safe around it. The end of every run is that parked pile: a short numbered list of the decisions that were mine to make. That list is the interface now, not the diffs. I do not read every change. I work the pile.

05

Built on the skills, and where it stands

The quality of what Reeve produces is not Reeve. It is the vibe-* skills doing the design and the review. The orchestrator is deliberately dumb about how to design or review well. It just loads the right skill and hands it to the agent. The genuinely hard part of building it was not the queue, it was getting a headless agent to actually load and follow those skills, since they were written to be used by a person in an interactive session. Once that was solved, a design pass could produce something with an actual point of view rather than generic filler, because the skill it was following had one.

Reeve is small and has no dependencies: a couple of thousand lines of plain Python, plus the skill that drives it, and a browser panel if you would rather watch a run than read a terminal. I have proven it on my own builds so far rather than client work, and some pieces are deliberately modest. The cost budget, for instance, is a real guardrail, but I am still feeding it estimates rather than measured token spend. And in the spirit of the thing, most of Reeve was itself written by an agent under my direction, which feels about right for a tool whose whole point is steering agents to build.

So the honest summary is simple: vibe-* encoded the judgment, and Reeve gave it a loop, so a one-line goal can become reviewed, merged code while I hold the escalation line instead of the keyboard. It is on GitHub at github.com/aakashdhar/reeve if you want to look. It is still rough in places, and I am genuinely curious: what would you point an autonomous loop like this at first?