Smith is an open source agent skill that runs several coding tasks at once with Claude Code, Codex, Cursor, OpenCode or any CLI agent, inside the Herdr terminal. Each task gets its own git worktree and Herdr tab, and nothing ships without your approval.

Links

On a normal day I have four or five small tasks waiting. A bug, a refactor, a follow-up from yesterday’s review. Any coding agent can handle one of them well. The trouble starts when I try to handle all of them at once.

I built Smith for that. It’s an agent skill that lets one agent session orchestrate several worker agents, each on its own task, its own branch and its own terminal tab. The name is a nod to Agent Smith from The Matrix, the agent who copies himself. The difference is that these copies stop and ask before anything important happens.

It started as a Claude Code skill. It now works with Codex, Cursor, OpenCode, Copilot, Pi, Antigravity, Grok and a long list of others, and a single run can mix them.

The code is on GitHub: aslamdoctor/skills.

Heads up: Smith only works inside Herdr. Herdr is a terminal multiplexer built for coding agents, and the skill relies on it to open worker tabs, send prompts, read each agent’s state and show gates in the sidebar. If your agent isn’t running in a Herdr pane, the skill stops at its first check and tells you so. It won’t run in a plain terminal, tmux or an IDE terminal. Install Herdr 0.9 or newer first.

Why “just open more terminals” stops working

The first thing everyone tries is three terminal tabs with three agent sessions, one issue pasted into each. I did this for weeks. It works for a while, then the same few things go wrong:

  • Branches collide. Two agents in one checkout means one of them is always on the wrong branch or stashing the other’s changes.
  • You miss the one that’s waiting. A tab sits on a permission prompt for twenty minutes while you watch a different one.
  • You repeat the setup. Base branch, commit format, test command, how PRs get opened. Again, in every tab.
  • Context dies overnight. Review comments arrive the next morning, the session that wrote the code is gone, and a fresh one has to rebuild everything.

I wanted one session to talk to, with workers doing the coding in isolated spaces and reporting back.

What Smith does

When you invoke the skill, your current agent session becomes the orchestrator. It doesn’t write task code. It creates workers, sends them prompts, watches them, and brings their questions to you with a recommendation.

Every task gets:

  1. A git worktree at ~/<repo>-wt/<branch>, so each worker has a real checkout on its own branch and never touches your main one.
  2. A background tab in Herdr. The worker’s pane shows its task and gate in the sidebar, like #102 · spec.
  3. A worker agent, started with a kickoff prompt that already includes your project’s rules.

You stay in one pane. The workers come to you.

Gates keep you in charge

I don’t want five agents opening five PRs while I’m making coffee. So every worker moves through gates and stops at each one.

GateWhat the worker has doneWhat you decide
specRead the issue and written a short spec and implementation plan. No code yet.Approve, change the approach, answer open questions
implementedWritten the code, run lint and tests, committed locally. Nothing pushed.Review the commit, ask for changes, or approve shipping
prPushed and opened the PR with your project’s ship flowCheck the PR and approve any replies to review bots

There’s also a question gate for when a worker needs a decision partway through.

At each gate the worker runs mark.sh, a small script that appends to a gate log, updates the run record and changes its sidebar label. Then it waits.

The orchestrator turns that into a short report. At the spec gate you get the approach in a few lines, any deviations from the task, and the open questions in a table with a recommended answer for each. It also flags when two workers are about to touch the same function or the same docs section. That happens more often than you’d expect.

You answer once. The orchestrator relays your exact words to the worker without paraphrasing them or adding decisions of its own.

Bring your own agent

This is the biggest change since the first version. Worker kinds come from Herdr, and the skill ships with a small agents.tsv file that maps each kind to its Herdr integration, its resume arguments and where its transcripts live.

Support comes in tiers:

TierAgentsWhat works
Aclaude, codex, agy, grok, pi, opencode, copilot, cursor, devin, droid, kimi, qodercli, qwen, letta, hermes, mastracode, kilo, omp, with the Herdr integration installedEverything, including resume with full context
BThe same agents without the integration, plus amp, kiro, maki, museStart, prompt, live state and watchers. Resume falls back to a catch-up prompt
Cgemini, clineLike B, but Herdr’s state detection is less tested
DAny CLI agent Herdr doesn’t recognizeLaunched with --cmd. Prompts are pasted into the pane and progress is tracked through gate lines

The one hard requirement for a worker is that it can run a shell command, since that’s how it marks gates.

Tasks in one run can use different kinds. You could put Claude Code on a tricky refactor, Codex on two small bugs and OpenCode on a docs task, all reporting to the same orchestrator.

The orchestrator can be any of these agents as well. If yours can run a command in the background and gets woken when it exits (Claude Code’s background shell does this), the watchers run that way. If it can’t, the watchers run detached and wake the orchestrator by sending it a Herdr prompt that starts with [smith]. If you happen to be answering a dialog when a report arrives, the watcher waits for it to close instead of typing into it.

A run, from start to finish

Start your agent inside a Herdr pane in the repo and ask for parallel work:

/smith 101 102 103

Plain language works too: “work on 101 and 102 in parallel”, “use codex workers for these issues”, or “pick 3 open issues that don’t depend on each other”. If you ask it to pick, it lists open candidates, reads their blockers, and shows a table of the ones it chose and why.

1. Setup detection

Before spawning anything, the orchestrator reads the repo’s instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, .cursor/rules and similar), its own memory or notes, and the last few PRs. From those it works out:

  • the base ref for new branches and the PR base
  • branch naming and commit format
  • the ship flow (a ship skill, or the project’s PR rules)
  • test and lint commands
  • the worker kind for each task

All of it goes into one table for you to confirm. It asks only about what it couldn’t find.

Two details here save a lot of pain later. If task B needs task A’s unmerged code, B is cut from A’s branch and shipped as a stacked PR. And if two tasks will both add to a shared list (an error-code registry, a changelog section), the orchestrator decides up front which worker owns it.

Skills and slash commands don’t carry over between agents, so when a step uses a skill like a ship flow, workers get the skill’s file path (“follow <path>/SKILL.md“) instead of a slash command.

2. Spawning workers

For each task, spawn.sh creates or reuses the worktree, opens a background tab, starts the worker and records it in run.json with its kind, pane, tab and session ID.

New worktree paths sometimes trigger a first-run dialog in the worker: folder trust, a login screen, a model picker. The orchestrator doesn’t click through those for you. It shows you what the dialog asks.

Then it sends the kickoff prompt through send.sh, which confirms the prompt was actually delivered and that the worker started working. A failed send is treated as “not delivered”, never assumed to have landed.

3. Watching without polling

Each worker gets a watcher that stays quiet until the worker needs you. It reports one of four things:

  • GATE: a new gate line
  • BLOCKED: a dialog that stays on screen
  • STOPPED: no gate and idle for 15 minutes (adjustable with --grace)
  • GONE: the agent exited

A Herdr toast with a sound goes along with it, so you notice even from another window. A worker that’s idle because its own background job is running (CI, a review bot) doesn’t set it off.

4. One status table

Every report includes a board built from run.json and the live agent states:

| Task | Worker         | Gate        | Status                          | PR |
|------|----------------|-------------|---------------------------------|----|
| #101 | i101 (idle)    | spec        | Spec written, 2 open questions  |    |
| #102 | i102 (working) | spec        | Approved, implementing          |    |
| #103 | i103 (idle)    | implemented | 3 files changed, tests pass     |    |

That table is what finally stopped me clicking through tabs to see what was going on.

5. Shipping

Shipping happens per task and only when you say so. The worker rebases on the PR base if it moved, runs your ship flow and opens the PR.

Ship flows ask the same routine questions every time: run local QA, deploy to staging, who reviews. The orchestrator asks once whether your standing answers apply to every worker in the run, and puts them in the ship prompt.

Before it tells you anything “shipped”, it checks the PR itself: base branch, reviewers, labels, resolved threads and the file list.

Resume: picking up the next day

Each run is saved in ~/.smith/runs/<repo>-<date>/. run.json holds one record per task (branch, worktree, tab, pane, agent kind, session ID, gate, PR), and gates/<id>.md is that worker’s gate log.

The next morning you can say:

resume yesterday's smith run

The orchestrator finds the run, shows the board, and restarts each worker in its original conversation using the resume arguments for its kind. The worker still has its spec, your decisions and its commits in context. Tell it “the reviewer left comments on PR 412, fix them and draft replies for my approval” and it carries on from there.

For tier B to D agents, which have no session to resume, the worker starts fresh and gets a short catch-up prompt with the issue, the PR, the current gate and your decisions so far.

Stacked PRs are handled in order. Once the parent’s fixes are pushed, the child rebases onto it. After the parent merges, the child rebases onto the real base and its PR base is switched.

Lessons from real runs

Most of the work in a skill like this goes into small failure cases. The repo keeps them in references/gotchas.md, and the orchestrator reads that file before its first run in a session. A few examples:

Prompts that silently didn’t send. An early version passed a timeout flag without the matching wait flag. Herdr rejected the command and sent nothing. send.sh now checks that the prompt arrived and the worker started.

Gates marked before anyone was watching. A fast worker could mark a gate before its watcher started, and the gate was missed. send.sh now records the gate-log baseline at send time.

Approval prompts that flash and vanish. Some agents show an approval prompt for a moment and then clear it themselves. The watcher re-checks after five seconds and only reports dialogs that stay up.

Late session IDs. OpenCode and Pi report their session ID only after the first turn, so mark.sh refreshes it at every gate.

Shared test databases. Worktrees of one repo often share a test database, so parallel test runs collide. Workers are told to use a separate DB name where the config allows it, and to rerun a one-off DB failure before debugging it.

The local site shows old code. A dev site that’s symlinked or mounted to the main checkout won’t see worktree changes. The standing answer for automated local QA from workers is “No”.

Symlinked vendor and node_modules. Worktrees borrow these from the main checkout as symlinks, and a /vendor/ ignore rule doesn’t match a symlink. Workers keep them out of commits, and the orchestrator checks the PR file list.

Closing keywords on stacked PRs. Fixes #123 doesn’t link an issue when the PR base isn’t the default branch, so stacked PRs get the issue linked directly.

The rules it won’t break

  • You approve every spec, every ship and every piece of text posted to GitHub or anywhere else. Workers don’t decide for you and neither does the orchestrator.
  • Workers ask questions in plain text. If one opens an interactive dialog anyway, send.sh presses Esc before sending your answer.
  • The orchestrator never focuses or closes tabs it didn’t create, and never stops the Herdr server.
  • At wrap-up it closes only this run’s tabs, stops its detached watchers, and keeps worktrees and run.json until the PRs merge. After that it offers to remove the worktrees, delete the local branches and archive the run.

Requirements

  • Herdr 0.9 or newer, with the orchestrator running inside a Herdr pane
  • git, jq and bash
  • gh if you want the orchestrator to pick issues or open PRs
  • At least one coding agent CLI. Install its Herdr integration (herdr integration install <name>) to get session resume

Install

From skills.sh:

npx skills add aslamdoctor/skills --skill smith

Or clone the repo and symlink the skill into your agent’s skills directory:

git clone https://github.com/aslamdoctor/skills ~/skills
ln -s ~/skills/skills/smith ~/.claude/skills/smith

Other agents read skills from other directories. Symlink the same folder there too, so there’s one copy to update:

ln -s ~/skills/skills/smith ~/.agents/skills/smith   # shared location (Codex, OpenCode, Pi and others)
ln -s ~/skills/skills/smith ~/.codex/skills/smith    # Codex

For an agent without skill support, tell it: “Read ~/skills/skills/smith/SKILL.md and follow it.”

Who it’s for

If you like working through one task at a time, you don’t need this. A single agent session is fine.

Smith is for days with a stack of small, independent tasks that you want moving at the same time, while you still read every spec and approve every PR. For me that was the whole point. More workers, same control.

Start with two small issues and see how the gates feel. Then try five, maybe with a different agent on each.

The skill is MIT licensed. Bugs and ideas go in the issue tracker.