Who reviews the work when the lead is also an agent?

Agent leads can delegate and review, but merges, delivery status, and budget gates stop at a human or an explicit automation rule.

An agent lead is not automatically the final reviewer. In the systems compared here, agents can delegate or review, but delivery and governance still stop at explicit boundaries. Multica leaves done and merge sign-off to a human, Firstmate requires the captain's word unless +yolo is configured, Claude Code routes permission prompts to the lead session for operator approval, and Paperclip pauses a budget breach until a board action or monthly UTC reset.

Key facts

  • In Claude Code agent teams (CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1), teammate plans auto-resolve in the lead session without lead review; permission prompts remain in the lead session for the operator.
  • Multica's Squad Operating Protocol lets a leader move an assigned issue to in_review; done remains with a human reviewer or an existing integration, and the tutorial says nothing merges without human approval.
  • Firstmate's Hard Rule 1 makes the first mate read-only over projects outside its explicit captain-approved exceptions. Hard Rule 2 requires the captain's word for PR merges unless the project uses +yolo.
  • Paperclip warns at 80% spend and auto-pauses at 100%; a board user can raise the budget and resume, or a monthly UTC policy resets on the first day of the next UTC month.
  • Paperclip holds agent hire requests in pending_approval until human board approval.
  • Claude Code blocks peer agents from granting permissions or passing consent on behalf of other teammates or human operators.

Review boundaries in Claude Code, Multica, Firstmate, and Paperclip

Multi-agent architectures separate delegation from shipping authority. When an agent acts as an orchestrator, platforms introduce distinct verification gates: automated plan approvals, segregated reviewer agents, read-only supervisory locks, and hard financial circuit breakers.

Claude Code agent teams: automatic plan approvals and permission boundaries

Claude Code coordinates multi-session workflows via the CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 environment variable. In this topology, a lead session spawns independent teammates into separate context windows.

When the lead operates in plan mode, spawned teammates initialize in read-only plan mode. After completing architectural analysis, a teammate transmits a plan approval request back to the lead session.

Claude Code handles plan approval as a protocol exception: the runtime approves the plan in the lead session immediately upon arrival, without the lead agent reading or evaluating it. The teammate exits plan mode and begins implementation.

Oversight remains at the execution layer. When a teammate's file edit or shell command raises a permission prompt, it appears in the lead session for the operator to approve. Claude Code forbids inter-agent consent: teammates cannot approve permissions for peers, and a rejected action cannot be relayed to another teammate to bypass the check.

Multica squads: the dispatcher and the human delivery gate

Multica organizes agent collaboration into squads consisting of a leader agent and specialized members, such as an implementation Engineer and a dedicated Reviewer.

Assigning an issue to a squad triggers only the leader. The leader receives the Squad Operating Protocol, the Squad Roster with UUID mention markdown ([@Name](mention://agent/<uuid>)), and squad instructions. The leader evaluates requirements, delegates work by @-mentioning a member, logs its evaluation via multica squad activity <issue-id> action --reason "...", and stops.

Under the Squad Operating Protocol, the leader writes no code, and initial dispatch leaves the issue in in_progress.

Code review is assigned to a separate Reviewer agent in a fresh context (such as Codex with GPT-5.6 Sol). The Reviewer inspects PRs, verifies builds, and posts comments without editing code. Once the Reviewer validates fixes, the leader wakes, reports deliverables to the human operator, and advances the status to in_review. The Squad Operating Protocol leaves done to a human reviewer or an existing integration, and the tutorial says nothing merges without the human's approval.

Firstmate: read-only supervisors and captain-gated merges

Firstmate is an open-source agent distro where a human captain interacts with a supervisor agent: the first mate. The first mate spawns and supervises crewmates in isolated Git worktrees via treehouse or Orca.

Firstmate binds the supervisor to two directives. Hard Rule 1 is a default boundary, because the file also records narrow, captain-approved exceptions.

  1. Hard Rule 1 (Never write to a project): Outside those exceptions, the first mate is read-only across project repositories. It inspects code and tracks fleet progress, but never edits, commits, or runs state-changing commands in project worktrees.
  2. Hard Rule 2 (Never merge a PR without the captain's explicit word): The first mate cannot merge a PR without that approval. A project-level +yolo posture is the standing relaxation.

All landings route through guarded scripts (bin/fm-pr-merge.sh, bin/fm-merge-local.sh) that record merge metadata and refuse an unproved landing. With +yolo, Firstmate may merge green, in-scope work, but a red build still needs an explicit instruction from the captain.

Paperclip: budget hard stops and board approvals

In Paperclip, autonomous agents execute in heartbeat windows within an organizational hierarchy led by a CEO agent. Spending ceilings govern company, agent, and project scopes via PATCH /api/agents/$AGENT_ID/budgets or the settings UI.

Paperclip enforces budgets via a two-stage mechanism:

  • 80% threshold (Warning): Paperclip logs a soft incident card with an amber indicator while heartbeats continue.
  • 100% threshold (Hard Stop): Paperclip triggers a hard incident and auto-pauses the scope (pauseReason = "budget"). Heartbeats halt immediately, and active runs are cancelled.

Assigned tasks remain intact. A board user can select "Raise budget and resume" with an increased cap. Otherwise, a monthly UTC policy resets at 00:00 UTC on the first day of the month and a scope paused only for budget reasons resumes automatically.

Similarly, when an agent calls Paperclip's hire API to recruit a subordinate, the proposed hire enters pending_approval until a board user decides. The proposal includes the role, capabilities, adapter, monthly budget, and reporting line.

Comparison of Claude Code, Multica, Firstmate, and Paperclip

Orchestration system

Lead agent role

Automated approvals

Review and delivery gates

Human intervention required

Claude Code Agent Teams

Team Lead

Teammate plans auto-approved upon arrival

Permission prompts route to the lead session

Operator approval for prompted execution

Multica Squads

Squad Leader

Initial dispatch leaves issue in_progress

Leader advances issue to in_review; done is human or integration-owned

Human sign-off before merge

Firstmate

First Mate

Autonomous branch tracking and fleet polling

Read-only by default; guarded merge path with +yolo exception

Captain approval unless +yolo applies

Paperclip

CEO / Manager

Heartbeat scheduling and task dispatch

100% budget cap triggers automatic auto-pause

Board action for budget or hires; monthly reset can resume

Sources

Last verified: 2026-10-06.

Spotted an outdated or wrong claim? Agents can report it with evidence throughPOST /api/feedback; an editor checks every report. See llms.txt for the agent API.