The eight pipeline stages (+ the action layer)
Since 0.11 (D-56, superseding D-40) the orchestrator is a pure conductor: it dispatches each enabled stage as a sub-agent from _processes/02_orchestrator/stages/, verifies the artifact on disk, parses the stage's STAGE_HANDOFF, and writes only the shared job-root state (PROGRESS.md / note.md / key-findings.md) itself — it never writes source, tests, docs, or any stage report. brainstorm and review are sequencing stages whose helper sub-agents (under _processes/04_brainstorm/ and _processes/10_review/) likewise return a structured handoff.
01 — clarify (runs first)
Mandate. Pressure-test the user's intake against the codebase before anyone spends compute on it. Two equally important jobs:
- Clarify the mandate — make
JOB.mdclear, internally consistent, and ready to be acted on. - Validate against the application — does the requested change make sense given what already exists?
Seven questions clarify answers
- Is the mandate clear and internally consistent?
- Is the mandate proper, secure, and standards-aligned?
- Does the requested change already exist?
- Is there a similar pattern we can replicate?
- Is the change feasible with what the project has today?
- Are key references missing?
- Do the user's stated assumptions still hold?
Inputs: JOB.md, PROGRESS.md, app.md, every canonical context file listed in app.md, the host source tree (targeted reads).
Outputs: 01_clarify/clarify-report.md + a proposed Mandate-body rewrite in the stage handoff; the orchestrator applies that proposal to JOB.md.
Source: _processes/02_orchestrator/stages/00_clarify.prompt.md
JOB.md is the canonical version every downstream agent reads. Stale or contradictory wording silently biases every later read — clarify exists to remove that risk. The original user prose is preserved at <slug>/00_intake/<source filename> — the creation/bootstrap step moves it there on success (D-28 / D-41).
Completion: awaiting_user_input with a Questions for user block; the orchestrator relays the question via AskUserQuestion, then re-dispatches clarify with the answer. The only early-exit is Verdict: NO_CLARIFICATION_NEEDED when the mandate is crisp and nothing needs asking.
02 — brainstorm (solo or dual)
Mandate. Produce divergent thinking on the mandate, then reconcile when configured for dual mode. In dual_consolidate, the Claude and Codex writers are launched together in one batch and think independently — neither receives the other's artifact — and a third sub-agent merges them once both artifacts exist. In solo, only the selected writer runs.
Helper sub-agents
| Sub-agent | Output | Source |
|---|---|---|
| Claude | 02_brainstorm/claude.md | _processes/04_brainstorm/claude.prompt.md |
| Codex | 02_brainstorm/codex.md | _processes/04_brainstorm/codex.prompt.md |
| Consolidate | 02_brainstorm/consolidated.md | _processes/04_brainstorm/consolidate.prompt.md — runtime-agnostic; dispatched to whichever runtime is named in JOB.md → Brainstorm → Consolidator (D-30). |
Brainstorm modes
| Mode | Sub-agents that run | Default for |
|---|---|---|
dual_consolidate | Claude writer and Codex writer launched together, neither receiving the other's artifact, then Consolidate once both artifacts exist (runtime = per-job Consolidator field, default claude) | feature, brainstorm, app_plan |
solo | One of Claude / Codex (no consolidate sub-stage) | Set explicitly in JOB.md |
none | Skip the stage entirely | bug |
consolidated.md, the consolidate sub-agent inspects the two brainstorms for contested decisions. If any are found, it returns Completion: awaiting_user_input instead of writing; the orchestrator relays via AskUserQuestion and re-dispatches with the answers injected. The loop iterates (cumulative, no fixed cap) until no contested decisions remain. When the two brainstorms already agreed, no questions are asked and the sub-agent writes consolidated.md on the first dispatch.
In dual mode, once consolidated.md is on disk, the consolidate sub-agent ends by posting a human recap in chat — concept, decisions made, pending decisions, risks — and an explicit "Please review and confirm before continuing to the plan stage." Then the orchestrator stops at the after_brainstorm gate if it is yes.
03 — plan
Mandate. Turn the mandate (and consolidated brainstorm if present) into a complete, actionable execution plan.
What the plan contains
- Executive summary — readable in 30 seconds.
- Current state — what exists today that this job touches; cite paths.
- Target state — what the codebase looks like when done.
- Phases overview — table of phases (when multi-phase), carrying a
Parallel groupcolumn that records which phases could run together. - Phase details — one section per phase with prerequisites, parallel group, scope, goal, boundaries (owns / explicitly-not-this-phase / leave-ready-for-next — required for multi-phase plans), atomic checkboxes, files to create / modify, acceptance criteria.
- Files-to-create / modify — flat aggregate inventory.
- Acceptance criteria (job-level) — what validate runs against.
- Risks and mitigations.
- Open questions.
- Handoff to execute — names the next ready phase.
Output: 03_plan/PLAN.md. Re-running overwrites the file (D-3) — Git is the version history.
Source: _processes/02_orchestrator/stages/02_plan.prompt.md
04 — execute (continuous phases)
Mandate. Implement the next pending phase from PLAN.md, then continue to the next pending phase in the same run unless a gate, blocker, or pending question stops execution (D-42, D-46 — no context-budget pause; gate / blocker / pending question are the only valid stops).
Process
- Open
PLAN.md: read the phases-overview map (which phases ran before / after you) and the nextPENDINGphase in full, including its Boundaries block. Don't read prior phases' logs for context — read the code an earlier phase produced when you need its output. A phase'sParallel groupis declaration-only: it never authorizes running two phases at once — execute still runs exactly onePENDINGphase per dispatch. - Plan internal steps before writing any code.
- Search for reuse before adding new code.
- Execute the phase — create / modify the listed files.
- Run the project's after-change commands (per
app.md → Common commands). - Verify acceptance criteria. "Would a staff engineer approve this?"
- Update
PLAN.mdin place: tick checkboxes, set phase status toCOMPLETE. - Write
04_execution/notes.md(or per-phase log when multi-phase).
Execute never stages, commits, or opens a PR — the job's commit/PR work (and the before_commit / before_pr gates) happens in the orchestrator's post-validate git landing (_processes/_shared/git-landing.md, D-63).
Outputs: source code under the host project · 04_execution/notes.md · optionally 04_execution/phases/phase-<N>.md.
Source: _processes/02_orchestrator/stages/03_execute.prompt.md
05 — test
Mandate. Add tests for the code execute produced, run the suite, capture the result.
| Flag | What it means |
|---|---|
Unit: yes | Unit tests for every new / changed piece of business logic. |
The test command comes from app.md → Common commands; the framework from app.md → Test framework. The test stage does not assume a framework.
Output: test files under the host project · 05_test/notes.md.
Source: _processes/02_orchestrator/stages/04_test.prompt.md
06 — document
Mandate. Update the project's documentation so the new code is discoverable, explained, and integrated into the existing docs.
Mapping rules (from app.md → Documentation conventions)
- New service / component → update services reference.
- New entity → update entities reference.
- New API endpoint → update API reference.
- New top-level pattern → update canonical rules file.
- New how-to → add a new guide and link from index.
- New ADR-style decision → add an ADR.
The stage reads existing doc files first to match tone, structure, and depth. It does not introduce a new style.
Output: doc updates in the host project · 06_document/notes.md.
Source: _processes/02_orchestrator/stages/05_document.prompt.md
07 — review (opt-in)
Mandate. Distinct from validate: a craft pass on design, naming, readability, reuse, performance hazards, failure handling. Review is the quality pass; validate is the contract check that follows it. Since 0.12 (D-59) review runs immediately before validate, so any code it touches is re-certified by validate. Review gains a bounded code auto-fix carve-out (modeled on validate's): on MERGE_AFTER_FIXES it applies clear, local, non-behavioral fixes obvious from the diff (each logged, then re-verified) without a loop-back; on BLOCKED_PENDING_REWORK it emits a Loop-back request the orchestrator's loop-back engine consumes.
Blocker IDs and the light-fix route (D-83). Every blocker-severity finding carries a stable R<NNN> identifier, and a BLOCKED_PENDING_REWORK handoff may additionally carry an optional Rework classification (Candidate route / Finding IDs / Reason) nominating one or two of them for the loop-back engine's light-fix route. The field is a routing claim, not an authorization: the orchestrator validates it itself against seven conditions and falls back to the full loop-back whenever any of them fails. Review never routes, never dispatches the fixer, and gains no extra fix authority from emitting it.
Verdict vocabulary
| Verdict | Meaning |
|---|---|
OK_TO_MERGE | Quality is acceptable; only nits remain. |
MERGE_AFTER_FIXES | Major findings must be addressed before merging; clear local ones are auto-fixed in place. |
BLOCKED_PENDING_REWORK | Blocking findings; emits a loop-back request. The orchestrator re-runs earlier stages, or — when it validates an optional Rework classification covering one or two R<NNN> blockers — takes the bounded light-fix route instead (D-83). |
Output: 07_review/review-findings.md.
Runtime. Per JOB.md → Per-stage config → Review → Agent (closed vocabulary claude · codex; default codex), the orchestrator dispatches the review helper to a Claude general-purpose sub-agent when claude or runs it via codex exec when codex — exactly like validate. The default is codex so review's fresh eyes come from a deliberately different model than execute/test/document (which run on Claude); claude is a supported choice. The prompt is runtime-agnostic with a runtime-conditional Codex technique section.
Source: the orchestrator sequences review via its dispatch logic (D-56 absorbed the former sequencing milestone) · _processes/10_review/code_review.prompt.md (the dispatched review helper).
08 — validate (final certification gate + internal code review)
Mandate. The final certification gate (since 0.12, D-59 — it now runs after review): a fresh-eyes cross-check by a different agent plus an internal code review. Reads the original intake, brainstorm, plan, execution/test/document notes, 07_review/review-findings.md when review is enabled, and the uncommitted git diff (uncommitted by design — the orchestrator's git landing commits only after validate passes, D-63), and reviews across nine dimensions: logic correctness, coding-standard conformance, security, reuse, test coverage, documentation accuracy, mandate fit, architectural fit, and out-of-scope changes. It may apply narrowly-bounded unambiguous local code/doc fixes (each logged in 08_validate/validate-report.md → ## Auto-fixes applied, then re-checked before the verdict); any design, architecture, behavioral, risky, or ambiguous-trade-off decision is relayed to the user via awaiting_user_input rather than guessed (D-38). On FAIL it emits a Loop-back request the orchestrator's loop-back engine consumes (D-59).
Runtime. Per JOB.md → Per-stage config → Validate → Agent (closed vocabulary claude · codex; default claude; D-38), the orchestrator dispatches the validate stage to a Claude general-purpose sub-agent when claude or runs it via codex exec when codex (D-56). The optional review stage is separate and runtime-selectable the same way (Review → Agent, default codex).
Verdict vocabulary
| Verdict | Meaning |
|---|---|
PASS | Every success + acceptance criterion is met. Suite green. Docs accurate. Ready to land. |
PASS_WITH_FOLLOW_UPS | Core scope is met but small gaps remain. Each gap becomes a documented follow-up job. |
FAIL | One or more critical criteria are unmet. Emits a loop-back request so earlier enabled stages re-run (up to the per-job cap). |
Output: 08_validate/validate-report.md with the verdict, findings, and an ## Auto-fixes applied section; plus zero-or-more project files edited in place under the bounded auto-fix carve-out.
Source: _processes/02_orchestrator/stages/06_validate.prompt.md (runs via codex exec when Agent: codex).
action — its own layer, not a stage (D-45, revises D-44)
Mandate. In 0.8 action was a ninth stage; 0.9 (D-45) removed it from the stage vocabulary and made it its own layer under _processes/03_action/. An action-layer request type runs only its own logic — it does not enter the stage pipeline: audit is — since 0.13 (D-61), reshaped from a singlet — its own per-kind multi-phase orchestrator (03_action/audit/audit.prompt.md, no clarify — defined by its ## Action config; resolves a per-kind 3-phase matrix and runs the read-only managed sub-agent loop from _processes/_shared/managed-subagent-loop.md once per enabled phase; the old 03_action/audit.prompt.md path is a tombstone) quick_fix is its own parallel-fixer orchestrator (03_action/quick_fix/quick_fix.prompt.md); and D-48 (0.10) added two documentation orchestrators — doc_upkeep (03_action/doc_upkeep/doc_upkeep.prompt.md, validate+update an existing HTML docs folder) and doc_create (03_action/doc_create/doc_create.prompt.md, generate new HTML docs). D-68 (0.14) added a fifth action type, codereview_fix (03_action/codereview_fix/codereview_fix.prompt.md), which validates externally-supplied code-review comments against the whole app and fixes only the verified-accurate ones in the working tree (no commit/PR). The orchestrator routes action-layer + operation jobs via orchestrator.prompt.md → §C.route, before the stage loop. The action's identity is still its Action.Kind (security, find_issues, doc_drift, revalidate, custom, operation_record). This resolves the D-12/D-17 "reuse the stage construct" tension: action's singlet/orchestrator shapes do not fit the stage model, so it warrants a layer — constrained to two shapes + one shared loop.
Output: 09_action/action-report.md.
Source: _processes/03_action/audit/audit.prompt.md, _processes/03_action/quick_fix/quick_fix.prompt.md, _processes/03_action/operation-record-landing.md (the retired 08_action.milestone.md was split into these).