Troubleshooting
Common workflow stops and the shortest safe recovery path.
Missing or placeholder app.md
Run init. It writes both app.md and config.json.
_processes/00_app/app-init.prompt.md
Missing runtime in config.json
Either install the runtime and re-probe with app-update, or edit the job to avoid that runtime. Missing Codex usually means using brainstorm solo mode as needed and setting Review → Agent: claude (or disabling review). Running app-update re-probes runtimes and overwrites config.json unconditionally (or copy config.example.json and edit the booleans by hand).
_processes/00_app/app-update.prompt.md
Stale .lock
Only remove a lock after confirming no orchestrator session is running. Then re-run the orchestrator.
rm jobs/3_ongoing/<slug>/.lock
A Codex stage fails
Codex runs as a detached codex exec process (not a plugin sub-agent) — no broker, no session-per-repo requirement, no watchdog. A failure surfaces as an honest non-zero exit whose stderr the orchestrator records as an ordinary blocker (removing its own .lock); a run of any duration is fine, since process exit — not a wall-clock deadline — signals completion. Re-trigger the orchestrator on the affected slug after addressing the cause; the run resumes from the canonical artifacts on disk. If the codex CLI itself is unavailable, the blocker says so. See D-75.
Workflow version mismatch
Current jobs must use workflow version 0.14. Follow the migration in the full playbook, then re-run the operation or orchestrator prompt.
Queue / progress conflict
Queue location is operational truth, PROGRESS.md is stage truth, and JOB.md is intended configuration. Reconcile the mismatch manually, then re-run the orchestrator.
Replay looks wrong
Rebuild the generated replay bundle with _tools/build_replay_fixtures.sh, then run node _docs/site/validate-replay.js. The replay fixtures currently follow an older (v0.5) job-folder contract and are pending a rebuild to the current 0.14 contract.