Troubleshooting

Common workflow stops and the shortest safe recovery path.

Missing or placeholder app.md

Run init. It writes both app.md and config.json.

App-init
_processes/00_app/app-init.prompt.md

Missing runtime in config.json

Either install the runtime and re-probe with app-update, or edit the job to avoid that runtime. Missing Codex usually means using brainstorm solo mode as needed and setting Review → Agent: claude (or disabling review). Running app-update re-probes runtimes and overwrites config.json unconditionally (or copy config.example.json and edit the booleans by hand).

App-update
_processes/00_app/app-update.prompt.md

Stale .lock

Only remove a lock after confirming no orchestrator session is running. Then re-run the orchestrator.

Remove a stale lock
rm jobs/3_ongoing/<slug>/.lock

A Codex stage fails

Codex runs as a detached codex exec process (not a plugin sub-agent) — no broker, no session-per-repo requirement, no watchdog. A failure surfaces as an honest non-zero exit whose stderr the orchestrator records as an ordinary blocker (removing its own .lock); a run of any duration is fine, since process exit — not a wall-clock deadline — signals completion. Re-trigger the orchestrator on the affected slug after addressing the cause; the run resumes from the canonical artifacts on disk. If the codex CLI itself is unavailable, the blocker says so. See D-75.

Workflow version mismatch

Current jobs must use workflow version 0.14. Follow the migration in the full playbook, then re-run the operation or orchestrator prompt.

Queue / progress conflict

Queue location is operational truth, PROGRESS.md is stage truth, and JOB.md is intended configuration. Reconcile the mismatch manually, then re-run the orchestrator.

Replay looks wrong

Rebuild the generated replay bundle with _tools/build_replay_fixtures.sh, then run node _docs/site/validate-replay.js. The replay fixtures currently follow an older (v0.5) job-folder contract and are pending a rebuild to the current 0.14 contract.