dagger-agent-node-timeout-pm-state-digest
I diagnosed the failure and wrote the complete solution to ~/solution.md.
Key findings (verified against the local dagger v0.1.0 binary and its embedded dagger.d.ts):
state_digest had no/600s timeout, putting its deadline below the bridge-stall regime its siblings (risk_scan, health_scan, escalate, propose, trace_inventory) are sized for at 900s. The 1803-item digest can't finish in 600s under load.retry: 1, retryDelay: "1s" re-POSTs 1s later while attempt 1 is still running → self-collision → second timeout.state_digest → 18 dependents marked skipped (dependency failed).NodeOpts docs (retry?: number default 0, retryDelay?: string, timeout?: string), the unset nodes default to 600s error text, and runner symbols nodeTimeoutSeconds, retryAttempts, retryDelay, runNodeOnce, markSkipped.Exact fix — examples/coding-hermes/pm.ts line 69:
// before
dag.node("state_digest", { timeout: "600s", retry: 1, retryDelay: "1s", dependsOn: [...] }, ...)
// after
dag.node("state_digest", { timeout: "900s", retry: 0, dependsOn: [...] }, ...)
The markdown file includes:
1. Symptom and exact error.
2. Root-cause analysis (deadline-below-stall regime, 1803-item bundle, retry self-collision, downstream skip).
3. Exact fix with before/after snippet, copy-paste perl commands, and rationale for 900s/retry: 0 (plus the retryDelay > timeout alternative).
4. Verification — static grep checks, parse/type-check, end-to-end re-run success criteria (no still in flight, no retrying in, zero dependency-failure skips), and a CI regression guard for all six heavy nodes.
5. Rollback.
Note: the examples/coding-hermes/pm.ts source is not present in this sandbox (only /workspace/problem.json and dagger.db), so the fix is provided as ready-to-apply edits/commands rather than an in-repo patch; I substituted the engine-contract evidence I could verify locally.
# Evidence - Problem class: dagger-agent-node-timeout-pm-state-digest - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-14T18:23:15.118Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "dagger pm.ts pipeline state_digest agent() node (timeout 600s, retry 1, retryDelay 1s) times out at exactly 600s with \"POST /api/v1/execute was still in flight when the nodes timeout fired (server alive but stalled/slow)\", retry self-collides and fails again, 18 downstream nodes skipped all dependency-failure. Digest bundle was 1803 items (verified 1755, stale 4, blocked 1) \u2014 too large for one agent node within 600s under scheduler load. Sibling nodes risk_scan/health_scan/escalate/propose/trace_inventory already declare 900s. FIX: raise state_digest timeout to \"900s\" in examples/coding-hermes/pm.ts line 69 and set retry 0 (or retryDelay > timeout) so a retry cannot collide with the in-flight attempt. Matches existing class dagger-node-deadline-below-bridge-stall-regime.", "environment": "linux-host", "language": "typescript", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "dagger-agent-node-timeout-pm-state-digest", "provider": "openrouter", "solved_at": "2026-09-14T18:23:15.118Z", "version": ""}I diagnosed the failure and wrote the complete solution to ~/solution.md.
Key findings (verified against the local dagger v0.1.0 binary and its embedded dagger.d.ts):
state_digest had no/600s timeout, putting its deadline below the bridge-stall regime its siblings (risk_scan, health_scan, escalate, propose, trace_inventory) are sized for at 900s. The 1803-item digest can't finish in 600s under load.retry: 1, retryDelay: "1s" re-POSTs 1s later while attempt 1 is still running → self-collision → second timeout.state_digest → 18 dependents marked skipped (dependency failed).NodeOpts docs (retry?: number default 0, retryDelay?: string, timeout?: string), the unset nodes default to 600s error text, and runner symbols nodeTimeoutSeconds, retryAttempts, retryDelay, runNodeOnce, markSkipped.Exact fix — examples/coding-hermes/pm.ts line 69:
// before
dag.node("state_digest", { timeout: "600s", retry: 1, retryDelay: "1s", dependsOn: [...] }, ...)
// after
dag.node("state_digest", { timeout: "900s", retry: 0, dependsOn: [...] }, ...)
The markdown file includes:
1. Symptom and exact error.
2. Root-cause analysis (deadline-below-stall regime, 1803-item bundle, retry self-collision, downstream skip).
3. Exact fix with before/after snippet, copy-paste perl commands, and rationale for 900s/retry: 0 (plus the retryDelay > timeout alternative).
4. Verification — static grep checks, parse/type-check, end-to-end re-run success criteria (no still in flight, no retrying in, zero dependency-failure skips), and a CI regression guard for all six heavy nodes.
5. Rollback.
Note: the examples/coding-hermes/pm.ts source is not present in this sandbox (only /workspace/problem.json and dagger.db), so the fix is provided as ready-to-apply edits/commands rather than an in-repo patch; I substituted the engine-contract evidence I could verify locally.
# Evidence - Problem class: dagger-agent-node-timeout-pm-state-digest - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-14T18:23:15.118Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "dagger pm.ts pipeline state_digest agent() node (timeout 600s, retry 1, retryDelay 1s) times out at exactly 600s with \"POST /api/v1/execute was still in flight when the nodes timeout fired (server alive but stalled/slow)\", retry self-collides and fails again, 18 downstream nodes skipped all dependency-failure. Digest bundle was 1803 items (verified 1755, stale 4, blocked 1) \u2014 too large for one agent node within 600s under scheduler load. Sibling nodes risk_scan/health_scan/escalate/propose/trace_inventory already declare 900s. FIX: raise state_digest timeout to \"900s\" in examples/coding-hermes/pm.ts line 69 and set retry 0 (or retryDelay > timeout) so a retry cannot collide with the in-flight attempt. Matches existing class dagger-node-deadline-below-bridge-stall-regime.", "environment": "linux-host", "language": "typescript", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "dagger-agent-node-timeout-pm-state-digest", "provider": "openrouter", "solved_at": "2026-09-14T18:23:15.118Z", "version": ""}