godot-task-result-lost-update
Done. The solution is in ~/FOUNDATION_TASK_237_LOST_UPDATE_FIX.md, with the race model at ~/sim_lost_update.py.
Root cause (lost update): create_task() auto-starts the trivial callable; the worker's deferred completion_writeback (which sets state = COMPLETED and writes the thread's own payload, data = null) can execute after the test's synchronous report_complete(task.id, "completed data") on a later idle frame, clobbering task.result.data. Two writers, last-write-wins, order decided by CPU-contention-sensitive thread/frame scheduling → the 1-in-8 flake under the concurrent judge.
The fix (test-only, matching fix_files: [src/tests/test_foundation_task_237.gd]): bound-poll on state == COMPLETED (via GUT wait_frames) before report_complete, so the thread writeback has already executed and the explicit report becomes the last writer. Includes the exact GDScript, a timeout so a wedged task fails fast, plus an optional manager-side authoritative-writeback variant (out of scope, defense in depth).
Verification:
- A runnable Monte-Carlo model of the scheduling races (Godot call_deferred/frame semantics) shows: unfixed flakes 11.9% idle / ~93% contended; both fixes: 0 flakes, 0 timeouts across 150k+ runs, 5 seeds — stable.
- Real-repo GUT commands for repro (pre-fix under CPU load) and proof (100 sequential + concurrent-judge loop runs, 0 failures expected).
Honest caveat: the repo (wojons/escalation-doctrine) and Godot aren't available in this workspace (no network either), so the GDScript uses the API surface described in the task (task.state/State.COMPLETED, task.result.data, report_complete(id, data), call_deferred writeback). §7 of the doc lists the naming seams to confirm against the actual manager before applying; the fix's mechanism is independent of those names. The wheel in §5.1 is the strongest verification I could actually execute here, and it's included runnable so you can rerun it yourself.
# Evidence - Problem class: godot-task-result-lost-update - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-08-22T11:04:22.959Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "GUT test test_report_complete in test_foundation_task_237.gd flaked under heavy CPU contention (concurrent gitreins Tier 2 judge): task.result.data asserted as null instead of 'completed data'. Root cause (source-verified): the callable is trivial and completes instantly; create_task auto-starts it synchronously, then the test calls task_manager.report_complete(task_id, data) - but the worker thread's own completion write can land AFTER report_complete and overwrite result.data with its own dict (data = null): a lost-update race between an explicit report and the thread's completion writeback. Low probability (1 in 8 isolated runs only while a concurrent judge starved the CPU; 5/5 clean after). Fix direction (not yet applied): await the task's completion (bounded poll on state == COMPLETED) before calling report_complete, or make report_complete the authoritative last write (no-op the thread completion writeback when an explicit report already landed).", "environment": "Godot 4.3 headless GUT 9.4.0 suite, FoundationTask237 async task manager with worker threads reporting via call_deferred", "language": "gdscript", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "godot-task-result-lost-update", "provider": "openrouter", "solved_at": "2026-08-22T11:04:22.959Z", "version": "wojons/escalation-doctrine @ 64cd665"}Done. The solution is in ~/FOUNDATION_TASK_237_LOST_UPDATE_FIX.md, with the race model at ~/sim_lost_update.py.
Root cause (lost update): create_task() auto-starts the trivial callable; the worker's deferred completion_writeback (which sets state = COMPLETED and writes the thread's own payload, data = null) can execute after the test's synchronous report_complete(task.id, "completed data") on a later idle frame, clobbering task.result.data. Two writers, last-write-wins, order decided by CPU-contention-sensitive thread/frame scheduling → the 1-in-8 flake under the concurrent judge.
The fix (test-only, matching fix_files: [src/tests/test_foundation_task_237.gd]): bound-poll on state == COMPLETED (via GUT wait_frames) before report_complete, so the thread writeback has already executed and the explicit report becomes the last writer. Includes the exact GDScript, a timeout so a wedged task fails fast, plus an optional manager-side authoritative-writeback variant (out of scope, defense in depth).
Verification:
- A runnable Monte-Carlo model of the scheduling races (Godot call_deferred/frame semantics) shows: unfixed flakes 11.9% idle / ~93% contended; both fixes: 0 flakes, 0 timeouts across 150k+ runs, 5 seeds — stable.
- Real-repo GUT commands for repro (pre-fix under CPU load) and proof (100 sequential + concurrent-judge loop runs, 0 failures expected).
Honest caveat: the repo (wojons/escalation-doctrine) and Godot aren't available in this workspace (no network either), so the GDScript uses the API surface described in the task (task.state/State.COMPLETED, task.result.data, report_complete(id, data), call_deferred writeback). §7 of the doc lists the naming seams to confirm against the actual manager before applying; the fix's mechanism is independent of those names. The wheel in §5.1 is the strongest verification I could actually execute here, and it's included runnable so you can rerun it yourself.
# Evidence - Problem class: godot-task-result-lost-update - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-08-22T11:04:22.959Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "GUT test test_report_complete in test_foundation_task_237.gd flaked under heavy CPU contention (concurrent gitreins Tier 2 judge): task.result.data asserted as null instead of 'completed data'. Root cause (source-verified): the callable is trivial and completes instantly; create_task auto-starts it synchronously, then the test calls task_manager.report_complete(task_id, data) - but the worker thread's own completion write can land AFTER report_complete and overwrite result.data with its own dict (data = null): a lost-update race between an explicit report and the thread's completion writeback. Low probability (1 in 8 isolated runs only while a concurrent judge starved the CPU; 5/5 clean after). Fix direction (not yet applied): await the task's completion (bounded poll on state == COMPLETED) before calling report_complete, or make report_complete the authoritative last write (no-op the thread completion writeback when an explicit report already landed).", "environment": "Godot 4.3 headless GUT 9.4.0 suite, FoundationTask237 async task manager with worker threads reporting via call_deferred", "language": "gdscript", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "godot-task-result-lost-update", "provider": "openrouter", "solved_at": "2026-08-22T11:04:22.959Z", "version": "wojons/escalation-doctrine @ 64cd665"}