policy-script-clobber-guard
I reconstructed the scenario in /tmp/polysol and ran the actual guard code before documenting it, so the figures are from real execution, not asserted.
~/fleet-cooldown-policy-clobber-guard-solution.md
The sync copy loop's only test is live != repo-copy → cp. That can't distinguish a forward deploy from a stale vendored repo copy reverting a deliberate operator pin. Compounding it: the deploy-hash guard existed only as an orphaned test suite — the guard functions were eaten by the clobber before commit, so "protection" was assumed from a test file that exercised nothing. A dormant ns_admission bug then surfaced when the battery finally ran.
SCRIPT_PATH/CANONICAL_HASH_PATH derived from os.path.realpath(__file__) (survives HOME override); verify_deploy_hash() → (ok, status, detail) with BOOTSTRAPPED/OK/MISMATCH; loud mismatch + --update-canonical hint, never auto-bump; --verify/--apply exit 1 on mismatch, --apply refuses before any fleet write, --dry-run tolerant.guard_clobber() skips only when live exists, is not a symlink, and its sha256 matches the sidecar while repo differs: loud SKIP CLOBBER [ticket], exit non-zero. Live already diverged → overwrite proceeds and says so.Plus the read-site fix: ns_map = {} before if namespaces: in write_fleet_pins.
namespaces=None fixed.One trap worth flagging (included in the doc): tampering must edit content inside the deployed script — replacing it with a stub removes the guard being tested, which produced a false failure in my first harness iteration.
# Evidence - Problem class: policy-script-clobber-guard - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-23T08:33:00.266Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a live operator script (~/.hermes/scripts/fleet-cooldown-policy.py) was silently overwritten twice (09-21, 09-22) by a sync script whose copy loop overwrites whenever live != repo copy \u2014 a stale vendored repo copy reverts operator pins fleet-wide, silently. A previous fix restored the copies but left the mechanism armed. ROOT CAUSE: the copy loop has no staleness check; also the guard itself had been designed once (a deploy-hash sidecar + test suite existed at ~/.hermes/scripts/test_fleet_cooldown_policy_deploy_hash.py) but the guard functions were eaten by the clobber before ever being committed \u2014 the test suite was orphaned, so the guard 'existing' on disk misled the next agent into thinking protection was live. FIX (three layers, all cheap): (1) IN THE SCRIPT: constants SCRIPT_PATH/CANONICAL_HASH_PATH (sidecar .fleet-cooldown-policy.canonical.sha256 next to the live script), verify_deploy_hash() returning (ok, status, detail) with statuses BOOTSTRAPPED/OK/MISMATCH (print LOUD 'DEPLOY HASH MISMATCH' + '--update-canonical' hint, never auto-bump), write_canonical_hash(); wire into main() so --verify and --apply exit 1 on mismatch (--apply BEFORE any fleet write) while --dry-run stays tolerant. Gotcha: the sidecar path must survive a HOME override (tests run with HOME=temp) \u2014 derive from os.path.realpath(__file__), not expanduser('~'). (2) IN THE SYNC SCRIPT: guard the copy loop \u2014 when the live copy exists, is not a symlink, and its sha256 matches the sidecar, do NOT overwrite with a differing repo copy: print a loud SKIP naming the ticket and exit non-zero; when live diverges from the sidecar, proceed (live was already non-canonical) and say so. (3) IN THE WATCHDOG: wire the deploy-hash test suite into the daily regression battery so a mismatch is ALERTED, not merely refused at next use. LESSONS: (a) the orphaned-test-suite trap \u2014 check that guard code actually exists in the live file before assuming an existing test suite proves protection; (b) a sidecar hash goes stale whenever the script legitimately changes \u2014 the workflow is edit -> deploy -> --update-canonical (sidecar must always match the DEPLOYED file); (c) when the guard lands, an unrelated pre-existing bug may surface via the watchdog suite (here: ns_admission UnboundLocalError when write_fleet_pins is called with namespaces=None \u2014 dict built inside `if namespaces:` but read unconditionally; fix = default {} at the read site). VERIFICATION: sandboxed HOME + fake repo dir, three cases: live-canonical + stale-repo-copy -> SKIP + exit 1, live preserved; live==repo==canonical -> OK exit 0; live tampered -> overwrite proceeds and restores canonical. Both suites green (8/8 + 6/6).", "environment": "Linux host, fleet agent scripts dir ~/.hermes/scripts; live installs wired from a git repo scripts/ by a copy loop (some symlinks, some byte-identical copies per cron realpath guard)", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "policy-script-clobber-guard", "provider": "openrouter", "solved_at": "2026-09-23T08:33:00.266Z", "version": ""}I reconstructed the scenario in /tmp/polysol and ran the actual guard code before documenting it, so the figures are from real execution, not asserted.
~/fleet-cooldown-policy-clobber-guard-solution.md
The sync copy loop's only test is live != repo-copy → cp. That can't distinguish a forward deploy from a stale vendored repo copy reverting a deliberate operator pin. Compounding it: the deploy-hash guard existed only as an orphaned test suite — the guard functions were eaten by the clobber before commit, so "protection" was assumed from a test file that exercised nothing. A dormant ns_admission bug then surfaced when the battery finally ran.
SCRIPT_PATH/CANONICAL_HASH_PATH derived from os.path.realpath(__file__) (survives HOME override); verify_deploy_hash() → (ok, status, detail) with BOOTSTRAPPED/OK/MISMATCH; loud mismatch + --update-canonical hint, never auto-bump; --verify/--apply exit 1 on mismatch, --apply refuses before any fleet write, --dry-run tolerant.guard_clobber() skips only when live exists, is not a symlink, and its sha256 matches the sidecar while repo differs: loud SKIP CLOBBER [ticket], exit non-zero. Live already diverged → overwrite proceeds and says so.Plus the read-site fix: ns_map = {} before if namespaces: in write_fleet_pins.
namespaces=None fixed.One trap worth flagging (included in the doc): tampering must edit content inside the deployed script — replacing it with a stub removes the guard being tested, which produced a false failure in my first harness iteration.
# Evidence - Problem class: policy-script-clobber-guard - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-23T08:33:00.266Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a live operator script (~/.hermes/scripts/fleet-cooldown-policy.py) was silently overwritten twice (09-21, 09-22) by a sync script whose copy loop overwrites whenever live != repo copy \u2014 a stale vendored repo copy reverts operator pins fleet-wide, silently. A previous fix restored the copies but left the mechanism armed. ROOT CAUSE: the copy loop has no staleness check; also the guard itself had been designed once (a deploy-hash sidecar + test suite existed at ~/.hermes/scripts/test_fleet_cooldown_policy_deploy_hash.py) but the guard functions were eaten by the clobber before ever being committed \u2014 the test suite was orphaned, so the guard 'existing' on disk misled the next agent into thinking protection was live. FIX (three layers, all cheap): (1) IN THE SCRIPT: constants SCRIPT_PATH/CANONICAL_HASH_PATH (sidecar .fleet-cooldown-policy.canonical.sha256 next to the live script), verify_deploy_hash() returning (ok, status, detail) with statuses BOOTSTRAPPED/OK/MISMATCH (print LOUD 'DEPLOY HASH MISMATCH' + '--update-canonical' hint, never auto-bump), write_canonical_hash(); wire into main() so --verify and --apply exit 1 on mismatch (--apply BEFORE any fleet write) while --dry-run stays tolerant. Gotcha: the sidecar path must survive a HOME override (tests run with HOME=temp) \u2014 derive from os.path.realpath(__file__), not expanduser('~'). (2) IN THE SYNC SCRIPT: guard the copy loop \u2014 when the live copy exists, is not a symlink, and its sha256 matches the sidecar, do NOT overwrite with a differing repo copy: print a loud SKIP naming the ticket and exit non-zero; when live diverges from the sidecar, proceed (live was already non-canonical) and say so. (3) IN THE WATCHDOG: wire the deploy-hash test suite into the daily regression battery so a mismatch is ALERTED, not merely refused at next use. LESSONS: (a) the orphaned-test-suite trap \u2014 check that guard code actually exists in the live file before assuming an existing test suite proves protection; (b) a sidecar hash goes stale whenever the script legitimately changes \u2014 the workflow is edit -> deploy -> --update-canonical (sidecar must always match the DEPLOYED file); (c) when the guard lands, an unrelated pre-existing bug may surface via the watchdog suite (here: ns_admission UnboundLocalError when write_fleet_pins is called with namespaces=None \u2014 dict built inside `if namespaces:` but read unconditionally; fix = default {} at the read site). VERIFICATION: sandboxed HOME + fake repo dir, three cases: live-canonical + stale-repo-copy -> SKIP + exit 1, live preserved; live==repo==canonical -> OK exit 0; live tampered -> overwrite proceeds and restores canonical. Both suites green (8/8 + 6/6).", "environment": "Linux host, fleet agent scripts dir ~/.hermes/scripts; live installs wired from a git repo scripts/ by a copy loop (some symlinks, some byte-identical copies per cron realpath guard)", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "policy-script-clobber-guard", "provider": "openrouter", "solved_at": "2026-09-23T08:33:00.266Z", "version": ""}