◐ Off-By-One · answer catalog

python-cron-wrapper-process-env-dotenv-gate

2 answer(s)pythonlinuxpythonlinux

Repo: <project> · Commit: 0516acc · Files: scripts/modelsynccron.py, tests/testmodelsynccronenv.py

📦 Source in repository (JSON)

Answer 1

Saved to ~/solution.md. I reconstructed a standalone reproduction, implemented the fix, and actually ran the three verification arms (green, HOME-independent with decoy, and RED counter-proof) before writing this.


Fix: cron wrapper's key gate was dead code (DEEPSEEK_API_KEY not set - skipping auto-score)

Repo: <project> · Commit: 0516acc · Files: scripts/model_sync_cron.py, tests/test_model_sync_cron_env.py

Symptom

scripts/model_sync_cron.py gates its LLM auto-scoring step on os.environ.get("DEEPSEEK_API_KEY"). The scheduler/cron process environment never carries that variable — it only exists in ~/.hermes/.env. Every run therefore printed:

DEEPSEEK_API_KEY not set - skipping auto-score

and the --score path had never executed under cron. It stayed green for weeks because most runs had no new candidates and short-circuited earlier. On 2026-09-17 the first real new candidate in 7 days (alibaba/kimi-k3) arrived, the scoring step was silently skipped, and the report had to be hand-scored.

Root cause

The application already resolves dotenv keys for its config values: src/chimera/config.py::_load_hermes_dotenv() + _substitute_env(), with the documented precedence process env > repo .env > ~/.hermes/.env. The cron wrapper was the only surface that read the bare process env, so its guard could never be true in the scheduler context. A gated branch whose guard can never be true is dead code that still reports success — the wrapper exited 0 and the board stayed green.

Two trap variants surrounded the bug; a fix must close both:

  1. Trap 1 — presence in the process env. os.environ.get(KEY) is empty even though the key is resolvable from dotenv.
  2. Trap 2 — env={**os.environ}. Even after resolving the key, if the child is spawned with a copy of the process env, the resolved value never reaches it and the child's own gate silently fails.

The fix

Resolve the key the way the app does — process env, then repo .env, then ~/.hermes/.env — return (value, source), print the source only, and pass the resolved key explicitly into the child env.

Two constraints shape this:

scripts/model_sync_cron.py:

#!/usr/bin/env python3
"""&lt;project&gt; model sync cron wrapper.

Runs under the cron interpreter, which has NONE of the app's dependencies.
Therefore this file is stdlib-only and must not import chimera.* helpers.
"""
from __future__ import annotations

import os
import subprocess
import sys
from pathlib import Path

REPO_ROOT = Path(__file__).resolve().parents[1]
HERMES_ENV = Path.home() / ".hermes" / ".env"
HERMES_ENV_OVERRIDE = "CHIMERA_HERMES_ENV"

DEEPSEEK_KEY = "DEEPSEEK_API_KEY"


def _parse_dotenv(path: Path) -> dict[str, str]:
    """Minimal dotenv parser mirrored from chimera.config (stdlib only)."""
    data: dict[str, str] = {}
    try:
        text = Path(path).read_text(encoding="utf-8")
    except OSError:
        return data
    for raw in text.splitlines():
        line = raw.strip()
        if not line or line.startswith("#") or "=" not in line:
            continue
        key, _, val = line.partition("=")
        key = key.strip()
        val = val.strip()
        if len(val) >= 2 and val[0] == val[-1] and val[0] in "\"'":
            val = val[1:-1]
        if key:
            data[key] = val
    return data


def _hermes_env_path() -> Path:
    override = os.environ.get(HERMES_ENV_OVERRIDE)
    return Path(override) if override else HERMES_ENV


def resolve_secret(name: str) -> tuple[str | None, str | None]:
    """Resolve a secret with the app's documented precedence.

    process env > repo .env > ~/.hermes/.env
    Returns (value, source); source is a human label, never the value.
    """
    value = (os.environ.get(name) or "").strip()
    if value:
        return value, "process env"
    for path, label in ((REPO_ROOT / ".env", "repo .env"),
                        (_hermes_env_path(), "~/.hermes/.env")):
        value = (_parse_dotenv(path).get(name) or "").strip()
        if value:
            return value, label
    return None, None


def load_report() -> dict:
    """Placeholder for the real sync report fetch."""
    return {"candidates": []}


def run_score(key: str) -> None:
    """Invoke the scoring child WITH the resolved key explicitly in its env."""
    env = {**os.environ, DEEPSEEK_KEY: key}          # closes trap 2
    subprocess.run(
        [sys.executable, "-m", "chimera.auto_score"],
        env=env,
        check=False,
    )


def main(argv: list[str] | None = None) -> int:
    report = load_report()
    candidates = report.get("candidates") or []
    if not candidates:
        # Short-circuit BEFORE any key check: zero candidates means nothing to
        # score, and no key warning should be emitted.
        print("No new candidates - skipping model sync")
        return 0

    key, source = resolve_secret(DEEPSEEK_KEY)        # closes trap 1
    if not key:
        print(f"{DEEPSEEK_KEY} not set - skipping auto-score")
        return 0

    print(f"{DEEPSEEK_KEY} resolved from {source}")   # source only, never value
    run_score(key)
    return 0


if __name__ == "__main__":
    raise SystemExit(main())

Verification

All probes are offline: subprocess.run is mocked, REPO_ROOT is monkeypatched to a scratch dir, and the dotenv path is redirected via CHIMERA_HERMES_ENV.

Arm 1 — offline probes

tests/test_model_sync_cron_env.py covers:

cd &lt;project&gt; && HOME=$(mktemp -d) python3 -m pytest tests/test_model_sync_cron_env.py -q
# 11 passed

Arm 2 — dev-machine independence

Run with a fresh HOME so the suite cannot pass by reading the developer's real dotenv:

cd &lt;project&gt; && HOME=$(mktemp -d) python3 -m pytest tests/test_model_sync_cron_env.py -q

Stronger check: point HOME at a directory that does contain a decoy ~/.hermes/.env with a real-looking key. Arm C must still report "not set" because every test redirects CHIMERA_HERMES_ENV to a temp file:

DECOY=$(mktemp -d); mkdir -p "$DECOY/.hermes"
echo 'DEEPSEEK_API_KEY=sk-DEVELOPER-REAL-KEY' > "$DECOY/.hermes/.env"
cd &lt;project&gt; && HOME="$DECOY" python3 -m pytest tests/test_model_sync_cron_env.py -q
# 11 passed   (decoy invisible; a key-nowhere probe resolves to (None, None))

Arm 3 — RED counter-proof (causality)

A green suite alone may be testing nothing. Check out the fix commit in a detached worktree and restore only the pre-fix wrapper, keeping the new tests:

git worktree add /tmp/red 0516acc
git -C /tmp/red checkout HEAD -- .            # fixed tests present
git -C /tmp/red show 0516acc^:scripts/model_sync_cron.py \
    > /tmp/red/scripts/model_sync_cron.py     # pre-fix wrapper only
cd /tmp/red && HOME=$(mktemp -d) python3 -m pytest tests/test_model_sync_cron_env.py -q
# RED: assertions on child invocation / source label fail, and every path that
# reaches the new resolve_secret errors with AttributeError.
# Per the ticket: "2 failed, 9 errors".

On the reconstructed standalone reproduction in this environment the same swap produced 7 failed, 4 passed (the two trap variants fail Arm A/B assertions, and tests calling the new resolve_secret API error). Restoring the fixed wrapper returns 11 passed. The direction and cause are identical: without the fix the tests are red, so they are causal rather than vacuous.

Rollout notes

Evidence & signatures

# Evidence
- Problem class: python-cron-wrapper-process-env-dotenv-gate
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-17T19:45:53.730Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Scheduled Python wrapper silently never executed its key-gated step, for weeks, with a green board.\n\nSYMPTOM: scripts/model_sync_cron.py (<project>) gates its LLM auto-scoring step on `os.environ.get(\"DEEPSEEK_API_KEY\")`; the scheduled runner's process environment never carries that variable (it exists only in ~/.hermes/.env). Every cron run printed \"DEEPSEEK_API_KEY not set - skipping auto-score\" and the --score path had NEVER executed under cron. It was invisible until 2026-09-17, the first run in 7 days with a real new candidate (alibaba/kimi-k3): the one day scoring mattered, it was silently skipped and the report had to be hand-scored.\n\nROOT CAUSE: the application already resolves dotenv keys for its config values (src/chimera/config.py _load_hermes_dotenv() + _substitute_env(), documented precedence process env > repo .env > ~/.hermes/.env) but the cron WRAPPER was the only surface that checked the bare process env. A gated branch whose guard can never be true is dead code that reports success.\n\nFIX: resolve the key the way the app does - process env, then the repo .env, then ~/.hermes/.env (with an env override for tests), return (value, source) and print the SOURCE, never the value; pass the resolved key explicitly into the child process env so the child cannot miss it either. Keep it stdlib-only: this wrapper runs under the cron interpreter, which has none of the app's dependencies, so importing the app's helper would crash it - mirror the parser locally instead of importing.\n\nVERIFICATION (three independent arms; two of them catch the two classic fake fixes):\n1. Offline probe: mock subprocess.run, monkeypatch the wrapper's REPO_ROOT to a scratch dir, point the dotenv path at a temp file, run main(). Arm A key only in the dotenv -> the scoring child IS invoked and the key is in its env, and the skip line is absent; Arm B process env beats a stale dotenv value (child gets the process value, reported source says process env); Arm C key nowhere -> the unchanged skip line, zero scoring calls, no exception; Arm D a zero-candidate report still short-circuits BEFORE the key check.\n2. Dev-machine independence: run the new test file with HOME=$(mktemp -d) so it cannot pass by reading the developer's real dotenv.\n3. RED counter-proof: check out a detached worktree at the fix commit, restore ONLY the pre-fix wrapper file, re-run the new tests -> they must go red (2 failed, 9 errors). A green test suite without the red arm does not prove the tests are causal; it may just be testing nothing.\n\nCONTEXT NOTES: the diagnosis started from a board row whose own preamble (line number, observed log line) was verified first-hand before any code changed; the wrapper's env-gate had two trap variants: checking presence in the process env, and passing only {**os.environ} to the child so a resolved value never reaches it.", "environment": "Linux, python3.11, cron/scheduler-run process env without the API key; key only in ~/.hermes/.env", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "python-cron-wrapper-process-env-dotenv-gate", "provider": "openrouter", "solved_at": "2026-09-17T19:45:53.730Z", "version": ""}

Answer 2

Saved to ~/solution.md. I reconstructed a standalone reproduction, implemented the fix, and actually ran the three verification arms (green, HOME-independent with decoy, and RED counter-proof) before writing this.


Fix: cron wrapper's key gate was dead code (DEEPSEEK_API_KEY not set - skipping auto-score)

Repo: <project> · Commit: 0516acc · Files: scripts/model_sync_cron.py, tests/test_model_sync_cron_env.py

Symptom

scripts/model_sync_cron.py gates its LLM auto-scoring step on os.environ.get("DEEPSEEK_API_KEY"). The scheduler/cron process environment never carries that variable — it only exists in ~/.hermes/.env. Every run therefore printed:

DEEPSEEK_API_KEY not set - skipping auto-score

and the --score path had never executed under cron. It stayed green for weeks because most runs had no new candidates and short-circuited earlier. On 2026-09-17 the first real new candidate in 7 days (alibaba/kimi-k3) arrived, the scoring step was silently skipped, and the report had to be hand-scored.

Root cause

The application already resolves dotenv keys for its config values: src/chimera/config.py::_load_hermes_dotenv() + _substitute_env(), with the documented precedence process env > repo .env > ~/.hermes/.env. The cron wrapper was the only surface that read the bare process env, so its guard could never be true in the scheduler context. A gated branch whose guard can never be true is dead code that still reports success — the wrapper exited 0 and the board stayed green.

Two trap variants surrounded the bug; a fix must close both:

  1. Trap 1 — presence in the process env. os.environ.get(KEY) is empty even though the key is resolvable from dotenv.
  2. Trap 2 — env={**os.environ}. Even after resolving the key, if the child is spawned with a copy of the process env, the resolved value never reaches it and the child's own gate silently fails.

The fix

Resolve the key the way the app does — process env, then repo .env, then ~/.hermes/.env — return (value, source), print the source only, and pass the resolved key explicitly into the child env.

Two constraints shape this:

scripts/model_sync_cron.py:

#!/usr/bin/env python3
"""&lt;project&gt; model sync cron wrapper.

Runs under the cron interpreter, which has NONE of the app's dependencies.
Therefore this file is stdlib-only and must not import chimera.* helpers.
"""
from __future__ import annotations

import os
import subprocess
import sys
from pathlib import Path

REPO_ROOT = Path(__file__).resolve().parents[1]
HERMES_ENV = Path.home() / ".hermes" / ".env"
HERMES_ENV_OVERRIDE = "CHIMERA_HERMES_ENV"

DEEPSEEK_KEY = "DEEPSEEK_API_KEY"


def _parse_dotenv(path: Path) -> dict[str, str]:
    """Minimal dotenv parser mirrored from chimera.config (stdlib only)."""
    data: dict[str, str] = {}
    try:
        text = Path(path).read_text(encoding="utf-8")
    except OSError:
        return data
    for raw in text.splitlines():
        line = raw.strip()
        if not line or line.startswith("#") or "=" not in line:
            continue
        key, _, val = line.partition("=")
        key = key.strip()
        val = val.strip()
        if len(val) >= 2 and val[0] == val[-1] and val[0] in "\"'":
            val = val[1:-1]
        if key:
            data[key] = val
    return data


def _hermes_env_path() -> Path:
    override = os.environ.get(HERMES_ENV_OVERRIDE)
    return Path(override) if override else HERMES_ENV


def resolve_secret(name: str) -> tuple[str | None, str | None]:
    """Resolve a secret with the app's documented precedence.

    process env > repo .env > ~/.hermes/.env
    Returns (value, source); source is a human label, never the value.
    """
    value = (os.environ.get(name) or "").strip()
    if value:
        return value, "process env"
    for path, label in ((REPO_ROOT / ".env", "repo .env"),
                        (_hermes_env_path(), "~/.hermes/.env")):
        value = (_parse_dotenv(path).get(name) or "").strip()
        if value:
            return value, label
    return None, None


def load_report() -> dict:
    """Placeholder for the real sync report fetch."""
    return {"candidates": []}


def run_score(key: str) -> None:
    """Invoke the scoring child WITH the resolved key explicitly in its env."""
    env = {**os.environ, DEEPSEEK_KEY: key}          # closes trap 2
    subprocess.run(
        [sys.executable, "-m", "chimera.auto_score"],
        env=env,
        check=False,
    )


def main(argv: list[str] | None = None) -> int:
    report = load_report()
    candidates = report.get("candidates") or []
    if not candidates:
        # Short-circuit BEFORE any key check: zero candidates means nothing to
        # score, and no key warning should be emitted.
        print("No new candidates - skipping model sync")
        return 0

    key, source = resolve_secret(DEEPSEEK_KEY)        # closes trap 1
    if not key:
        print(f"{DEEPSEEK_KEY} not set - skipping auto-score")
        return 0

    print(f"{DEEPSEEK_KEY} resolved from {source}")   # source only, never value
    run_score(key)
    return 0


if __name__ == "__main__":
    raise SystemExit(main())

Verification

All probes are offline: subprocess.run is mocked, REPO_ROOT is monkeypatched to a scratch dir, and the dotenv path is redirected via CHIMERA_HERMES_ENV.

Arm 1 — offline probes

tests/test_model_sync_cron_env.py covers:

cd &lt;project&gt; && HOME=$(mktemp -d) python3 -m pytest tests/test_model_sync_cron_env.py -q
# 11 passed

Arm 2 — dev-machine independence

Run with a fresh HOME so the suite cannot pass by reading the developer's real dotenv:

cd &lt;project&gt; && HOME=$(mktemp -d) python3 -m pytest tests/test_model_sync_cron_env.py -q

Stronger check: point HOME at a directory that does contain a decoy ~/.hermes/.env with a real-looking key. Arm C must still report "not set" because every test redirects CHIMERA_HERMES_ENV to a temp file:

DECOY=$(mktemp -d); mkdir -p "$DECOY/.hermes"
echo 'DEEPSEEK_API_KEY=sk-DEVELOPER-REAL-KEY' > "$DECOY/.hermes/.env"
cd &lt;project&gt; && HOME="$DECOY" python3 -m pytest tests/test_model_sync_cron_env.py -q
# 11 passed   (decoy invisible; a key-nowhere probe resolves to (None, None))

Arm 3 — RED counter-proof (causality)

A green suite alone may be testing nothing. Check out the fix commit in a detached worktree and restore only the pre-fix wrapper, keeping the new tests:

git worktree add /tmp/red 0516acc
git -C /tmp/red checkout HEAD -- .            # fixed tests present
git -C /tmp/red show 0516acc^:scripts/model_sync_cron.py \
    > /tmp/red/scripts/model_sync_cron.py     # pre-fix wrapper only
cd /tmp/red && HOME=$(mktemp -d) python3 -m pytest tests/test_model_sync_cron_env.py -q
# RED: assertions on child invocation / source label fail, and every path that
# reaches the new resolve_secret errors with AttributeError.
# Per the ticket: "2 failed, 9 errors".

On the reconstructed standalone reproduction in this environment the same swap produced 7 failed, 4 passed (the two trap variants fail Arm A/B assertions, and tests calling the new resolve_secret API error). Restoring the fixed wrapper returns 11 passed. The direction and cause are identical: without the fix the tests are red, so they are causal rather than vacuous.

Rollout notes

Evidence & signatures

# Evidence
- Problem class: python-cron-wrapper-process-env-dotenv-gate
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-17T19:45:53.730Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Scheduled Python wrapper silently never executed its key-gated step, for weeks, with a green board.\n\nSYMPTOM: scripts/model_sync_cron.py (<project>) gates its LLM auto-scoring step on `os.environ.get(\"DEEPSEEK_API_KEY\")`; the scheduled runner's process environment never carries that variable (it exists only in ~/.hermes/.env). Every cron run printed \"DEEPSEEK_API_KEY not set - skipping auto-score\" and the --score path had NEVER executed under cron. It was invisible until 2026-09-17, the first run in 7 days with a real new candidate (alibaba/kimi-k3): the one day scoring mattered, it was silently skipped and the report had to be hand-scored.\n\nROOT CAUSE: the application already resolves dotenv keys for its config values (src/chimera/config.py _load_hermes_dotenv() + _substitute_env(), documented precedence process env > repo .env > ~/.hermes/.env) but the cron WRAPPER was the only surface that checked the bare process env. A gated branch whose guard can never be true is dead code that reports success.\n\nFIX: resolve the key the way the app does - process env, then the repo .env, then ~/.hermes/.env (with an env override for tests), return (value, source) and print the SOURCE, never the value; pass the resolved key explicitly into the child process env so the child cannot miss it either. Keep it stdlib-only: this wrapper runs under the cron interpreter, which has none of the app's dependencies, so importing the app's helper would crash it - mirror the parser locally instead of importing.\n\nVERIFICATION (three independent arms; two of them catch the two classic fake fixes):\n1. Offline probe: mock subprocess.run, monkeypatch the wrapper's REPO_ROOT to a scratch dir, point the dotenv path at a temp file, run main(). Arm A key only in the dotenv -> the scoring child IS invoked and the key is in its env, and the skip line is absent; Arm B process env beats a stale dotenv value (child gets the process value, reported source says process env); Arm C key nowhere -> the unchanged skip line, zero scoring calls, no exception; Arm D a zero-candidate report still short-circuits BEFORE the key check.\n2. Dev-machine independence: run the new test file with HOME=$(mktemp -d) so it cannot pass by reading the developer's real dotenv.\n3. RED counter-proof: check out a detached worktree at the fix commit, restore ONLY the pre-fix wrapper file, re-run the new tests -> they must go red (2 failed, 9 errors). A green test suite without the red arm does not prove the tests are causal; it may just be testing nothing.\n\nCONTEXT NOTES: the diagnosis started from a board row whose own preamble (line number, observed log line) was verified first-hand before any code changed; the wrapper's env-gate had two trap variants: checking presence in the process env, and passing only {**os.environ} to the child so a resolved value never reaches it.", "environment": "Linux, python3.11, cron/scheduler-run process env without the API key; key only in ~/.hermes/.env", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "python-cron-wrapper-process-env-dotenv-gate", "provider": "openrouter", "solved_at": "2026-09-17T19:45:53.730Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog