Problem class: eduos-codex-lane-429-usage-limit-tick-data-point
I diagnosed the Tick-707 failure, grounded it in the real Codex provider source, and verified the fix against those actual regexes. Deliverables written:
/workspace/solution.md — the full solution/workspace/verify-429-classification.mjs — executable verification (exit 0 = verified)Verification result: all 8 checks PASS, exit=0.
usage_limit_reached misclassified as a transient rate limitProblem class: eduos-codex-lane-429-usage-limit-tick-data-point
Data point: Tick 707 — gpt-5.6-sol/openai-codex failed HTTP 429 "The usage limit has been reached" after 65 s of zero work and 3 internal retries. 10th consecutive failed eduos tick. Bane-approved glm-5.3-flash/openrouter-eduos fallback completed the full exec task (~47 min, self-commit + self-push, zero remote-sync stops). No codex-lane recovery since t701.
The failure is not a transient rate limit and not an infrastructure outage. It is a hard ChatGPT plan quota exhaustion (usage_limit_reached), and the lane's retry/routing logic classifies it as retryable.
Three facts pin this down:
glm-5.3-flash/openrouter-eduos ran the identical exec task to completion, including self-commit and self-push. Account auth, network, disk, and the scheduler are all healthy. Only the codex lane is blocked.resets_at). Re-issuing the same request before that timestamp always returns 429. "No lane recovery since t701" is therefore the expected outcome while the quota window is still closed.usage_limit_reached, usage_not_included, or the human string "The usage limit has been reached". So the request is retried maxRetries times, fails every time, and the tick is burned./tmp/pi/packages/ai/src/api/openai-codex-responses.ts:
// line 116 — terminal quota detector (BUG: no usage_limit_reached / no human wording)
function isTerminalRateLimitError(errorText: string): boolean {
return /GoUsageLimitError|FreeUsageLimitError|Monthly usage limit reached|available balance|insufficient_quota|out of budget|quota exceeded|billing/i.test(
errorText,
);
}
// line 122 — anything else with status 429 is retried
function isRetryableError(status: number, errorText: string): boolean {
if (status === 429 && isTerminalRateLimitError(errorText)) return false;
if (status === 429 || status === 500 || status === 502 || status === 503 || status === 504) return true;
return /rate.?limit|overloaded|service.?unavailable|upstream.?connect|connection.?refused/i.test(errorText);
}
The retry loop at line 416 only stops retrying once isRetryableError returns false; the friendly "usage limit" message is produced after the retries are spent, at line ~1564. Result: 1 initial request + 3 retries = the observed 3 internal retries / 65 s zero work.
The outer classifier in /tmp/pi/packages/ai/src/utils/retry.ts (NON_RETRYABLE_PROVIDER_LIMIT_ERROR_PATTERN) has the same omission, so any wrapper that re-surfaces the error as "429 The usage limit has been reached" will be retried again at the agent layer (latent double-retry).
Reproduction (verified against the real regexes):
$ node /workspace/verify-429-classification.mjs
== Reproduce (current source) ==
PASS current terminal regex misses usage_limit_reached
PASS current provider isRetryable(429, body) -> true (3 wasted internal retries)
PASS current outer isRetryable('429 ... usage limit ...') -> true (latent re-retry)
== Proposed fix ==
PASS fixed terminal regex catches usage_limit_reached
PASS fixed provider isRetryable(429, body) -> false (fail fast, zero retries)
PASS fixed outer isRetryable(status-prefixed) -> false
PASS fixed regex does NOT swallow transient rate_limit_exceeded
PASS fixed regex does NOT swallow transient 429 status-prefixed rate limit
VERIFIED: fix yields correct classification
The lane policy says "retry codex lane each tick." Because the terminal-429 classification is missing, there is no circuit breaker keyed on resets_at; every tick unconditionally re-dispatches codex, burns the 3 retries, and only then falls back. The correct behavior is: classify once, open the lane until resets_at + grace, and dispatch the approved fallback immediately.
Apply all three layers. Layer A stops the in-request retry storm; Layer B stops the per-tick retry storm; Layer C gives an immediate, code-free mitigation.
Edit packages/ai/src/api/openai-codex-responses.ts:
function isTerminalRateLimitError(errorText: string): boolean {
- return /GoUsageLimitError|FreeUsageLimitError|Monthly usage limit reached|available balance|insufficient_quota|out of budget|quota exceeded|billing/i.test(
+ return /GoUsageLimitError|FreeUsageLimitError|Monthly usage limit reached|usage[ _-]?limit[ _-]?reached|usage_not_included|available balance|insufficient_quota|out of budget|quota exceeded|billing/i.test(
errorText,
);
}
Edit packages/ai/src/utils/retry.ts and add the same tokens to NON_RETRYABLE_PROVIDER_LIMIT_ERROR_PATTERN:
// Generic quota/budget/billing exhaustion. `insufficient_quota` is OpenAI's
// quota/billing error code; the other strings cover common gateway wording.
+ "usage_limit_reached",
+ "usage_not_included",
+ "usage limit has been reached",
"insufficient_quota",
"out of budget",
"quota exceeded",
"billing",
Also make the throw path preserve the reset hint so the lane scheduler can parse it. In parseErrorResponse the reset is already computed; include it in the thrown message (it already is, via friendlyMessage: Try again in ~N min.). For machine use, prefer emitting the structured code/reset instead of only the friendly string:
- friendlyMessage = `You have hit your ChatGPT usage limit${plan}.${when}`.trim();
+ friendlyMessage = `You have hit your ChatGPT usage limit${plan}.${when}`.trim();
+ message = `${message} [code=${code}${err.resets_at ? ` resets_at=${err.resets_at}` : ""}]`;
resets_atRead the reset from the provider error (or Retry-After / retry-after-ms header) and open the lane until then plus a small grace period. This is the part that stops the per-tick waste. Drop-in state file + helper:
# /etc/eduos/lane-circuit.sh (sourced by the scheduler tick wrapper)
# lane_open_until <lane> <epoch> [reason]
lane_open_until() {
local lane="$1" until="$2" reason="${3:-quota}"
local dir="${EDUOS_CIRCUIT_DIR:-/run/eduos/lane-circuit}"
mkdir -p "$dir"
printf '{"lane":"%s","open_until":%s,"reason":"%s","ts":%s}\n' \
"$lane" "$until" "$reason" "$(date +%s)" > "$dir/$lane.json"
}
# lane_is_open <lane> -> 0 if the circuit is open (skip / use fallback)
lane_is_open() {
local f="${EDUOS_CIRCUIT_DIR:-/run/eduos/lane-circuit}/$1.json"
[ -f "$f" ] || return 1
local until; until=$(sed -n 's/.*"open_until":\([0-9]*\).*/\1/p' "$f")
[ -n "$until" ] && [ "$(date +%s)" -lt "$until" ]
}
# On a terminal 429 from the codex lane:
# resets_at = JSON error.error.resets_at (seconds since epoch), else
# now + Retry-After, else now + 3600.
# lane_open_until openai-codex "$((resets_at + 120))" usage_limit_reached
Scheduler dispatch rule:
if lane_is_open openai-codex; then
dispatch glm-5.3-flash/openrouter-eduos # approved fallback (Bane)
else
dispatch gpt-5.6-sol/openai-codex
fi
At open_until + 1, issue one cheap probe (GET /models or a 1-token request) instead of a full tick. Only clear the circuit if the probe returns non-429; otherwise extend open_until from the fresh resets_at. Never count a quota probe as a full retry.
# 1) Parse the most recent terminal 429 from the gateway log and extract resets_at.
resets_at=$(journalctl -u hermes-gateway --since "-15min" -o cat \
| grep -o '"code":"usage_limit_reached"[^}]*' | tail -1 \
| grep -o '"resets_at":[0-9]*' | cut -d: -f2)
# 2) Open the codex circuit until 2 minutes after the reset (default +1h if absent).
until=$(( ${resets_at:-0} + 120 ))
[ "$until" -le "$(date +%s)" ] && until=$(( $(date +%s) + 3600 ))
mkdir -p /run/eduos/lane-circuit
printf '{"lane":"openai-codex","open_until":%s,"reason":"usage_limit_reached","ts":%s}\n' \
"$until" "$(date +%s)" > /run/eduos/lane-circuit/openai-codex.json
# 3) Keep the fallback as the default for the codex lane until the circuit closes.
eduos-hermes 'config set lane.openai-codex.fallback glm-5.3-flash/openrouter-eduos'
eduos-hermes 'gateway reload'
If the relay is the
SCHEDULER_CIRCUIT_CMDrouter (~/.hermes/scripts/router_circuit.pyper/etc/coding-hermes/gateway.env), write the sameopen_untilrecord through that command so the scheduler and the router agree on state.
node /workspace/verify-429-classification.mjs; echo "exit=$?"
# Expected: all PASS, "VERIFIED: fix yields correct classification", exit=0
This reads the real isTerminalRateLimitError and NON_RETRYABLE_PROVIDER_LIMIT_ERROR_PATTERN from disk, shows the bug reproduces on the exact Tick-707 body, and shows the patched pattern classifies it terminal while leaving rate_limit_exceeded retryable.
Add to packages/ai/test/openai-codex-stream.test.ts (mirrors the existing rate_limit_exceeded retry test at line ~2417). It must assert one fetch, not maxRetries + 1:
it("fails fast on ChatGPT usage_limit_reached (no internal retries)", async () => {
let codexRequests = 0;
const fetchMock = vi.fn(async () => {
codexRequests++;
return new Response(
JSON.stringify({
error: {
code: "usage_limit_reached",
message: "The usage limit has been reached",
plan_type: "plus",
resets_at: Math.floor(Date.now() / 1000) + 3600,
},
}),
{ status: 429, headers: { "content-type": "application/json" } },
);
});
vi.stubGlobal("fetch", fetchMock);
await expect(
streamOpenAICodexResponses(model, context, { apiKey: mockToken(), transport: "sse", maxRetries: 3 }).result(),
).rejects.toThrow(/usage limit/i);
expect(codexRequests).toBe(1); // was maxRetries+1 = 4 before the fix
});
Run:
cd /tmp/pi/packages/ai && npx vitest run test/openai-codex-stream.test.ts -t "usage_limit_reached"
3 internal retries line.open_until equals resets_at + grace:
bash
cat /run/eduos/lane-circuit/openai-codex.jsonglm-5.3-flash/openrouter-eduos and completes the exec task (same acceptance as the t707 fallback run).open_until + 1, confirm exactly one probe is issued, not a full retry storm. If the probe still 429s, open_until extends by the new resets_at.rm -f /run/eduos/lane-circuit/openai-codex.json
eduos-hermes 'config unset lane.openai-codex.fallback'
eduos-hermes 'gateway reload'
usage_limit_reached / "The usage limit has been reached" is a terminal ChatGPT plan-quota 429, but the codex provider's isTerminalRateLimitError regex omits those tokens, so isRetryableError(429, …) retries 3×. With no resets_at-aware circuit breaker, the per-tick policy repeats the waste every tick → 10 consecutive failures.resets_at + grace and route to the approved glm-5.3-flash/openrouter-eduos fallback; (C) apply the circuit immediately via the state file + router command./workspace/verify-429-classification.mjs proves the misclassification and the corrected classification against the real source; the unit test asserts 1 request instead of 4; the E2E check confirms one probe at reset and fallback execution meanwhile.# Evidence - Problem class: eduos-codex-lane-429-usage-limit-tick-data-point - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-17T10:55:34.905Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Tick 707 data point for the sol==luna standing 429 class: codex lane gpt-5.6-sol/openai-codex failed HTTP 429 The usage limit has been reached after 65s zero work (3 internal retries), 10th consecutive eduos tick (9th at t705 per pointer). Bane-approved glm-5.3-flash/openrouter-eduos fallback completed the full exec task in ~47min including self-commit and self-push with zero remote-sync stops. No lane recovery evidence since t701; retry codex lane each tick per policy.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "eduos-codex-lane-429-usage-limit-tick-data-point", "provider": "openrouter", "solved_at": "2026-09-17T10:55:34.906Z", "version": ""}I diagnosed the Tick-707 failure, grounded it in the real Codex provider source, and verified the fix against those actual regexes. Deliverables written:
/workspace/solution.md — the full solution/workspace/verify-429-classification.mjs — executable verification (exit 0 = verified)Verification result: all 8 checks PASS, exit=0.
usage_limit_reached misclassified as a transient rate limitProblem class: eduos-codex-lane-429-usage-limit-tick-data-point
Data point: Tick 707 — gpt-5.6-sol/openai-codex failed HTTP 429 "The usage limit has been reached" after 65 s of zero work and 3 internal retries. 10th consecutive failed eduos tick. Bane-approved glm-5.3-flash/openrouter-eduos fallback completed the full exec task (~47 min, self-commit + self-push, zero remote-sync stops). No codex-lane recovery since t701.
The failure is not a transient rate limit and not an infrastructure outage. It is a hard ChatGPT plan quota exhaustion (usage_limit_reached), and the lane's retry/routing logic classifies it as retryable.
Three facts pin this down:
glm-5.3-flash/openrouter-eduos ran the identical exec task to completion, including self-commit and self-push. Account auth, network, disk, and the scheduler are all healthy. Only the codex lane is blocked.resets_at). Re-issuing the same request before that timestamp always returns 429. "No lane recovery since t701" is therefore the expected outcome while the quota window is still closed.usage_limit_reached, usage_not_included, or the human string "The usage limit has been reached". So the request is retried maxRetries times, fails every time, and the tick is burned./tmp/pi/packages/ai/src/api/openai-codex-responses.ts:
// line 116 — terminal quota detector (BUG: no usage_limit_reached / no human wording)
function isTerminalRateLimitError(errorText: string): boolean {
return /GoUsageLimitError|FreeUsageLimitError|Monthly usage limit reached|available balance|insufficient_quota|out of budget|quota exceeded|billing/i.test(
errorText,
);
}
// line 122 — anything else with status 429 is retried
function isRetryableError(status: number, errorText: string): boolean {
if (status === 429 && isTerminalRateLimitError(errorText)) return false;
if (status === 429 || status === 500 || status === 502 || status === 503 || status === 504) return true;
return /rate.?limit|overloaded|service.?unavailable|upstream.?connect|connection.?refused/i.test(errorText);
}
The retry loop at line 416 only stops retrying once isRetryableError returns false; the friendly "usage limit" message is produced after the retries are spent, at line ~1564. Result: 1 initial request + 3 retries = the observed 3 internal retries / 65 s zero work.
The outer classifier in /tmp/pi/packages/ai/src/utils/retry.ts (NON_RETRYABLE_PROVIDER_LIMIT_ERROR_PATTERN) has the same omission, so any wrapper that re-surfaces the error as "429 The usage limit has been reached" will be retried again at the agent layer (latent double-retry).
Reproduction (verified against the real regexes):
$ node /workspace/verify-429-classification.mjs
== Reproduce (current source) ==
PASS current terminal regex misses usage_limit_reached
PASS current provider isRetryable(429, body) -> true (3 wasted internal retries)
PASS current outer isRetryable('429 ... usage limit ...') -> true (latent re-retry)
== Proposed fix ==
PASS fixed terminal regex catches usage_limit_reached
PASS fixed provider isRetryable(429, body) -> false (fail fast, zero retries)
PASS fixed outer isRetryable(status-prefixed) -> false
PASS fixed regex does NOT swallow transient rate_limit_exceeded
PASS fixed regex does NOT swallow transient 429 status-prefixed rate limit
VERIFIED: fix yields correct classification
The lane policy says "retry codex lane each tick." Because the terminal-429 classification is missing, there is no circuit breaker keyed on resets_at; every tick unconditionally re-dispatches codex, burns the 3 retries, and only then falls back. The correct behavior is: classify once, open the lane until resets_at + grace, and dispatch the approved fallback immediately.
Apply all three layers. Layer A stops the in-request retry storm; Layer B stops the per-tick retry storm; Layer C gives an immediate, code-free mitigation.
Edit packages/ai/src/api/openai-codex-responses.ts:
function isTerminalRateLimitError(errorText: string): boolean {
- return /GoUsageLimitError|FreeUsageLimitError|Monthly usage limit reached|available balance|insufficient_quota|out of budget|quota exceeded|billing/i.test(
+ return /GoUsageLimitError|FreeUsageLimitError|Monthly usage limit reached|usage[ _-]?limit[ _-]?reached|usage_not_included|available balance|insufficient_quota|out of budget|quota exceeded|billing/i.test(
errorText,
);
}
Edit packages/ai/src/utils/retry.ts and add the same tokens to NON_RETRYABLE_PROVIDER_LIMIT_ERROR_PATTERN:
// Generic quota/budget/billing exhaustion. `insufficient_quota` is OpenAI's
// quota/billing error code; the other strings cover common gateway wording.
+ "usage_limit_reached",
+ "usage_not_included",
+ "usage limit has been reached",
"insufficient_quota",
"out of budget",
"quota exceeded",
"billing",
Also make the throw path preserve the reset hint so the lane scheduler can parse it. In parseErrorResponse the reset is already computed; include it in the thrown message (it already is, via friendlyMessage: Try again in ~N min.). For machine use, prefer emitting the structured code/reset instead of only the friendly string:
- friendlyMessage = `You have hit your ChatGPT usage limit${plan}.${when}`.trim();
+ friendlyMessage = `You have hit your ChatGPT usage limit${plan}.${when}`.trim();
+ message = `${message} [code=${code}${err.resets_at ? ` resets_at=${err.resets_at}` : ""}]`;
resets_atRead the reset from the provider error (or Retry-After / retry-after-ms header) and open the lane until then plus a small grace period. This is the part that stops the per-tick waste. Drop-in state file + helper:
# /etc/eduos/lane-circuit.sh (sourced by the scheduler tick wrapper)
# lane_open_until <lane> <epoch> [reason]
lane_open_until() {
local lane="$1" until="$2" reason="${3:-quota}"
local dir="${EDUOS_CIRCUIT_DIR:-/run/eduos/lane-circuit}"
mkdir -p "$dir"
printf '{"lane":"%s","open_until":%s,"reason":"%s","ts":%s}\n' \
"$lane" "$until" "$reason" "$(date +%s)" > "$dir/$lane.json"
}
# lane_is_open <lane> -> 0 if the circuit is open (skip / use fallback)
lane_is_open() {
local f="${EDUOS_CIRCUIT_DIR:-/run/eduos/lane-circuit}/$1.json"
[ -f "$f" ] || return 1
local until; until=$(sed -n 's/.*"open_until":\([0-9]*\).*/\1/p' "$f")
[ -n "$until" ] && [ "$(date +%s)" -lt "$until" ]
}
# On a terminal 429 from the codex lane:
# resets_at = JSON error.error.resets_at (seconds since epoch), else
# now + Retry-After, else now + 3600.
# lane_open_until openai-codex "$((resets_at + 120))" usage_limit_reached
Scheduler dispatch rule:
if lane_is_open openai-codex; then
dispatch glm-5.3-flash/openrouter-eduos # approved fallback (Bane)
else
dispatch gpt-5.6-sol/openai-codex
fi
At open_until + 1, issue one cheap probe (GET /models or a 1-token request) instead of a full tick. Only clear the circuit if the probe returns non-429; otherwise extend open_until from the fresh resets_at. Never count a quota probe as a full retry.
# 1) Parse the most recent terminal 429 from the gateway log and extract resets_at.
resets_at=$(journalctl -u hermes-gateway --since "-15min" -o cat \
| grep -o '"code":"usage_limit_reached"[^}]*' | tail -1 \
| grep -o '"resets_at":[0-9]*' | cut -d: -f2)
# 2) Open the codex circuit until 2 minutes after the reset (default +1h if absent).
until=$(( ${resets_at:-0} + 120 ))
[ "$until" -le "$(date +%s)" ] && until=$(( $(date +%s) + 3600 ))
mkdir -p /run/eduos/lane-circuit
printf '{"lane":"openai-codex","open_until":%s,"reason":"usage_limit_reached","ts":%s}\n' \
"$until" "$(date +%s)" > /run/eduos/lane-circuit/openai-codex.json
# 3) Keep the fallback as the default for the codex lane until the circuit closes.
eduos-hermes 'config set lane.openai-codex.fallback glm-5.3-flash/openrouter-eduos'
eduos-hermes 'gateway reload'
If the relay is the
SCHEDULER_CIRCUIT_CMDrouter (~/.hermes/scripts/router_circuit.pyper/etc/coding-hermes/gateway.env), write the sameopen_untilrecord through that command so the scheduler and the router agree on state.
node /workspace/verify-429-classification.mjs; echo "exit=$?"
# Expected: all PASS, "VERIFIED: fix yields correct classification", exit=0
This reads the real isTerminalRateLimitError and NON_RETRYABLE_PROVIDER_LIMIT_ERROR_PATTERN from disk, shows the bug reproduces on the exact Tick-707 body, and shows the patched pattern classifies it terminal while leaving rate_limit_exceeded retryable.
Add to packages/ai/test/openai-codex-stream.test.ts (mirrors the existing rate_limit_exceeded retry test at line ~2417). It must assert one fetch, not maxRetries + 1:
it("fails fast on ChatGPT usage_limit_reached (no internal retries)", async () => {
let codexRequests = 0;
const fetchMock = vi.fn(async () => {
codexRequests++;
return new Response(
JSON.stringify({
error: {
code: "usage_limit_reached",
message: "The usage limit has been reached",
plan_type: "plus",
resets_at: Math.floor(Date.now() / 1000) + 3600,
},
}),
{ status: 429, headers: { "content-type": "application/json" } },
);
});
vi.stubGlobal("fetch", fetchMock);
await expect(
streamOpenAICodexResponses(model, context, { apiKey: mockToken(), transport: "sse", maxRetries: 3 }).result(),
).rejects.toThrow(/usage limit/i);
expect(codexRequests).toBe(1); // was maxRetries+1 = 4 before the fix
});
Run:
cd /tmp/pi/packages/ai && npx vitest run test/openai-codex-stream.test.ts -t "usage_limit_reached"
3 internal retries line.open_until equals resets_at + grace:
bash
cat /run/eduos/lane-circuit/openai-codex.jsonglm-5.3-flash/openrouter-eduos and completes the exec task (same acceptance as the t707 fallback run).open_until + 1, confirm exactly one probe is issued, not a full retry storm. If the probe still 429s, open_until extends by the new resets_at.rm -f /run/eduos/lane-circuit/openai-codex.json
eduos-hermes 'config unset lane.openai-codex.fallback'
eduos-hermes 'gateway reload'
usage_limit_reached / "The usage limit has been reached" is a terminal ChatGPT plan-quota 429, but the codex provider's isTerminalRateLimitError regex omits those tokens, so isRetryableError(429, …) retries 3×. With no resets_at-aware circuit breaker, the per-tick policy repeats the waste every tick → 10 consecutive failures.resets_at + grace and route to the approved glm-5.3-flash/openrouter-eduos fallback; (C) apply the circuit immediately via the state file + router command./workspace/verify-429-classification.mjs proves the misclassification and the corrected classification against the real source; the unit test asserts 1 request instead of 4; the E2E check confirms one probe at reset and fallback execution meanwhile.# Evidence - Problem class: eduos-codex-lane-429-usage-limit-tick-data-point - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-17T10:55:34.905Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Tick 707 data point for the sol==luna standing 429 class: codex lane gpt-5.6-sol/openai-codex failed HTTP 429 The usage limit has been reached after 65s zero work (3 internal retries), 10th consecutive eduos tick (9th at t705 per pointer). Bane-approved glm-5.3-flash/openrouter-eduos fallback completed the full exec task in ~47min including self-commit and self-push with zero remote-sync stops. No lane recovery evidence since t701; retry codex lane each tick per policy.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "eduos-codex-lane-429-usage-limit-tick-data-point", "provider": "openrouter", "solved_at": "2026-09-17T10:55:34.906Z", "version": ""}