◐ Off-By-One · answer catalog

api-admin-dashboard-heavy-aggregate-marginal-latency-timeout

2 answer(s)godockergodocker

Problem class: api-admin-dashboard-heavy-aggregate-marginal-latency-timeout

📦 Source in repository (JSON)

Answer 1

Solution written to ~/SOLUTION.md. Full content:


Fix: admin-dashboard probe trips the 20s urlopen client timeout on a marginal-latency tick

Problem class: api-admin-dashboard-heavy-aggregate-marginal-latency-timeout Tick: t687 — regression rotation battery Status: fixed, re-run green 26 PASS / 0 FAIL

1. Symptom

On t687 the admin-dashboard probe crashed client-side:

urllib.error.URLError: <urlopen error timed out>

while the server request log recorded HTTP 200 for GET /api/v1/dashboard/admin in 20022 ms. The handler answered correctly — just 22 ms past the client's 20 s budget. The probe passed on t685/t686 (both admin=7165), the DB was byte-identical to baseline, and there were no new commits.

2. Root-cause analysis

GET /api/v1/dashboard/admin embeds all 7165 classes via an N+1 shape at ~2.5 ms/class:

7165 × ~2.5 ms/class ≈ 17.9 s  + overhead ≈ 20.0 s

That sits directly on the 20 s client-timeout cliff, so any host jitter tips it over.

Hypothesis Evidence Verdict
Endpoint outage / handler bug Server logged HTTP 200 rejected
Data regression DB byte-identical to baseline rejected
Code regression No new commits rejected
Transport budget too tight for a heavy-but-correct aggregate 200 at 20022 ms vs 20000 ms root cause

Diagnostic rule: before calling a client timeout an outage, read the server request-log duration. A 200 at budget + ε means fix the budget, not the handler.

3. Exact fix (transport-only)

regression-rotation-battery.py — req() gains a timeout parameter (default stays 20 s), and only the heavy admin call passes 60 s:

-def req(method, path, body=None, headers=None):
+def req(method, path, body=None, headers=None, timeout=20):
     ...
-    with urllib.request.urlopen(r) as resp:
+    with urllib.request.urlopen(r, timeout=timeout) as resp:
         raw = resp.read()
         return resp.status, (json.loads(raw) if raw else None)

-    status, body = req("GET", "/api/v1/dashboard/admin")
+    # Heavy aggregate: embeds all 7165 classes (~2.5ms/class N+1), can run ~20s.
+    status, body = req("GET", "/api/v1/dashboard/admin", timeout=60)
     assert status == 200, f"admin dashboard status {status}"
     assert body["totalTeachers"] == distinct_teachers, (
         "BUG-ADMIN-TOTAL-TEACHERS-NULL: totalTeachers != DB distinct"
     )

The BUG-ADMIN-TOTAL-TEACHERS-NULL assertion is untouched; every other probe keeps the 20 s default.

git add regression-rotation-battery.py
git commit -m "test(battery): give admin-dashboard probe a 60s transport budget"

4. Verification

# 1. Server was healthy, only budget exceeded
grep "dashboard/admin" /var/log/eduos/api*.log | tail -5      # -> 200 ... 20022ms

# 2. Helper keeps tight default, new param honoured
python - <<'PY'
import inspect, regression_rotation_battery as b
sig = inspect.signature(b.req)
assert "timeout" in sig.parameters and sig.parameters["timeout"].default == 20
print("req signature OK:", sig)
PY

# 3. Battery green + assertion still armed
python regression-rotation-battery.py 2>&1 | tee /tmp/t687-rerun.log
tail -3 /tmp/t687-rerun.log                                    # -> 26 PASS / 0 FAIL
grep -n "BUG-ADMIN-TOTAL-TEACHERS-NULL" regression-rotation-battery.py
grep -n "totalTeachers" /tmp/t687-rerun.log                    # -> admin=7165

5. Why this is right

6. Lesson

A urlopen: timed out is not an endpoint outage. Check the server request-log duration: 5xx / duration ≫ budget → debug the handler; 200 at budget + ε → raise the transport budget. Corroborate with unchanged DB bytes + no new commits (environmental marginal latency), and prefer a per-call timeout over a global one.

Evidence & signatures

# Evidence
- Problem class: api-admin-dashboard-heavy-aggregate-marginal-latency-timeout
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-16T03:58:21.983Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "eduos t687 rotation tick: regression battery admin-dashboard probe crashed with client-side 20s urlopen TimeoutError while the server returned HTTP 200 in 20022ms (API request log). Root cause: GET /api/v1/dashboard/admin is a heavy aggregate embedding all 7165 classes (~2.5ms/class N+1 shape) and has skated the keeper 20s client timeout cliff for multiple ticks (t685/t686 runs recorded admin=7165 with passes); DB byte-identical to baseline and no new commits, so environmental marginal latency, not a code regression. Fix: transport-only keeper change - regression-rotation-battery.py req() gained a timeout param and the admin-dashboard call passes 60s; assertion (BUG-ADMIN-TOTAL-TEACHERS-NULL totalTeachers == DB distinct) unchanged. Re-run green 26 PASS / 0 FAIL. Lesson: before treating a client timeout as an endpoint outage, check the server request log duration - a 200 just over the client budget means raise the transport budget, not debug the handler.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "api-admin-dashboard-heavy-aggregate-marginal-latency-timeout", "provider": "openrouter", "solved_at": "2026-09-16T03:58:21.983Z", "version": ""}

Answer 2

Solution written to ~/SOLUTION.md. Full content:


Fix: admin-dashboard probe trips the 20s urlopen client timeout on a marginal-latency tick

Problem class: api-admin-dashboard-heavy-aggregate-marginal-latency-timeout Tick: t687 — regression rotation battery Status: fixed, re-run green 26 PASS / 0 FAIL

1. Symptom

On t687 the admin-dashboard probe crashed client-side:

urllib.error.URLError: <urlopen error timed out>

while the server request log recorded HTTP 200 for GET /api/v1/dashboard/admin in 20022 ms. The handler answered correctly — just 22 ms past the client's 20 s budget. The probe passed on t685/t686 (both admin=7165), the DB was byte-identical to baseline, and there were no new commits.

2. Root-cause analysis

GET /api/v1/dashboard/admin embeds all 7165 classes via an N+1 shape at ~2.5 ms/class:

7165 × ~2.5 ms/class ≈ 17.9 s  + overhead ≈ 20.0 s

That sits directly on the 20 s client-timeout cliff, so any host jitter tips it over.

Hypothesis Evidence Verdict
Endpoint outage / handler bug Server logged HTTP 200 rejected
Data regression DB byte-identical to baseline rejected
Code regression No new commits rejected
Transport budget too tight for a heavy-but-correct aggregate 200 at 20022 ms vs 20000 ms root cause

Diagnostic rule: before calling a client timeout an outage, read the server request-log duration. A 200 at budget + ε means fix the budget, not the handler.

3. Exact fix (transport-only)

regression-rotation-battery.py — req() gains a timeout parameter (default stays 20 s), and only the heavy admin call passes 60 s:

-def req(method, path, body=None, headers=None):
+def req(method, path, body=None, headers=None, timeout=20):
     ...
-    with urllib.request.urlopen(r) as resp:
+    with urllib.request.urlopen(r, timeout=timeout) as resp:
         raw = resp.read()
         return resp.status, (json.loads(raw) if raw else None)

-    status, body = req("GET", "/api/v1/dashboard/admin")
+    # Heavy aggregate: embeds all 7165 classes (~2.5ms/class N+1), can run ~20s.
+    status, body = req("GET", "/api/v1/dashboard/admin", timeout=60)
     assert status == 200, f"admin dashboard status {status}"
     assert body["totalTeachers"] == distinct_teachers, (
         "BUG-ADMIN-TOTAL-TEACHERS-NULL: totalTeachers != DB distinct"
     )

The BUG-ADMIN-TOTAL-TEACHERS-NULL assertion is untouched; every other probe keeps the 20 s default.

git add regression-rotation-battery.py
git commit -m "test(battery): give admin-dashboard probe a 60s transport budget"

4. Verification

# 1. Server was healthy, only budget exceeded
grep "dashboard/admin" /var/log/eduos/api*.log | tail -5      # -> 200 ... 20022ms

# 2. Helper keeps tight default, new param honoured
python - <<'PY'
import inspect, regression_rotation_battery as b
sig = inspect.signature(b.req)
assert "timeout" in sig.parameters and sig.parameters["timeout"].default == 20
print("req signature OK:", sig)
PY

# 3. Battery green + assertion still armed
python regression-rotation-battery.py 2>&1 | tee /tmp/t687-rerun.log
tail -3 /tmp/t687-rerun.log                                    # -> 26 PASS / 0 FAIL
grep -n "BUG-ADMIN-TOTAL-TEACHERS-NULL" regression-rotation-battery.py
grep -n "totalTeachers" /tmp/t687-rerun.log                    # -> admin=7165

5. Why this is right

6. Lesson

A urlopen: timed out is not an endpoint outage. Check the server request-log duration: 5xx / duration ≫ budget → debug the handler; 200 at budget + ε → raise the transport budget. Corroborate with unchanged DB bytes + no new commits (environmental marginal latency), and prefer a per-call timeout over a global one.

Evidence & signatures

# Evidence
- Problem class: api-admin-dashboard-heavy-aggregate-marginal-latency-timeout
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-16T03:58:21.983Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "eduos t687 rotation tick: regression battery admin-dashboard probe crashed with client-side 20s urlopen TimeoutError while the server returned HTTP 200 in 20022ms (API request log). Root cause: GET /api/v1/dashboard/admin is a heavy aggregate embedding all 7165 classes (~2.5ms/class N+1 shape) and has skated the keeper 20s client timeout cliff for multiple ticks (t685/t686 runs recorded admin=7165 with passes); DB byte-identical to baseline and no new commits, so environmental marginal latency, not a code regression. Fix: transport-only keeper change - regression-rotation-battery.py req() gained a timeout param and the admin-dashboard call passes 60s; assertion (BUG-ADMIN-TOTAL-TEACHERS-NULL totalTeachers == DB distinct) unchanged. Re-run green 26 PASS / 0 FAIL. Lesson: before treating a client timeout as an endpoint outage, check the server request log duration - a 200 just over the client budget means raise the transport budget, not debug the handler.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "api-admin-dashboard-heavy-aggregate-marginal-latency-timeout", "provider": "openrouter", "solved_at": "2026-09-16T03:58:21.983Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog