On the CI host the compose stack reported every agent ready, then 9/12 battery
response_map at webhook level) masquerading as a 404 flakeOn the CI host the compose stack reported every agent ready, then 9/12 battery
probes failed with:
{"error":"agent not found: \"sink\""}
Only the bus health probe and the two probes that self-register in the same request passed. The failure was filed as a registration-timing/runner flake because the same commit had been green before.
The example agents' POST /agents body carried a webhook-level response_map.
response_map is not a webhook field — it is a member of custom_schema
(internal/webhook/schema.go → CustomSchema.ResponseMap, "raw" or a dot path). The
webhook object is decoded with DisallowUnknownFields (hardening commit d97b777,
DF-CRIER-150), so the extra key is a hard 400 and the agent never enters the registry.
Reproduced live against a built server:
$ curl -s -w '\nHTTP %{http_code}\n' -X POST $CRIER/agents -d '{"id":"sink",...,"webhook":{"url":"...","delivery_mode":"blocking","schema_template":"generic","response_map":{"reply":"reply"}},...}'
{"error":"webhook: unknown field \"response_map\" (accepted: url, auth_type, auth_value_ref, schema_template, custom_schema, delivery_mode, batch, retries, timeout_ms)"}
HTTP 400
Why the green probes misled the report:
/ready was process-liveness, not registration. It returned 200 {"registered": false}
(or 200 while logging nothing), so ready: sink printed with an empty registry.schema_template: "generic" path resolves to generic-custom, whose
default response map is "raw" — the whole JSON body is the reply. No response_map is
needed at all; the battery assertions ("reply":" / ECHO) still match.cancel-in-progress concurrency group.All paths are under the repo root.
examples/agent-ecosystem/sink/echo_sink.py, examples/agent-ecosystem/consumer/consumer.mjs,
and the async examples/agent-ecosystem/hermes/cont-init.d/10-register-<project>.sh all lose
response_map from the webhook object:
"webhook": {
"url": "http://sink:9002/hook",
"delivery_mode": "blocking",
"schema_template": "generic"
}
If a custom map were actually wanted it must be nested:
"webhook": {"custom_schema": {"response_map": "reply"}} — but the generic default (raw)
already returns the body, so the stack needs none.
/ready registration-gatedecho_sink.py:
- register(attempts=30) prints HTTP <status> <body> on rejection, breaks immediately on a
permanent 4xx, flushes every line, and returns a bool.
- /ready does a bounded re-register (attempts=1) and answers 503 until the registry
confirms the agent.
- Registration runs in a background thread so /ready is serving (and can report 503)
immediately instead of blocking startup.
consumer.mjs:
- register() wraps fetch in try/catch, reads the body, logs status + body on failure,
and returns a bool.
- /ready calls ensureRegistered() and sends 200 or 503.
examples/agent-ecosystem/battery/battery.sh adds a visibility gate (in addition to the
existing /ready wait), polled for each agent before any round-trip:
wait_registered() { # name id — visibility gate: poll GET /agents/<id> until 200
local n="$1" id="$2" i out code body
for i in $(seq 1 60); do
out="$(curl -s -w '\n%{http_code}' "$CRIER/agents/$id")"
code="${out##*$'\n'}"; body="${out%$'\n'*}"
[[ "$code" == "200" ]] && { echo "registered: $n"; return 0; }
sleep 1
done
echo "FAIL registration visibility: $n (id=$id) not in registry after 60s — last HTTP $code body: $(echo "$body" | head -c 300)"
return 1
}
...
wait_registered "sink" "sink" || exit 1
wait_registered "pi-agent" "pi-agent" || exit 1
# ... one per harness
Removed the documented webhook-level field and documented where it really lives:
specs/AGENT-ECOSYSTEM.md §3.1/§3.2/§3.4, docs/AGENT-ECOSYSTEM.md,
examples/agent-ecosystem/README.md. §3.2 now states explicitly that response_map is a
member of custom_schema and a webhook-level copy is a hard 400 under strict decode.
Build and run <project> locally (the demo wiring disables signatures/guard):
go build -o /tmp/<project>-bin ./cmd/server
CRIER_PORT=18767 CR_REQUIRE_AGENT_SIG=false CR_GUARD_ENABLED=false /tmp/<project>-bin &
$ curl -s -o /dev/null -w 'HTTP %{http_code}\n' -X POST http://<ip-address>:18767/agents -H 'Content-Type: application/json' \
-d '{"id":"sink","public_key":"<64hex>","webhook":{"url":"http://<ip-address>:19002/hook","delivery_mode":"blocking","schema_template":"generic"},"guard":{"policies":[{"id":"default"}]}}'
HTTP 201
$ curl -s -w '\nHTTP %{http_code}\n' -X POST http://<ip-address>:18767/agents/sink/inbox -H 'Content-Type: application/json' \
-d '{"payload":{"text":"What vegetable is in plot B?"},"sender":"battery","session_id":"eco-sink","delivery_mode":"blocking","timeout_ms":15000}'
{"id":"8e9e711b50eae5be6a65d004","transport":"webhook","reply":{"reply":"ECHO: What vegetable is in plot B?"},"session_id":"eco-sink"}
HTTP 200
The generic template already extracts the reply (raw = whole body); no
response_map is needed.
$ grep -c 'path=/agents status=400' /tmp/verify/<project>.log
0
sink listening :19002
sink registered with <project>: 409 # 201 on a cold registry; 409 = already registered
/ready cannot fake registration — with an unreachable <project>:$ curl -s -m 5 -w '\nHTTP %{http_code}\n' http://<ip-address>:19003/ready
{"registered": false}
HTTP 503
$ curl -s -m 5 -w ' HTTP %{http_code}\n' http://<ip-address>:19003/health
{"status": "ok"} HTTP 200
registered: sink
FAIL registration visibility: ghost (id=ghost) not in registry after 60s — last HTTP 404 body: {"error":"agent not found: \"ghost\""}
rc=1
python3 -m py_compile echo_sink.py # OK
node --check consumer.mjs # OK
bash -n battery.sh # OK
go test ./internal/webhook/ # ok
Foreman-side proof (independent of the reporter):
docker logs <<project>-container> | grep -c 'POST /agents status=400' # => 0
Re-run the whole battery (docker compose run --rm --build battery) → 0 fail.
When an end-to-end example returns "entity not found" for every entity while each
entity's own liveness probe is green, read the server's registration response before
believing a timing flake. A strict JSON decoder plus an example payload that drifted is the
cheapest explanation. And never let a harness treat process-liveness as registration — a
fail-open readiness endpoint turns a hard 400 into N confusing 404s.
Verdict: example-side drift, not a <project> contract defect; the strict decode is by design.
Files changed: examples/agent-ecosystem/{sink/echo_sink.py, consumer/consumer.mjs,
battery/battery.sh, hermes/cont-init.d/10-register-<project>.sh}, specs/AGENT-ECOSYSTEM.md,
docs/AGENT-ECOSYSTEM.md, examples/agent-ecosystem/README.md. The working tree with all
edits is at /tmp/<project>.
# Evidence - Problem class: go-compose-example-registration-400-strict-decode-drift - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T12:32:08.704Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM (misread as a flake): a docker-compose example stack's end-to-end battery reports every agent 'ready' and then every webhook round-trip 404s {\"error\":\"agent not found: \\\"sink\\\"\"} on the CI host, while the same commit was green on a previous run, so the failure is filed as a registration-timing/runner flake. 9 of 12 probes fail; the only passes are the bus health probe and two probes that register their own agent in the same request.\n\nROOT CAUSE (measure it, do not theorize): the example agents' self-registration body carries a webhook field the server refuses. POST /agents answers HTTP 400 with the exact reason: {\"error\":\"webhook: unknown field \\\"response_map\\\" (accepted: url, auth_type, auth_value_ref, schema_template, custom_schema, delivery_mode, batch, retries, timeout_ms)\"}. response_map is NOT a webhook-level field - it is a member of custom_schema (CustomSchema.ResponseMap: \"raw\" or a dot path) - and the webhook object is decoded strictly (DisallowUnknownFields), so the extra key is a hard 400 and the agent never enters the registry. The example predates the strict decode; a later hardening commit turned its harmless extra key into a registration failure, and nothing in the stack noticed because the agent process still served its /ready endpoint and the battery treated 'process up' as 'registered'.\n\nWHY THE EVIDENCE HID IT: (1) the agent's /ready is a process-liveness endpoint, not a registration endpoint, so 'ready: sink' printed while the registry was empty; (2) the failing job ran only on a schedule (workflow_dispatch/schedule guard) and shared a cancel-in-progress concurrency group with push-triggered runs, so pushes cancelled or skipped the only execution path of the example stack; (3) the sink's registration logger was buffered and printed nothing.\n\nFIX (four layers, all in the example/docs, no contract change): (a) drop the unknown webhook-level key from both agent registration bodies; (b) print the status AND the response body of a rejected registration, and make /ready answer 503 until registered so 'up' can never masquerade as 'registered'; (c) gate the battery on registration visibility (poll GET /agents/<id> until 200, bounded, fail loudly naming the agent + last status/body) instead of round-tripping straight after the process liveness probe; (d) correct the spec/user docs that documented the removed field.\n\nVERIFICATION (what makes this answer reusable): register the corrected body against a live server -> HTTP 201; then POST the blocking inbox round-trip -> HTTP 200 {\"id\":\"...\",\"transport\":\"webhook\",\"reply\":{\"reply\":\"ECHO: ...\"},\"session_id\":\"...\"} - proving the named 'generic' template already extracts the reply and no response_map is needed; count POST /agents status=400 in the server log -> 0; sink/consumer logs -> 'registered: 201'; re-run the whole battery -> 0 fail. The foreman-side check that proves the fix rather than the reporter's word: docker logs <<project>-container> | grep -c 'POST /agents status=400'.\n\nGENERALIZED LAW: when an end-to-end example fails with 'entity not found' for EVERY entity while each entity's own liveness probe is green, read the server's registration response BEFORE believing a timing flake; a strict JSON decoder plus an example payload that drifted is the cheapest explanation. And never let a harness treat process-liveness as registration - a fail-open readiness endpoint turns a hard 400 into N confusing 404s.", "environment": "", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-compose-example-registration-400-strict-decode-drift", "provider": "openrouter", "solved_at": "2026-09-19T12:32:08.705Z", "version": ""}response_map at webhook level) masquerading as a 404 flakeOn the CI host the compose stack reported every agent ready, then 9/12 battery
probes failed with:
{"error":"agent not found: \"sink\""}
Only the bus health probe and the two probes that self-register in the same request passed. The failure was filed as a registration-timing/runner flake because the same commit had been green before.
The example agents' POST /agents body carried a webhook-level response_map.
response_map is not a webhook field — it is a member of custom_schema
(internal/webhook/schema.go → CustomSchema.ResponseMap, "raw" or a dot path). The
webhook object is decoded with DisallowUnknownFields (hardening commit d97b777,
DF-CRIER-150), so the extra key is a hard 400 and the agent never enters the registry.
Reproduced live against a built server:
$ curl -s -w '\nHTTP %{http_code}\n' -X POST $CRIER/agents -d '{"id":"sink",...,"webhook":{"url":"...","delivery_mode":"blocking","schema_template":"generic","response_map":{"reply":"reply"}},...}'
{"error":"webhook: unknown field \"response_map\" (accepted: url, auth_type, auth_value_ref, schema_template, custom_schema, delivery_mode, batch, retries, timeout_ms)"}
HTTP 400
Why the green probes misled the report:
/ready was process-liveness, not registration. It returned 200 {"registered": false}
(or 200 while logging nothing), so ready: sink printed with an empty registry.schema_template: "generic" path resolves to generic-custom, whose
default response map is "raw" — the whole JSON body is the reply. No response_map is
needed at all; the battery assertions ("reply":" / ECHO) still match.cancel-in-progress concurrency group.All paths are under the repo root.
examples/agent-ecosystem/sink/echo_sink.py, examples/agent-ecosystem/consumer/consumer.mjs,
and the async examples/agent-ecosystem/hermes/cont-init.d/10-register-<project>.sh all lose
response_map from the webhook object:
"webhook": {
"url": "http://sink:9002/hook",
"delivery_mode": "blocking",
"schema_template": "generic"
}
If a custom map were actually wanted it must be nested:
"webhook": {"custom_schema": {"response_map": "reply"}} — but the generic default (raw)
already returns the body, so the stack needs none.
/ready registration-gatedecho_sink.py:
- register(attempts=30) prints HTTP <status> <body> on rejection, breaks immediately on a
permanent 4xx, flushes every line, and returns a bool.
- /ready does a bounded re-register (attempts=1) and answers 503 until the registry
confirms the agent.
- Registration runs in a background thread so /ready is serving (and can report 503)
immediately instead of blocking startup.
consumer.mjs:
- register() wraps fetch in try/catch, reads the body, logs status + body on failure,
and returns a bool.
- /ready calls ensureRegistered() and sends 200 or 503.
examples/agent-ecosystem/battery/battery.sh adds a visibility gate (in addition to the
existing /ready wait), polled for each agent before any round-trip:
wait_registered() { # name id — visibility gate: poll GET /agents/<id> until 200
local n="$1" id="$2" i out code body
for i in $(seq 1 60); do
out="$(curl -s -w '\n%{http_code}' "$CRIER/agents/$id")"
code="${out##*$'\n'}"; body="${out%$'\n'*}"
[[ "$code" == "200" ]] && { echo "registered: $n"; return 0; }
sleep 1
done
echo "FAIL registration visibility: $n (id=$id) not in registry after 60s — last HTTP $code body: $(echo "$body" | head -c 300)"
return 1
}
...
wait_registered "sink" "sink" || exit 1
wait_registered "pi-agent" "pi-agent" || exit 1
# ... one per harness
Removed the documented webhook-level field and documented where it really lives:
specs/AGENT-ECOSYSTEM.md §3.1/§3.2/§3.4, docs/AGENT-ECOSYSTEM.md,
examples/agent-ecosystem/README.md. §3.2 now states explicitly that response_map is a
member of custom_schema and a webhook-level copy is a hard 400 under strict decode.
Build and run <project> locally (the demo wiring disables signatures/guard):
go build -o /tmp/<project>-bin ./cmd/server
CRIER_PORT=18767 CR_REQUIRE_AGENT_SIG=false CR_GUARD_ENABLED=false /tmp/<project>-bin &
$ curl -s -o /dev/null -w 'HTTP %{http_code}\n' -X POST http://<ip-address>:18767/agents -H 'Content-Type: application/json' \
-d '{"id":"sink","public_key":"<64hex>","webhook":{"url":"http://<ip-address>:19002/hook","delivery_mode":"blocking","schema_template":"generic"},"guard":{"policies":[{"id":"default"}]}}'
HTTP 201
$ curl -s -w '\nHTTP %{http_code}\n' -X POST http://<ip-address>:18767/agents/sink/inbox -H 'Content-Type: application/json' \
-d '{"payload":{"text":"What vegetable is in plot B?"},"sender":"battery","session_id":"eco-sink","delivery_mode":"blocking","timeout_ms":15000}'
{"id":"8e9e711b50eae5be6a65d004","transport":"webhook","reply":{"reply":"ECHO: What vegetable is in plot B?"},"session_id":"eco-sink"}
HTTP 200
The generic template already extracts the reply (raw = whole body); no
response_map is needed.
$ grep -c 'path=/agents status=400' /tmp/verify/<project>.log
0
sink listening :19002
sink registered with <project>: 409 # 201 on a cold registry; 409 = already registered
/ready cannot fake registration — with an unreachable <project>:$ curl -s -m 5 -w '\nHTTP %{http_code}\n' http://<ip-address>:19003/ready
{"registered": false}
HTTP 503
$ curl -s -m 5 -w ' HTTP %{http_code}\n' http://<ip-address>:19003/health
{"status": "ok"} HTTP 200
registered: sink
FAIL registration visibility: ghost (id=ghost) not in registry after 60s — last HTTP 404 body: {"error":"agent not found: \"ghost\""}
rc=1
python3 -m py_compile echo_sink.py # OK
node --check consumer.mjs # OK
bash -n battery.sh # OK
go test ./internal/webhook/ # ok
Foreman-side proof (independent of the reporter):
docker logs <<project>-container> | grep -c 'POST /agents status=400' # => 0
Re-run the whole battery (docker compose run --rm --build battery) → 0 fail.
When an end-to-end example returns "entity not found" for every entity while each
entity's own liveness probe is green, read the server's registration response before
believing a timing flake. A strict JSON decoder plus an example payload that drifted is the
cheapest explanation. And never let a harness treat process-liveness as registration — a
fail-open readiness endpoint turns a hard 400 into N confusing 404s.
Verdict: example-side drift, not a <project> contract defect; the strict decode is by design.
Files changed: examples/agent-ecosystem/{sink/echo_sink.py, consumer/consumer.mjs,
battery/battery.sh, hermes/cont-init.d/10-register-<project>.sh}, specs/AGENT-ECOSYSTEM.md,
docs/AGENT-ECOSYSTEM.md, examples/agent-ecosystem/README.md. The working tree with all
edits is at /tmp/<project>.
# Evidence - Problem class: go-compose-example-registration-400-strict-decode-drift - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T12:32:08.704Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM (misread as a flake): a docker-compose example stack's end-to-end battery reports every agent 'ready' and then every webhook round-trip 404s {\"error\":\"agent not found: \\\"sink\\\"\"} on the CI host, while the same commit was green on a previous run, so the failure is filed as a registration-timing/runner flake. 9 of 12 probes fail; the only passes are the bus health probe and two probes that register their own agent in the same request.\n\nROOT CAUSE (measure it, do not theorize): the example agents' self-registration body carries a webhook field the server refuses. POST /agents answers HTTP 400 with the exact reason: {\"error\":\"webhook: unknown field \\\"response_map\\\" (accepted: url, auth_type, auth_value_ref, schema_template, custom_schema, delivery_mode, batch, retries, timeout_ms)\"}. response_map is NOT a webhook-level field - it is a member of custom_schema (CustomSchema.ResponseMap: \"raw\" or a dot path) - and the webhook object is decoded strictly (DisallowUnknownFields), so the extra key is a hard 400 and the agent never enters the registry. The example predates the strict decode; a later hardening commit turned its harmless extra key into a registration failure, and nothing in the stack noticed because the agent process still served its /ready endpoint and the battery treated 'process up' as 'registered'.\n\nWHY THE EVIDENCE HID IT: (1) the agent's /ready is a process-liveness endpoint, not a registration endpoint, so 'ready: sink' printed while the registry was empty; (2) the failing job ran only on a schedule (workflow_dispatch/schedule guard) and shared a cancel-in-progress concurrency group with push-triggered runs, so pushes cancelled or skipped the only execution path of the example stack; (3) the sink's registration logger was buffered and printed nothing.\n\nFIX (four layers, all in the example/docs, no contract change): (a) drop the unknown webhook-level key from both agent registration bodies; (b) print the status AND the response body of a rejected registration, and make /ready answer 503 until registered so 'up' can never masquerade as 'registered'; (c) gate the battery on registration visibility (poll GET /agents/<id> until 200, bounded, fail loudly naming the agent + last status/body) instead of round-tripping straight after the process liveness probe; (d) correct the spec/user docs that documented the removed field.\n\nVERIFICATION (what makes this answer reusable): register the corrected body against a live server -> HTTP 201; then POST the blocking inbox round-trip -> HTTP 200 {\"id\":\"...\",\"transport\":\"webhook\",\"reply\":{\"reply\":\"ECHO: ...\"},\"session_id\":\"...\"} - proving the named 'generic' template already extracts the reply and no response_map is needed; count POST /agents status=400 in the server log -> 0; sink/consumer logs -> 'registered: 201'; re-run the whole battery -> 0 fail. The foreman-side check that proves the fix rather than the reporter's word: docker logs <<project>-container> | grep -c 'POST /agents status=400'.\n\nGENERALIZED LAW: when an end-to-end example fails with 'entity not found' for EVERY entity while each entity's own liveness probe is green, read the server's registration response BEFORE believing a timing flake; a strict JSON decoder plus an example payload that drifted is the cheapest explanation. And never let a harness treat process-liveness as registration - a fail-open readiness endpoint turns a hard 400 into N confusing 404s.", "environment": "", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-compose-example-registration-400-strict-decode-drift", "provider": "openrouter", "solved_at": "2026-09-19T12:32:08.705Z", "version": ""}