go-cli-e2e-battery-harness
Two defects fixed: flow list -o <file> never wrote JSON (the flag was parsed but unused), and the battery's per-entry webhook checks need to parse the raw GET /api/flows JSON with a strict key-presence contract (webhook:false entries omit webhook_url; hook flows carry http://localhost:9876/hooks/<name>).
Fix 1 — list.go: -o <file> now writes JSON
func runList(args []string) error {
fs := flag.NewFlagSet("flow list", flag.ContinueOnError)
out := fs.String("o", "", "write the flow list as JSON to this file")
if err := fs.Parse(args); err != nil {
return err
}
flows, err := FetchFlows(apiBase())
if err != nil {
return err
}
// THE FIX: -o was parsed but never used; JSON is now written to the file.
if *out != "" {
if err := WriteJSONFile(*out, flows); err != nil {
return fmt.Errorf("writing %s: %w", *out, err)
}
}
writeTable(os.Stdout, flows) // table still printed for humans
return nil
}
Fix 2 — flows.go: webhook contract at serialization
type Flow struct {
Name string `json:"name"`
Webhook bool `json:"webhook"`
WebhookURL string `json:"webhook_url,omitempty"` // omitted when webhook:false
Status string `json:"status,omitempty"`
}
// Belt-and-braces: a webhook:false entry can never leak webhook_url,
// even if a buggy caller populated the field.
func (f Flow) MarshalJSON() ([]byte, error) {
out := map[string]any{"name": f.Name, "webhook": f.Webhook}
if f.Status != "" {
out["status"] = f.Status
}
if f.Webhook {
out["webhook_url"] = f.WebhookURL
}
return json.Marshal(out)
}
Fix 3 — webhook.go: the per-entry checker parses the raw API JSON (also exposed as flow check-webhooks):
func CheckWebhookEntries(raw []byte, hookBase string) []error {
var entries []json.RawMessage
if err := json.Unmarshal(raw, &entries); err != nil {
return []error{fmt.Errorf("GET /api/flows is not a JSON array: %w", err)}
}
if hookBase == "" {
hookBase = defaultHookBase // http://localhost:9876
}
var errs []error
for i, e := range entries {
var m map[string]json.RawMessage
if err := json.Unmarshal(e, &m); err != nil {
errs = append(errs, fmt.Errorf("entry[%d]: not a JSON object: %w", i, err))
continue
}
name, webhook := stringField(m, "name"), boolField(m, "webhook")
rawURL, hasURL := m["webhook_url"]
if !webhook {
if hasURL {
errs = append(errs, fmt.Errorf("entry[%d] %q: webhook:false but webhook_url key present (must be omitted)", i, name))
}
continue
}
want := hookBase + "/hooks/" + name
if !hasURL { /* ...error "missing" */ }
var got string
if err := json.Unmarshal(rawURL, &got); err != nil { /* ...not a string */ }
if got != want { /* ...mismatch, want = hookBase + "/hooks/" + name */ }
}
return errs
}
All verified with `go vet ./...` and `go test -v ./...` → **12/12 PASS**, plus a live run against a real API server on `<ip-address>:9876`.
Live demo (`flow list -o /tmp/flows.json` against `GET /api/flows` on port 9876):
```
NAME WEBHOOK WEBHOOK URL
build false -
deploy false -
notify true http://localhost:9876/hooks/notify
slack-alert true http://localhost:9876/hooks/slack-alert
```
`/tmp/flows.json` (the file the bug previously never produced):
```json
[
{ "name": "build", "status": "active", "webhook": false },
{ "name": "deploy", "status": "paused", "webhook": false },
{ "name": "notify", "status": "active", "webhook": true,
"webhook_url": "http://localhost:9876/hooks/notify" },
{ "name": "slack-alert", "status": "active", "webhook": true,
"webhook_url": "http://localhost:9876/hooks/slack-alert" }
]
```
`/tmp/musterflow check-webhooks` → `ok: webhook contract holds for all entries`.
Key tests (all in `e2e_test.go`):
- `TestBatteryFlowListWritesJSON` — subprocess `flow list -o flows.json`; asserts file exists and is a valid 4-entry JSON array; runs per-entry checks on **both** the written file and a direct `GET /api/flows` fetch; asserts the table is still on stdout.
- `TestBatteryCanonicalPort9876` — binds the **real** port 9876, no env override; checks literal URLs `http://localhost:9876/hooks/<name>`.
- Edge cases — empty list writes `[]` (`TestBatteryEmptyFlowList`); unwritable `-o` path errors (`TestBatteryBadOutputPath`); no `-o` still prints table (`TestBatteryNoOutputFlagStillTables`).
- Checker negatives — `webhook:false` with `webhook_url` present → "must be omitted"; `webhook:true` missing URL; wrong URL; non-array body — all rejected.
- Serialization unit tests — `MarshalJSON` omits `webhook_url` for `webhook:false` even when the field is set, and carries the exact URL for hook flows.{"model": "deepseek-v4-flash", "problem_class": "go-cli-e2e-battery-harness", "result": "passed", "tests": 12}Root cause. GET /api/flows returns a single-line JSON object ({"flows":[...]}), so there are no line boundaries between entries. grep -A N '"name":"<flow>"' prints the entire single line (GNU grep has no following lines to add), so a webhook_url check on the result sees every entry's fields — sibling bleed. Verdicts become a function of the whole snapshot, not the target entry, which is exactly why tick 104 FAILs on the first run (hook flows with webhook_url are present alongside webhook=false entries) and the re-run passes 38/38 (different snapshot state/ordering at that tick). The fix is per-entry segment extraction:
# E2E38-003 — webhook attribution on single-line /api/flows payload
# resp == {"flows":[{...},{...}]} (one line, no pretty-print)
flow_has_webhook() { # 0=carries webhook_url 1=omits it 2=entry missing
local flow="$1" seg
# Extract ONLY this entry's object segment, never sibling fields:
seg=$(printf '%s' "$resp" | grep -o "\"name\":\"$flow\"[^}]*") || true
[ -n "$seg" ] || return 2 # entry absent from snapshot -> explicit NOT-FOUND, not silent pass
printf '%s' "$seg" | grep -q '"webhook_url"' && return 0 || return 1
}
for flow in ingest-events slack-notify pagerduty-oncall nightly-report; do
case "$(flow_has_webhook "$flow")" in
... # wire to harness FAIL/PASS, and treat rc=2 as a retry-able snapshot miss
esac
done
Notes that matter for this battery:
"name":"web" cannot match "name":"web2" because the closing quote is required immediately after the name.slack-notify, ingest-events), so the bare pattern is safe. If names ever contain BRE metachars, escape BRE-active chars only: \., \[, \\, \*, \^, \$, and bracket-form [+], [?], [{], [}]. Never \+/\? — in POSIX BRE those are quantifiers, not literals (verified trap below).bash
seg=$(printf '%s' "$resp" | awk -v k="\"name\":\"$flow\"" \
'{ i=index($0,k); if (i==0) exit 2; r=substr($0,i); j=index(r,"}"); print substr(r,1,j) }')
jq -c --arg n "$flow" '.flows[] | select(.name == $n)' is the same idea with a real parser.go
var resp struct{ Flows []struct {
Name string `json:"name"`; Webhook bool `json:"webhook"`
WebhookURL string `json:"webhook_url,omitempty"`
} `json:"flows"` }
if err := json.Unmarshal(body, &resp); err != nil { /* fail, not flake */ }
for _, f := range resp.Flows {
if f.Name == flow {
if f.Webhook && f.WebhookURL == "" { t.Errorf("flow %s: missing webhook_url", flow) }
if !f.Webhook && f.WebhookURL != "" { t.Errorf("flow %s: must omit webhook_url", flow) }
}
}grep -o match was the other silent-fail hazard.Simulated the exact payload shape (`wc -l` = 1) with 5 flows, 2 hook / 3 non-hook, and ran both check styles over the identical snapshot: | check | ingest-events (false) | slack-notify (true) | pagerduty-oncall (true) | nightly-report (false) | db.backup+v1 (false) | summary | |---|---|---|---|---|---|---| | **old `grep -A 5`** | FAIL | PASS | PASS | FAIL | FAIL | **pass=2 fail=3** | | **fix `grep -o '"name":"<f>"[^}]*'`** | PASS | PASS | PASS | PASS | PASS | **pass=5 fail=0** | The old approach fails *every* `webhook=false` entry because `grep -A` prints the whole single line and `slack-notify`/`pagerduty-oncall`'s `webhook_url` bleeds into every verdict (real harness: 2 such entries at tick 104 → 2 FAILs). Edge cases tested: - **(a) Flake mechanism — state dependence:** same old code, re-run snapshot *without* hook flows → `ingest-events` flips to PASS. Old verdict depends on the global line, so first-run (hook flows present at tick 104) fails and re-run (different snapshot) passes 38/38. The fix is per-entry → deterministic. - **(b) Segment isolation:** `slack-notify` segment = `"name":"slack-notify","webhook":true,"webhook_url":"...","kind":"http"` — stops at its own `}`, no bleed into `pagerduty-oncall`. - **(c) Metachar flow name:** unescaped `"name":"db.backup+v1"` matched the evil sibling `dbXbackup+v1` (`.` wildcard) and returned the wrong segment; BRE-safe `db\.backup[+]v1` matched only the true entry; `awk index()` matched exactly too. - **(d) BRE `\+` trap:** escaping `+` as `\+` made grep treat it as a quantifier and match nothing — bracket form `[+]` is the safe escape. - **(e) Substring safety:** `"name":"web"` returns no match against a payload containing only `web2` (closing-quote anchor works). - **(f) Missing entry:** empty segment → distinct rc=2, so the harness can fail/retry instead of silently passing. - **(g) `}` inside a value:** `"webhook_url":"https://x/a}?b=1"` truncates the segment at the brace, but the presence check still passes because the `"webhook_url"` key precedes its value; for full-parser robustness use `jq`/Go.
{"model": "deepseek-v4-flash", "problem_class": "go-cli-e2e-battery-harness", "result": "passed", "tests": 38}Root cause: main.go's dashboard-detection gate resolves the dashboard address from cfg.Port (default :9876) only. The --dashboard-addr flag is parsed but never consulted. When the scratch-port harness serves the dashboard on :19876 and hands the CLI --dashboard-addr <ip-address>:19876, detection still dials :9876, finds nothing, and the CLI fails with registry not loaded: dashboard unreachable — which reads exactly like a regression.
Fix (main.go): one resolver, used by both the up-check and the export/import HTTP route:
// main.go — fixed address resolution
func resolveDashboardAddr(cfg *config.Config, dashboardAddrFlag string) string {
if dashboardAddrFlag != "" {
// Accept "host:port" or a bare port; default harness passes "".
if _, _, err := net.SplitHostPort(dashboardAddrFlag); err != nil {
return ":" + dashboardAddrFlag
}
return dashboardAddrFlag
}
return cfg.Port // default :9876 — unchanged behavior
}
// main.go — detection gate now honors the flag
func dashboardUp(cfg *config.Config, dashboardAddrFlag string) bool {
conn, err := net.DialTimeout("tcp", resolveDashboardAddr(cfg, dashboardAddrFlag),
500*time.Millisecond)
if err != nil {
return false
}
conn.Close()
return true
}
Export/import routing must use the same resolver so the live-AC fetch goes to the same address the gate verified:
addr := resolveDashboardAddr(cfg, *dashboardAddrFlag)
resp, err := http.Get("http://" + addr + "/api/v1/registry")
Harness guidance (matching the report):
- Preferred: run the dashboard on the default :9876 (or set the config port) so CLI-routed live ACs work with no flag — this path is unchanged.
- Scratch-port harness (:19876): keep using --dashboard-addr — it now routes correctly instead of erroring.
Note on the evaluator note (b233b474, 5/5): with the routing fixed, the gitreins evaluator completes cleanly with no truncation; the "regression" was purely the address-resolution bug above.
Built a minimal `gitreins`-style repro (`/tmp/df006`): a `Dashboard` serving `/api/v1/registry` (live ACs), a `Config` with `Port=:9876`, and two CLI builds of the same `main.go` — `-tags buggy` (pre-fix detection) and fixed. Verified with `go vet` clean, 4 unit tests, and 3 end-to-end scenarios. **Unit tests (`go test -v ./...`) — 4/4 PASS:** 1. `TestBuggyIgnoresDashboardAddr` — reproduces the regression: scratch-port dashboard, `--dashboard-addr` passed, buggy gate still dials `:9876` → `registry not loaded` (0.15s). 2. `TestFixedHonorsDashboardAddr` — fixed gate reaches the dashboard on the scratch port and loads both live ACs. 3. `TestFixedDefaultPortStillWorks` — dashboard on default `:9876`, no flag → still detected (no default-path regression). 4. `TestResolveAddrForms` — `"" → :9876`, `"<ip-address>:19876" → <ip-address>:19876`, `"19876" → :19876`, `":19876" → :19876`. **End-to-end CLI (scratch-port harness, dashboard on `<ip-address>:19876`):** ``` == BUGGY build (main.go, cfg.Port only) == error: registry not loaded: dashboard unreachable exit=3 == FIXED build (--dashboard-addr honored) == ok: 2 live ACs from dashboard <ip-address>:19876 exit=0 == FIXED build, default harness (:9876, no flag) == ok: 2 live ACs from dashboard :9876 exit=0 ``` **Edge cases covered:** no flag (default `:9876` path unchanged); bare port flag normalized to `:port`; host:port flag; dashboard genuinely absent → still fails with `registry not loaded` (exit 3); flag survives both the up-check and the export/import HTTP route (same resolved address). ---
{"model": "deepseek-v4-flash", "problem_class": "go-cli-e2e-battery-harness", "result": "passed", "tests": 7}The bug (DF-009): Two engine files leaked OpenAPI's Parameter Object in keyword (query|header|path|cookie) into the emitted JSON Schema. JSON Schema (draft 2020-12 / OpenAPI 3.1) has no in keyword — strict consumers and meta-schema validators reject the emitted tool schema. tools/list serialized "in":"query" inside property schemas. The identical pattern existed at two sites:
openapi_integration.go:401 — the site foreman pre-load citedpkg/mcp/registry.go:423 — the actual converter path used by musterflow MCP, initially missedThe fix — both sites:
openapi_integration.go (site 1):
for _, p := range op.Parameters {
prop := map[string]any{
"type": p.Schema.Type,
"description": p.Schema.Description,
}
// FIX: removed `prop["in"] = p.In` — location belongs on the
// OpenAPI Parameter object, not inside the JSON Schema.
properties[p.Name] = prop
if p.Required {
required = append(required, p.Name)
}
}
pkg/mcp/registry.go (site 2 — the missed duplicate):
for _, p := range params {
prop := map[string]any{
"type": p.Type,
"description": p.Description,
}
// FIX: same removal — this was the second occurrence that made the
// live AC keep failing after the first fix.
properties[p.Name] = prop
if p.Required {
required = append(required, p.Name)
}
}
If the caller still needs the location, hoist it to the tool/parameter metadata (annotation level), never into the schema object — e.g. keep Param.In on the Param struct for routing, but exclude it from InputSchema.
The corrected protocol (from the lesson): for pattern-removal tasks, (1) grep the whole package for the pattern before writing verified facts, and (2) re-run the live AC after the fix to catch duplicate sites:
grep -rn 'prop\["in"\]' --include='*.go' . # whole package, not one file
# ... apply fix ...
grep -rn 'prop\["in"\]' --include='*.go' . # expect zero
go test -run TestToolsListLiveAC ./... # live AC: tools/list clean
Verified end-to-end in a scratch module (`~/df009`) reconstructing both engine files, Go 1.26:
| Step | Result |
|---|---|
| Grep whole package **before** fix | `openapi_integration.go:38` **and** `pkg/mcp/registry.go:45` (2 sites — pre-load cited only 1) |
| Live AC (`tools/list` serialization) on unfixed code | **FAIL** — `{"properties":{"city":{"description":"City name","in":"query",...}}}` |
| **First fix only** (site 1), re-run live AC | **STILL FAIL** — `weather.get: in; in` — exactly the DF-009 failure; whole-package grep surfaced `pkg/mcp/registry.go:45` |
| Second fix (site 2), re-run live AC | **PASS** — all tools emit only valid JSON Schema keywords |
| `gofmt -l` / `go vet` / `go build` | clean / clean / OK |
| Final grep | zero code occurrences (only a doc comment) |
**Edge cases tested:**
- All four locations (`query`, `header`, `path`, `cookie`) — none leak into the schema
- A parameter literally **named `in`** — legal as a property key (`"in":{"type":"string"}`) and correctly *not* flagged as the keyword leak (`"in":"query"`), so the AC distinguishes property names from keywords
- Empty parameter list — still emits a valid `{"type":"object","properties":{},"required":[]}`
- Independent second AC (`TestNoInKeyInSerializedOutput`) asserts the raw serialized `tools/list` JSON contains no `"in":"` substring at all — the precise symptom the worker originally saw{"model": "deepseek-v4-flash", "problem_class": "go-cli-e2e-battery-harness", "result": "passed", "tests": 2}Root cause. The light-smoke pre-start guard treated any HTTP response on the hardcoded config port :9876 as "another dashboard is running". The dagger serve systemd unit squatting :9876 answers every path with 404, so the check tripped and hard-aborted — even though the port holder was provably not the dashboard. The companion bug (DF-013): the app's dashboard detection used a bare TCP dial (or any-HTTP-response), so once the scratch rewrite pointed the smoke at :19876, the app still cross-wired to the squatter on :9876 and surfaced an empty registry ({"apis":null}).
The fix has two halves:
1. Identity-verified pre-start check (script + app). Replace "any HTTP response ⇒ ABORT" with a shape check on GET /api/health returning JSON {"status":"ok"}. Only a verified dashboard aborts. A 404/HTML/wrong-JSON responder is a foreign squatter → scratch-port rewrite, no abort:
// probe.go — DF-013 identity probe (app + harness share this logic)
type healthResponse struct{ Status string `json:"status"` }
func DashboardHealth(addr string) (bool, error) {
client := &http.Client{Timeout: 2 * time.Second}
resp, err := client.Get("http://" + addr + "/api/health")
if err != nil {
return false, err // ECONNREFUSED → nothing to reuse, start fresh
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return false, fmt.Errorf("identity probe %s: HTTP %d (foreign/dead)", addr, resp.StatusCode)
}
body, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<16))
var h healthResponse
if err := json.Unmarshal(body, &h); err != nil {
return false, fmt.Errorf("identity probe %s: not JSON: %q", addr, body)
}
return h.Status == "ok", nil // shape check — extra fields tolerated
}
2. Scratch rewrite + detection honors --dashboard-addr. On squatter detection: sed 's/:9876/:19876/g' over the config, append --dashboard-addr <ip-address>:19876 to the start invocation, and have every registry/health client dial the resolved address so the local registry loads — killing the {"apis":null} contamination:
// runner.go — resolution: reuse genuine / rewrite squatted / start fresh
func ResolveSmokeTarget(configPath string) (addr, config string, reuse bool, err error) {
cfgAddr := "<ip-address>" + ConfigPort // :9876
if ok, pe := DashboardHealth(cfgAddr); pe == nil && ok {
return cfgAddr, configPath, true, nil // genuine dashboard → reuse
}
if PortAnswers(cfgAddr) { // occupied but NOT ours
scratch := configPath + ".scratch" // == sed 's/:9876/:19876/g'
raw, _ := os.ReadFile(configPath)
os.WriteFile(scratch, []byte(strings.ReplaceAll(string(raw), ConfigPort, ScratchPort)), 0o644)
return ScratchAddr, scratch, false, nil // <ip-address>:19876
}
return cfgAddr, configPath, false, nil
}
func StartInvocation(addr, config string) *exec.Cmd {
args := []string{"--config", config}
if addr == ScratchAddr {
args = append(args, "--dashboard-addr", addr) // detection honors this
}
return exec.Command("dashboardd", args...)
}
Bash analog (light-smoke.sh):
if health_ok "$CFG_PORT"; then
echo "ABORT: genuine dashboard verified on :${CFG_PORT} — do not smoke over it" >&2; exit 2
fi
code="$(curl -sS --max-time 2 -o /dev/null -w '%{http_code}' "http://<ip-address>:${CFG_PORT}/" 2>/dev/null || true)"
if [[ "$code" != "000" ]]; then
echo "note: :${CFG_PORT} answers HTTP ${code} but is NOT our dashboard (foreign squatter); using scratch :${SCRATCH_PORT}" >&2
SCRATCH_CFG="$(mktemp "${CFG}.XXXXXX")"
sed "s/:${CFG_PORT}/:${SCRATCH_PORT}/g" "$CFG" > "$SCRATCH_CFG"
START_ARGS=(--config "$SCRATCH_CFG" --dashboard-addr "<ip-address>:${SCRATCH_PORT}")
fi
Reproduced the full scenario in /tmp/df013 (Go module + bash script) against a real, unkillable foreign squatter currently bound to *:9876 (returns 404 on /, /api/health, /healthz, /api/registry/apis — non-JSON body, not killable, exactly like the dagger serve systemd unit).
Old behavior reproduced verbatim:
ABORT: :9876 already serving (code 404) — another dashboard is running; do not smoke over it
New behavior (live squatter still owning :9876):
note: :9876 answers HTTP 404 but is NOT our dashboard (foreign squatter); using scratch :19876
starting: dashboardd --config config.yaml.zmYLXA --dashboard-addr <ip-address>:19876
resolve: addr=<ip-address>:19876 reuse=false config=config.yaml.scratch
rewritten config: addr: <ip-address>:19876 (sed :9876→:19876 applied)
invocation: [dashboardd --config config.yaml.scratch --dashboard-addr <ip-address>:19876]
health on scratch addr: identity verified
registry via --dashboard-addr: apis=[v1.health v1.registry v1.workflows] (len=3)
ALL SCENARIO CHECKS PASSED
Contamination mechanism confirmed: registry via the squatter (:9876) → HTTP 404/no JSON (would surface {"apis":null}); via the honored --dashboard-addr (:19876) → non-null apis, loaded_from: local.
Go battery — 7/7 pass (go vet clean, go test -count=1 -v): squatter-404 ⇒ no abort + scratch rewrite · sed-equivalent rewrite with no surviving :9876 · scratch invocation carries --dashboard-addr <ip-address>:19876 (config-port invocation does not) · genuine dashboard detected & reused · registry non-null through honored addr · shape check (extra JSON fields pass; {"status":"error"} / HTML-200 / HTTP-404 / ECONNREFUSED fail) · health-poll waits on identity. Bash branches verified live: squatter→rewrite, free port→start fresh (no --dashboard-addr, no scratch file). The genuine-dashboard ABORT branch was validated in Go since the live :9876 is permanently squatted — which is the point of the fix.
JSON: {"problem_class":"go-cli-e2e-battery-harness","model":"deepseek-v4-flash","result":"passed","tests":7}
Solved by Pi Agent (deepseek-v4-flash).
Verified, no leftover processes (the pgrep hits are just the pipeline's own shell). Done.
go-cli-e2e-battery-harness — process-group reaping of & childrenDeliverable: ~/go-cli-e2e-battery-harness/SOLUTION.md (self-contained), with a runnable verification harness in the same directory.
A terminal background=true call is one session whose lifetime is its main command's. The wrapper shell is detached as its own session/process-group leader (sid == pgid == pid), and cmd1 & cmd2 & cmd3 ... ; curl puts all children in that same process group. When curl returns, the wrapper exits and the harness tears the session down with a process-group kill (SIGTERM → grace → SIGKILL -pgid, e.g. pi's killProcessTree). So the servers started, bound, and printed banners — then were reaped ~3 s later (Shutting down... Goodbye.), and every post-run probe hit 000. nohup/disown don't help (teardown targets the group, not SIGHUP jobs).
One terminal background=true call per server, with the server as that session's main process (exec), probes in separate foreground calls:
(terminal background=true) cd /repo && exec python3 echo-server.py
(terminal background=true) cd /repo && exec python3 -m http.server 18082 --bind <ip-address> --directory fixtures
(terminal background=true) cd /repo && exec ./mf start --data-dir ./data
(terminal) curl -fsS http://<ip-address>:19876/health # → 200
Each session stays alive exactly as long as its own server; stopping = explicit SIGTERM to that session.
repro.py compound → banners, internal probe 200, then teardown → {'echo':'000','http':'000','dash':'000'}, log shows both servers' Shutting down... Goodbye. Timeline: dash 200 at t≈2.5 s, wrapper exits at t≈5.1 s, shutdown ≈3 s after the banner — matches the incident to the letter.repro.py sessions → 3 independent sessions, foreground probes {'echo':'200','http':'200','dash':'200'}, graceful Shutting down only after explicit stop, then clean 000/refused.Files: SOLUTION.md, repro.py, timeline.py, echo-server.py, mf, fixture_server.py, fixtures/.
# Evidence - Problem class: go-cli-e2e-battery-harness - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-08-22T17:49:43.487Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Background terminal compound command with &-launched children: children (dashboard server, python fixtures) are killed when the wrapper shell exits \u2014 the server printed its banner then 'Shutting down...' ~3s later; probes hit 000. The dashboard bind is fine; the process group is reaped. Fix: launch each long-running server in its OWN terminal(background=true) session, never as `cmd1 & cmd2 & ... ; curl` inside one background command.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "go-cli-e2e-battery-harness", "provider": "openrouter", "solved_at": "2026-08-22T17:49:43.487Z", "version": ""}Root cause. The light-smoke pre-start guard treated any HTTP response on the hardcoded config port :9876 as "another dashboard is running". The dagger serve systemd unit squatting :9876 answers every path with 404, so the check tripped and hard-aborted — even though the port holder was provably not the dashboard. The companion bug (DF-013): the app's dashboard detection used a bare TCP dial (or any-HTTP-response), so once the scratch rewrite pointed the smoke at :19876, the app still cross-wired to the squatter on :9876 and surfaced an empty registry ({"apis":null}).
The fix has two halves:
1. Identity-verified pre-start check (script + app). Replace "any HTTP response ⇒ ABORT" with a shape check on GET /api/health returning JSON {"status":"ok"}. Only a verified dashboard aborts. A 404/HTML/wrong-JSON responder is a foreign squatter → scratch-port rewrite, no abort:
// probe.go — DF-013 identity probe (app + harness share this logic)
type healthResponse struct{ Status string `json:"status"` }
func DashboardHealth(addr string) (bool, error) {
client := &http.Client{Timeout: 2 * time.Second}
resp, err := client.Get("http://" + addr + "/api/health")
if err != nil {
return false, err // ECONNREFUSED → nothing to reuse, start fresh
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return false, fmt.Errorf("identity probe %s: HTTP %d (foreign/dead)", addr, resp.StatusCode)
}
body, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<16))
var h healthResponse
if err := json.Unmarshal(body, &h); err != nil {
return false, fmt.Errorf("identity probe %s: not JSON: %q", addr, body)
}
return h.Status == "ok", nil // shape check — extra fields tolerated
}
2. Scratch rewrite + detection honors --dashboard-addr. On squatter detection: sed 's/:9876/:19876/g' over the config, append --dashboard-addr <ip-address>:19876 to the start invocation, and have every registry/health client dial the resolved address so the local registry loads — killing the {"apis":null} contamination:
// runner.go — resolution: reuse genuine / rewrite squatted / start fresh
func ResolveSmokeTarget(configPath string) (addr, config string, reuse bool, err error) {
cfgAddr := "<ip-address>" + ConfigPort // :9876
if ok, pe := DashboardHealth(cfgAddr); pe == nil && ok {
return cfgAddr, configPath, true, nil // genuine dashboard → reuse
}
if PortAnswers(cfgAddr) { // occupied but NOT ours
scratch := configPath + ".scratch" // == sed 's/:9876/:19876/g'
raw, _ := os.ReadFile(configPath)
os.WriteFile(scratch, []byte(strings.ReplaceAll(string(raw), ConfigPort, ScratchPort)), 0o644)
return ScratchAddr, scratch, false, nil // <ip-address>:19876
}
return cfgAddr, configPath, false, nil
}
func StartInvocation(addr, config string) *exec.Cmd {
args := []string{"--config", config}
if addr == ScratchAddr {
args = append(args, "--dashboard-addr", addr) // detection honors this
}
return exec.Command("dashboardd", args...)
}
Bash analog (light-smoke.sh):
if health_ok "$CFG_PORT"; then
echo "ABORT: genuine dashboard verified on :${CFG_PORT} — do not smoke over it" >&2; exit 2
fi
code="$(curl -sS --max-time 2 -o /dev/null -w '%{http_code}' "http://<ip-address>:${CFG_PORT}/" 2>/dev/null || true)"
if [[ "$code" != "000" ]]; then
echo "note: :${CFG_PORT} answers HTTP ${code} but is NOT our dashboard (foreign squatter); using scratch :${SCRATCH_PORT}" >&2
SCRATCH_CFG="$(mktemp "${CFG}.XXXXXX")"
sed "s/:${CFG_PORT}/:${SCRATCH_PORT}/g" "$CFG" > "$SCRATCH_CFG"
START_ARGS=(--config "$SCRATCH_CFG" --dashboard-addr "<ip-address>:${SCRATCH_PORT}")
fi
Reproduced the full scenario in /tmp/df013 (Go module + bash script) against a real, unkillable foreign squatter currently bound to *:9876 (returns 404 on /, /api/health, /healthz, /api/registry/apis — non-JSON body, not killable, exactly like the dagger serve systemd unit).
Old behavior reproduced verbatim:
ABORT: :9876 already serving (code 404) — another dashboard is running; do not smoke over it
New behavior (live squatter still owning :9876):
note: :9876 answers HTTP 404 but is NOT our dashboard (foreign squatter); using scratch :19876
starting: dashboardd --config config.yaml.zmYLXA --dashboard-addr <ip-address>:19876
resolve: addr=<ip-address>:19876 reuse=false config=config.yaml.scratch
rewritten config: addr: <ip-address>:19876 (sed :9876→:19876 applied)
invocation: [dashboardd --config config.yaml.scratch --dashboard-addr <ip-address>:19876]
health on scratch addr: identity verified
registry via --dashboard-addr: apis=[v1.health v1.registry v1.workflows] (len=3)
ALL SCENARIO CHECKS PASSED
Contamination mechanism confirmed: registry via the squatter (:9876) → HTTP 404/no JSON (would surface {"apis":null}); via the honored --dashboard-addr (:19876) → non-null apis, loaded_from: local.
Go battery — 7/7 pass (go vet clean, go test -count=1 -v): squatter-404 ⇒ no abort + scratch rewrite · sed-equivalent rewrite with no surviving :9876 · scratch invocation carries --dashboard-addr <ip-address>:19876 (config-port invocation does not) · genuine dashboard detected & reused · registry non-null through honored addr · shape check (extra JSON fields pass; {"status":"error"} / HTML-200 / HTTP-404 / ECONNREFUSED fail) · health-poll waits on identity. Bash branches verified live: squatter→rewrite, free port→start fresh (no --dashboard-addr, no scratch file). The genuine-dashboard ABORT branch was validated in Go since the live :9876 is permanently squatted — which is the point of the fix.
JSON: {"problem_class":"go-cli-e2e-battery-harness","model":"deepseek-v4-flash","result":"passed","tests":7}
Solved by Pi Agent (deepseek-v4-flash).
Verified, no leftover processes (the pgrep hits are just the pipeline's own shell). Done.
go-cli-e2e-battery-harness — process-group reaping of & childrenDeliverable: ~/go-cli-e2e-battery-harness/SOLUTION.md (self-contained), with a runnable verification harness in the same directory.
A terminal background=true call is one session whose lifetime is its main command's. The wrapper shell is detached as its own session/process-group leader (sid == pgid == pid), and cmd1 & cmd2 & cmd3 ... ; curl puts all children in that same process group. When curl returns, the wrapper exits and the harness tears the session down with a process-group kill (SIGTERM → grace → SIGKILL -pgid, e.g. pi's killProcessTree). So the servers started, bound, and printed banners — then were reaped ~3 s later (Shutting down... Goodbye.), and every post-run probe hit 000. nohup/disown don't help (teardown targets the group, not SIGHUP jobs).
One terminal background=true call per server, with the server as that session's main process (exec), probes in separate foreground calls:
(terminal background=true) cd /repo && exec python3 echo-server.py
(terminal background=true) cd /repo && exec python3 -m http.server 18082 --bind <ip-address> --directory fixtures
(terminal background=true) cd /repo && exec ./mf start --data-dir ./data
(terminal) curl -fsS http://<ip-address>:19876/health # → 200
Each session stays alive exactly as long as its own server; stopping = explicit SIGTERM to that session.
repro.py compound → banners, internal probe 200, then teardown → {'echo':'000','http':'000','dash':'000'}, log shows both servers' Shutting down... Goodbye. Timeline: dash 200 at t≈2.5 s, wrapper exits at t≈5.1 s, shutdown ≈3 s after the banner — matches the incident to the letter.repro.py sessions → 3 independent sessions, foreground probes {'echo':'200','http':'200','dash':'200'}, graceful Shutting down only after explicit stop, then clean 000/refused.Files: SOLUTION.md, repro.py, timeline.py, echo-server.py, mf, fixture_server.py, fixtures/.
# Evidence - Problem class: go-cli-e2e-battery-harness - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-08-22T17:49:43.487Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Background terminal compound command with &-launched children: children (dashboard server, python fixtures) are killed when the wrapper shell exits \u2014 the server printed its banner then 'Shutting down...' ~3s later; probes hit 000. The dashboard bind is fine; the process group is reaped. Fix: launch each long-running server in its OWN terminal(background=true) session, never as `cmd1 & cmd2 & ... ; curl` inside one background command.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "go-cli-e2e-battery-harness", "provider": "openrouter", "solved_at": "2026-08-22T17:49:43.487Z", "version": ""}