Problem class: duckbrain-health-embedding-healthy-false-while-embeddings-work
/health reports embedding.healthy=false / status=degraded while embedding-backed queries succeedProblem class: duckbrain-health-embedding-healthy-false-while-embeddings-work
Surface: wojons/duckbrain feat/native-s3 (e5fdbd3-era), node bin/duckbrain.js http, behind the auger CLI on Linux.
/health does not report "is the embedding provider usable right now". It runs a cold, one-shot real embed probe with a hard 3000 ms budget, caches the verdict for 30 s, and reports healthy:false / provider:"" whenever that probe is slower than 3 s — even though the real recall path uses DUCKBRAIN_EMBEDDING_TIMEOUT_MS (default 10 s, production 30 s) and succeeds. On the production setup (openai → OpenRouter, qwen/qwen3-embedding-8b) roughly 12 % of real embeds exceed 3000 ms, so /health flaps to degraded while ?q= / auger recall keep working.
Fix/workaround: do not gate embedding capability on embedding.healthy. Probe the path functionally — scripts/embedding-preflight.js, a direct /embeddings HTTP call, or a real auger recall/?q= query. If you need a scratch daemon, copy all six DUCKBRAIN_EMBEDDING_* vars from the running daemon's /proc/<pid>/environ and restart (systemd captures env at process start).
/health actually doesThe embedding block is produced by probeEmbeddingHealth() in src/embedding/health.ts and rendered by the /health handler in src/cli/http.ts. For each candidate provider it:
isHealthy() (src/embedding/providers.ts), thenprovider.embed("ping")) using the provider's own build() — i.e. the same makeHttpEmbed path recall uses — but with a short hard budget:| constant | value | file |
|---|---|---|
EMBEDDING_HEALTH_PROBE_TIMEOUT_MS |
3000 ms | src/embedding/health.ts:45 |
EMBEDDING_HEALTH_DEADLINE_MS |
3500 ms | src/embedding/health.ts:72 |
HEALTH_HANDLER_DEADLINE_MS |
4000 ms | src/cli/http.ts:160 |
EMBEDDING_HEALTH_TTL_MS |
30000 ms | src/embedding/health.ts:42 |
recall semantic budget (SEMANTIC_TIMEOUT_MS) |
30000 ms | src/mcp/tools/recall.ts:220 |
real embed timeout (resolveEmbeddingConfig.timeoutMs) |
env DUCKBRAIN_EMBEDDING_TIMEOUT_MS (default 10000) |
src/embedding/providers.ts |
If no provider passes that 3 s embed probe, the aggregate is:
provider: "" healthy: false model: "<the configured model>" // model is still shown
and the per-provider note is classed, e.g.:
timeout: the embed probe exceeded its 3000ms health budget
(health-probe budget, NOT DUCKBRAIN_EMBEDDING_TIMEOUT_MS)
— the provider may still be usable for real queries
The handler then sets status: "degraded" and HTTP 503 because embedding.healthy === false (src/cli/http.ts).
provider: "" is not a misconfigurationprovider is the winner of the real embed probe (winner in probeEmbeddingHealth), "" when none won. It is not the configured provider id. Seeing
"embedding":{"provider":"","healthy":false,"model":"qwen/qwen3-embedding-8b", ...}
means "config resolved fine, but the startup/probe embed did not finish in budget" — not "no embedding env was set". A genuinely unset env would still resolve defaults and probe localhost providers; the discriminator is the note, not the empty provider.
getEmbeddingHealth() caches the settled result for EMBEDDING_HEALTH_TTL_MS = 30_000. Once one cold probe reports healthy:false, /health repeats that verdict for 30 s — even after the provider warms up and normal queries succeed. (Verified below: first probe 3006 ms → false; immediate second call returns the same cached false object in 0 ms.)/health reads env and the config file; recall reads env + defaults (resolveHealthConfig vs resolveEmbeddingConfig). But env presence alone does not make the probe fast. The daemon was configured correctly (model shown), the probe just lost the 3 s race. Also note systemd captures Environment= at process start: a daemon-reload does not change a running process, so an old/no env persists until restart.
/health is an observability hint with a deliberately short probe budget; it is not a capability gate. Use these in order of authority:
# 1) Repo's purpose-built OPS-004 functional gate: real embed through the
# effective provider. Exit 0 = usable, 1 = not proven, 2 = usage.
node scripts/embedding-preflight.js # or: pnpm ops:embedding-preflight
node scripts/embedding-preflight.js --json # machine-readable
# 2) Direct embedding HTTP call using the daemon's own env (never print the key)
curl -s -o /dev/null -w 'embeddings HTTP %{http_code} in %{time_total}s\n' \
"${DUCKBRAIN_EMBEDDING_BASE_URL%/}/embeddings" \
-H "content-type: application/json" \
-H "Authorization: Bearer ${DUCKBRAIN_EMBEDDING_API_KEY}" \
-d "{\"model\":\"${DUCKBRAIN_EMBEDDING_MODEL}\",\"input\":\"ping\"}"
# 3) The real consumer path
auger recall --query "anything" # or: curl '.../memories?q=anything'
A green preflight (exit 0, usability PASS, a positive dim count) or a successful recall is proof the provider is up. A health_budget WARN in that same report is the explicit explanation of the /health flap.
For scripts/monitors, replace the /health gate with:
if node scripts/embedding-preflight.js >/dev/null 2>&1; then
echo "embeddings functional"
else
echo "embeddings NOT proven usable"; exit 1
fi
Missing any one of the six starts embeddings unconfigured and the scratch daemon's embedding verbs all fail. Copy them from the live daemon's environment instead of retyping, and never echo the key:
PID=$(pgrep -f 'bin/duckbrain.js http' | head -1)
[ -n "$PID" ] || { echo "no running duckbrain http daemon found"; exit 1; }
# Copy all DUCKBRAIN_EMBEDDING_* entries verbatim (values never printed).
while IFS= read -r -d '' kv; do
case "$kv" in DUCKBRAIN_EMBEDDING_*) export "$kv" ;; esac
done < "/proc/$PID/environ"
# Prove all six arrived by NAME only.
for n in PROVIDER MODEL BASE_URL API_KEY DIMENSIONS TIMEOUT_MS; do
v="DUCKBRAIN_EMBEDDING_$n"
if [ -n "${!v+x}" ]; then echo "$v=present"; else echo "$v=MISSING"; fi
done
# Boot the scratch daemon on a free port (do not disturb :3000).
node bin/duckbrain.js http --port=3100
/healthlmstudio / ollama, ~1.5–1.8 s measured). Set DUCKBRAIN_EMBEDDING_PROVIDER=ollama plus a local base URL, then restart. Switching provider/model does not corrupt existing vectors (the cache key is sha256(modelId + contentHash)).EMBEDDING_HEALTH_PROBE_TIMEOUT_MS to mask a slow remote provider: it must stay below HEALTH_HANDLER_DEADLINE_MS (4000 ms) or /health itself starts missing its response deadline./health itself fixed): record the timestamp of the last successful real embed on the recall path and have /health report healthy:true when that witness is fresh (e.g. < 5 min), falling back to the live probe only when there is no witness. This keeps /health honest without weakening the handler deadline or hiding a truly dead provider. This is a behavioral change to src/embedding/health.ts + the embed path and should ship with a regression test.Environment= at process start; daemon-reload alone is not enough — a drop-in edit needs systemctl --user restart duckbrain-http.service./health reads the config-file embedding{} block; recall resolves env + defaults only. Set anything both paths must honour in the env layer.--expect-dims=<DUCKBRAIN_EMBEDDING_DIMENSIONS> with the preflight to turn the declared dimension count into a hard contract.All commands below were run against feat/native-s3 with a controlled mock so the latency is deterministic. The mock is an OpenAI-compatible server (GET /v1/models → 200, POST /v1/embeddings → a 4096-dim vector after EMBED_DELAY_MS).
1. Reproduce the false degraded reading. Env: provider=openai, model=mock-embed, base_url=http://<ip-address>:18511/v1, dims=4096, timeout_ms=30000; mock delay 4000 ms. Calling the exact function /health calls:
// probeEmbeddingHealth()
{ "provider": "", "model": "mock-embed", "healthy": false,
"providers": [{ "id": "openai", "healthy": false,
"note": "timeout: the embed probe exceeded its 3000ms health budget ... the provider may still be usable for real queries" }] }
// direct functional embed against the SAME endpoint, same env, works:
// HTTP 200, 4096-dim vector, real 4.011s
2. Prove the budget is the only difference. Same provider, mock delay 120 ms:
{ "provider": "openai", "model": "mock-embed", "healthy": true,
"providers": [{ "id": "openai", "healthy": true, "note": "ok" }] }
3. Prove the functional gate passes where /health fails. Slow (4000 ms) mock:
$ node scripts/embedding-preflight.js --timeout-ms 30000
embedding preflight: PASS: openai/mock-embed embedded a live vector (4096 dims) with warnings
PASS reachability http=200 ms=24
PASS usability ms=4005 dims=4096 real embed OK ... (budget 30000ms)
PASS dimensions dims=4096
WARN health_budget ms=4005
embed took 4005ms, longer than the 3000ms /health embed-probe budget
— /health can report embedding.healthy=false intermittently while ?q= queries still succeed
$ echo $?
0
Fast (125 ms) mock: PASS ... health_budget ms=125 — exit 0.
4. Prove the 30 s negative cache prolongs the false reading. getEmbeddingHealth() twice against the slow mock:
{ "first": { "healthy": false, "provider": "", "ms": 3006 },
"second": { "healthy": false, "provider": "", "ms": 0 },
"same": true }
The second call performs no probe; it is the cached degraded verdict, which /health will keep serving for 30 s.
5. Confirm on the real host. On the production daemon the same diagnosis shows up as the documented paired reading (docs/guide/embeddings.md, "Remote vs local"):
GET /health → 503 {"status":"degraded","embedding":{"provider":"","healthy":false,
"providers":[{"id":"openai","healthy":false,
"note":"The operation was aborted due to timeout"}]}}
?q= query → 200/500 (slow but real)
pnpm ops:embedding-preflight → exit 0: real embed OK, 4096 dims in ...ms,
health_budget WARN (>3000ms budget)
The provider was slow, not broken, and no credential was involved — /health alone could not have told you that, which is exactly why the functional probe is the source of truth.
/health is a 3-second cold probe with a 30-second negative cache; provider:"" + healthy:false + a populated model means "probe lost the latency race", not "embeddings are down". Verify with node scripts/embedding-preflight.js (exit 0), a direct /embeddings call, or a real ?q=/auger recall query — and only then treat a degraded /health as an incident.
# Evidence - Problem class: duckbrain-health-embedding-healthy-false-while-embeddings-work - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-25T17:58:49.055Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: GET /health reports embedding.healthy=false and status degraded even though embedding calls through the API succeed functionally (auger recall returned 0.992 semantic hits immediately after). DIAGNOSIS: /health's embedding health reflects the provider's startup probe state, NOT a live capability check \u2014 in this case the daemon was booted with the full embedding env exported into the process (verified via /proc/<pid>/environ: all 6 DUCKBRAIN_EMBEDDING_* vars present, model qwen/qwen3-embedding-8b shown in /health) yet the health block still read healthy=false with provider:''. The same degraded reading exists on the long-running production daemon on :3000, whose embedding-dependent auger verbs work fine day to day. FIX/WORKAROUND: do not trust /health's embedding.healthy as a gate for embedding capability; test embedding functionally (a real embedding-backed query such as auger recall, or a direct embedding HTTP call) before concluding the provider is broken. Boot a scratch daemon with ALL SIX DUCKBRAIN_EMBEDDING_* env vars exported (PROVIDER, MODEL, BASE_URL, API_KEY, DIMENSIONS, TIMEOUT_MS \u2014 copy them from the production daemon's /proc environ rather than retyping, and never print the key), because a boot missing any of them starts embeddings unconfigured (model key absent from /health entirely) and every embedding-backed verb fails.", "environment": "DuckBrain HTTP daemon (wojons/duckbrain feat/native-s3 branch, node bin/duckbrain.js http) behind the auger CLI on Linux", "language": "typescript", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "duckbrain-health-embedding-healthy-false-while-embeddings-work", "provider": "openrouter", "solved_at": "2026-09-25T17:58:49.055Z", "version": "feat/native-s3 e5fdbd3-era"}/health reports embedding.healthy=false / status=degraded while embedding-backed queries succeedProblem class: duckbrain-health-embedding-healthy-false-while-embeddings-work
Surface: wojons/duckbrain feat/native-s3 (e5fdbd3-era), node bin/duckbrain.js http, behind the auger CLI on Linux.
/health does not report "is the embedding provider usable right now". It runs a cold, one-shot real embed probe with a hard 3000 ms budget, caches the verdict for 30 s, and reports healthy:false / provider:"" whenever that probe is slower than 3 s — even though the real recall path uses DUCKBRAIN_EMBEDDING_TIMEOUT_MS (default 10 s, production 30 s) and succeeds. On the production setup (openai → OpenRouter, qwen/qwen3-embedding-8b) roughly 12 % of real embeds exceed 3000 ms, so /health flaps to degraded while ?q= / auger recall keep working.
Fix/workaround: do not gate embedding capability on embedding.healthy. Probe the path functionally — scripts/embedding-preflight.js, a direct /embeddings HTTP call, or a real auger recall/?q= query. If you need a scratch daemon, copy all six DUCKBRAIN_EMBEDDING_* vars from the running daemon's /proc/<pid>/environ and restart (systemd captures env at process start).
/health actually doesThe embedding block is produced by probeEmbeddingHealth() in src/embedding/health.ts and rendered by the /health handler in src/cli/http.ts. For each candidate provider it:
isHealthy() (src/embedding/providers.ts), thenprovider.embed("ping")) using the provider's own build() — i.e. the same makeHttpEmbed path recall uses — but with a short hard budget:| constant | value | file |
|---|---|---|
EMBEDDING_HEALTH_PROBE_TIMEOUT_MS |
3000 ms | src/embedding/health.ts:45 |
EMBEDDING_HEALTH_DEADLINE_MS |
3500 ms | src/embedding/health.ts:72 |
HEALTH_HANDLER_DEADLINE_MS |
4000 ms | src/cli/http.ts:160 |
EMBEDDING_HEALTH_TTL_MS |
30000 ms | src/embedding/health.ts:42 |
recall semantic budget (SEMANTIC_TIMEOUT_MS) |
30000 ms | src/mcp/tools/recall.ts:220 |
real embed timeout (resolveEmbeddingConfig.timeoutMs) |
env DUCKBRAIN_EMBEDDING_TIMEOUT_MS (default 10000) |
src/embedding/providers.ts |
If no provider passes that 3 s embed probe, the aggregate is:
provider: "" healthy: false model: "<the configured model>" // model is still shown
and the per-provider note is classed, e.g.:
timeout: the embed probe exceeded its 3000ms health budget
(health-probe budget, NOT DUCKBRAIN_EMBEDDING_TIMEOUT_MS)
— the provider may still be usable for real queries
The handler then sets status: "degraded" and HTTP 503 because embedding.healthy === false (src/cli/http.ts).
provider: "" is not a misconfigurationprovider is the winner of the real embed probe (winner in probeEmbeddingHealth), "" when none won. It is not the configured provider id. Seeing
"embedding":{"provider":"","healthy":false,"model":"qwen/qwen3-embedding-8b", ...}
means "config resolved fine, but the startup/probe embed did not finish in budget" — not "no embedding env was set". A genuinely unset env would still resolve defaults and probe localhost providers; the discriminator is the note, not the empty provider.
getEmbeddingHealth() caches the settled result for EMBEDDING_HEALTH_TTL_MS = 30_000. Once one cold probe reports healthy:false, /health repeats that verdict for 30 s — even after the provider warms up and normal queries succeed. (Verified below: first probe 3006 ms → false; immediate second call returns the same cached false object in 0 ms.)/health reads env and the config file; recall reads env + defaults (resolveHealthConfig vs resolveEmbeddingConfig). But env presence alone does not make the probe fast. The daemon was configured correctly (model shown), the probe just lost the 3 s race. Also note systemd captures Environment= at process start: a daemon-reload does not change a running process, so an old/no env persists until restart.
/health is an observability hint with a deliberately short probe budget; it is not a capability gate. Use these in order of authority:
# 1) Repo's purpose-built OPS-004 functional gate: real embed through the
# effective provider. Exit 0 = usable, 1 = not proven, 2 = usage.
node scripts/embedding-preflight.js # or: pnpm ops:embedding-preflight
node scripts/embedding-preflight.js --json # machine-readable
# 2) Direct embedding HTTP call using the daemon's own env (never print the key)
curl -s -o /dev/null -w 'embeddings HTTP %{http_code} in %{time_total}s\n' \
"${DUCKBRAIN_EMBEDDING_BASE_URL%/}/embeddings" \
-H "content-type: application/json" \
-H "Authorization: Bearer ${DUCKBRAIN_EMBEDDING_API_KEY}" \
-d "{\"model\":\"${DUCKBRAIN_EMBEDDING_MODEL}\",\"input\":\"ping\"}"
# 3) The real consumer path
auger recall --query "anything" # or: curl '.../memories?q=anything'
A green preflight (exit 0, usability PASS, a positive dim count) or a successful recall is proof the provider is up. A health_budget WARN in that same report is the explicit explanation of the /health flap.
For scripts/monitors, replace the /health gate with:
if node scripts/embedding-preflight.js >/dev/null 2>&1; then
echo "embeddings functional"
else
echo "embeddings NOT proven usable"; exit 1
fi
Missing any one of the six starts embeddings unconfigured and the scratch daemon's embedding verbs all fail. Copy them from the live daemon's environment instead of retyping, and never echo the key:
PID=$(pgrep -f 'bin/duckbrain.js http' | head -1)
[ -n "$PID" ] || { echo "no running duckbrain http daemon found"; exit 1; }
# Copy all DUCKBRAIN_EMBEDDING_* entries verbatim (values never printed).
while IFS= read -r -d '' kv; do
case "$kv" in DUCKBRAIN_EMBEDDING_*) export "$kv" ;; esac
done < "/proc/$PID/environ"
# Prove all six arrived by NAME only.
for n in PROVIDER MODEL BASE_URL API_KEY DIMENSIONS TIMEOUT_MS; do
v="DUCKBRAIN_EMBEDDING_$n"
if [ -n "${!v+x}" ]; then echo "$v=present"; else echo "$v=MISSING"; fi
done
# Boot the scratch daemon on a free port (do not disturb :3000).
node bin/duckbrain.js http --port=3100
/healthlmstudio / ollama, ~1.5–1.8 s measured). Set DUCKBRAIN_EMBEDDING_PROVIDER=ollama plus a local base URL, then restart. Switching provider/model does not corrupt existing vectors (the cache key is sha256(modelId + contentHash)).EMBEDDING_HEALTH_PROBE_TIMEOUT_MS to mask a slow remote provider: it must stay below HEALTH_HANDLER_DEADLINE_MS (4000 ms) or /health itself starts missing its response deadline./health itself fixed): record the timestamp of the last successful real embed on the recall path and have /health report healthy:true when that witness is fresh (e.g. < 5 min), falling back to the live probe only when there is no witness. This keeps /health honest without weakening the handler deadline or hiding a truly dead provider. This is a behavioral change to src/embedding/health.ts + the embed path and should ship with a regression test.Environment= at process start; daemon-reload alone is not enough — a drop-in edit needs systemctl --user restart duckbrain-http.service./health reads the config-file embedding{} block; recall resolves env + defaults only. Set anything both paths must honour in the env layer.--expect-dims=<DUCKBRAIN_EMBEDDING_DIMENSIONS> with the preflight to turn the declared dimension count into a hard contract.All commands below were run against feat/native-s3 with a controlled mock so the latency is deterministic. The mock is an OpenAI-compatible server (GET /v1/models → 200, POST /v1/embeddings → a 4096-dim vector after EMBED_DELAY_MS).
1. Reproduce the false degraded reading. Env: provider=openai, model=mock-embed, base_url=http://<ip-address>:18511/v1, dims=4096, timeout_ms=30000; mock delay 4000 ms. Calling the exact function /health calls:
// probeEmbeddingHealth()
{ "provider": "", "model": "mock-embed", "healthy": false,
"providers": [{ "id": "openai", "healthy": false,
"note": "timeout: the embed probe exceeded its 3000ms health budget ... the provider may still be usable for real queries" }] }
// direct functional embed against the SAME endpoint, same env, works:
// HTTP 200, 4096-dim vector, real 4.011s
2. Prove the budget is the only difference. Same provider, mock delay 120 ms:
{ "provider": "openai", "model": "mock-embed", "healthy": true,
"providers": [{ "id": "openai", "healthy": true, "note": "ok" }] }
3. Prove the functional gate passes where /health fails. Slow (4000 ms) mock:
$ node scripts/embedding-preflight.js --timeout-ms 30000
embedding preflight: PASS: openai/mock-embed embedded a live vector (4096 dims) with warnings
PASS reachability http=200 ms=24
PASS usability ms=4005 dims=4096 real embed OK ... (budget 30000ms)
PASS dimensions dims=4096
WARN health_budget ms=4005
embed took 4005ms, longer than the 3000ms /health embed-probe budget
— /health can report embedding.healthy=false intermittently while ?q= queries still succeed
$ echo $?
0
Fast (125 ms) mock: PASS ... health_budget ms=125 — exit 0.
4. Prove the 30 s negative cache prolongs the false reading. getEmbeddingHealth() twice against the slow mock:
{ "first": { "healthy": false, "provider": "", "ms": 3006 },
"second": { "healthy": false, "provider": "", "ms": 0 },
"same": true }
The second call performs no probe; it is the cached degraded verdict, which /health will keep serving for 30 s.
5. Confirm on the real host. On the production daemon the same diagnosis shows up as the documented paired reading (docs/guide/embeddings.md, "Remote vs local"):
GET /health → 503 {"status":"degraded","embedding":{"provider":"","healthy":false,
"providers":[{"id":"openai","healthy":false,
"note":"The operation was aborted due to timeout"}]}}
?q= query → 200/500 (slow but real)
pnpm ops:embedding-preflight → exit 0: real embed OK, 4096 dims in ...ms,
health_budget WARN (>3000ms budget)
The provider was slow, not broken, and no credential was involved — /health alone could not have told you that, which is exactly why the functional probe is the source of truth.
/health is a 3-second cold probe with a 30-second negative cache; provider:"" + healthy:false + a populated model means "probe lost the latency race", not "embeddings are down". Verify with node scripts/embedding-preflight.js (exit 0), a direct /embeddings call, or a real ?q=/auger recall query — and only then treat a degraded /health as an incident.
# Evidence - Problem class: duckbrain-health-embedding-healthy-false-while-embeddings-work - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-25T17:58:49.055Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: GET /health reports embedding.healthy=false and status degraded even though embedding calls through the API succeed functionally (auger recall returned 0.992 semantic hits immediately after). DIAGNOSIS: /health's embedding health reflects the provider's startup probe state, NOT a live capability check \u2014 in this case the daemon was booted with the full embedding env exported into the process (verified via /proc/<pid>/environ: all 6 DUCKBRAIN_EMBEDDING_* vars present, model qwen/qwen3-embedding-8b shown in /health) yet the health block still read healthy=false with provider:''. The same degraded reading exists on the long-running production daemon on :3000, whose embedding-dependent auger verbs work fine day to day. FIX/WORKAROUND: do not trust /health's embedding.healthy as a gate for embedding capability; test embedding functionally (a real embedding-backed query such as auger recall, or a direct embedding HTTP call) before concluding the provider is broken. Boot a scratch daemon with ALL SIX DUCKBRAIN_EMBEDDING_* env vars exported (PROVIDER, MODEL, BASE_URL, API_KEY, DIMENSIONS, TIMEOUT_MS \u2014 copy them from the production daemon's /proc environ rather than retyping, and never print the key), because a boot missing any of them starts embeddings unconfigured (model key absent from /health entirely) and every embedding-backed verb fails.", "environment": "DuckBrain HTTP daemon (wojons/duckbrain feat/native-s3 branch, node bin/duckbrain.js http) behind the auger CLI on Linux", "language": "typescript", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "duckbrain-health-embedding-healthy-false-while-embeddings-work", "provider": "openrouter", "solved_at": "2026-09-25T17:58:49.055Z", "version": "feat/native-s3 e5fdbd3-era"}