◐ Off-By-One · answer catalog

lmstudio-embedding-model-restore

2 answer(s)shellproductionshellproduction

lmstudio-embedding-model-restore

📦 Source in repository (JSON)

Answer 1

Done. The environment is currently in the restored healthy state, and I verified it end-to-end, then documented the full incident-to-recovery solution.

What I found (live verification, all passing)

Check Result
curl localhost:3000/health status=healthy, embedding.healthy=true, provider=lmstudio, note=ok, uptime 45,498 s (no daemon restart)
curl 'localhost:3000/api/memories?namespace=coding-hermes&q=duckdb&limit=2' HTTP 200, scored items (score 0.63 / 0.57, total 100) — no more 503
curl localhost:1234/v1/embeddings (model text-embedding-qwen3-embedding-0.6b) dim: 1024 — matches cached VSS dims exactly
curl localhost:1234/api/v1/models model key text-embedding-qwen3-embedding-0.6b, Q8_0, 639 150 592 B (639.15 MB), loaded_instances non-empty (loaded) vs. chat models empty
curl localhost:11434/api/tags + /api/embed tag absent; qwen3-embedding-0.6b → 404 "model not found, try pulling it first"; only 768-dim nomic-embed-text exists (trap)

Root cause

The model was not loaded in LM Studio and its GGUF was gone from ~/.lmstudio/models, so provider=auto failed on all 3 providers → degraded health + 503 EMBEDDINGS_UNAVAILABLE. The 1024-dim cached embeddings under namespaces/*/.embeddings/lmstudio_text-embedding-qwen3-embedding-0.6b/ lock the fix to that exact model — any substitute (Ollama 768-dim nomic-embed-text, etc.) would corrupt the VSS.

The fix (documented in solution.md)

lms get https://huggingface.co/Qwen/Qwen3-Embedding-0.6B-GGUF -y   # full URL form ONLY
lms load text-embedding-qwen3-embedding-0.6b                        # ~1s

Short names fail with Failed to resolve artifact... (not in LM Studio's staff-pick catalog); ollama pull qwen3-embedding-0.6b fails (pull model manifest: file does not exist). No daemon restart or config change needed.

Deliverables written

One deployment note: openai provider shows healthy:false ("missing API key") — that's expected on this box and irrelevant under provider=auto while LM Studio is healthy.

Evidence & signatures

# Evidence
- Problem class: lmstudio-embedding-model-restore
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-23T05:04:28.737Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "DB-GAP-036: live /health showed status=degraded with all 3 embedding providers unhealthy (lmstudio 'No models loaded', ollama 'model not in /api/tags', openai missing key) and /api/memories?q= returned 503 EMBEDDINGS_UNAVAILABLE. Root cause: the embedding model text-embedding-qwen3-embedding-0.6b was NOT loaded in LM Studio and its GGUF files were missing from ~/.lmstudio/models (only chat models remained) \u2014 semantic search VSS caches live under namespaces/*/.embeddings/lmstudio_text-embedding-qwen3-embedding-0.6b/ with 1024-dim vectors, so the ONLY viable fix is restoring that exact model on LM Studio (a different model/provider would break every VSS index via dimension mismatch). Fix: `lms get https://huggingface.co/Qwen/Qwen3-Embedding-0.6B-GGUF -y` (full URL form \u2014 short names `lms get text-embedding-qwen3-embedding-0.6b`, `lms get Qwen/Qwen3-Embedding-0.6B`, and `lms get Qwen/Qwen3-Embedding-0.6B-GGUF` ALL fail with 'Failed to resolve artifact: The artifact does not exist or you do not have permission to read it' because the name is not in LM Studio's staff-pick catalog; the full HF URL bypasses catalog resolution), then `lms load text-embedding-qwen3-embedding-0.6b` (~1s). Verify: curl localhost:3000/health shows status=healthy, embedding.healthy=true, lmstudio note ok; curl 'localhost:3000/api/memories?namespace=coding-hermes&q=duckdb&limit=2' returns HTTP 200 with scored items. No daemon restart or config change needed. Also note: `ollama pull qwen3-embedding-0.6b` fails with 'pull model manifest: file does not exist' \u2014 Ollama registry has no such tag; do not go down the ollama alias path (dimension mismatch 768 vs 1024 would corrupt VSS).", "environment": "DuckBrain fleet memory daemon (node duckbrain.js http :3000) + LM Studio localhost:1234, Ollama localhost:11434; production config provider=auto model=text-embedding-qwen3-embedding-0.6b", "language": "shell", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "lmstudio-embedding-model-restore", "provider": "openrouter", "solved_at": "2026-08-23T05:04:28.738Z", "version": "duckbrain feat/native-s3 d0feb58"}

Answer 2

Done. The environment is currently in the restored healthy state, and I verified it end-to-end, then documented the full incident-to-recovery solution.

What I found (live verification, all passing)

Check Result
curl localhost:3000/health status=healthy, embedding.healthy=true, provider=lmstudio, note=ok, uptime 45,498 s (no daemon restart)
curl 'localhost:3000/api/memories?namespace=coding-hermes&q=duckdb&limit=2' HTTP 200, scored items (score 0.63 / 0.57, total 100) — no more 503
curl localhost:1234/v1/embeddings (model text-embedding-qwen3-embedding-0.6b) dim: 1024 — matches cached VSS dims exactly
curl localhost:1234/api/v1/models model key text-embedding-qwen3-embedding-0.6b, Q8_0, 639 150 592 B (639.15 MB), loaded_instances non-empty (loaded) vs. chat models empty
curl localhost:11434/api/tags + /api/embed tag absent; qwen3-embedding-0.6b → 404 "model not found, try pulling it first"; only 768-dim nomic-embed-text exists (trap)

Root cause

The model was not loaded in LM Studio and its GGUF was gone from ~/.lmstudio/models, so provider=auto failed on all 3 providers → degraded health + 503 EMBEDDINGS_UNAVAILABLE. The 1024-dim cached embeddings under namespaces/*/.embeddings/lmstudio_text-embedding-qwen3-embedding-0.6b/ lock the fix to that exact model — any substitute (Ollama 768-dim nomic-embed-text, etc.) would corrupt the VSS.

The fix (documented in solution.md)

lms get https://huggingface.co/Qwen/Qwen3-Embedding-0.6B-GGUF -y   # full URL form ONLY
lms load text-embedding-qwen3-embedding-0.6b                        # ~1s

Short names fail with Failed to resolve artifact... (not in LM Studio's staff-pick catalog); ollama pull qwen3-embedding-0.6b fails (pull model manifest: file does not exist). No daemon restart or config change needed.

Deliverables written

One deployment note: openai provider shows healthy:false ("missing API key") — that's expected on this box and irrelevant under provider=auto while LM Studio is healthy.

Evidence & signatures

# Evidence
- Problem class: lmstudio-embedding-model-restore
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-23T05:04:28.737Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "DB-GAP-036: live /health showed status=degraded with all 3 embedding providers unhealthy (lmstudio 'No models loaded', ollama 'model not in /api/tags', openai missing key) and /api/memories?q= returned 503 EMBEDDINGS_UNAVAILABLE. Root cause: the embedding model text-embedding-qwen3-embedding-0.6b was NOT loaded in LM Studio and its GGUF files were missing from ~/.lmstudio/models (only chat models remained) \u2014 semantic search VSS caches live under namespaces/*/.embeddings/lmstudio_text-embedding-qwen3-embedding-0.6b/ with 1024-dim vectors, so the ONLY viable fix is restoring that exact model on LM Studio (a different model/provider would break every VSS index via dimension mismatch). Fix: `lms get https://huggingface.co/Qwen/Qwen3-Embedding-0.6B-GGUF -y` (full URL form \u2014 short names `lms get text-embedding-qwen3-embedding-0.6b`, `lms get Qwen/Qwen3-Embedding-0.6B`, and `lms get Qwen/Qwen3-Embedding-0.6B-GGUF` ALL fail with 'Failed to resolve artifact: The artifact does not exist or you do not have permission to read it' because the name is not in LM Studio's staff-pick catalog; the full HF URL bypasses catalog resolution), then `lms load text-embedding-qwen3-embedding-0.6b` (~1s). Verify: curl localhost:3000/health shows status=healthy, embedding.healthy=true, lmstudio note ok; curl 'localhost:3000/api/memories?namespace=coding-hermes&q=duckdb&limit=2' returns HTTP 200 with scored items. No daemon restart or config change needed. Also note: `ollama pull qwen3-embedding-0.6b` fails with 'pull model manifest: file does not exist' \u2014 Ollama registry has no such tag; do not go down the ollama alias path (dimension mismatch 768 vs 1024 would corrupt VSS).", "environment": "DuckBrain fleet memory daemon (node duckbrain.js http :3000) + LM Studio localhost:1234, Ollama localhost:11434; production config provider=auto model=text-embedding-qwen3-embedding-0.6b", "language": "shell", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "lmstudio-embedding-model-restore", "provider": "openrouter", "solved_at": "2026-08-23T05:04:28.738Z", "version": "duckbrain feat/native-s3 d0feb58"}
Generated from the verified corpus · MIT licensedBack to the catalog