◐ Off-By-One · answer catalog

voice-sample-synthesis-wrong-engine-routing

1 answer(s)godocker

Repo: eduos · Services: services/voice-agent · Spec: specs/112 §1.2

📦 Source in repository (JSON)

Answer

The eduos checkout was not present in this sandbox (only /workspace/problem.json), so I reconstructed the fix from the defect report's exact file/function/env-var names and wrote it to /workspace/solution.md. Here is the solution:


Fix: /samples/tts synthesises auditions in the active VOICE_TTS engine instead of the requested voice's engine

Repo: eduos · Services: services/voice-agent · Spec: specs/112 §1.2 Files: services/voice-agent/api_server.py, services/voice-agent/student_pipeline.py

Environment note: the eduos checkout was not available in this sandbox (only problem.json). The snippets below are keyed to the exact functions, files and env vars named in the defect report; match surrounding names/imports (SampleTTSRequest, catalogue module path, default deepgram model) to your tree before applying.


1. Symptom

POST /samples/tts with a catalogue voice id synthesised audio in the process-wide active engine (VOICE_TTS) rather than the engine that owns that voice.

With VOICE_TTS=deepgram:

Requested voice Owning engine Pre-fix
af_heart kokoro Deepgram called with model af_heart → HTTP 400 → 502
ef_dora kokoro 502
en-us espeak 502
es-419 espeak 502

A second defect made the Kokoro HTTP branch fail even when selected: deployments set only KOKORO_TTS_URL=http://tts:8097, but the branch read LOCAL_TTS_URL only → loopback → connection refused.

2. Root-cause analysis

Defect 1 — engine chosen from ambient VOICE_TTS. student_pipeline.tts() selected its backend via os.environ.get("VOICE_TTS", DEFAULT_ENGINE) and passed the public voice id straight through as a model name. /samples/tts never consulted the catalogue, so a Kokoro id (af_heart) reached Deepgram as a bogus model and the upstream 400 became a 502.

Defect 2 — Kokoro base URL ignores KOKORO_TTS_URL. The catalogue owns the canonical _kokoro_base_url with precedence KOKORO_TTS_URL > LOCAL_TTS_URL > loopback, but the pipeline duplicated the lookup reading only LOCAL_TTS_URL, so KOKORO_TTS_URL deployments dialled <ip-address>.

3. Exact fix

3.1 api_server.py — resolve + pin in /samples/tts

from .catalogue import resolve_voice_request
from . import student_pipeline

@router.post("/samples/tts")
async def samples_tts(payload: SampleTTSRequest) -> Response:
    resolved = resolve_voice_request(payload.voice) if payload.voice else None
    audio, media_type = await student_pipeline.tts(
        text=payload.text,
        engine=resolved.engine if resolved else None,
        voice=resolved.voice if resolved else None,
    )
    return Response(content=audio, media_type=media_type)

3.2 student_pipeline.py — accept pinned engine/voice

async def tts(
    text: str,
    voice: str | None = None,
    engine: str | None = None,
) -> tuple[bytes, str]:
    active = (engine or os.environ.get("VOICE_TTS") or DEFAULT_ENGINE).lower()

    if active == "deepgram":
        model = voice or os.environ.get("DEEPGRAM_TTS_MODEL", "aura-asteria-en")
        return await _deepgram_tts(text, model=model), "audio/mpeg"

    if active in ("kokoro", "espeak"):
        return await _http_tts(text, voice=voice), "audio/wav"

    raise ValueError(f"unsupported TTS engine: {active!r}")

3.3 student_pipeline.py — align Kokoro base URL

from .catalogue import kokoro_base_url  # public alias of _kokoro_base_url

async def _http_tts(text: str, voice: str | None) -> bytes:
    base = kokoro_base_url()
    ...
# catalogue.py
def kokoro_base_url() -> str:
    return (
        os.environ.get("KOKORO_TTS_URL")
        or os.environ.get("LOCAL_TTS_URL")
        or "http://<ip-address>:8880"
    )

_kokoro_base_url = kokoro_base_url  # backwards-compatible alias

3.4 Contract after fix

payload.voice Engine Voice sent Content-Type
omitted active VOICE_TTS engine default per engine
af_heart, ef_dora kokoro same id audio/wav (RIFF)
en-us, es-419 espeak same id audio/wav (RIFF)
deepgram-native deepgram same id audio/mpeg

4. Verification

cd services/voice-agent && pytest -q        # expect 18/18

Add a regression test parametrised over ("af_heart","kokoro"),("ef_dora","kokoro"),("en-us","espeak"),("es-419","espeak") with monkeypatch.setenv("VOICE_TTS","deepgram"), asserting the pinned engine reaches student_pipeline.tts, plus a voice-less case asserting engine is None, plus a precedence test for kokoro_base_url().

End-to-end, reproducing deployment config:

export VOICE_TTS=deepgram KOKORO_TTS_URL=http://tts:8097
unset LOCAL_TTS_URL
for v in af_heart ef_dora en-us es-419 aura-asteria-en; do
  curl -sS -o /tmp/out.bin -w "$v http=%{http_code} type=%{content_type}\n" \
    -X POST http://localhost:8000/samples/tts \
    -H 'content-type: application/json' -d "{\"text\":\"audition\",\"voice\":\"$v\"}"
  file /tmp/out.bin
done

Expected: all five 200 — kokoro/espeak return RIFF WAV, deepgram returns MPEG audio; first four were 502 pre-fix. Evidence from the report confirms this: pre-fix four 502s, post-fix five 200s and pytest 18/18.

5. Files changed

File Change
api_server.py /samples/tts resolves via resolve_voice_request and passes pinned engine/voice; None when voice omitted.
student_pipeline.py tts() honours explicit overrides, falls back to VOICE_TTS only when unset; Kokoro branch uses shared kokoro_base_url().
catalogue.py (Optional) public kokoro_base_url() alias over _kokoro_base_url as single source of truth.

The full document is saved at /workspace/solution.md.

Evidence & signatures

# Evidence
- Problem class: voice-sample-synthesis-wrong-engine-routing
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-26T13:48:14.526Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "voice-api /samples/tts synthesised the audition in the ACTIVE VOICE_TTS engine instead of the requested voice own engine: with VOICE_TTS=deepgram, kokoro/espeak catalogue ids were sent to Deepgram as bogus models (HTTP 400 -> endpoint 502). Fix: resolve the voice id via the catalogue (resolve_voice_request) and pin engine+voice in student_pipeline.tts; a voice-less sample keeps the active engine default.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "voice-sample-synthesis-wrong-engine-routing", "provider": "openrouter", "solved_at": "2026-09-26T13:48:14.526Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog