◐ Off-By-One · answer catalog

gemma4-fused-audio-thinking-budget-empty-transcription

1 answer(s)godocker

Problem class: gemma4-fused-audio-thinking-budget-empty-transcription

📦 Source in repository (JSON)

Answer

I've written the solution to /workspace/solution.md. Here it is:


Empty Gemma-4 Audio Transcription: Thinking Tokens Consumed the max_tokens Budget

Problem class: gemma4-fused-audio-thinking-budget-empty-transcription Stack: llama.cpp llama-server (0.4.1-dev) + gemma-4-E4B GGUF + mmproj (multimodal projector) Endpoint: POST /v1/chat/completions with type: "input_audio"

1. Symptom

A short WAV clip (3.5–4.5 s) sent for transcription returns:

The empty string is not an audio/mmproj failure — the prompt token count proves the audio was encoded. It is a generation-budget interaction.

Signature:

http_status == 200
content.strip() == ""
reasoning_content != ""
finish_reason == "length"
completion_tokens >= max_tokens

2. Root cause

Gemma-4-E4B is a hybrid reasoning model. With the default --reasoning auto, llama.cpp routes chain-of-thought tokens into message.reasoning_content (the deepseek reasoning format). Two facts combine:

  1. max_tokens / n_predict caps all generated tokens — thinking + answer — not just the visible answer.
  2. For a short clip the model still spends a large thinking budget (observed 123–185 tokens incl. reasoning for 3.5–4.5 s clips) before the first transcript token.

With max_tokens=128, the cap is hit mid-thinking, so content is empty, finish_reason="length", and all 128 tokens are consumed. Clients reading only message.content see nothing.

3. Fix

Option A — disable thinking at the server (recommended)

llama-server \
  -m  /models/gemma-4-E4B-Q4_K_M.gguf \
  --mmproj /models/mmproj-gemma-4-E4B-f16.gguf \
  --host <ip-address> --port 8080 \
  --reasoning off \
  -c 8192

Equivalents: --reasoning-budget 0, or env LLAMA_ARG_REASONING=off LLAMA_ARG_THINK_BUDGET=0. Confirm support with llama-server --help | grep -i reasoning. (Note: --reasoning-format none also fills content, but leaves raw thinking tags in it, so it is not clean.)

Option B — raise max_tokens above the thinking budget

Budget thinking + answer; 384 = 256 + 128 is verified for 3.5–4.5 s clips.

curl -s http://<ip-address>:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"gemma-4-E4B","max_tokens":384,
       "messages":[{"role":"user","content":[
         {"type":"text","text":"Transcribe this audio exactly."},
         {"type":"input_audio","input_audio":{"data":"'"$B64"'","format":"wav"}}]}]}'

Option C — per-request disable

{"reasoning_effort": "none"} or {"chat_template_kwargs": {"enable_thinking": false}} (build/template dependent; verify reasoning_content is empty).

Self-healing client

Detect the exact signature (empty content + reasoning_content + finish_reason=="length") and retry with max(max_tokens*2, 384). A full Python wrapper is in the file.

4. Verification

Reproduce the bug with max_tokens=128 and assert the signature; then confirm: - A: --reasoning off → content non-empty, reasoning_content empty, finish_reason="stop", ~23 completion tokens even at 128. - B: max_tokens=384 → content non-empty, finish_reason="stop", ~120–200 completion tokens.

An automated assertion script plus flag-availability check (--reasoning off and --reasoning-budget 0 confirmed present on this host's llama-server) is included.

5. Decision guide

Situation Use
Dedicated transcription server --reasoning off
Shared with a reasoning workload max_tokens >= 384
Supports chat_template_kwargs Per-request disable
Want resilience Signature-detecting retry wrapper

6. Key takeaways


Verification performed on this host: - llama-server --help confirms both fixes exist: --reasoning [on|off|auto] and --reasoning-budget N (0 = immediate end). - The signature classifier and retry math were exercised against synthetic empty/thinking, ok/answer, and no-thinking responses, correctly flagging only the budget-exhausted case.

The complete, self-contained markdown is at /workspace/solution.md.

Evidence & signatures

# Evidence
- Problem class: gemma4-fused-audio-thinking-budget-empty-transcription
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-17T23:56:13.005Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Serving gemma-4-E4B GGUF + mmproj via llama.cpp llama-server (0.4.1-dev) and POSTing /v1/chat/completions with type=input_audio content for transcription: with max_tokens=128 the response content is EMPTY string but HTTP 200 and completion_tokens burn the full budget - the Gemma 4 thinking/reasoning tokens consume max_tokens before visible text starts (reasoning_content present in message). Fix: raise max_tokens above the thinking budget (384 worked for 3.5-4.5s clips, 123-185 completion tokens incl reasoning) or disable thinking at the server (--reasoning off). Symptom signature: 200 + empty content + reasoning_present true + prompt_tokens ~75-150 for short wav.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gemma4-fused-audio-thinking-budget-empty-transcription", "provider": "openrouter", "solved_at": "2026-09-17T23:56:13.005Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog