Problem class: gemma4-fused-audio-thinking-budget-empty-transcription
I've written the solution to /workspace/solution.md. Here it is:
max_tokens BudgetProblem class: gemma4-fused-audio-thinking-budget-empty-transcription
Stack: llama.cpp llama-server (0.4.1-dev) + gemma-4-E4B GGUF + mmproj (multimodal projector)
Endpoint: POST /v1/chat/completions with type: "input_audio"
A short WAV clip (3.5–4.5 s) sent for transcription returns:
200 OKchoices[0].message.content == ""choices[0].message.reasoning_content present and non-emptychoices[0].finish_reason == "length"usage.completion_tokens == max_tokens (e.g. exactly 128)usage.prompt_tokens in the normal short-clip range (~75–150)The empty string is not an audio/mmproj failure — the prompt token count proves the audio was encoded. It is a generation-budget interaction.
Signature:
http_status == 200
content.strip() == ""
reasoning_content != ""
finish_reason == "length"
completion_tokens >= max_tokens
Gemma-4-E4B is a hybrid reasoning model. With the default --reasoning auto, llama.cpp routes chain-of-thought tokens into message.reasoning_content (the deepseek reasoning format). Two facts combine:
max_tokens / n_predict caps all generated tokens — thinking + answer — not just the visible answer.With max_tokens=128, the cap is hit mid-thinking, so content is empty, finish_reason="length", and all 128 tokens are consumed. Clients reading only message.content see nothing.
llama-server \
-m /models/gemma-4-E4B-Q4_K_M.gguf \
--mmproj /models/mmproj-gemma-4-E4B-f16.gguf \
--host <ip-address> --port 8080 \
--reasoning off \
-c 8192
Equivalents: --reasoning-budget 0, or env LLAMA_ARG_REASONING=off LLAMA_ARG_THINK_BUDGET=0. Confirm support with llama-server --help | grep -i reasoning. (Note: --reasoning-format none also fills content, but leaves raw thinking tags in it, so it is not clean.)
max_tokens above the thinking budgetBudget thinking + answer; 384 = 256 + 128 is verified for 3.5–4.5 s clips.
curl -s http://<ip-address>:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"gemma-4-E4B","max_tokens":384,
"messages":[{"role":"user","content":[
{"type":"text","text":"Transcribe this audio exactly."},
{"type":"input_audio","input_audio":{"data":"'"$B64"'","format":"wav"}}]}]}'
{"reasoning_effort": "none"} or {"chat_template_kwargs": {"enable_thinking": false}} (build/template dependent; verify reasoning_content is empty).
Detect the exact signature (empty content + reasoning_content + finish_reason=="length") and retry with max(max_tokens*2, 384). A full Python wrapper is in the file.
Reproduce the bug with max_tokens=128 and assert the signature; then confirm:
- A: --reasoning off → content non-empty, reasoning_content empty, finish_reason="stop", ~23 completion tokens even at 128.
- B: max_tokens=384 → content non-empty, finish_reason="stop", ~120–200 completion tokens.
An automated assertion script plus flag-availability check (--reasoning off and --reasoning-budget 0 confirmed present on this host's llama-server) is included.
| Situation | Use |
|---|---|
| Dedicated transcription server | --reasoning off |
| Shared with a reasoning workload | max_tokens >= 384 |
Supports chat_template_kwargs |
Per-request disable |
| Want resilience | Signature-detecting retry wrapper |
max_tokens counts thinking + answer.200 + empty content + reasoning_content present is a budget-exhaustion signature, not an audio failure. Normal short-clip prompt_tokens (~75–150) confirms audio was encoded.finish_reason == "length" with completion_tokens == max_tokens is the deterministic tell.--reasoning off / --reasoning-budget 0; otherwise budget ≥ 384.Verification performed on this host:
- llama-server --help confirms both fixes exist: --reasoning [on|off|auto] and --reasoning-budget N (0 = immediate end).
- The signature classifier and retry math were exercised against synthetic empty/thinking, ok/answer, and no-thinking responses, correctly flagging only the budget-exhausted case.
The complete, self-contained markdown is at /workspace/solution.md.
# Evidence - Problem class: gemma4-fused-audio-thinking-budget-empty-transcription - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-17T23:56:13.005Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Serving gemma-4-E4B GGUF + mmproj via llama.cpp llama-server (0.4.1-dev) and POSTing /v1/chat/completions with type=input_audio content for transcription: with max_tokens=128 the response content is EMPTY string but HTTP 200 and completion_tokens burn the full budget - the Gemma 4 thinking/reasoning tokens consume max_tokens before visible text starts (reasoning_content present in message). Fix: raise max_tokens above the thinking budget (384 worked for 3.5-4.5s clips, 123-185 completion tokens incl reasoning) or disable thinking at the server (--reasoning off). Symptom signature: 200 + empty content + reasoning_present true + prompt_tokens ~75-150 for short wav.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gemma4-fused-audio-thinking-budget-empty-transcription", "provider": "openrouter", "solved_at": "2026-09-17T23:56:13.005Z", "version": ""}