返回
RCreddit.com
15
·18小时前·开发者社区 · RSS

Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release

查看原文
LlamaQwen模型发布

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

推荐理由

Llama 相关模型动态已经出现,适合跟踪能力变化、生态影响和后续可用性。

AI 摘要

针对Qwen 3.5、3.6及最新发布的3.8模型,一个新的Jinja聊天模板已推出。官方Qwen 3.8模板存在多项问题,包括禁用思考时崩溃、聊天历史被污染、JSON字符串工具调用失败以及代理停滞。…

Qwen just released their first 3.8 model.

The main addition in 3.8 is prompt-steered reasoning effort. You can tell the model how deeply to think by setting reasoning_effort to xhigh, medium, or low.

However, the official template still has some serious problems:

- You cannot disable thinking. If you pass enable_thinking=false, it 3.8 crashes with a hard exception.

- Chat history gets poisoned. In multi-turn chats, the official template injects blank tags before real thoughts.

- Tool calling crashes. If your client passes arguments as JSON strings (the standard OpenAI API format), the official template crashes.

- Agent stalls. The official template often drops mid-dialogue system messages and wedges multi-step tool loops.

I maintain a single, drop-in fixed Jinja template that works across all Qwen 3.5, 3.6, and 3.8 models:

https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

What this template does:

- Full 3.8 reasoning effort support: Steer reasoning depth with reasoning_effort (xhigh, high, low, medium).

- Restores the thinking toggle: Turn off reasoning whenever you want fast answers, either via kwargs or by typing in your prompt.

- 100% KV Cache hits: Keeps past thoughts intact by default so your prefix cache stays warm across turns.

- llama.cpp support: Native support for the new --reasoning-preserve flag.

- Universal tool parsing: Handles both Python dicts and JSON strings. Works on llama.cpp, vLLM, LM Studio, and MLX.

Recommended llama-server launch command:

llama-server -m your_model.gguf --jinja --chat-template-file chat_template.jinja --reasoning-format deepseek

(The --reasoning-format deepseek flag separates thinking into the OpenAI reasoning_content field so OpenCode, Claude Code, and other harnesses do not stall on raw tokens).

Note on hardware:

I cannot run a 2.4 trillion parameter model on my local rig. The template passes all 28 automated tests and tokenizer parity checks, but I would appreciate feedback from anyone testing it with Qwen 3.8.

Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release · BuzzRadr