Qwen3.8-27B MTP head — Chinese-tuned

A Multi-Token-Prediction (MTP) draft head for Qwen3.8-27B, fine-tuned on Chinese text so that speculative decoding accepts more of its drafts on Chinese output. Only the head was trained; the 27B trunk is untouched.

Because MTP drafts are always verified by the trunk (verify-then-accept), a different head can only change speed, never the model's output.

Files

File What it is
qwen3.8-27b-mtp-zh-4bit.safetensors Deployed head, quantized to match a 4-bit MLX trunk. Drop-in replacement for the stock head in KaedeTai/mlx-mtp-graft (qwen3.8-27b-mtp-4bit.safetensors).
fp16/mtp_ce.safetensors, fp16/mtp_ov.safetensors The two fp16 training outputs (out_ce, out_ov runs) before quantization.

Used locally with orcarouter/Qwen3.8-27B-Uncensored-MLX (4-bit).

How it was trained

Hard self-distillation: the frozen 4-bit trunk is run forward on Chinese text, and the head is trained to predict the trunk's own argmax (not the corpus token). The goal is to imitate this specific trunk, which is exactly what speculative-decoding acceptance measures. 512-token blocks, gradient accumulation 4, lr 1e-5.

Draft acceptance rate (accepted / drafted tokens)

Text Stock head Chinese-tuned head
zh_web 56.8% 62.6%
zh_news 56.5% 62.4%
english 70.0% 69.6%
code 81.5% 81.8%

Chinese acceptance rises about 6 points; English and code are essentially unchanged.

中文說明

Qwen3.8-27B 的 MTP(多 token 預測)草稿頭,用中文文本微調,讓推測解碼在中文輸出時接受率更高(約 +6 個百分點),英文與程式碼不受影響。只訓練草稿頭、主幹不動;由於草稿一律由主幹驗證後才接受,換草稿頭只影響速度,不會改變模型輸出。

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KaedeTai/Qwen3.8-27B-MTP-zh

Base model

Qwen/Qwen3.8-27B
Finetuned
(511)
this model