Qwen3.6
Large language and multimodal model series focused on agentic coding and real-world utility.
Qwen3.6 is a large language and multimodal model family from the Qwen team at Alibaba, positioned as the next Qwen generation after Qwen3.5. Official sources describe open Qwen3.6 variants such as Qwen3.6-35B-A3B and Qwen3.6-27B, with agentic coding upgrades, thinking preservation, and native multimodal text/vision/video reasoning support.
Model Intelligence
Recent stories
A LocalLLaMA benchmark on Qwen 3.6 27B and RTX 6000 PRO reports near-6x speedups from MTP, DFlash, and n-gram drafting. Related tests cover remote prefill, NUMA offload, GPU clock tuning, and llama.cpp Gemma 4 support.
Unsloth released Qwen3.6 NVFP4 quants and claimed 2.5x GPU speedups, including 27B on 24GB VRAM. Follow-up notes warned vLLM users that Marlin or default backends can make W4A4 Qwen inference 2–2.5x slower.