Published · Phase 3
Qwen3.8 vs Qwen3.6 on DGX Spark
Choose Qwen3.6-35B-A3B for coding agents, tool use, and long-context workflows; choose Qwen3.8-27B for a simpler dense deployment and predictable broad-runtime support. Test both on your tasks before assuming the newer model wins.
Editorial review: complete · Updated 2026-08-30
Quick verdict
Quick verdict guidance: Choose Qwen3.6-35B-A3B for coding agents, tool use, and long-context workflows; choose Qwen3.8-27B for a simpler dense deployment and predictable broad-runtime support. Test both on your tasks before assuming the newer model wins.
Comparison table
| Configuration | Evidence state | Metric | Source |
|---|---|---|---|
| Qwen3.6-27B · NVIDIA DGX Spark · vLLM | Source-reported · Grade C | 31 tokens/second | Qwen3.6-27B-NVFP4 vLLM Performance Report |
| Qwen3.8-27B · NVIDIA DGX Spark public lab · vLLM | Measured · Grade A | 30.009033471349664 tokens/second | StackBench controlled lab approved source qwen38-dgx-spark-owner-archive-20260817 |
Memory / capacity
Memory / capacity guidance: At 4 bit, Qwen3.8 is roughly 16–19 GB and Qwen3.6 roughly 20–24 GB. Both fit a 32 GB GPU, but Qwen3.8 leaves more cache headroom and Qwen3.6's MoE may deliver more capability per active token.
Observed LLM performance
Observed LLM performance guidance: Qwen3.6 activates about 3B of 35B parameters, so optimized inference can be fast; Qwen3.8 is dense and simpler to reason about. Use task latency and correctness, not parameter count alone.
Prefill vs decode
Prefill vs decode guidance: Qwen3.6 advertises very long context; begin at 16K–32K. Qwen3.8 should likewise use a measured operational context rather than its maximum.
Power
Power guidance: MoE reduces active compute but still moves/routs full expert weights. Efficiency depends on kernel quality and batch size.
Current market cost
Current market cost guidance: Both avoid a professional GPU at 4 bit. The meaningful cost difference is engineering time around a newer MoE runtime versus the simpler dense model.
Which models fit
Which models fit guidance: At 4 bit, Qwen3.8 is roughly 16–19 GB and Qwen3.6 roughly 20–24 GB. Both fit a 32 GB GPU, but Qwen3.8 leaves more cache headroom and Qwen3.6's MoE may deliver more capability per active token.
Who each option suits
Who each option suits guidance: Agent and coding teams should favor 3.6; general chat, document analysis, and users who value deployment simplicity should consider 3.8.
What stands out
Start with Qwen3.6 for an agent product and Qwen3.8 for a robust general local model. Keep the one that passes a shared 20–50 task evaluation at acceptable latency.
Qwen3.6 activates about 3B of 35B parameters, so optimized inference can be fast; Qwen3.8 is dense and simpler to reason about. Use task latency and correctness, not parameter count alone.
Qwen3.5-35B-A3B has a more established deployment history; smaller 14B models reduce latency; larger Mistral/gpt-oss options use higher-memory hardware.
Evidence limitations
- Model version labels are not quality guarantees across every domain.
- Compare the same precision, prompt, sampler, and runtime before drawing a speed conclusion.
Planner CTA
Compare this recommendation against your model, context, concurrency, latency, and budget in the ComputeSage Planner.
Sources / methodology
Recommendations combine official specifications, public model/runtime documentation, adjacent benchmark observations, and clearly labeled engineering estimates. Review the ComputeSage evidence methodology.