Published · Phase 3

Qwen3.8 vs Qwen3.6 on DGX Spark

Choose Qwen3.6-35B-A3B for coding agents, tool use, and long-context workflows; choose Qwen3.8-27B for a simpler dense deployment and predictable broad-runtime support. Test both on your tasks before assuming the newer model wins.

Editorial review: complete · Updated 2026-08-30

Quick verdict

Quick verdict guidance: Choose Qwen3.6-35B-A3B for coding agents, tool use, and long-context workflows; choose Qwen3.8-27B for a simpler dense deployment and predictable broad-runtime support. Test both on your tasks before assuming the newer model wins.

Comparison table

Qwen3.8 vs Qwen3.6 on DGX Spark: eligible evidence
ConfigurationEvidence stateMetricSource
Qwen3.6-27B · NVIDIA DGX Spark · vLLMSource-reported · Grade C31 tokens/secondQwen3.6-27B-NVFP4 vLLM Performance Report
Qwen3.8-27B · NVIDIA DGX Spark public lab · vLLMMeasured · Grade A30.009033471349664 tokens/secondStackBench controlled lab approved source qwen38-dgx-spark-owner-archive-20260817

Memory / capacity

Memory / capacity guidance: At 4 bit, Qwen3.8 is roughly 16–19 GB and Qwen3.6 roughly 20–24 GB. Both fit a 32 GB GPU, but Qwen3.8 leaves more cache headroom and Qwen3.6's MoE may deliver more capability per active token.

Observed LLM performance

Observed LLM performance guidance: Qwen3.6 activates about 3B of 35B parameters, so optimized inference can be fast; Qwen3.8 is dense and simpler to reason about. Use task latency and correctness, not parameter count alone.

Prefill vs decode

Prefill vs decode guidance: Qwen3.6 advertises very long context; begin at 16K–32K. Qwen3.8 should likewise use a measured operational context rather than its maximum.

Power

Power guidance: MoE reduces active compute but still moves/routs full expert weights. Efficiency depends on kernel quality and batch size.

Current market cost

Current market cost guidance: Both avoid a professional GPU at 4 bit. The meaningful cost difference is engineering time around a newer MoE runtime versus the simpler dense model.

Which models fit

Which models fit guidance: At 4 bit, Qwen3.8 is roughly 16–19 GB and Qwen3.6 roughly 20–24 GB. Both fit a 32 GB GPU, but Qwen3.8 leaves more cache headroom and Qwen3.6's MoE may deliver more capability per active token.

Who each option suits

Who each option suits guidance: Agent and coding teams should favor 3.6; general chat, document analysis, and users who value deployment simplicity should consider 3.8.

What stands out

Start with Qwen3.6 for an agent product and Qwen3.8 for a robust general local model. Keep the one that passes a shared 20–50 task evaluation at acceptable latency.

Qwen3.6 activates about 3B of 35B parameters, so optimized inference can be fast; Qwen3.8 is dense and simpler to reason about. Use task latency and correctness, not parameter count alone.

Qwen3.5-35B-A3B has a more established deployment history; smaller 14B models reduce latency; larger Mistral/gpt-oss options use higher-memory hardware.

Evidence limitations

  • Model version labels are not quality guarantees across every domain.
  • Compare the same precision, prompt, sampler, and runtime before drawing a speed conclusion.

Planner CTA

Compare this recommendation against your model, context, concurrency, latency, and budget in the ComputeSage Planner.

Sources / methodology

Recommendations combine official specifications, public model/runtime documentation, adjacent benchmark observations, and clearly labeled engineering estimates. Review the ComputeSage evidence methodology.