Published · Phase 1
Qwen3.6 35B on DGX Spark
The cited NVIDIA forum configuration reports 171.64 aggregate output tokens per second for Qwen3.6-35B-A3B-NVFP4 on one DGX Spark under vLLM nightly, a 65,536-token context, concurrency 16, an 8,000-token prompt, and 1,000 generated tokens. It is third-party service-throughput evidence, not a ComputeSage-controlled single-user decode measurement.
Editorial review: complete · Updated 2026-08-30
Measured / sourced results
| Configuration | Evidence state | Metric | Source |
|---|---|---|---|
| Qwen3.6-35B-A3B-NVFP4 · NVIDIA DGX Spark · vLLM | Source-reported · Grade C / unverified | 171.64 tokens/second | NVIDIA Developer Forums Qwen3.6 NVFP4 cross-platform report |
What the numbers mean
The 171.64 value is aggregate output throughput for concurrency 16, an 8,000-token prompt, and 1,000 generated tokens. It is a group service rate for the exact vLLM tuple, not one user's decode speed, TTFT, or prefill rate.
What stands out
The row identifies Qwen3.6-35B-A3B, a mixture-of-experts variant, and NVFP4 explicitly. That prevents this 35B article from inheriting values reported for the separate 27B model or a dense representation.
Concurrency 16 makes 171.64 aggregate output tokens per second a shared-serving result. Relabeling it as an individual request's decode rate would erase scheduler and batching effects.
The result belongs to vLLM nightly at the recorded 65,536-token context and exact prompt/output tuple. A different runtime build, context, or concurrency needs its own observation.
Evidence quality / source
The visible result comes only from the named NVIDIA Developer Forums report for Qwen3.6-35B-A3B-NVFP4 on DGX Spark. The approved source-rights record permits publication, while the result remains source-reported Grade C and unverified by a ComputeSage-controlled run.
What is unknown
- The NVIDIA forum row is a source-reported Grade C observation, not a ComputeSage-controlled run; its source records the tuple but this site has not reproduced it locally.
- No selected row establishes single-request decode, TTFT, prefill, power draw, or device-residency telemetry for this 35B-A3B DGX configuration.
Workload fit
Use this row only as a bounded concurrency-16 serving signal for the recorded Qwen3.6-35B-A3B-NVFP4 tuple. Batch or shared-service planning can consider the aggregate rate; interactive chat planning still needs exact per-request latency and decode observations.
Alternatives
Use the Planner to inspect a separate model or hardware option as a new decision. Treat every alternative as its own model identity, runtime, and workload protocol; no conclusion on this page transfers to another setup.
Planner CTA
Carry this workload into the ComputeSage Planner. It keeps missing evidence visible rather than inventing a recommendation.
Sources / methodology
Review the ComputeSage evidence methodology. Bound selector: `smfworks-qwen36-dgx`; NInfer Qwen3.6 C4/C8; NVIDIA forum cross-platform. Eligible source: NVIDIA Developer Forums Qwen3.6 NVFP4 cross-platform report.