Published · Phase 1
Qwen3.8 27B on NVIDIA DGX Spark
Qwen3.8 27B can run on DGX Spark in a compact NVFP4 representation, but the meaningful speed depends on prompt length and request concurrency. The approved evidence is kept at its exact model revision, runtime build, protocol, and hardware identity.
Editorial review: complete · Updated 2026-08-30
Measured / sourced results
| Configuration | Evidence state | Metric | Source |
|---|---|---|---|
| Qwen3.8-27B · NVIDIA DGX Spark public lab · vLLM | Measured · Grade A | 30.009033471349664 tokens/second | StackBench controlled lab approved source qwen38-dgx-spark-owner-archive-20260817 |
What the numbers mean
Decode tokens per second describes generation after the first token for the stated request, while aggregate output rate describes the serving group. Effective prefill is retained only with its proxy label and cannot be read as strict model prompt-processing throughput.
What stands out
The capacity question is comparatively straightforward here: the approved NVFP4 representation is intended for the Spark memory envelope. The harder question is whether the serving stack and latency profile match the workload.
Concurrency changes the meaning of throughput. A C4 or C16 group can improve aggregate service output while an individual request experiences a different decode rate, so these columns are never treated as interchangeable.
The prefill proxy can be distorted by endpoint and cache behavior. For that reason, the page treats strict prefill as unavailable where the collection method did not isolate it, even when a derived prompt-per-TTFT value exists.
Evidence quality / source
The eligible rows come from the owner-approved sanitized controlled-lab projection, bound to the NVFP4 revision and the recorded vLLM build. Raw host identifiers and the held unsanitized archive are not publication inputs.
What is unknown
- The held private owner archive is excluded; only the separately approved controlled-lab projection may support visible claims.
- Serving-derived effective prefill is not strict model prefill, and decode timing retains the documented first-non-empty-chunk approximation.
Workload fit
The configuration is relevant to long-context analysis and interactive generation when a 27B dense model is desired on one compact system. Size alone does not prove acceptable latency; use the protocol nearest the intended prompt and concurrency.
Alternatives
Use the Planner to evaluate alternatives only after pinning the exact model revision, representation, runtime, prompt, output, and concurrency. A nearby hardware label or model family is not substitute evidence.
Planner CTA
Carry this workload into the ComputeSage Planner. It keeps missing evidence visible rather than inventing a recommendation.
Sources / methodology
Review the ComputeSage evidence methodology. Bound selector: NInfer Qwen3.8 NVFP4 C1/C4/C8; public Qwen3.8 DGX evidence. Eligible source: StackBench controlled lab approved source qwen38-dgx-spark-owner-archive-20260817.