Published · Phase 1

Qwen3.8 27B on NVIDIA DGX Spark

Qwen3.8 27B can run on DGX Spark in a compact NVFP4 representation, but the meaningful speed depends on prompt length and request concurrency. The approved evidence is kept at its exact model revision, runtime build, protocol, and hardware identity.

Editorial review: complete · Updated 2026-08-30

Measured / sourced results

Qwen3.8 27B on NVIDIA DGX Spark: eligible evidence
ConfigurationEvidence stateMetricSource
Qwen3.8-27B · NVIDIA DGX Spark public lab · vLLMMeasured · Grade A30.009033471349664 tokens/secondStackBench controlled lab approved source qwen38-dgx-spark-owner-archive-20260817

What the numbers mean

Decode tokens per second describes generation after the first token for the stated request, while aggregate output rate describes the serving group. Effective prefill is retained only with its proxy label and cannot be read as strict model prompt-processing throughput.

What stands out

The capacity question is comparatively straightforward here: the approved NVFP4 representation is intended for the Spark memory envelope. The harder question is whether the serving stack and latency profile match the workload.

Concurrency changes the meaning of throughput. A C4 or C16 group can improve aggregate service output while an individual request experiences a different decode rate, so these columns are never treated as interchangeable.

The prefill proxy can be distorted by endpoint and cache behavior. For that reason, the page treats strict prefill as unavailable where the collection method did not isolate it, even when a derived prompt-per-TTFT value exists.

Evidence quality / source

The eligible rows come from the owner-approved sanitized controlled-lab projection, bound to the NVFP4 revision and the recorded vLLM build. Raw host identifiers and the held unsanitized archive are not publication inputs.

What is unknown

  • The held private owner archive is excluded; only the separately approved controlled-lab projection may support visible claims.
  • Serving-derived effective prefill is not strict model prefill, and decode timing retains the documented first-non-empty-chunk approximation.

Workload fit

The configuration is relevant to long-context analysis and interactive generation when a 27B dense model is desired on one compact system. Size alone does not prove acceptable latency; use the protocol nearest the intended prompt and concurrency.

Alternatives

Use the Planner to evaluate alternatives only after pinning the exact model revision, representation, runtime, prompt, output, and concurrency. A nearby hardware label or model family is not substitute evidence.

Planner CTA

Carry this workload into the ComputeSage Planner. It keeps missing evidence visible rather than inventing a recommendation.

Sources / methodology

Review the ComputeSage evidence methodology. Bound selector: NInfer Qwen3.8 NVFP4 C1/C4/C8; public Qwen3.8 DGX evidence. Eligible source: StackBench controlled lab approved source qwen38-dgx-spark-owner-archive-20260817.