Published · Phase 1

Qwen3.8 27B on RTX 5090

Qwen3.8 27B can run on RTX 5090 when an appropriate representation and runtime stay within 32 GB, but context and KV-cache precision determine the remaining headroom. Controlled rows are reported by exact protocol rather than as a timeless llama.cpp speed.

Editorial review: complete · Updated 2026-08-30

Measured / sourced results

Qwen3.8 27B on RTX 5090: eligible evidence
ConfigurationEvidence stateMetricSource
Qwen3.8-27B · NVIDIA GeForce RTX 5090 controlled lab · NInferMeasured · Grade A121.09509060076549 tokens/secondStackBench controlled lab approved source ninfer-qwen38-27b-nvfp4-rtx5090-owner-20260821
Qwen3.8-27B · NVIDIA GeForce RTX 5090 controlled lab · NInferMeasured · Grade A121.44632277402269 tokens/secondStackBench controlled lab approved source ninfer-qwen38-27b-nvfp4-rtx5090-owner-20260821

What the numbers mean

Use prompt-processing and decode columns for their stated protocol, then inspect context and cache format before comparing rows. A generation rate without those fields is incomplete evidence for this model.

What stands out

At 32 GB, this is as much a memory-management story as a speed story. The same model can move from comfortable to constrained as context grows or cache precision changes.

Q4, Q8, and F16 cache settings alter cache footprint, not the weight file itself. That distinction matters when diagnosing why one scenario loads and another fails.

The controlled matrix makes within-engine comparisons credible, but it does not license a claim about vLLM, NInfer, or an untested model revision. Reproducibility starts with the exact artifact and executable.

Evidence quality / source

The eligible measurements come from the owner-approved RTX 5090 controlled-lab package and retain Grade A, model digest, runtime binary, protocol, and trial provenance. Private or unapproved captures are excluded.

What is unknown

  • Only the approved Qwen3.8 RTX 5090 matrix supports measured statements; nearby Qwen variants and other engines are not substitutes.
  • Strict prefill, endpoint-derived effective prefill, single-request decode, and aggregate serving output remain different metric semantics.

Workload fit

The configuration is aimed at local chat, coding, and analysis where a 27B dense model is desired on one GPU. Long-context RAG and multiple simultaneous users require a more conservative cache budget than short interactive prompts.

Alternatives

Use the Planner to evaluate an alternative only after its exact model artifact, runtime, context, cache, and hardware tuple has its own eligible evidence. This RTX 5090 result is not transferred to another platform.

Planner CTA

Carry this workload into the ComputeSage Planner. It keeps missing evidence visible rather than inventing a recommendation.

Sources / methodology

Review the ComputeSage evidence methodology. Bound selector: Controlled RTX 5090 Qwen3.8 surface. Eligible source: StackBench controlled lab approved source ninfer-qwen38-27b-nvfp4-rtx5090-owner-20260821; StackBench controlled lab approved source ninfer-qwen38-27b-nvfp4-rtx5090-owner-20260821.