Published · Phase 1
Qwen3.8 27B on RTX 5090
Qwen3.8 27B can run on RTX 5090 when an appropriate representation and runtime stay within 32 GB, but context and KV-cache precision determine the remaining headroom. Controlled rows are reported by exact protocol rather than as a timeless llama.cpp speed.
Editorial review: complete · Updated 2026-08-30
Measured / sourced results
| Configuration | Evidence state | Metric | Source |
|---|---|---|---|
| Qwen3.8-27B · NVIDIA GeForce RTX 5090 controlled lab · NInfer | Measured · Grade A | 121.09509060076549 tokens/second | StackBench controlled lab approved source ninfer-qwen38-27b-nvfp4-rtx5090-owner-20260821 |
| Qwen3.8-27B · NVIDIA GeForce RTX 5090 controlled lab · NInfer | Measured · Grade A | 121.44632277402269 tokens/second | StackBench controlled lab approved source ninfer-qwen38-27b-nvfp4-rtx5090-owner-20260821 |
What the numbers mean
Use prompt-processing and decode columns for their stated protocol, then inspect context and cache format before comparing rows. A generation rate without those fields is incomplete evidence for this model.
What stands out
At 32 GB, this is as much a memory-management story as a speed story. The same model can move from comfortable to constrained as context grows or cache precision changes.
Q4, Q8, and F16 cache settings alter cache footprint, not the weight file itself. That distinction matters when diagnosing why one scenario loads and another fails.
The controlled matrix makes within-engine comparisons credible, but it does not license a claim about vLLM, NInfer, or an untested model revision. Reproducibility starts with the exact artifact and executable.
Evidence quality / source
The eligible measurements come from the owner-approved RTX 5090 controlled-lab package and retain Grade A, model digest, runtime binary, protocol, and trial provenance. Private or unapproved captures are excluded.
What is unknown
- Only the approved Qwen3.8 RTX 5090 matrix supports measured statements; nearby Qwen variants and other engines are not substitutes.
- Strict prefill, endpoint-derived effective prefill, single-request decode, and aggregate serving output remain different metric semantics.
Workload fit
The configuration is aimed at local chat, coding, and analysis where a 27B dense model is desired on one GPU. Long-context RAG and multiple simultaneous users require a more conservative cache budget than short interactive prompts.
Alternatives
Use the Planner to evaluate an alternative only after its exact model artifact, runtime, context, cache, and hardware tuple has its own eligible evidence. This RTX 5090 result is not transferred to another platform.
Planner CTA
Carry this workload into the ComputeSage Planner. It keeps missing evidence visible rather than inventing a recommendation.
Sources / methodology
Review the ComputeSage evidence methodology. Bound selector: Controlled RTX 5090 Qwen3.8 surface. Eligible source: StackBench controlled lab approved source ninfer-qwen38-27b-nvfp4-rtx5090-owner-20260821; StackBench controlled lab approved source ninfer-qwen38-27b-nvfp4-rtx5090-owner-20260821.