Published · Phase 2
Qwen3.5 on RTX 5090
Qwen3.5-35B-A3B is a practical RTX 5090 model at 4 bit and can deliver very high throughput with optimized kernels. Use it for coding and multimodal agents, but keep context and concurrency bounded by the 32 GB card.
Editorial review: complete · Updated 2026-08-30
Measured / sourced results
| Configuration | Evidence state | Metric | Source |
|---|---|---|---|
| Qwen3.5-35B-A3B · NVIDIA RTX 5090 · llama.cpp | Source-reported · Grade C | 194 tokens/second | llama.cpp discussion #19890 Qwen3.5 GPU comparison |
What the numbers mean
What the numbers mean guidance: Optimized source-reported results can approach roughly 190 aggregate tok/s, while single-user decode will be lower. A conservative planning band is 80–160 tok/s for a tuned 4-bit deployment. Use 4 bit, 8K–16K context, and concurrency one or two. If you need 32K+ context for multiple users, move to 48–96 GB rather than squeezing the cache.
What stands out
Use 4 bit, 8K–16K context, and concurrency one or two. If you need 32K+ context for multiple users, move to 48–96 GB rather than squeezing the cache.
Optimized source-reported results can approach roughly 190 aggregate tok/s, while single-user decode will be lower. A conservative planning band is 80–160 tok/s for a tuned 4-bit deployment.
Qwen3.8-27B is simpler and leaves more cache headroom; Qwen3.6-35B-A3B is the newer coding-oriented choice; DGX Spark supports 8-bit/BF16 experimentation.
Evidence quality / source
Evidence quality / source guidance: Architecture and repository size come from Qwen; the throughput band is an estimate anchored by source-reported RTX 5090 results. Aggregate throughput can include multiple sequences and should not be advertised as one user's token rate.
What is unknown
- Aggregate throughput can include multiple sequences and should not be advertised as one user's token rate.
- Retest output quality when changing quantization or enabling experimental attention/MoE kernels.
Workload fit
Workload fit guidance: Recommended for developers who want a capable agentic/multimodal model on one consumer GPU. It is less suitable for very long-context multi-user service on the same card. Budget about 20–24 GB for a reviewed 4-bit checkpoint. That leaves enough space for a useful cache and one active service; 8-bit or BF16 versions belong on 48–96 GB hardware.
Alternatives
Alternatives guidance: Qwen3.8-27B is simpler and leaves more cache headroom; Qwen3.6-35B-A3B is the newer coding-oriented choice; DGX Spark supports 8-bit/BF16 experimentation.
Planner CTA
Compare this recommendation against your model, context, concurrency, latency, and budget in the ComputeSage Planner.
Sources / methodology
Recommendations combine official specifications, public model/runtime documentation, adjacent benchmark observations, and clearly labeled engineering estimates. Review the ComputeSage evidence methodology.