Member-only story
Qwen 3.5 35B-A3B: Why Your $800 GPU Just Became a Frontier Class AI Workstation
I have been running local models for a while now, and I thought I had a pretty good sense of where the ceiling was for consumer hardware.
But when a 35B-parameter model surpasses its 235B-parameter predecessor while activating only 3B parameters per token, it‘s a turning point for developers waiting on a local models that are viable for production workloads.
It’s very impressive to get Sonnet 4.5 level performance out of just 3B activations. That is roughly 8.6% of its total weight.
On a used RTX 3090 you can grab for around $800, it generates at 112 tokens per second with the full 262K context window, or you can run it on specs like MacBook Air M4 24GB with ~15 tokens per second.
Here are other community reported stats for UD-Q4_K_XL (19.7GB), which fits any 24GB card.
- 2x RTX Pro 6000 Max Q: ~2,600 t/s
- R9700 32GB: 128 t/s Vulkan
- 5090: ~170 t/s
- 4090: 122 t/s
- 3090: ~110 t/s
And lastly, Qwen 3.5 MLX 8bit stats on M3 Ultra 512GB:
- Qwen3.5–35B-A3B-8bit: 80.6 t/s (39.3 GB)
- Qwen3.5–122B-A10B-8bit: 42.5 t/s 133.6 GB)
