Sitemap

Member-only story

Qwen 3.5 35B-A3B: Why Your $800 GPU Just Became a Frontier Class AI Workstation

10 min readMar 1, 2026

--

I have been running local models for a while now, and I thought I had a pretty good sense of where the ceiling was for consumer hardware.

But when a 35B-parameter model surpasses its 235B-parameter predecessor while activating only 3B parameters per token, it‘s a turning point for developers waiting on a local models that are viable for production workloads.

It’s very impressive to get Sonnet 4.5 level performance out of just 3B activations. That is roughly 8.6% of its total weight.

Press enter or click to view image in full size

On a used RTX 3090 you can grab for around $800, it generates at 112 tokens per second with the full 262K context window, or you can run it on specs like MacBook Air M4 24GB with ~15 tokens per second.

Here are other community reported stats for UD-Q4_K_XL (19.7GB), which fits any 24GB card.

  • 2x RTX Pro 6000 Max Q: ~2,600 t/s
  • R9700 32GB: 128 t/s Vulkan
  • 5090: ~170 t/s
  • 4090: 122 t/s
  • 3090: ~110 t/s

And lastly, Qwen 3.5 MLX 8bit stats on M3 Ultra 512GB:

  • Qwen3.5–35B-A3B-8bit: 80.6 t/s (39.3 GB)
  • Qwen3.5–122B-A10B-8bit: 42.5 t/s 133.6 GB)

--

--

Agent Native
Agent Native

Written by Agent Native

Hyperscalers, open-source developments, startup activity and the emerging enterprise patterns shaping agentic AI.