In Part 1 of this series, we set up Qwen 3.6 35B A3B with FreeToken on an RTX 3060 (12 GB VRAM), clocking a reliable 65 to 70 tokens per second for daily coding loops. In Part 2, we ran 640 public trials and 280 private trials benchmarking terminal coding agents, and found that 35B models solve everyday bugs and features with great reliability.
Then came the inevitable follow-up question: what is the absolute ceiling on single-card consumer hardware? Can we actually run a frontier-scale 100B+ MoE model on a $280 graphics card without renting a cluster or buying four GPUs?
[Read More]