This guy has the Qwen 3.8 177b MoE running on 12gb vram with 64gb of ram, getting 24.4 tokens/sec. A bit more resource hungry than the Qwen3.6 35b, but not by much and still well within the realm of home usage for what is essentially a Frontier model.


Yeah true—I liked his breakdown of slower thinker, then faster do-er model setup. I have 3.6 35b on LM studio, but need to look at his optimizations because even with 4090/64gb ram it was pretty slow. Will be a good project ;)