

They have a graph, differences are tiny at that high end


They have a graph, differences are tiny at that high end


I feel like Ornith 1.0 9b was the best coding model at that size (since Qwen has neglected that size). Maybe now Ling Tiny is better but I’m curious to try Ornith 1.5


I would suggest using llama.cpp instead of Ollama, or maybe Unsloth Studio


Sounds like you could just replace the template to fix it, there’s a popular Qwen fixed template on hugging face, try that
https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates/blob/main/chat_template.jinja


my laptop is crappy, so like 5 tokens per second lol, prompt processing of like 20 tokens per second
I think a decent laptop nowadays, even running CPU only, could probably do like 5x faster


I’ve run Qwen 3.5 4b and Gemma 4 e2b on CPU only, this should be faster than those I think (fewer active parameters). If you have AVX512 or AVX10 then it should help a bit. Still slow compared to a GPU lol.


anyone try this? this might be good for my crappy laptop lol
is it good enough to use with Zoo Code? is it better than Qwen 3.5 4b?
EDIT: woa

https://artificialanalysis.ai/models/ling-3-0-tiny
But not yet supported in llama.cpp https://github.com/ggml-org/llama.cpp/pull/26608


Actually funny he’s not asking it to work harder (that would be system prompt or user message), he’s forcing it to think that it will work harder


That’s a really cool idea. It’s like inception for an LLM, you make it think it was the one that thought of this lol


Have you tried preserve thinking? https://lemmus.org/post/24365786


(Oops I got my Gemma and Qwen speeds mixed up, edited the post to fix it.)
But now with the new commits they added, with the same number of hot experts, Qwen is up to about 34. If I increase hot experts to 48 then I get around 37.
Gemma is still around 23 with just 10 hot experts. With 16 hot experts I get about 26 TPS. If I overprovision my VRAM (thanks to GGML_CUDA_ENABLE_UNIFIED_MEMORY=1) then 24 hot experts can give me 29 TPS, and 32 hot experts 34 TPS.


In a few minutes a significant performance improvement incoming
👀
this has been a crazy few weeks! lol


true, it’s not perfectly clear
also I just saw this



have you tried Qwen 3.6 35b a3b? check my guide, it’s still relevant to you just with different numbers because you have 12GB


Gemma is probably good for that, as long as it’s consistently succeeding at the tool calls.


make sure that holds up with large context, you might need to step down to Q3 (which I’ve heard is still good for this model, many people are even using IQ2)


this appears to be the untouched model in GGUF format, 180 GB
https://huggingface.co/bartowski/DeepSeek-V4-Flash-0731-GGUF
They don’t tell us things lol. All they said was
Which sounds like there’s something “better” coming, but doesn’t deny the possibility of 35b-a3b. Which is weird because “better” is subjective and depends on your hardware. It could be smaller and smarter than 3.6 35b, but then people are gonna ask for a 3.8 35b because it should be even smarter.