It is much faster than MTP (around 2x). Also if you have read their paper which contains lots of really cool ideas, they clearly documented their limitations and did a good comparison with existing methods. You can run it easily on a strix halo (albeit quantized) and the peformance is fast given how quantized it is.
At least they published their research and went into great length for their novel discoveries, much bettter than any other “frontier” lab did. Them, GLM and Kimi are doing great, but deepseek hands down have the most cool and fun innovative papers in this field.
Breakthrough hardly. Amazing how the article specifies cloud inference while the OP mentions local
Similar or slightly better than MTP using the same tech
It is much faster than MTP (around 2x). Also if you have read their paper which contains lots of really cool ideas, they clearly documented their limitations and did a good comparison with existing methods. You can run it easily on a strix halo (albeit quantized) and the peformance is fast given how quantized it is.
At least they published their research and went into great length for their novel discoveries, much bettter than any other “frontier” lab did. Them, GLM and Kimi are doing great, but deepseek hands down have the most cool and fun innovative papers in this field.
“Novel” and “groundbreaking” comment. You’ve made your 5 cents
people are saying 60-80% speed up?
Deepseek reports this number