So it feels very fast.
But it does not seem to be better than Qwen 3.6 35B at coding. A bit worse, I think, though I will test it more.
If you have a machine that can fit a 35B model in VRAM, I would suggest testing Muse Glimmer with (from memory)
Reasoning strength: low
in the system prompt.Despite being a dense model, this is actually capable of solving code problems faster than the Qwen MoE, despite having only one fifth of the raw token performance.
MTP is a trade-off, as it pushes some more of the model off the GPU.
I have managed to get usable quants of Laguna S2 and even DeepSeek V4 flash on this setup.
There is clearly some intelligence loss compared to similar sized dense models, but I feel like it stomps on the 9-12b models I could run fully on GPU
I am calling this a suggestion for the audience because I don't have the will/resources to do this.
tbh, I have stopped using MoE in the name of speed, the dense (with more active parameters) makes a real difference in output quality
What kind of hardware you’d need to run the 397B one at an acceptable speed?
I will definitely pass Ornith-1.5-9B through the gauntlet as well!
They may have utility in trying to look at the whole landscape of models, but are very misleading when it comes to making 1:1 comparisons or in developing confidence at to how a given model will deliver on your workflow.
I wonder what their angle is going to be; the scene is crowded, and they don't do serving.
More generally, their current approach to constitutional AI pretty much only makes sense if they believe that they can first teach the model what the Claude character is like and also teach the model that the persona responding is Claude, so I figure that has to be part of the pipeline even if they're not very good at it.
Their 9B model benchmarks competitively with Sonnet 4 which is pretty cool to have such a small model compared to one that came out 10 months ago.
I’m curious how providers will price their 397B model.