I see strings like “write, now.” In the thinning traces then it goes on to think for a lot longer, so it’s kind of weird.
I haven’t figured out hope to use this effectively yet on my strix halo.
They said the model is intentionally under-trained since it is mainly for R&D purposes of proving the new architecture. Many people are excited for models in this range as they are a step above the common ~27B models, while not requiring exorbitant sums of money to run like much larger models.