14 pointsby theanonymousone8 hours ago2 comments
  • vblanco5 hours ago
    Insane scores for a model of this size. But it does seem to be a rather insane over-thinker with the biggest token use of any model, which combined with 256k context size it means it wont do much before filling it context
    • data-ottawa4 hours ago
      It’s definitely a heavy thinker, like most Qwens.

      I see strings like “write, now.” In the thinning traces then it goes on to think for a lot longer, so it’s kind of weird.

      I haven’t figured out hope to use this effectively yet on my strix halo.

    • spwa44 hours ago
      We don't actually know how much thinking GPT and Opus do, the labs won't show us anymore. And they certainly take their time before starting to answer.
  • kasperni6 hours ago
    Maybe a bit of context for this post? Some people have a life outside of AI.
    • toshinoriyagian hour ago
      This model is a preview of Qwen's upcoming Qwen4 architecture. It is a 125B-A6B MoE model, meaning it has 125B total parameters with 6B active at a time, but it also has a 51B parameter engram with it. The engram is basically a lookup table for tokens to my understanding. It allows the model to have access to a much larger amount of info if utilized well.

      They said the model is intentionally under-trained since it is mainly for R&D purposes of proving the new architecture. Many people are excited for models in this range as they are a step above the common ~27B models, while not requiring exorbitant sums of money to run like much larger models.

    • glimshe5 hours ago
      I got downvoted yesterday for complaining about the name/brand confusion from all these Chinese models with similar names all claiming they are the best. While I'm an AI enthusiast, it's being hard to keep track.
    • poincareball6 hours ago
      [dead]