310 pointsby anana_9 hours ago33 comments
  • beltsazar8 hours ago
    As a comparison, Qwen3.6 27B scores 38, which was the highest in its small model category (4B–40B).

    Qwen3.8 27B beats all medium models (40B–150B). It has the same score as DeepSeek V4 Flash 0731, which ranks #5 in large model category (> 150B).

    Sources:

    - https://artificialanalysis.ai/models/open-source/small

    - https://artificialanalysis.ai/models/open-source/medium

    - https://artificialanalysis.ai/models/open-source/large

    • phsource8 hours ago
      Simon Willison's post about this gives a good context on why exactly this is happening. While it doesn't mention this in the Artificial Analysis page, this is likely with Max reasoning, which has extremely long reasoning traces:

      https://simonwillison.net/2026/Aug/16/qwen-38-27b/

      It seems like the token usage is 2.3x GPT Luna Max and almost 2x Kimi K3!

      https://imgur.com/a/dDSyhr2

      I'm curious if they can make up for this with insanely high tokens-per-second especially when served from hosted providers, though, given how tiny it is (37B!)

      • thousand_nights3 hours ago
        i feel like Simon omitted an important part of how Qwen's "reasoning" levels work. they are just one sentence additions/omissions to the system prompt

        xhigh -> "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer."

        medium -> no mention of effort (sentence omitted)

        low -> "Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration."

        in my testing this doesn't seem to produce exactly deterministic thinking levels, because it's just a system prompt nudge. i had instances where medium thought longer than xhigh

        • anon3738393 hours ago
          The models are post-trained on these prompt additions so they’re more structural than thinking of them as “system prompts” suggests. (All LLMs ever see is tokens going in, so even the concept of a system prompt is just formatting they’ve seen in post-training.)

          You can also apply fixed token budgets for the reasoning blocks, though it will decrease quality in some cases.

          • nixon_why69an hour ago
            Why not invent a few magic token values for reasoning level instead? It would be like 4 out of a vocabulary of 200k and save like 30 tokens in every prompt
          • thousand_nightsan hour ago
            yes of course, I understand that. but I feel like it would've been nice to include in the article because the main point of it is the effort and overthinking
      • kees997 hours ago
        Qwen models are slower in tokens/s, compared to similarly sized gemma4 and others, and they use more tokens per task, in part thanks to that xhigh default.

        On the other hand, there are some of us who are stuck with hardware that has plenty of compute, but limited (V)RAM. The new 27B is just perfect for that.

        • petu7 hours ago
          > Qwen models are slower in tokens/s, compared to similarly sized gemma4 and others

          No? Gemma 31B and Qwen 27B are about the same speed. Gemma 26B-A4B and Qwen 35B-A3B are about the same speed.

          • trouve_search7 hours ago
            What configuration are you using? On both vllm and llama-cpp, I get significantly higher speeds from gemma4 than qwen3.6 (with their respective speculative decoding methods).

            Output TPS in vllm for instance:

            - Gemma4 26B-A4B: 200-300TPS

            - Qwen3.6 35B-A3B: 120-180TPS

            - Gemma4 31B: 80-120TPS

            - Qwen3.6 27B: 60-80TPS

            This is for a first request on a dual 5090 setup, with their respective speculative decoding methods.

            • mirekrusin2 hours ago
              Dual 4090, getting 85-113 t/s depending on task (draft seems to speed up quite a lot, disproportionately more for content like svg etc):

                ./llama.cpp/llama-server \
                      -hf unsloth/Qwen3.8-27B-GGUF:UD-Q8_K_XL \
                      --webui-mcp-proxy \
                      --no-mmproj \
                      --parallel 1 \
                      --kv-unified \
                      --flash-attn on \
                      --fit off \
                      --split-mode tensor \
                      -ngl 999 \
                      --cache-type-k q8_0 \
                      --cache-type-v q8_0 \
                      -ub 256 \
                      --no-context-shift \
                      --host 0.0.0.0 \
                      --tools all \
                      --jinja \
                      --ctx-size 262144 \
                      --spec-type draft-mtp \
                      --spec-draft-n-max 3 \
                      --reasoning on \
                      --chat-template-kwargs '{"reasoning_effort":"medium"}' \
                      --reasoning-preserve \
                      --temp 1.0 \
                      --top-p 0.95 \
                      --top-k 20 \
                      --min-p 0.0 \
                      --presence-penalty 0.0 \
                      --repeat-penalty 1.0
              
              Use claude/codex/whatever with /goal to optimize params for you.

              IMHO draft model support on dense models is great alternative to MoE on GPUs (high bandwidth, less memory) – more intelligence, speed somewhere mid way there which is usually sufficient.

            • petu6 hours ago
              Single 3090 under llama.cpp:

                | model               |    size |   test |  t/s |
                | ------------------- | ------- | ------ | ---- |
                | gemma4 31B Q4_0     | 16.1 GB | pp2048 | 1248 |
                | gemma4 31B Q4_0     | 16.1 GB |  tg512 |   40 |
                | qwen35 27B Q4_K     | 15.9 GB | pp2048 | 1248 |
                | qwen35 27B Q4_K     | 15.9 GB |  tg512 |   39 |
                | gemma4 26B.A4B Q4_0 | 13.3 GB | pp2048 | 4304 |
                | gemma4 26B.A4B Q4_0 | 13.3 GB |  tg512 |  160 |
                | qwen35 35B.A3B Q3_K | 15.7 GB | pp2048 | 3329 |
                | qwen35 35B.A3B Q3_K | 15.7 GB |  tg512 |  144 |
              
              > with their respective speculative decoding methods

              You're benchmarking drafter acceptance rate, then. Which is real life values, yes, but attributing worse drafter performance to the other 95% of the model being inherently slower.

            • xfalcox4 hours ago
              Have you tried running it on a single 5090? Dual 5090 require https://github.com/aikitoria/open-gpu-kernel-modules for higher perf. Are you using TP?
          • stymaar7 hours ago
            There's no Qwen3.8-35B-A3B though.
            • hadlock6 hours ago
              I benched Qwen 3.6 35B-A3B against Qwen 3.8 27B with the same parameters, thinking set to low. Despite 35B having 9x fewer active parameters, it benched only 2.34x slower. The 35B got only 50% more agentic tasks done per hour.
      • skohan8 hours ago
        I'm running 3.8 27B locally, and the results from the past few days have been excellent. I find raw speed is less of an issue when you can trust the model more to reach the right result.
      • stymaar7 hours ago
        > insanely high tokens-per-second especially when served from hosted providers, though, given how tiny it is (37B!)

        It's a dense model so it will use all of its parameters per token. 37B active parameters isn't tiny at all, it's almost what Deepseek R1 had, and it's 2/3 of what Kimi k3 uses, so it's not going to be “insanely high” tps: it's going to be three times slower than Deepseek Flash (Prefil speed is going to be quite high though, but not token generation).

        • Azantys6 hours ago
          Its 27B not 37B and having just 27B in total and 3T and like 30B active of those is still totally different. A 120B with 5B active is still much slower than a proper 5B. Just like the new Ling 3.0 Tiny with 8B and 1B active only gets around 120tk/s compared to 250tk/s which a real 1B one gets on my hardware.
          • jakswa5 hours ago
            Thanks for mentioning Ling 3 Tiny. This model has completely bypassed me and seems promising for how small it is.
      • 2001zhaozhao5 hours ago
        On the other hand, it used about the same tokens as GLM 5.2 and got 1 point lower score.

        The fact that we have a GLM 5.2-class model that can run on two 3090's comfortably at Q8 is absolutely insane. It wasn't long ago that GLM 5.2 was considered amazing for open weight models.

      • drob5187 hours ago
        It’s still going to chew up context quickly. Surely, some of the added tokens are helping the model, but does it require as many as it generates? What happens on long, multi step tasks as it pushes old tokens out of context? I’m not sure we know the answers to those.
      • ArvidSu8 hours ago
        A ThinkingCap variant of Qwen 3.8 27b would be extremely interesting.

        https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B

        And then a Bonsai ternary on top of that model.

        • kees997 hours ago
          Re: bonsai - unsloth's quants have Q2 (UD-IQ2) variants, which are more or less same in size.

          ...or did Prism do something special with their "bonsai" releases? I didn't notice anything like QAT being mentioned.

          • drob5187 hours ago
            There is some special sauce that they have. It’s not just a simple quant of another release. Or so they imply. I don’t have any insight into how it works or what the Bonsai special sauce is.
    • kzrdude7 hours ago
      It's fun to see "test time" scaling work out so well, maybe the best example of all.
    • bermudi4 hours ago
      This only makes me understand how flawed AAII is. This Qwen model is nowhere close to the other models in that score range.
  • Balinares6 hours ago
    And once again, Qwen 3.8 27B beats Opus 4.6, what the hell.

    It's both funny and a bit terrifying and I still can't quite believe it. It runs decently on a gaming PC! Opus 4.6 came out only 6 months ago and was then broadly considered the new SOTA by a comfortable margin! How in hell did they package capability in the ballpark of a Feb 2026 frontier SOTA into 27B?!

    More importantly, what's the point of building monster-scale data centers on unprecedented amounts of debt when a more than good enough model runs on a GPU from a couple years ago?

    The coming months are going to be exciting, that's for sure...

    • fluoridation4 hours ago
      I'm reminded of that paradox from sci-fi that says that starting an interstellar journey as soon as the technology is capable of it is uneconomical, because the trip will take so long that newer technology will arrive at the destination first, despite departing at a later date.
    • boc2 hours ago
      > More importantly, what's the point of building monster-scale data centers on unprecedented amounts of debt when a more than good enough model runs on a GPU from a couple years ago?

      Probably because the future "monster" models will be insane. 100T+ param models might be the type of things that can independently run a small business, which means anyone not using them is at a distinct disadvantage to their competitors.

      The top model from 2025 looks silly compared to the top model of the first half of 2026. Do you feel like progress has stalled?

    • 4 hours ago
      undefined
  • x3136 hours ago
    I used this a lot over the weekend, and it's a really intelligent and strange model.

    It gets really agentic at the higher reasoning levels. It does the basics like goal tracking and tool calling well, but more than that, it gets obsessed with solving problems and will do insane/unusual things to get to the solution. It actually reminds me of GPT-5.6-Sol-max which is similarly obsessive.

    It doesn't surprise me at all that it outscores Opus 4.6. Opus had way better world knowledge but was more "human" with agent stuff - sort of lazy and uncreative, basically giving up once the obvious solutions failed. These newer models work more like magic, they are so creative and persistent at finding ways to get to the solution.

    • graceful68005 hours ago
      Obsessive is the right word. Over the weekend I had to stop it multiple times deep into a multi-hour long turn to ask what the hell it was doing. It was like a dog with a bone and would NOT let go of its current work to talk to me. I had to interrupt it three times with increasingly aggressive instructions to STOP and answer my questions before proceeding. In another session it straight up told me it was in the middle of debugging something important and to ask later.

      I'm running an RTX 6000 Blackwell. It regularly spent over an hour per turn thinking. Every time I looked at it, the thinking trace seemed coherent, sensible, appropriate. But it could never settle on a solution.

      Granted, I was trying to have it solve a hard problem that 5.6 Sol couldn't solve, but still.

      Either way, I'm still impressed. It genuinely feels better than Sonnet 5

      • tandr25 minutes ago

            > Granted, I was trying to have it solve a hard problem that 5.6 Sol couldn't solve, but still.
        
        Did it solve it?
    • culi6 hours ago
      I have the same reaction reading the internal "thinking" monologues of Kimi K3. When I sent a message that was basically "Nope, I'll just do XYZ instead. Thanks for your help", Kimi basically had an identity crisis. Like there was two wolves inside. One that deeply wanted to help more and go above and beyond and one that was trying to tame the other and make a graceful exit. Here's an excerpt of it

      > Should I verify their README changes? They didn't ask me to. "I've added some notes in the README. Thanks" — that's a closing statement, not a request. Reading the README unprompted to check their notes could be seen as helpful diligence, but they didn't ask for review. Keep it simple: acknowledge, brief close.

    • celrod4 hours ago
      I think I'd rather have the model stop once the obvious solutions failed and ask me. It can suggest more creative ideas, but I don't necessarily want it to try implementing them.
  • K0IN7 hours ago
    I used Qwen 3.6 27B extensively (>1B tokens) and DeepSeek V4 Flash (the older one also 2B+ tokens).

    And I just can't fathom that the new 3.8 beats the new DeepSeek V4 Flash (which, in my eyes, is one of the best everyday coding models).

    What an insane release, and convenient size to use every day/locally.

    but i will test this model extensivly.

    • drob5187 hours ago
      I’ve been using v4 Flash 0731 a lot lately and you can’t beat the price performance. That said, it sometimes takes my prompts as more of a suggestion than a directive. I’ve found that introducing a reviewer subagent (even with the same model) helps push it back to what I’ve asked for. But makes every coding session a back and forth: “do X” -> “use a reviewer subagent to analyze whether you really did X as I asked”.
      • 0xc1333 hours ago
        deepseek-v4-flash-0731 has been awfully prone to infinite looping output in reasoning for me, and once it hallucinated in the middle of going in circles that I had instructed it to start using Yoda-speak, which… I don’t have any idea where that came from.
      • Saris7 hours ago
        What model do you normally run the subagent on? You mentioned flash as well for that, but I wonder if a more 'strict' model would do a better job at pushing the main back on track.
        • drob5186 hours ago
          For cost reasons, I’ve been using Flash for the reviewer, too, but I plan on trying to use Pro for that. Thus far, however, Flash has been doing well at reviewing. I’m cheap as I’m paying for all the tokens myself.
    • f311a7 hours ago
      How is the general knowledge of Qwen 3.6? Do you need to explain things outside of algorithms to it? Since the size is so small, I guess you need more explanations to it. General knowledge helps with coding when your don't specify a lot of details and ask for big changes.
      • SwellJoe4 hours ago
        It researches what it doesn't know, just give it a web search tool. It searches unprompted, if it can. It's impressive.
    • algo_trader6 hours ago
      > I used Qwen 3.6 27B extensively (>1B tokens) and DeepSeek V4 Flash

      Were your opinions effected by the harness ?

      DS is an amazing combo. It probably could only happen in China, not in current USA or EU (for different reasons)

      • ignoramous6 hours ago
        > It probably could only happen in China, not in current USA or EU (for different reasons)

        Per Artificial Analysis benchmarks, Meta's Muse Glimmer 30b (open weight) holds its own (for agentic code workloads) against models 5x to 10x its size, too.

        • skohan5 hours ago
          Glimmer benchmarks around Qwen 3.6 27B levels no?
        • kube-system6 hours ago
          and Muse Glimmer uses a lot fewer tokens than Qwen 3.8. I found it more usable on my hardware because I can get an answer quicker.
          • skohan5 hours ago
            I found Glimmer underwhelming in terms of coding - I tried it as a drop-in replacement for 3.6, and the output was noticeably worse. 3.8 has been a significant step up so far from early testing.
    • JacobAsmuth7 hours ago
      It has double the active params.
      • K0IN7 hours ago
        yeah but using the rule of thumb (I see floating in the Internet) which is sqrt(#params * #active params) which would give sqrt(284B*13B) = 60B, so deepseek should perform better.
        • apitman6 hours ago
          Seems like that equation really should weight total params and active params differently
  • kmike847 hours ago
    I have an internal automated benchmark, which roughly follows my workflow, and I've been testing various models on it, local and cloud. Qwen 3.8 27B did awesome. Its understanding is correct, research is better than e.g. glm's (and I like glm), and implementation is good and careful.

    Qwen 3.8 27B doesn't look benchmaxxed. These "52 AA score" numbers feel real, which is surprising. I've been using it locally for a few days for other tasks as well. If not the speed, I'd be totally happy to use it as a daily driver instead of cloud models, it is that good.

    --- (benchmark, to get an idea):

    1. First, initial prompt which is not super precise - similar to how I'd write a task when talking e.g. to Opus. I'm describing an idea, and asking model to come up with some plan, and also to criticize the approach. Task is about implementing a particular pi extension. I'm checking if a model actually understands what I'm asking.

    2. Then, as a follow-up, I ask to research alternative implementations, research UX of similar extensions, etc. It needs to do web searches, inspect open source codebases, read articles and papers, etc. I don't prompt to do this exactly, but I expect good models to figure out they need to do it.

    3. Then, implementation.

    Also, one finding: Q4 and Q8 seem to have very different behavior in this benchmark. Q4 produces 2-3x thinking in the end, and makes more turns - it seems it makes more mistakes, and needs effort to recover from them, while Q8 gets more things right in a first try. In the end, quality is roughly similar, but Q8 gets there much faster, especially the implementation (tried it several times). Could be a difference between concrete artifacts, or between runtimes, I don't know, but be careful - it seems the real-world experience with qwen 3.8 27B can be vastly different, depending on how it's set up.

    Regarding DeepSeek 0731 vs Qwen 3.8 27B. On this benchmark, Qwen understand my intent better, it's better at research, and I also liked its implementation more. But: if you're more precise in what you ask, 0731 is also very good, and it's quite a lot faster on mac; raw speed is better, and it needs less thinking to get there. So, I'd say it's a tie in practice, both are awesome :)

    • algo_trader6 hours ago
      > So, I'd say it's a tie in practice, both are awesome :)

      Which harness for the benchmark ?

      You have previously commented on using OC/GLM. R u going to stock with it?

      • kmike846 hours ago
        > Which harness for the benchmark ?

        pi, with a plugin to do web search / web fetch.

        > You have previously commented on using OC/GLM. R u going to stock with it?

        For personal use - probably yes, z.ai + kimi + opencode go subscriptions, with some share of local models now. For work - claude code, codex.

    • ignoramous6 hours ago
      > I have an internal automated benchmark ... I've been testing various models on it, local and cloud

      Once you send your benchmark to "cloud", I don't think you can rely on it being secret/private any longer.

      • kmike846 hours ago
        Heh, a good point.
  • padolsey7 hours ago
    The smaller these frontier-nearing models get, the more I'm reminded of https://en.wikipedia.org/wiki/Lottery_ticket_hypothesis
    • keeganpoppen7 hours ago
      i think there definitely is some truth to this in terms of embeddings spaces, which is why i believe they are implemented by OpenAI/Anthropic in roughly highest import => least import bit order-- an overwhelming majority of the variance is in the first few hundred vector bits. i haven't actually tested this myself by manually truncating vectors, but it is my understanding that they generally speaking have this property.
  • anana_9 hours ago
    For more context, this puts it on par with models like GLM 5.2 and GPT 5.6 Luna, which are far larger
    • anana_8 hours ago
      And to read the tea leaves a little:

      3.8 actually performs slightly worse than 3.6 on AA-Omniscience Accuracy, which could imply that they traded out world knowledge for capability in other areas.

      It also produces nearly twice as many tokens per task as 3.6 (and by extension, time), which may be a tradeoff required to achieve correctness at this parameter size.

      • skohan8 hours ago
        Imo it makes sense for things to move in the direction of small, focused models that excel in one area. I use LLMs for technical work 99% of the time, I could care less about general world knowledge, or if the model is good at creative writing.

        With good orchestration and delegation you can get surprisingly far with small models running on consumer hardware.

        • tancop7 hours ago
          The biggest untapped market is pure agentic models that are built for tool calling and non hallucination instead of memorizing facts. You need some world knowledge (as in common sense) to build a useful model, but I don't think perfect recall on general QA is a good use of space when you have web search and structured knowledge in Wikidata or Wolfram Alpha.

          Training should focus on tasks that require real intelligence instead of memory. Creative writing is actually good for this if you score it on coherence instead of getting random real life details right. Basic level of coding (simple prompt to code, don't need to one shot complex projects) is also great because writing a small script is more efficient than 20 separate tool calls.

          • CamperBob24 hours ago
            With respect to Wolfram Alpha, it's worth noting that VibeThinker-3B is basically a match for the larger frontier models -- hundreds of times larger -- in the narrow domain of mathematical and logical reasoning problems. It doesn't seem necessary to resort to external models or tools for that, at least in principle.
        • anana_7 hours ago
          Agreed. Luckily, this model also scores high in AA non-hallucination, so it knows what it doesn't know -- perfect for situations where it can just tool call a web search.
        • drob5187 hours ago
          Yep, exactly. I keep saying that I want the “coding expert” extracted from these multi-T parameter models to run locally on reasonable hardware (large laptops, not servers). Yea, I know there’s no single “coding expert” that you can actually extract in these models, but you get what I mean. Like you said, when I’m coding, I don’t care about world knowledge, and I’m fine with consulting another model when I need that.
    • bertili8 hours ago
      And more context:

      Same score as the latest DeepSeek Flash 0731 which has 284B parameters! (13B active)

      Its also the second best Qwen model, much better than Qwen 3.7 Max, but significantly below Qwen 3.8 Max.

      • anthonypasq7 hours ago
        isnt the active parameter count more relevant than the total? qwen is a dense model no?
        • kzrdude6 hours ago
          For some tasks yes, and we don't know how many active parameters Luna is using..probably less than 27B
    • nsingh28 hours ago
      Also with Qwen 3.8 being more token hungry than Luna, using around 2.3x tokens. Which hurts for local deployment.
      • sottol8 hours ago
        I'm torn on this - on the one hand performance matters, on the other so does capability.

        I could run Qwen 3.6 27B on my laptop, but at 5 tok/s it was too slow even without overthinking - I never used it. OTOH, Qwen 3.6 35B A3B ran at 20 tok/s but it just could not get done what I asked of it. It sort of got close but you had to repeat and retry so much that it might have been faster to run 27B dense... maybe?

        So that said, I might take a much better model that runs 2-3x slower (total time per task) but that's more capable over a faster, less capable one.

        I'd also like to try a proper "plan-then-execute" type execution where thinking is entirely disabled (or low) during the execution stage but enabled/max during the planning stage.

        I will definitely give 3.8 27B a better shot than 3.6 though.

      • skohan7 hours ago
        Depends on your use-case. Over the past couple days, I've found 192k context more than enough for coding. There's more thinking for sure compared to comparably sized models (running on xhigh), but I've found the results are so much better that the entire session consumes less tokens on average since weaker models need more review passes.
    • 8 hours ago
      undefined
    • johnnyApplePRNG8 hours ago
      We don't actually know how large they are, actually.
      • halJordan8 hours ago
        Well, ackshually. E do know exactly how big glm 5.2 is. And there's more than enough data to draw conclusions about luna. Or are you one of the guys who says "big bang is just a theory"?
        • knicholes7 hours ago
          The big bang is a theory. It's not JUST a theory, however.
    • catigula7 hours ago
      Which should tell you how useful these benchmarks are.
  • f311a7 hours ago
    Why is it so small, but expensive?

    Open Router

    Input /M $0.45

    Output /M $3.20

    Cache read /M $0.05

    Throughput 27 tps

    It would be a very nice model at 200-300 tps and if it was dirt cheap. What's the limiting factor of optimizing speed and price for inference providers?

    • AgentLemon7 hours ago
      It's a dense model, so 27B active parameters to compute. Compare that to DeepSeek V4 Flash, which has only 13B active parameters (MoE).
    • freakynit7 hours ago
      I read it somewhere recently that it's architecture does not allow serving as many concurrent requests as the deepseek models allow. Maybe that's why.
      • FuckButtons7 hours ago
        From first principles, 27b dense vs 13b moe, means you spend ~2x more memory bandwidth per request amortized over the whole server (ie, you assume all experts are being concurrently used by some user during the forward pass, then on average the bandwidth required for one forward pass for any individual request is just the size of one expert). deepseek also have some innovations around kv cache and compressed attention which allow for further reductions in memory bandwidth which means that the thing that’s actually bottlenecking inference, (memory bandwidth) is significantly lower than for qwen 3.8, which has been optimized for running 1 instance ~= 1 user.
        • meatmanek2 hours ago
          When processing multiple users in parallel, don't you end up having to load in multiple experts? Not every session is going to use each expert at exactly the same time.
      • Grimblewald2 hours ago
        where'd you read that? sounds like total bs but i could be wrong and would like to learn more.
    • theanonymousone7 hours ago
      That's my questions as well. DeepSeek v4 0731 is served dirt cheap and it needs 10 times more RAM.
      • kmike847 hours ago
        DeepSeek needs more RAM for weights, Qwen requires more compute.

        Also, DeepSeek's KV cache requires less RAM than Qwen's. In concurrent situations (on servers) you load model weights once, but you have different context in each parallel session. So, it can also need less RAM than Qwen to serve, even if it's a larger model.

      • petu7 hours ago
        > and it needs 10 times more RAM.

        More like 3-6.

        Qwen 27B full quality is FP16. So 54GB. In practice most inference providers would serve FP8, so 27GB.

        DeepSeek V4 Flash in full quality is mostly FP4. ~167GB official release.

        So Deepseek has 140GB model size overhead... which is shared between 100s of users single inference node serves, so not even a gigabyte of VRAM per user.

        Memory required for 200K of context per user:

          V4 Flash: 1GB. 
          Qwen 27B: 13GB.
  • ComplexSystems5 hours ago
    China cleaned house these past few months. Kudos to them.

    I would really like to see some open source US companies out there.

    • chr15man hour ago
      Meta's Glimmer is cracked for local inference!
  • josephcooney6 hours ago
    Why are hosting providers charging to much to host it, compared to much larger models? https://openrouter.ai/compare/qwen/qwen3.8-27b/deepseek/deep...
    • sleepyeldrazi6 hours ago
      2 things, 1st: Alibaba's official endpoint pricing. they don't want to undercut too much as there is profit to be made to be close to it but not too low

      2nd, and maybe more importantly: KV is not as efficient (vram usage-wise) as something like deepseek v4 flash. for 256k, fp8 kv is 9.3gb (full precision ~17.3gb). deepseek v4 flash is ~2.5b for the same size at full precision (which is fp4/8, if you are interested in it, read the paper, its pretty cool).

      Doing the math, hosting 27B at NVFP4 (~23gb) with 2.3M total ctx (9 agents) matches the vram usage of ds v4 flash for the same 2.3M ctx (2.3 agents). the break point is 1.5M (6 27B agents) if you use full precision 27B.

      To be clear, the qwen3.5 architecture (what 3.8 uses) is still considered decent in terms of KV efficiency, its just that dsv4f's architecture is SOTA in that space, and with the lower active params, you get better max kv scaling and higher speed serving that, if you have a lot of gpus.

      • josephcooney2 hours ago
        Thanks for the detailed response. I saw some details later in the thread that I had somehow overlooked that touched on this....but anti-procrastination settings prevented me from changing my question.
  • RachelFan hour ago
    I have a bad feeling about this.

    US companies have spent hundreds of billions on their models, and they are not much better than the cheaper open Chinese models.

    Perhaps they can out-compete them. If not there will be increasing calls to limit access to open models on the grounds of "safety".

    Basically, if you can't beat 'em, ban 'em.

    • chr15man hour ago
      I have a good feeling about this.

      Unbelievably cheap and remarkable technology is coming. History repeats.

  • chr15man hour ago
    The leveraged US labs are cooked. Debt's coming home. There will be bailouts.
  • sp19828 hours ago
    Perhaps model size and reasoning length trade off to some extent, similar to CPU vs. RAM. A smaller model with a longer reasoning trace has more intermediate structure to latch onto and build on.
    • deflator7 hours ago
      Makes sense to me.

      We will see, since if true then it is likely the other makers of small, dense models will copy it and include high reasoning by default.

      If that also makes the other dense open source models better, then you are probably correct.

  • colingauvin8 hours ago
    It's 7th (!!!) overall on the agentic index, above Terra.
    • euazOn5 hours ago
      That is actually insane. In my opinion, the future is local AI: for most daily tasks, you absolutely don't need Fable level intelligence - you need Fable level agentic capabilities. And this model has (almost) just that. If we get a similarly capable MoE model in a few months (yes we will), it's going to be an utterly wild ride.
    • hadlock8 hours ago
      Strangely Qwen 3.8 Max isn't on their list, at all.
  • hrmon6 hours ago
    I want to highlight its (1-hallucation rate) at 70%. BRAVO! For me, this is its most wonderful score. GPT-5.6-Sol sits at 8%.
  • bertili8 hours ago
    I can't shake this the existential feeling that this compact series of 27G bytes represent something profound and universal.
    • kzrdude7 hours ago
      Raw model size is around 27x2 GB since it's in BF16 format
  • jakswa4 hours ago
    I've been waiting on this model to show up on the Deep SWE benchmark results and treat its absence/delay as an indication of how slow and unusable it is for good results. I bet it thinks to the moon on some of those complex challenges.
  • JV007 hours ago
    Why is it not included in the Pareto line intelligence/cost chart?
    • leprials7 hours ago
      Theres no official API yet. So theres nothing to price against.
      • ignoramous6 hours ago
        > So theres nothing to price against.

        To my surprise, providers on OpenRouter (io/akash/chutes) are serving Qwen3.8 27B at ~ $0.4 (in) / $3 (out) / $0.25 (cache), more expensive than DeepSeek v4 Flash.

        https://openrouter.ai/qwen/qwen3.8-27b / https://archive.vn/RrDGO

        • JV005 hours ago
          Ah so luna still wins, unless you deploy on your own hardware
      • culi7 hours ago
        Most are running it locally
      • WithinReason7 hours ago
        Just count the cost of electricity then :)
  • dethos7 hours ago
    I'm impressed with the score. This is a model that runs on a good, but still regular, desktop PC.
  • sottol8 hours ago
    A lot of the benchmarks seem often near meaningless these days - really bench-maxxed to the hilt. I tend to still look at the Artificial Analysis rankings to get at least an idea on relative performance of models, is that still warranted?

    What or other opinions on how representative the AA rankings are of real-world performance? Any better indicators?

    • re5i5tor8 hours ago
      Have you tried it? I’d recommend doing so, it’s impressive in real use cases.
    • aqme287 hours ago
      Do you have any evidence that this model is bench-maxxed? I know that's particularly difficult to quantify. If there is an indicator of bench-maxxing, that just becomes the new benchmark to benchmax.
      • deaux7 hours ago
        Here, filtered down for you. [0] Look at the individual benchmarks, not the combined one. You can tell that this model is much more benchmaxxed as its relative ranking swings between benchmarks is much larger. This is a hallmark.

        [0] https://artificialanalysis.ai/models/qwen3-8-27b?models=deep...

        • throwa3562627 hours ago
          I dont think this is benchmaxing.

          They have simply decided to not train the model in some areas such as world physics

      • achrono7 hours ago
        Sounds obvious but just try using the models for anything outside the evals. Take something arcane from Greek history, use it to create a masked linguistic puzzle, which you then ask the model to solve mathematically, all wrapped as an ask to generate ASCII art. Yes, all these elements exist in some form in the evals but the key is in how utterly unconventional the elements are that you pick and in how you combine them.

        I have consistently noticed Opus 4.8 and GPT-5.6 far outshine the Chinese models. Gemini is sort of middle of the road, Grok is better than Gemini but not really close to Opus/GPT. OAI & Anthropic still remain unbeaten by a wide margin in my eyes.

        • skohan7 hours ago
          At that point aren't you just edge-case testing?

          Surely most of your use-cases are not novel tasks that combine obscure domains.

          It seems to me the real way to evaluate the value of a model is how it performs in your real-life workflows.

          • achrono20 minutes ago
            Certainly, real-life is the ultimate benchmark. But for various reasons that isn't always immediately possible to go evaluate a model on.

            Maybe my Greek idea sounded too high falutin' or simply seemingly clever (I give an example below -- try it out!).

            So here's something way simpler that Qwen3.8 27B does not get; only GLM-5.2 and K3 do.

            **

            Analyze the 2 structural (not semantic) patterns in this text:

            Morning revient. Birds saluent Morgenlicht. We suivons Waldwege toward maison. Rain tombe plötzlich; we cherchons Schutz beneath sapins. Night vient langsam; we trouvons Wärme near le Feuer, sharing quelques Geschichten together.

            **

            The answer should get not just the obvious E-F-G pattern but also the word counts being Fibonacci. Surprisingly few models get this. The only way I got Qwen to do this was on the 2.4T model, with extensive prompt scaffolding.

        • CamperBob24 hours ago
          That problem sounds reminiscent of one I like to use as a benchmark, which is to request that the model create an .SVG of a logarithmic spiral of 50 numbered stones. Qwen 3.8 27B absolutely knocked that one out of the park, where a lot of larger models have failed outright or otherwise performed suboptimally.

          Can you share an example of the Greek-history puzzle prompts you're talking about?

    • drob5187 hours ago
      IMO, ELO rating from The Intelligence company and arena.ai are more representative of rankings since they use humans to judge a head-to-head comparison between a couple models at a time. https://www.intelligence.ai/ http://arena.ai
    • Iolaum8 hours ago
      Yea and we are reaching the point where this benchmaxing is visible in the model's reported overthinking.
      • logicchains8 hours ago
        It's not overthinking, it's the right amount of thinking necessary for such a small model to get good results. The dumber the model, the more it has to think to be smart. There's no easy way to reduce the thinking without reducing the model quality.
        • zdragnar7 hours ago
          Qwen doom loops were amusing to watch the first time or two, but it's incredibly vexing to have it waffle over the same decision over and over and over and over again. I can get more done with a faster model by correcting it, and it feels better to babysit them than it does to babysit qwen to see if I need to intervene or if it will actually finish.

          I do like the output from qwen when I get it, but honestly I haven't been impressed enough with it to put up with the downsides.

          • skohan7 hours ago
            It's only been a couple days, but I haven't seen looping issues with 3.8 so far, compared to 3.6 which did occasionally have this problem.
    • deaux7 hours ago
      This one is very benchmaxxed, and you can tell from this page alone. Look at the huge variance in ranking per benchmark. Most models, including at that size, are much more consistent.
  • prakashbuilds7 hours ago
    Interesting to see where local models are going to be in the coming days. I am already starting to believe open source models are the way to go in the coming days. With Qwen 3.8 Max, Kimi K3 etx already delivering at part perf with frontier models, the future is going to be exciting.
  • IronWolve7 hours ago
    Anyone try the 9B/2B distills yet? Wondering how they do for local tools
  • apitman9 hours ago
    Very interesting. I was not expecting anything close to this.
  • manofmanysmiles7 hours ago
    Imagine this, and sucesor models on Cerebras or other silicon...
    • WASDx6 hours ago
      That might actually compensate for the overthinking, if it can think really fast. Dense models are easier than MoE to put on silicon. https://chatjimmy.ai/ is getting 16k tps with an 8B model. Extrapolating that gives nearly 5k tps for 27B. And we're still early in this technology.

      If tps is so high, a compaction step could be performed over every thinking turn to keep context size down.

    • Moduke6 hours ago
      Very exciting indeed. It is in the works. Their current dense offering, Gemma 4 31B, sits at ~1800t/s

      https://news.ycombinator.com/item?id=49308715

  • johnnyApplePRNG8 hours ago
    Unbelievable. Bravo Qwen team.
  • cardboard99268 hours ago
    Where's GLM 5.3 score?
    • jakswa7 hours ago
      I'm waiting for this comparison too. I was impressed by a 1-shot GLM 5.3 did for me the other day.
  • armcat7 hours ago
    So it's effectively on-par with GLM 5.2 and GPT 5.6 Luna?
  • 8 hours ago
    undefined
  • matheusmoreira7 hours ago
    It tied with Luna/max. Simply incredible.
  • marcfrommelious7 hours ago
    [flagged]
  • manunicholasjac8 hours ago
    [flagged]
  • kessler99 hours ago
    [flagged]
  • Lynnr7 hours ago
    [flagged]