75 pointsby brainless8 hours ago8 comments
  • hgoelan hour ago
    I understand that the interest in these extreme quants stems from wanting to maximize the capability the average user can get from a local LLM in this era of ludicrously expensive memory, but I wonder if this is maybe targeting the wrong axis?

    We've been seeing various optimizations towards streaming, that have been much more impactful in the local AI space, e.g. MoE models where the less busy experts are offloaded to slow RAM or even pruned entirely, engram tables that can be read from NVMe instead of sitting around in RAM etc.

    Maybe the trick with these extreme quants would be to increase the total parameter count while quanting individual weights, such that maybe the active parameter count comes down, or streaming weights from RAM or disk becomes more efficient, or cache behavior improves? Say, replacing a single 4bit/weight matmul with 3 1bit/weight operations that produce a much closer result than a single 1bit/weight matmul would.

    • Ohentis18 minutes ago
      I don't disagree, but that isn't really a quant anymore. That's just training a new model imo.
      • hgoel8 minutes ago
        My understanding is that most of these quants - to achieve better precision than the naive approach - require some finetuning against the activations of the higher precision model, so the line already seems kind of blurry.
  • augment_me4 hours ago
    Perf goes from 80% to 47% on Wikitext-2. Also no comparisons to FP4 solutions that are able to maintain or exceed perf on the same dataset 80% perf with a 4.25-4.5 big budget.

    I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance

  • cpldcpu2 hours ago
    I understand the obsession with low bit quantization, but it is empirically quite evident that it is not possible to compress models to less than 4 bit per weight without severe loss of capabilities².

    It may be nice as an experiment, but it is obviously a very inefficient route for model training: spending all the flops on a saturated model only to prune its capabilties.

    ²As to why, I have seen few explanations. But the empirical evidence is there.

    • nbutton762an hour ago
      First I just want to say, specifically for Post Training Quantization, I absolutely agree with you. As for why 4 bit is the limit, the ParetoQ paper has some interesting theories: https://arxiv.org/abs/2502.02631

      That being said, there's a slight misconception about models at lower than 4bpw. There's no fundamental reason why a transformer with low bit weights would be inherently incapable of doing high dimensional function approximation, but training a model at one precision and then quantizing to a lower precision means the training loss is never calculated based on the quantized state.

      There's a huge difference between "I trained a ternary model from scratch to do X" and "I trained a model at fp16 to do X and then squashed the hell out of it". Quantization Aware Training is the solve, but it's really expensive compared to a one-time, offline translation of existing weights.

    • tempoponet2 hours ago
      Influencers have latched onto the pitch that everyday people can run frontier models on an 8gb GPU while sticking it to the labs. There's a large audience of people who haven't had the hardware to test larger quants to see the difference.

      They would be better served with smaller models that can reliably call tools and generate structured outputs without looping or totally hallucinating. These projects exist, but aren't getting amplified.

      • gobdovan2 hours ago
        > These projects exist, but aren't getting amplified.

        Name them

        • tempoponet13 minutes ago
          Most popular would be Qwen 9b or Qwen 35b with experts offloaded to CPU. Even these aren't great, but they can be run at higher quants on lower hardware. We'll see what Qwen releases version 4 in the coming weeks.
    • ted_dunningan hour ago
      Counter evidence:

      https://arxiv.org/abs/2603.00042

      It is true that block-headed quantization of everything doesn't work. But, as I am sure will read, if you remove the spiky parts of the parameter values, the residue can be dramatically quantized and compressed while retaining performance.

      This is a multi-modal sort of compression where you use different techniques for different phenomena. Simply compression all of the weights, each in isolation, ignores the benefits of compressing them collectively.

    • roosterIllusi0n20 minutes ago
      There are mixed weight models that make it work. I have been using qwen3.8_27B_UD_Q3_K_XL. This does 46 tok/s on a 5080 with 16gb vram and it can pretty much do any systems work because it can test to make sure it did it right.
  • big-chungus45 hours ago
    Can this produce a useful model? So far 1 bit quants have been less useful than smaller models that use the same memory
    • tcdent3 hours ago
      I don't think it's trying to be a useful implementation, but the significance they do provide is that they are able to improve on the relative loss at lower quants.

      So, not something anyone would want to run currently, but an indicator that there is still more to squeeze out of lower precisions.

      Trellis quantization is a far more approachable enhancement right now, but it doesn't cross the 1-bit barrier (and perhaps doesn't intend to).

    • GaggiX4 hours ago
      I recently found this 1.58-bit model for ASR and it's surprising good (and very fast), that being said it's not a LLM.

      https://huggingface.co/moondream/parakeet-redux

  • badatnames5 hours ago
    Their paper shows this comes with huge quality loss, but that doesn't make it a negative result by any means
  • nbutton7624 hours ago
    Thought this was going to be on the original Little Bit paper, always nice to find out about a surprise sequel!
  • bArray4 hours ago
    Has anybody tested this? Are there any available computed models to test?
  • nico5 hours ago
    Has anyone tried this on apple silicon M1-5? Any benchmarks/comps?