103 pointsby matt_d3 hours ago7 comments
  • infogulch23 minutes ago
    So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.

    If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.

    • kadushka16 minutes ago
      By “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.
  • om8an hour ago
    Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
    • janalsncm28 minutes ago
      PTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU.

      If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.

      • mitxela5 minutes ago
        which is important though since sending it across the wire over and over and over is actually the main bottleneck.
    • om8an hour ago
      If you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS
  • plqbfbvan hour ago
    Very interesting, I was just exploring this to hopefully fit one of the latest quantized models in 16GB of VRAM.
  • NooneAtAll3an hour ago
    This is the only time "1.58 bit" phrase makes more sense than "1 trit"

    Who knew that if you actually look at information entropy you can pack stuff better!

  • Kevcmkan hour ago
    Woah. Good science.
  • kadushka41 minutes ago
    [dead]
  • kadushkaan hour ago
    [flagged]