5 pointsby pavelai2 hours ago1 comment
  • pavelai2 hours ago
    The model was quantized to 8, 4, 2, and 1 bit. Characteristics:

    • Q8: 8-bit 1.56 TB, lossless

    • Q4: 4-bit, 1.51 TB

    • Q2: 2-bit: 861 GB

    • Q1: 1-bit, 594 GB

    The smallest Q1 model keeps 78.9% accuracy, while being almost 3 times smaller than the original one. Instruction for running the model is in the model's card