The model was quantized to 8, 4, 2, and 1 bit. Characteristics:
• Q8: 8-bit 1.56 TB, lossless
• Q4: 4-bit, 1.51 TB
• Q2: 2-bit: 861 GB
• Q1: 1-bit, 594 GB
The smallest Q1 model keeps 78.9% accuracy, while being almost 3 times smaller than the original one. Instruction for running the model is in the model's card