It’s 12B active so should hopefully be pretty fast at 2 bit quant if it fits, as in the ratio of total to active params is big compared to e.g. Qwen 3.5 122A10 which seems more commom
Edit: and it’s out - excited to try https://huggingface.co/unsloth/Inkling-Small-GGUF