my hypothesis is they are using flash and quant'd models to keep up with demand in the hardware crunch
You see the same pattern across open weight models and homelab setups, as we try to fit the models on devices while preserving capability and getting reasonable throughput