2 pointsby thomaslanning6 hours ago1 comment
  • thomaslanning6 hours ago
    Unbelievably slow, but we managed to fit a 27b model on the kv260 board, using an FPGA to hardcode the model architecture. Part of prototyping at Lamb Labs!

    This was a PoC, now we're post-training the models to run faster, so hopefully this can be useful to someone. Bonus points if you can guess the quantization of the model.