3 pointsby antonellofan hour ago1 comment
  • antonellofan hour ago
    Run a 27B reasoning model locally on a 16GB M2 Mac. Ferrox + ternary quantization delivers 5.9GB models without sacrificing much quality.