4 pointsby Pragmata2 hours ago1 comment
  • Pragmata2 hours ago
    I've been using it for the past few days, and it runs really well!

    I usually get 7 token/s using llama or lm studio, but this inference recipe runs at a smooth 80 tokens per second.

    Genuinely very usable, and fully local!