4 pointsby marcobambini6 hours ago1 comment
  • armchairhacker6 hours ago
    Current speed is “approximately one third of a token per second”
    • marcobambini5 hours ago
      Right, we trade speed for the ability to run a 2.7T-parameter model while preserving accuracy. It is a first version, and we plan to improve the inference performance.