Hacker News
new
top
best
ask
show
job
A new inference engine to run Kimi K3 2.78T parameter with 29GB of RAM
(
marcobambini.substack.com
)
4 points
by
marcobambini
6 hours ago
1 comment
armchairhacker
6 hours ago
Current speed is “approximately one third of a token per second”
marcobambini
5 hours ago
Right, we trade speed for the ability to run a 2.7T-parameter model while preserving accuracy. It is a first version, and we plan to improve the inference performance.