Hacker News
new
top
best
ask
show
job
Making models smarter and cheaper at the same time
2 points
by
Wetime
18 hours ago
1 comment
Wetime
18 hours ago
Byte-exact KV grafting stores verified reasoning on disk. Replaying it lets a frozen 12B LLM beat 31B models at 8,700x less energy.