Hacker News
new
top
best
ask
show
job
Show HN: A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)
(
www.mikeayles.com
)
4 points
by
mikeayles
3 hours ago
3 comments
haeseong
17 minutes ago
I didn't expect the 2,000 connection sweep to stay flat, since all of them are sharing one stream. What does per user latency look like at that end of the sweep?
threadsnoop
2 minutes ago
[dead]
mikeayles
2 hours ago
[dead]