Hacker News
new
top
best
ask
show
job
Show HN: LLM Inference Calculator – Estimate VRAM, Latency, and Throughput
(
llm-inference-calculator-delta.vercel.app
)
6 points
by
popopanda
6 hours ago
2 comments
maestroquirk
4 hours ago
Need one for vision models too tbh. Token/s doesn't really map easily
popopanda
2 hours ago
Yeah, tokens/s varies a lot based on workload. However, I’ve calibrated the estimator against public benchmarks, so it stays within a 30% error margin!
I'll definitely explore how to estimate vision models next!
popopanda
6 hours ago
[flagged]