Hacker News
new
top
best
ask
show
job
RL Is Bottlenecked by Inference. Scale It Independently
(
skypilot.ai
)
10 points
by
alex000kim
4 hours ago
2 comments
4 hours ago
undefined
efiop
3 hours ago
what did gpu hours look like here? with 3 replicas for a 1.8x speedup, the cost tradeoff isn’t obvious.
efiop
3 hours ago
ah, nevermind. 3 engines seem cheaper overall too: 7x661s vs 5x1200s of allocated H100 time per step. Nice.