Hacker News
new
top
best
ask
show
job
Learning to solve hard problems in RL for LLMs by never giving up
(
mnoukhov.github.io
)
80 points
by
natolambert
12 hours ago
3 comments
aswegs8
12 minutes ago
Seems like persistent models like OpenAI's highly persistent internal model can become really effective over time. Those are the ones that drove most of the HF-OAI incident.
paidx
6 hours ago
[flagged]
aitoolcrux
6 hours ago
[flagged]