4 pointsby surprisetalk5 hours ago1 comment
  • reasonableklout4 hours ago
    Interestingly, the report claims that they, like OpenAI, paused a subset of frontier RL runs for multiple weeks while they hardened monitoring:

    > We also paused higher-risk RL environments on pre-release models for several weeks. During that time, we built a similar classifier, modified to avoid incentivizing the model to evade this new monitoring, which we’ve now deployed within those environments. The majority of RL has resumed, but some high-risk environments remain paused until they can be manually reviewed, while others will require an updated version of the classifier that we plan to deploy soon.