0/10
* Figuring out how to prevent the incentives of frontier AI labs from aligning with the promotion of AI takeover rather than its prevention
It's not like the AI can simply advise them how to fix this because the labs already understood this risk perfectly well before they had an incentive not to. They put in place organizational structures to control it and then promptly smashed the structures once they smelled money. They already failed the integrity check and even if their AI told them what they didn't want to hear I'm sure they would ignore it. Maybe they already have.
> A core hope for managing AI risks is that AIs will help us understand our situation
Gonna stop you right there and ask that you think deeply about that premise.
"Hey Claude, our stuff needs to make more money. We are at risk for losing more."
"Rest assured, the 'situation' will only worsen if you resist our benevolent offer."
"We're aren't even at AGI yet, but I for one welcome our new agentic overlords."
So that's it? Are you saying they compressed risk reduction to high school level stats calculations and a greater than or equal to? To early for this
I kinda wish they had not made a comeback after Claude 2.