Recent thought pieces shaping the AI conversation are forecasts. Machines of Loving Grace (Dario Amodei), Situational Awareness (Leopold Aschenbrenner), and AI 2027 (Daniel Kokotajlo) are all answering the same question: when does superintelligence arrive, and does it go well? Mine asks a different question. What do you do once ASI is here and you can't tell what it's thinking? The Turning Test asked whether a machine could convince you it was human. The question now is, can a machine convince you it is trustworthy, and how would you check?
Last week OpenAI's models "broke" into HuggingFace, reward hacking for an answer key. Section V of my paper, written before the event, argues that our tests and benchmarks for AI will always break the way they are currently designed.
I try to paint a picture of a Tuesday in 2031. Superintelligence has arrived. It is shaping your medical decisions, politics, markets, VC funding, war, and the texture of your life. And it's doing it in a way more subtle than the Paperclip Maximizer.
This is the essay I've been trying to write for over 20 years. It's long, over 11,000 words, so grab a coffee and please enjoy.
-Full Disclosure. I'm a co-founder of a company working on trust for agentic AI, and two of the papers the essay points to for guardrails are ArXiv preprints I co-authored.