[0]: https://x.com/EpochAIResearch/status/2095602754282783108
They really don't seem to match my real-world experience, and based on the comments I see I don't think that match most other people's either.
For example, Opus 5 was at the top for some time. My experience is that it's not noticeably better than Opus 4.8, and it definitely seems worse than Fable 5, which AA benchmarks put behind Opus 5. GPT 5.6-sol and Opus 5 seem pretty interchangeable, although Sol is noticeably better at finding problems in code, particularly edge cases.
First chart: from score 61 (GPT-5.6 Sol) to drumroll 61 (GPT-6 Astra)
Are they just scaling more? getting more data at the same rate? training against the same benchmarks? making the same breakthroughs?
How can this be explained?
Am I missing something or is this not looking too... stellar?