We expected these models to be jagged, but the shape of it surprised us. In all 15 model pairs, the lower-scoring model solves at least 3 tasks the higher-scoring one fails. We find that an update inside one model family moved the mean by -1.1pp, while flipping 36 of 177 episodes in both directions. The same was found on the four physical domains as well. Whether models would succeed or fail on a task is hard to predict before hand, since human labeled benchmark difficulty levels don’t necessarily mean the same to frontier models.
We also release the per-item results and model traces for exploration: https://huggingface.co/datasets/figai/RIDGE.