I’m honestly surprised this is better benchmark wise than the text only model. I figured the addition of vision would take away from some of the text capabilities.
Tangentially this makes me wonder how large Opus really is. Perhaps Opus is a lot smaller than most of the 1T+ assumptions, just a lot more post-training/ finetuning on a 300-400B sized MoE model.