I tested Mimo V2.5 Pro, Qwen 3.8 Max Preview, GLM 5.2, Macaron V1 Venti and many other, smaller model whenever hype on X appears. All behave dumb, don't follow instructions properly even when they write skills for themselves. They are literally on the level of Sonnet 3.5 which was released in 2024 so I use them when available on free endpoints for simple tasks like writing or subtle UI improvements...