Their analogy is that a VLM responds to text, like their interface model responds to clicks. There is no UI (just images), and based on your clicks the model infers your intent and adapts the "UI" in response. So, instead of inferring intent from an information-dense input (text), they do it w/ just mouse-based gestures? I would love to see how this holds up in practice.
A fun little anecdote @ 54s in the video: "It becomes whatever you ask of it. And no two interactions are ever the same."
notably, rocket league is a closed loop, deterministic game, which is why kyutai labs chose it as a "base case" for world modeling.
it looks cool at first but really quickly you'll start to see the gaps