> Thinking, tool calls, code, etc. remain completely unconstrained.
I wonder how much of these tricks actually work? As a side effect does it make the model produce thinking tokens to remember not to do that, “ohhh wait the user instructions says I should not use PG style in my thinking switching back to my …”