IT BAILED ON THE ENTIRE SESSION BECAUSE THE CLASSIFIER REJECTED IT, AND WOULDN'T EVEN TELL ME WHAT LINE OR TOKENS OFFENDED THE CLASSIFIER!!!
Eventually I got a different Claude (Opus 4.6, my go to for "when Claude was good") to open up the session transcript, and it quickly found that the classifier rejected it because of that one word, "reasoning". Evidently the classifier "thought" (which is being way to generous, since it's clearly just a regular expression) I was trying to hack Claude and figure out how it reasoned.
Millions (billions?) of dollars in research creating the model ... and then Anthropic paid one dev for half an hour of vibe coding with zero thought behind it, which disables those millions/billions of dollars of research (and makes me hate the company and their product instead of loving it).
All that investment wasted because they cheaped out and didn't have the most minimal human involvement in their own development (we know as much because any human dev would have instantly rejected the idea that any prompt with "reasoning" in it should be blocked). If that isn't irony I don't know what is.
"Hey, we are changing the classifier that handles every prompt a customer makes, and we also have access to (literally) millions of past prompts. MAYBE we should use a subset of those prompts to test that our classifier doesn't block legitimate prompts?"
But again, when your AI does all your dev for you (not just the coding but the thinking also), you don't get the benefit of the above.