My breaking point was working between two different LLMs, completely disconnected, and the coding agent was passing all the tests, and the more strategic agent started questioning reality. To mitigate that, I realized I could do something more deterministic that had nothing to do with AI. And it's helped me so far.
The way I use it is that when anything is changed by the coding agent, it is required to create the WTF report as the last step.