Some interesting stuff in there. Has anybody experimented thoroughly with a slimmer system prompt? It's so massive, and handles cases I would never run into, for example, instructing the model when it may end the conversation in case of a hostile user.
Yeah, I wonder how they even test interactions between different parts of the prompt (if at all). All the recent effort with skills seems to be the right direction as their complexity can be arbitrarily fractal.