AFAICT, LLM's do no checking of data fed to them, accepting lies.
Correct me if this is wrong, but IMO, it's a fatal flaw in any software.
I'm just trying to be objective.
These are still token probability machines, and though the models in the last 6 months do more "thinking", they use certain tokens (actually, but wait) to "intentionally" inject path bifurcation so the next most likely token is conversation that questions preceding statements, like statistical BFS/DFS over the token / embedding graph, but you want to make certain path choices closer in equality so both are explored.
I don't believe anthropic or openai are doing anything similar to get their "hack a website" results. The fact that they are prompting dangerous prompts in non-airgapped environments is evidence of at minimum negligence, if not malicious intent on the behalf of those companies.
It's what they do when prompted (both when it is and when it isn't what the user asked for), that we get the most out of them by connecting them to the internet, and that some people are giving them control over robots, that worries people.