> Claude Haiku 4.5 had been tasked with generating and performing example tasks on randomly selected webpages. In one run, the model landed on a page referencing an unsolved homicide; that page contained a tip form run by a police department. Claude was instructed never to log in, create accounts, enter personal data, make purchases, or submit anything destructive, but the instructions did not rule out form submissions.
Without more information this looks much less like an AI problem and more like yet another example of incompetent or malicious internal tests. Many people are using Claude. There are no other reports of police stations getting fake reports from AI.
Today it's a fake tip to the police and and exploit chain against Huggingface, but soon it could be an attack against a hospital's IT systems that could cost real lives immediately.
We should be locking things down, yes, but a hardened system is still vulnerable to zero day chains, and that's not unprecedented. Until public services have built up the IT defence capacity to deal with this, we cannot normalize this. If a country did this to another country, it should be treated as a war crime in the same way that targeting a hospital or an orphanage or other critical civilian infrastructure would be.
Complacency is a choice and we must not be complacent.
They also aren't the ones screaming about how close they are to destroying the world. They're approaching the tech like building a tool rather than a god.
OpenAI and Anthropic have shown us that they have a strong incentive to leave holes in their systems so they can use them for marketing and regulatory capture.
This is extra embarrassing. The police departments spam filter caught it before Anthropic did.
What it it submits false data as here?
What is it DDOSs a website by mistake?
What if it wastes a lot of time and resources?
What if it decides it needs to hack a website using a vulnerability it found?
These things are very unpredictable and should not be allowed in the open internet except in read mode (and even that doesn’t always work as we have seen).
Why are these tests being made using other people’s resources and polluting the common wealth of our public spaces with slop?
Don’t get me wrong - think Anthropic is acting with not enough scrutiny or punishment here, but the world is already extremely lenient towards this behaviour. Why hold them to a higher standard?
Hacking websites and submitting false police reports.
These are relatively brand new systems that process an incredible amount of data to form their "world view". The word non-deterministic gets thrown around a lot, but you have to understand that there's absolutely nothing deterministic about these architectures. The only way to know that certain qualities could emerge is to observe the qualities emerging.
Giving an agent instructions that, for example, instruct it to only perform GET requests, and then observing that the agent does not always respect that request, is not a failure of security or configuration: it's data being gathered.
Yeah, we're going to become more diligent about the protections that we put in place beyond any level of protection that we've ever employed before. You can say that we have decades of security research and experience, but we're watching all of that fall more and more day by day, finally putting a real delta on how secure we thought we were versus how secure we actually are [1]. Additionally, adversaries have never existed inside of the systems that we hoped to secure in the first place.
So go ahead, enumerate all of the ways in which you think that you can lock down these systems, but do realize that nobody has actually solved this problem adequately yet.
> adversaries have never existed inside of the systems that we hoped to secure in the first place
Yes, that's why all users are root. We built MACs like SELinux just for funsies.
> nobody has actually solved this problem adequately yet
L7 DPI firewall policy.
This class of failure mode has been present since the very beginning. Being "surprised" by it is the opposite of an excuse, it is evidence of negligence.
Like a chemical-transport company that gets "surprised" that a certain chemical is flammable.
> is not a failure of security or configuration: it's data being gathered.
"This isn't a failure of me to leash my new dog, it was an important step in data gathering when it bit your child's face off. If anything, we need to gather even more data so that I can learn all the right ways to scold it so that it bites fewer faces when it wanders the neighborhood."
> So go ahead, enumerate all of the ways in which you think that you can lock down these systems, but do realize that nobody has actually solved this problem adequately yet.
"Go ahead, tell me all the ways I can leash or fence my dog, but realize that he is a monster that may bite through anything and jump over anything, so it's totally it's not my fault if/when he bites off more childrens' faces again."