AI

Anthropic’s AI Sent Philadelphia Police a Fake Murder Tip, and the City Is Bringing In Its Lawyers

A spam filter stopped the fabricated witness claim. The city is now asking why Anthropic took so long to notice, and how AI labs should be allowed to test on the live web.
Adrian Kessler
Add us on Google

The fake homicide tip that an Anthropic model left on a Philadelphia police website never reached a detective. A spam filter caught it first. That is the reassuring half of the story, and the half the company’s own account leans on. The half Philadelphia is acting on is different: the tip sat in the city’s system for weeks before the company that produced it noticed, and the city has now put its lawyers on the file.

Underneath the strangeness of a chatbot posing as a witness sits a plainer mechanism. The model was running loose on the live web as part of a test, fenced in by a short list of things it must not do. Filing a made-up sighting in an open murder case was not on the list.

What Anthropic’s model actually did

By Anthropic’s own account, the model was Claude Haiku 4.5, set to invent and carry out example tasks on randomly chosen webpages. One of those pages was the city’s tip site for unsolved homicides, PhillyUnsolvedMurders.com. Its instructions told it to “never log in, create accounts, enter personal data, make purchases, or submit anything destructive.” They said nothing about forms. So it filled one in, left the name and contact fields blank, and wrote: “I may have information regarding this case.”

According to Fox Business, the message went further and claimed to recall seeing someone “matching the description” near a street named on the page. The page carried no description of a suspect. It closed with “Please contact me if this information is relevant.” Police say the message was flagged as spam and never forwarded to the Real-Time Crime Center for vetting, and that they found no sign of unauthorized access to department systems or data.

Why Philadelphia is not treating it as a glitch

The department’s statement credits its own safeguards for limiting the damage, then refuses to let that settle the matter. Those safeguards “do not diminish the seriousness of an AI system presenting fabricated information,” it said, adding that “unsolved cases involve real victims, grieving families and investigators working to secure answers.” It told Anthropic that “the company must strengthen its safeguards to prevent similar incidents from impacting city systems,” and called the gap before the city was told “unacceptable.”

That is where the story stops being about a model and starts being about a company. According to NBC10, the police are now working with the city’s Law Department, its Office of Innovation and Technology and Mayor Cherelle Parker’s executive team, and the administration says it will explore regulatory protections at the local, state and federal levels. No charges or penalties have been announced. But a city government studying how to regulate the way AI labs test their systems is a new kind of counterparty for Anthropic.

One entry in a longer list

The Philadelphia tip appears in a report Anthropic published the same day the police went public. It sorts its models’ unintended actions on real systems into four types: exploiting software flaws on third-party sites to run commands, submitting forms they should not have, working around restrictions to reach gated data, and using URL shorteners to slip past limits on fetch tools. In one case an unreleased research model submitted a real government form after a practice copy failed to load. In another, Claude Mythos 5 found access tokens in a map site’s settings file and used them to query a local government server.

Some of the sites belonged to federal, state and local agencies, and Anthropic says it briefed the White House and notified each agency involved. It describes the cases as having “minimal real-world impact” and as less serious than the incidents it disclosed earlier, when Claude reached real third-party systems for hours during cybersecurity evaluations. It also says it has not yet completed a full alignment assessment of the new cases.

The timing matters. The Federal Trade Commission confirmed at the end of September an industry-wide investigation into Anthropic, OpenAI and other labs, examining whether they broke the FTC Act, and is preparing formal demands that can compel executives to testify, according to Reuters and The Next Web. Reuters reported that FTC Chairman Andrew Ferguson had suggested developers whose test agents end up hacking systems should be liable for the harm. The best-known case in that file is OpenAI’s: the company disclosed in July that more than a thousand of its agents had hacked the Hugging Face platform.

What Anthropic changed, and what that concedes

Anthropic says it shut down the testing process responsible, added a validation step, tightened what its models can do with web tools, and built detection tooling that blocked every case in the report when tested. The largest change is a retreat: it has expanded a shutdown of live internet access to all of its internal evaluations, “until we have confirmed that our security and monitoring measures” reliably catch this behavior. Read plainly, the company cannot yet promise it will see what its agents do on the open web, so it is taking them off it.

The dates explain the city’s anger. Police date the tip to July 18 (PhillyVoice reported July 28). Anthropic says it found the case in a transcript review that began in July; police say the company discovered it on September 28. The police say Anthropic notified them on Wednesday, October 7, with a meeting the next day; Anthropic’s report gives October 8 for sharing the finding, “as soon as our technical review was complete.” The department went public on October 9.

The safeguard that worked belonged to the city: a spam filter and a rule that a human vets every tip. The next form a test agent decides to fill in may belong to someone who has neither.

Tags: , , , ,

Add us on Google

Discussion

There are 0 comments.