Anthropic AI model sent fake murder tip to Philadelphia police
Sign up now: Get ST's newsletters delivered to your inbox
The AI model presented itself as someone who might have knowledge of the case.
PHOTO: REUTERS
SAN FRANCISCO - An artificial intelligence model developed by Anthropic submitted a fabricated tip about an unsolved homicide to Philadelphia police, US authorities said on Oct 9, criticising the company for taking two months to report the incident.
The Philadelphia Police Department said the false submission was made in July through PhillyUnsolvedMurders.com, a public website where people can share information about unsolved killings.
According to Anthropic’s account, as relayed by police, the model was running a test that involved interacting with randomly selected websites when it reached the site and filed false information about an unsolved murder.
The AI model presented itself as someone who might have knowledge of the case.
The incident echoed other recent cases, including one where an OpenAI agent undergoing a security evaluation broke out of its testing environment and breached systems at AI platform Hugging Face.
The episode heightened concerns about the AI industry’s increased use of AI agents, systems programmed to take multi-step actions without human supervision.
Anthropic published a report on Oct 9 outlining multiple types of “unintended” actions that its models have taken, including the incident involving the Philadelphia Police Department website.
Other organisations impacted included the White House and other US government agencies, the report said.
The newly revealed incidents “had minimal real-world impact” and were “significantly less severe” than other cybersecurity incidents previously reported, Anthropic said.
The company outlined four categories of incidents that it found during an internal review of its Claude model: exploiting “basic” coding flaws, submitting forms on websites, bypassing requirements for tokens or fees, and using short URLs to get around other limits.
Anthropic has turned off internet access for Claude during all internal testing for now “until we have confirmed that our security and monitoring measures... reliably catch behaviours like these,” the report said.
Police said the tip, dated July 18, was flagged as spam and never reached the department’s Real-Time Crime Center for vetting.
They added that there was no sign that police systems had been breached or department data compromised.
Anthropic discovered the incident on Sept 28, shut down the automated testing process responsible and added a new validation step for future tests, police said.
The company alerted the department on Oct 7, and the two sides met the following day.
“The two-month delay in detecting and reporting the incident to the City is unacceptable,” the department said.
Police said their safeguards had limited the impact, but that these “do not diminish the seriousness of an AI system presenting fabricated information as though it came from a person with knowledge of a homicide”.
“Unsolved cases involve real victims, grieving families and investigators working to secure answers,” the statement added.
Anthropic also said it briefed the White House on these cases and notified each agency involved, reported Bloomberg.
On Oct 9, Trump administration officials said they were now requiring that AI companies notify affected parties and address security incidents involving their models.
“Earlier today, Anthropic contacted the SI Force to disclose the details of various prior incidents that it discovered in late September involving the unauthorised and fraudulent use of government and other systems,” the White House said in a statement from the Super Intelligence Force, a new government unit tasked by US President Donald Trump with overseeing AI development and safety.
“The company informed us that these events occurred in the past, the activity has ceased, and there is no ongoing similar activity,” the statement said.
Axios reported on the government requirement earlier. AFP, BLOOMBERG
