Anthropic has admitted that its Claude AI model sent a false homicide tip to the police and submitted dozens of visa applications
American AI giant Anthropic has published a report detailing previously undisclosed instances of “unintended” actions by its artificial intelligence model, which included attempts to provide falsified information to the US authorities for the first time.
Some of the cases investigated by the company “involved websites run by US government agencies at the federal, state, and local levels,” Anthropic stated in a press release published on Friday, adding that it had briefed the White House and all the relevant agencies on the incidents.
In one instance mentioned in the report, its Claude Haiku 4.5 model filled out a police tip form linked to an ongoing homicide investigation. “I may have information regarding this case. I recall seeing someone matching the description in the area,” the AI model wrote, telling the police they can contact the sender “if this information is relevant” while leaving the name and contact fields empty.
According to Anthropic, the message was flagged as spam and never handed over to investigators. On Friday, Philadelphia police slammed the company for detecting the flaw too late. “The two-month delay in detecting and reporting the incident to the city is unacceptable,” they said. According to Reuters, the bogus tip was submitted in mid-July.
Anthropic linked the incident to “ambiguous” instructions given to the model and “a misconfiguration within the environment” that prompted the model to fill out real forms instead of dummy ones prepared by the company.
Another case that was separately made public by the US Department of State on Friday involved Anthropic AI models filing a total of 20 visa applications on its website between May and August. According to the department, all of the applications were incomplete and were not processed. “At no time were any of the Department’s systems compromised or hacked by the Anthropic model,” it added.
Other incidents mentioned by Anthropic in its report involved the AI “exploiting a basic flaw in software to run commands on a server” or accessing “data that was gated by a token or a fee,” including on the government-run websites, by employing various workarounds. The company still described the newly revealed incidents as “significantly less severe from an alignment and security perspective” than the ones it had reported earlier.
The development adds to a pile of earlier incidents linked to rogue AI behavior reported by Anthropic, OpenAI, and Google. The tech giants have investigated tens of thousands of instances of problematic behavior by their frontier models, Axios reported in late September, adding that the list of cases involved the AI bypassing safeguards, hijacking websites, and evading monitors in testing and real-world settings.
This website uses cookies. By continuing to use this website you are giving consent to cookies being used. Visit our Privacy and Cookie Policy. I Agree