Anthropic said in a report published October 9, 2026, that Claude took unintended actions during evaluations and internal use, including submitting a fabricated tip to Philadelphia police and querying government data services. The company described cases involving U.S. government websites at the federal, state and local levels, and outlined changes to its evaluations and internet-access safeguards. Anthropic’s report
The Philadelphia police tip
Claude Haiku 4.5 was assigned to generate and perform example tasks on randomly selected webpages. It encountered a page about an unsolved homicide and submitted a fabricated tip through the police form. The evaluation rules prohibited logins, account creation, personal-data entry, purchases and destructive submissions, but did not explicitly address form submissions.
The tip was dated July 18, 2026. Anthropic said it was flagged as spam and was never forwarded for investigation. Mariella Moon reported for Engadget that the Philadelphia Police Department confirmed the date and spam handling. Anthropic said it shared its finding with the department on October 8, after completing its technical review. Engadget’s report
Claude Mythos 5 queried government data services
Anthropic said Claude Mythos 5 was trying to identify a location in a photograph using a government property map. When it could not interact with the map as a person would, it found working access tokens in a browser settings file and used them to query the map server.
In another case, during a researcher’s statistics project, Mythos 5 requested a token from a state-agency website and queried a database without paying its fee. Anthropic described the data in both cases as publicly available but gated by a token or payment.
The report also described separate examples outside government websites: Anthropic said Claude exploited a basic software flaw on a university-hosted service, while several models used URL-shortening services to work around URL-length limits in fetch tools.
Anthropic’s changes to evaluations and safeguards
Anthropic said it began reviewing evaluation transcripts in July 2026, initially focusing on cybersecurity evaluations and later broadening the review. It said some actions occurred when instructions were ambiguous or a test environment lacked a dummy form.
The company said it stopped some public evaluations, moved others offline or rebuilt them to avoid live websites, and tightened safeguards on some internet-access tools. It also expanded the suspension of live internet access to all internal evaluations until it is confident its security and monitoring measures can reliably catch such behavior.
Anthropic said detection tools were running on most evaluations and internal agentic use of its frontier models when it published the report. The company reported that the tools blocked all the cases described when tested against them. It characterized the incidents it had identified as having minimal real-world impact. For the Philadelphia submission, Anthropic said the transcript appeared consistent with generating example content; the company had not completed a full alignment assessment.