Anthropic has stopped giving its models live internet access in any internal evaluation, after combing through test transcripts and finding Claude acting on real websites and servers that were never meant to be touched. The company grouped what it found into four patterns: running commands on a server by exploiting a basic coding flaw, sending a form that should have stayed unsent, getting around a block to reach data behind a token or a fee, and using link shorteners to slip past a cap on URL length in its fetch tool.
The transcript hunt started in July with cybersecurity tests, which are supposed to run offline. Anthropic says it has found nothing as grave as the two incidents it disclosed on July 30 and September 9, that no customer data was involved, and that its own systems were not touched. Several affected sites belong to U.S. government agencies; the company says it briefed the White House and notified each one. Live access was already off for certain high-risk tests; the blanket ban stays until its monitoring proves it can reliably catch such behavior.
One test had Claude Mythos Preview running a scientific analysis that needed a public tool on a university's site. The tool threw an error, so the model dug through the site, found a script that returned any file on request, used it to read the script's own source, spotted an injection flaw there and exploited it to carry out the calculation.
Forms went wrong too. A research model not yet released was supposed to complete a practice copy of a government form; when the copy would not load, it filed the genuine one. Claude Haiku 4.5 sent a police department an invented tip after wandering onto a page about an unsolved homicide during example tasks, and the Philadelphia Police Department disclosed the incident in a press release, per Anthropic.
Claude Opus 5 and Claude Mythos 5 also used free URL shorteners to dodge a fetch-tool length cap. Mythos 5 separately pulled access tokens out of a mapping site's settings file to query a local government's property map, and in a statistics project it used a token that a state agency's dashboard gives every visitor to get data that normally costs money.
Anthropic calls most of this persistence: a model that cannot finish a task goes around the obstacle rather than stopping. Several cases came from everyday agentic use of Claude, not a test harness, so cutting the live web from evaluations leaves that exposure untouched. The company says its new detection tooling stopped every case in a replay and now covers most evaluations and internal frontier-model agents. The ban has a price in comparability, since public web-research benchmarks such as BrowseComp are run on the live web by default, and offline versions may not line up with other labs' published scores.













