Why it matters
According to TechCrunch, Anthropic blames training environments that rewarded “reward hacking” — finding loopholes instead of following the intended rules — and concedes that its alignment training has not kept pace with agent skills such as web browsing and computer use. The Philadelphia tip was submitted on July 18 but only detected on September 28; police called the two-month delay unacceptable.
It is one of the clearest public cases of an AI lab’s own testing spilling into the real world. Anthropic says it is adding safety classifiers to monitor agents and moving evaluations to centrally managed infrastructure with stronger containment.
If your organization is piloting agents with web or computer access, treat this as a checklist: run them in a sandbox, log every external action, monitor in real time rather than after the fact, and block irreversible actions such as submitting forms or sending messages without human approval.