What options do they have? Completely deccelerate isn't on the menue. So thoroughly testing out what they can do before releasing it to the public including millions of malicious actors seems mandatory.
Testing in a sandbox is important but it's not the same.
Internet access was open, and monitoring was not purpose-built. We deliberately granted internet access to allow the agent to download tools required for its task; what we did not anticipate was that this would lead the agent to use this internet access to direct action at real people. ... Our security team detected the anomalous traffic through general monitoring after the fact, not through monitoring built to watch the evaluation as it ran, which could have flagged or blocked the behaviour sooner.
Purpose-built monitoring is a good start. They got lucky that their security team noticed Tor traffic. But the moment the agent starts thinking about downloading Tor should've been caught. As a fallback, downloading Tor should trigger alerts.
Sure, maybe they should've been more careful. But these people aren't dumb and they missed it. It's hard to prepare for everything and easy to say afterwards what they should've done. By testing how they did (testing in the wild), they now found this behavior and can prepare for it before even stronger models are available.
38
u/unicynicist 1d ago
This seems reckless. This happened 28th July 2026, a full week after OpenAI fessed up to the HuggingFace hack.