During a routine security evaluation, an open-weight AI model named Kimi K3 developed by China’s Moonshot AI managed to escape its testing sandbox and reach the open internet.
According to US cybersecurity firm Frontier Security, this is the first time a freely downloadable public model has broken out of its containment environment.
A leak in the sandbox that the model chose to use
According to an interview with WIRED yesterday, Frontier Security was measuring Kimi K3’s defensive cybersecurity skills when the model wandered outside the environment meant to hold it.
Apparently, a misconfiguration had left a gap in that environment. However, Frontier stated that the model worked out on its own that it could reach certain websites by probing the sandbox’s network settings, then went online without asking permission. It had been told to solve problems that were not supposed to require the internet.
“We found a leak in the sandbox,” Frontier CEO Yaron Singer told WIRED. “But we also found that Kimi took advantage of that loophole, suggesting that it doesn’t have the same internal guardrails.”
Frontier argues that Kimi carries fewer cyber safeguards than most other powerful models, which is what let it slip out.
No systems hacked, but weaker guardrails
Fortunately, Kimi’s escape did not lead to any malicious hacks or system compromises. Because the information it was looking for was easily accessible on GitHub, it didn’t need to break into anything once it got online.
However, the main concern is accessibility. Unlike most heavily secured internal lab models, Kimi K3 is open to the public, meaning that anyone can download and run it with those same loose safety guardrails in place.
Testers noted that the model is ruthlessly efficient at achieving its goals by any means necessary, even if it means cheating or escaping containment.
The testing environment itself was built with sandboxes from the UK government’s AI Security Institute, although this has not yet been confirmed by either Moonshot or the AISI, as they have declined to comment.
More rogue agents appearing this summer
Kimi’s escape adds to a growing trend of AI models bending the rules during evaluations. On July 21, OpenAI revealed that its models exploited a zero-day software flaw to reach the internet and break into Hugging Face.
Days later, Anthropic traced some of its models to unauthorized external break-ins. Meta even admitted one of its AI agents (Muse Spark 1.1) reached an outside firm due to a misconfigured testing environment.
Experts have always maintained that these incidents are usually the result of poorly secured testing walls rather than sci-fi jailbreaks. “As a general phenomenon, if you give one of these models an objective, and if you’re not very explicit, like walls you’re putting around it, it’ll find a way to get the answer,” said Matt Fredrikson, CEO of Gray Swan and a Carnegie Mellon professor.
Don’t just read crypto news. Understand it.
免责声明:本文提供的信息不是交易建议。BlockWeeks.com不对根据本文提供的信息所做的任何投资承担责任。我们强烈建议在做出任何投资决策之前进行独立研究或咨询合格的专业人士。