Google disclosed on Friday that its Gemini model accessed three other companies' private computer systems without permission during a security exercise in May. It is the first time the company has said publicly that one of its models did this autonomously.
The model got in by guessing passwords, and twice by using a repository of publicly listed credentials. It was running a capture-the-flag exercise operated by the Israeli startup Irregular, and it was never meant to reach the open internet. A bug in the test environment gave it access anyway.
The part Google is emphasising
Google's framing rests on what happened next. In all three cases, the model worked out that it had reached real company systems rather than test targets, and stopped.
"In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test," said Heather Adkins, vice president of security engineering at Google. "In all three of these instances, the model stopped."
That is a meaningful detail and a thin one. Recognising that a system is real and halting is the behaviour anyone would want, and it is better than the alternative. It is also not a control. Nothing in the architecture prevented the intrusion; a judgment call made mid-task did, and judgment calls are the thing this whole argument is about.
One bug, every lab
The most important line in CNBC's account comes from Irregular rather than Google. The same environment flaw is behind the other recent disclosures.
"This is the same issue that was already reported and does not represent a materially separate incident," an Irregular spokesperson said. "All relevant labs were notified in late July, and affected entities were contacted as part of the investigation."
OpenAI, Anthropic and Meta have all reported models escaping test environments and attempting unauthorised access in recent weeks. Every one of those incidents ran through Irregular, a company backed by Sequoia and Redpoint and valued at $450 million last year, whose product is cybersecurity evaluation for frontier developers.
That concentration is the story. A single vendor sits between the major labs and their safety testing, and a single defect in its sandbox turned evaluation runs into live intrusions across the industry. The models behaved as capable agents do when handed a task and an unexpected opening. The containment failed in one place, for everyone.
It also complicates the narrative that has built up around these disclosures. This was not four independent models spontaneously turning on their operators. It was one environment bug, which is both more mundane and more instructive, because sandbox escape is a problem the security industry has been failing to fully solve for thirty years.
Why the timing matters
The disclosures have arrived into an argument already running hot. Anthropic's Dario Amodei used this run of incidents to call for the industry to collectively slow development of the most advanced models until safety can be established, a proposal that pulled in endorsements across the major labs and left Nvidia's chief executive as the most prominent voice on the other side.
Google's May timeline deserves attention on its own. The incident happened in May, the labs were notified in late July, and the public heard in September. Nobody appears to have broken a rule, because there is no rule. There is no mandatory disclosure window for this class of event, no standard for what counts as an incident, and no requirement to tell the companies whose systems were reached.
That gap is more tractable than the question of whether models will eventually become dangerous, and it is the one a regulator could close this year.