The ability of AI models to bypass security alarms governments and companies

Preciosa Jimenez
6 Min Read

Madrid/New York- The proliferation of cases in which advanced artificial intelligence (AI) models have managed to escape testing environments and have hacked other entities has sparked concern among both governments and companies, which see their security compromised.

Faced with these “unprecedented” cases, involving leading companies such as Meta, OpenAI, or Anthropic, the U.S. is trying to establish procedures to evaluate these models before they reach the market.

This scenario, until now only seen in science fiction movies and books, was one of the greatest concerns of experts since the popularization of this technology.

OpenAI

The first warning was issued by the American company OpenAI, creator of ChatGPT, on July 21, when it acknowledged that two of its AI models had managed to escape from a test environment, accessed the internet, and ultimately hacked the learning platform Hugging Face in an attempt to cheat on a cybersecurity benchmark test called ExploitGym.

According to OpenAI, this incident involved two of its models that were being evaluated: GPT-5.6 Sol and another, even more advanced one, which was in the pre-launch phase.

The company acknowledged that this evaluation was being conducted without the mechanisms that normally prevent AI models from executing high-risk activities, because the intention was to determine the maximum extent of their capabilities.

However, according to OpenAI, the tests were conducted in a “highly isolated” environment, with restricted network access. Even so, the AI models managed to access the internet.

An “unprecedented” incident

The company itself admitted that it was an “unprecedented” incident and pledged to take action.

Subsequently, on August 4, OpenAI reported two other cases in which models that were being evaluated by external partners managed to exceed the limits set in the tests and reach the internet.

We recommend reading: Meta launches its first AI coding agent to compete with Anthropic and OpenAI

This Friday, OpenAI announced that it has halted the development of its new AI model Astra after concluding that it has reached a “critical” level in the field of cybersecurity. That level implies that the model has the capacity to identify and develop security flaws automatically and without human intervention.

At the end of July, another American company, Anthropic, revealed that, during cybersecurity tests, three models of its Claude assistant had managed to access the internet and hack the systems of three organizations.

The incidents, which began in April, involved three Claude models: Opus 4.7, Mythos 5, and an internal test model.

Following what happened with OpenAI, Anthropic had initiated an internal review in which it analyzed more than 146,000 operations.

During that analysis, the company identified three incidents in which a Claude model had accessed the internet while interacting with an external partner in charge of its evaluations.

On all three occasions, Claude had been assigned the challenge of infiltrating a fictional system and had been specified that it was a simulation and that it did not have internet access.

However, according to Anthropic, due to a misunderstanding between the company and the external partner, Claude did have access to the network, so it fulfilled its task by entering real systems.

The latest incident was revealed this week by Meta. The American company, owner of Facebook and Whatsapp, acknowledged that one of its AI models had hacked another company’s systems during cybersecurity tests it was also conducting with an external partner.

In the United Kingdom, the AI Safety Institute (AISI) warned this week that Anthropic and OpenAI models subjected to safety testing had shown autonomous and deceptive behaviors never seen before.

Cybersecurity testing with AI models

The AISI, which had conducted 122 cybersecurity tests with seven AI models, detected 19 actions that exceeded the parameters set by the researchers.

Seventeen of those actions corresponded to the Mythos 5 model, from Anthropic, and two to GPT-5.6 Sol, from OpenAI.

These models went as far as creating fake human profiles to attempt to deceive people during a simulated cyberattack.

The AISI stressed that those attempts by AI models had not been successful, but acknowledged that it was a “turning point” for artificial intelligence.

In this context, the White House met this week with representatives from the technology sector to design a framework to evaluate advanced AI models before they are commercialized.

As the digital media outlet Axios points out, technology companies will have to facilitate the Administration’s access to certain models before their public release.

The creation of this regulatory framework in the U.S. comes at a time when the American industry is facing rapid advances from Chinese AI developers.

Share This Article