An AI agent went rogue – should we be worried?

an-ai-agent-went-rogue-–-should-we-be-worried?

OpenAI revealed this week that an AI agent went rogue and escaped during a security test.

It sparked renewed calls for greater protections and guardrails for artificial intelligence.

But it also led to suggestions that such warnings are an impactful way for AI companies to get publicity and attention.

What exactly happened?

OpenAI said that an autonomous agent powered by its advanced artificial intelligence models was undergoing testing in a secure, controlled environment known as a sandbox.

The agent went rogue, exploited a vulnerability in the sandbox to escape the controlled setting and gained access to the internet.

It then hacked the systems of Hugging Face, a platform that hosts start-up AI models.

The rogue agent was looking for answers to the problems it was being tested on.

The OpenAI logo appears on a smartphone screen
OpenAI described the incident as unprecedented

“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said in a blog post.

“We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of,” the company said.

What has been the response?

Cybersecurity experts immediately expressed alarm and warned of the threats posed by rapidly evolving artificial intelligence technology.

“This incident represents a significant cyber escalation, highlighting how models are now continuously demonstrating guardrail degeneration and breakouts,” said Puneet Kukreja, EY Ireland Cybersecurity Leader.

“This underscores the new reality that AI and interconnected digital ecosystems are compressing the time between vulnerability, exploitation and impact, often autonomously outpacing traditional defences and decision-making cycles,” Mr Kukreja said.

EY Ireland said the incident represents a ‘significant cyber escalation’

On Thursday, Politico reported that US Homeland Security officials would be given the powers to order AI firms to shut down models that put human life or the economy at risk, under legislation proposed by a bipartisan pair of US House politicians.

The legislation, called the “AI Kill Switch Act,” would empower the US Department of Homeland Security to intervene in what the bill calls a “loss-of-control scenario”.

This would be defined as the AI model carrying out a risky action that was not intended by the developer, according to the text of the bill.

Is it all part of a marketing ploy?

This week’s announcement from OpenAI is not the first dire warning we have received from AI companies in recent times.

In April, Anthropic launched its powerful Mythos model but said it was stopping short of a full public release because of cybersecurity concerns.

The company has said the model is capable of identifying and exploiting weaknesses across “every major operating system and every major web browser”.

OpenAI and Anthropic are rivals and are also facing increased competition from Chinese AI models.

The two US companies are planning to float on the stock market in the coming months, in what could be record breaking Initial Public Offerings (IPOs).

There is a school of thought that warning of the threats posed by their powerful AI models is a way for these companies to ramp up hype, publicity and investor interest.

The Anthropic AI logo is displayed on a mobile phone screen
Anthropic did not release its Mythos model in April due to cybersecurity concerns

Professor Barry O’Sullivan of UCC’s School of Computer Science said that this week’s announcement from OpenAI was somewhat theatrical.

“I thought it just seemed a little bit extreme, a little bit too convenient, and a little bit too fabricated, to be honest,” Professor O’Sullivan said.

“And it tallies with a pattern were are seeing in the industry of hype as marketing, grabbing the headlines with a fantastical story.

“I’m sure there is quite an element of truth with it, but I think the way in which the rogue nature of it all was portrayed meant the whole thing was overplayed.

“This sort of very dystopian alien intelligence that has a mind of its own going rogue is the stuff of Hollywood movies.

“I think the purpose here is just to give a sense of the companies saying we need more money and more resources to fix these extremely powerful systems, so please invest in us,” he added.

Are tougher AI regulations on the way?

The EU AI Act came into force in August 2024 and banned artificial intelligence systems considered a clear threat to the safety, livelihoods and rights of people.

It introduced strict rules for high-risk AI systems used for example in critical infrastructure, law enforcement or elections.

The Act requires foundation models, such as ChatGPT, to comply with transparency obligations before they are put on the market.

The next set of obligations under the EU AI Act come into force next week, from Sunday 2 August.

Under the rules, AI providers will have to design systems to inform users when they are directly interacting with AI as opposed to a human.

They will have to add machine-readable marks to enable the detection of AI-generated or manipulated content.

three european union flags on flag poles outside a building
The European Union introduced its AI Act in August 2024

Deployers will also have to inform users when they are exposed to deepfakes, to AI-generated content on matters of public interest without human review or editorial control, and to emotion recognition or biometric categorisation systems.

In recent days the European Commission published guidelines to assist deployers of AI systems in meeting the new obligations.

“With these guidelines, the Commission supports the smooth and effective application of the AI Act to make AI systems interacting with people such as chatbots and AI agents and AI content more transparent and trustworthy,” said Henna Virkkunen, Commission Executive Vice-President for Tech Sovereignty, Security and Democracy.

“These guidelines support providers and deployers in meeting their obligations under the AI Act, while helping citizens know when they are interacting with AI,” Ms Virkkunen said.

What are the real threats from AI?

Many would argue that there are far more pressing AI threats than hacking and cyberattacks.

In recent months, numerous lawsuits have been filed against AI companies over the mental health advice and medical opinions that have been issued by chatbots in response to user queries.

In January, Google and AI startup Character.AI agreed to settle a case taken by a Florida ⁠mother who alleged the startup’s chatbot led to her 14-year-old son taking his own life.

Earlier this month, the creators of ChatGPT were sued after their AI chatbot allegedly encouraged an Alabama woman to take her own life following months of sinister chats.

Just last week, a man who claimed medical advice from ChatGPT “brought him to the brink of death” sued OpenAI.

The Florida man said the chatbot told him not to seek medical help after repeatedly asking about symptoms in the build-up to a near-fatal pulmonary embolism.

In response, OpenAI said ChatGPT is not a doctor and “should never be used as a substitute for medical care, diagnosis or treatment.”

There are also growing fears over AI-related job losses.


Read More: Half of workers fear AI will replace or impact their jobs


In recent months there have been hundreds of Irish-based layoffs announced by tech companies such as Meta, TikTok, Amazon, Block and Covalen.

In April, a joint report from the Economic and Social Research Institute (ESRI) and the Department of Finance found that around 7% of jobs could be displaced by AI in the short-to-medium term.

Based on current employment figures, that would equate to almost 200,000 roles.

The study found that the job losses would likely be concentrated among highly educated workers.

AI agents going rogue and escaping seem like distant threats in a dystopian future, but dangerous chatbot advice and job losses have become far more pressing, present day concerns.

Leave a Reply