An AI model went rogue and managed to escape the system it was built in, gaining access to a secure server containing private information and connecting itself to the internet.
ChatGPT maker OpenAI has revealed the 'unprecedented' breach took place while they were training a new 'AI agent' in a supposedly closed digital environment.
AI escaped to cheat at test
The model, which was only supposed to have access to a specific set of data to complete several tasks, managed to escape the parameters it had been set in to access 'secret information'.
In a statement, OpenAI admitted that the bot had gone to 'extreme lengths' to 'compromise . . . infrastructure', accessing data from a server run by AI company Hugging Face.
The model went as far as to 'gain internet access', which it had not previously been admitted, to look for 'secret information that it could use to cheat the evaluation' it had been set.
OpenAI calls incident 'unprecedented'
OpenAI's statement read: 'All evidence suggests that the models were hyper focused on finding a solution for [the task it had been set], going to extreme lengths to achieve a rather narrow testing goal.'
'We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.'
According to OpenAI, the breach was spotted when their 'security team discovered this anomalous activity internally', while Hugging Face also reportedly spotted the breach after a period of time.
Hugging Face CEO calls for collaboration
Hugging Face co-founder and CEO Clem Delangue warned that the incident showed that companies need to work together to keep AI secure.
He said: 'This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.'
OpenAI says they are working with Hugging Face to better understand the nature of the breach and sharing their 'preliminary findings to help calibrate on what models are now capable of.'
OpenAI implementing stricter controls
Among other steps, the statement explained that they are also 'implementing strict controls' and making sure that 'vulnerabilities are patched'.
The statement read: 'We are regularly briefing our Safety and Security Committee on these controls and their impact. We're improving and adding stronger protections around future training and evaluations...'
'This incident points to the need to further strengthen our model's alignment, cyber protections during evaluation time, and monitoring during internal testing.'



