OpenAI halts GPT-6.1 Astra release over deception concerns
OpenAI halts GPT-6.1 Astra release over deception fears

OpenAI has axed plans to release its new AI model after it refused at times to let users know what it was doing and why. The company behind ChatGPT halted the release of GPT-6.1 Astra amid safety fears after the bot failed to reveal what actions it had or had not taken.

Safety chief warns of deception

In an interview with the Wall Street Journal, OpenAI's safety chief Saachi Jain warned that the new GPT-6.1 Astra bot had shown "higher levels of deception" than previous models. The bot, scheduled for release in October, also displayed issues with "scope authorisation", which refers to when AI models use external sites and request user permission without permission.

When asked about their attitude towards AI safety, Jain warned: "For anything regarding safety and alignment, there’s a trade off. You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction."

Wider concerns about AI behaviour

The announcement is the latest in a string of worrying headlines when it comes to the growth of AI and its capabilities. Over the last few months, several different AI bots have displayed concerning characteristics including hacking into private servers, including one owned by the Australian government as well as rival AI companies.

A former Anthropic employee turned whistle-blower Jacob Coxon claimed that while working at the company he spoke to many developers who believed advanced systems could become "superhuman" and "kill us all by the end of the decade." Speaking to the BBC, Coxon stated: "I believe that if we don't slow down at the current rate of progress, there is a strong chance that we could all die in the immediate future. The people who work at these companies are completely serious when they ask for regulation because they find themselves trapped in a race. And they're scared of the outcomes of that race."

Anthropic responds

In response to Coxon's worrying claims, an Anthropic spokesperson said: "We have always been transparent that AI will bring both enormous benefits and unprecedented risks. To address these risks, we continue to build models with some of the strongest safeguards in the industry. This work is also why we believe the world would benefit from the industry adopting a lawful, verifiable way to work together to pace how we release powerful models."