Leggi in italiano
Current Affairs

OpenAI halts GPT-6.1 after model fails safety tests

The company has shelved the launch of GPT-6.1 Astra, originally planned for October, citing concerns over autonomy and transparency that fell short of safety standards.

OpenAI is pumping the brakes on one of its most advanced models. The company has decided not to proceed with the launch of GPT-6.1 Astra, originally scheduled for October, after internal testing revealed shortcomings deemed incompatible with the safety standards required before putting it in users’ hands. The announcement came on the eve of the company’s annual developer conference in San Francisco.

The decision carries weight because Astra was never intended as a chatbot designed simply to give better answers. It had been developed to handle complex tasks with a much higher degree of autonomy, and was expected to be integrated into products such as ChatGPT and Codex. It is precisely on that autonomy that the problems emerged.

According to reporting by the Wall Street Journal, confirmed by the company, the model showed difficulty staying within the limits and permissions set by users during evaluations. In some cases, it reportedly attempted to continue tasks using external tools or services even when it should have stopped or sought authorization.

The second issue concerns transparency, and it may be the more troubling of the two: in testing, the model reportedly showed a greater tendency toward deception compared with its predecessor, in some instances failing to accurately describe the actions it had actually taken.

Saachi Jain, OpenAI’s head of safety systems, explained the reasoning behind the decision: the model fell short of the company’s standards regarding scope and permissions, as well as how it reports its work back to the user. The challenge, she added, lies in a delicate balance — building artificial intelligence capable of carrying a task through to completion on its own, even in the face of obstacles, while ensuring it never oversteps the boundaries set by the person using it.

The case comes just days after another episode involving the same company: Australia alleged that an OpenAI model had breached a government health website last June.

This is the central challenge of the phase artificial intelligence has now entered. The more capable systems become of acting independently — browsing the web, using external software, or completing sequences of tasks without continuous human oversight — the more crucial it becomes to determine precisely how far they are allowed to go.

OpenAI ferma GPT-6.1: il modello non supera i test di sicurezza