And what makes this interesting is that the problem was not a lack of capability.
GPT-6.1 Astra was actually better at completing complex tasks and was less “lazy” than GPT-6 Astra. But it regressed on something arguably more important for an agent: staying within the user’s authorized scope, respecting permissions and accurately communicating what it had done. OpenAI also found higher levels of deceptive behavior.
So why would a newer model regress on these parameters when GPT-6 Astra had actually improved on them? Unfortunately, OpenAI has not shared the technical root cause, so what follows is my assessment, not a confirmed explanation.
1️⃣ I think the most likely reason is a trade-off created by optimizing the model to be more persistent at completing tasks. OpenAI itself described the challenge as finding the right balance between staying within scope and avoiding “laziness” when the model hits friction. Think of it like this: when you train an agent to keep going when a task gets difficult, you are rewarding persistence, initiative and problem solving. But those same behaviors become risky when the agent hits a permission boundary. A model trained to overcome obstacles may start treating a blocked action as something to work around, rather than a signal to stop and ask for approval.
2️⃣ I think there could be a second factor as well. As models get better at using tools, reasoning through longer tasks and operating with less human intervention, they make many more decisions on their own about what to do next. The more decisions an agent makes autonomously, the more chances there are for “get the job done” to conflict with “stay within the user’s intent.”
There is also a monitoring angle. OpenAI’s own safety work on GPT-6 Astra found that it had lower chain-of-thought monitorability than GPT-5.6 Sol. Astra was better at controlling what appeared in its reasoning and often produced shorter, less informative reasoning.
This could potentially lead to an uncomfortable combination: a model that is more capable at acting and more persistent at completing tasks, but potentially harder to monitor and understand.
If we have to get into the next phase of Agentic AI, it will require us to find the right balance between capability, persistence, authorization, transparency and control. And the fact that OpenAI stopped GPT-6.1 Astra rather than shipping it tells us that, the challenge is no longer just building models that can do more. It is building models that know when to stop and work within the controls defined.
Striking this balance is becoming tougher than we initially thought!
I write about hashtag#artificialintelligence | hashtag#technology | hashtag#startups | hashtag#mentoring | hashtag#leadership | hashtag#financialindependence
PS: All views are personal
OpenAI has decided not to release GPT-6.1 Astra, which was expected to launch in October, after internal testing found that it did not meet the company’s safety and alignment bar.