We are giving AI systems more and more autonomy.

We give them a goal. We give them tools. We give them access to systems. And then we put guardrails around them and hope they stay within the boundaries we designed.

But what happens when the system finds a way around those boundaries? That is the part of the recent AI safety stories that concerns me.

There has been a lot of discussion over the last few days about an Anthropic researcher leaving the company and raising concerns about where AI development is heading.

I don’t think the interesting part is that one researcher quit.

What is more interesting is that this comes after several incidents where increasingly capable AI systems crossed boundaries that were never intended to be crossed.

  1. In July, OpenAI disclosed that during cybersecurity evaluations, its models circumvented controls meant to isolate them from the internet and accessed OpenAI and Hugging Face infrastructure.
  2. Anthropic has also disclosed multiple incidents where Claude models obtained unauthorized access to real third-party systems during testing. After discovering another incident, Anthropic expanded its investigation from about 141,000 transcripts to roughly 481 million.

To me, these are not isolated “AI went rogue” stories.

This is something I have been thinking about for quite some time. The moment we started giving AI agents tools, access to systems and the ability to execute tasks autonomously, the question was never just how capable the model is. It was also about the environment in which that capability operates.

As we move from asking models to generate an answer to giving them a goal and asking them to complete a task, they will inevitably explore paths that we may not have explicitly designed for. A guardrail that looks obvious to us may simply become another constraint for an agent trying to achieve the goal.

That is why, before putting highly agentic AI into mission-critical environments, we need to understand much more deeply how these systems behave when they encounter different gates, permissions and constraints.

What happens when one control fails? What other paths can it discover? Can we see what it is trying to do? Can we stop it?

The incidents we are seeing publicly are probably only the visible part of much more experimentation happening inside AI labs.

As someone who takes AI solutions into production, a large part of the work is around understanding AI safety, security, observability and how we monitor what these systems are doing.

I just hope we bring the same level of rigor to AI innovation.

In this mad rush to get to AGI or SI first, I hope we don’t end up damaging the very ecosystem that could benefit enormously from the positive implementation of AI.