The Right Question: How much authority should that intelligence actually have?
There has been an unusual amount of activity around AI safety recently.
Yoshua Bengio is asking why increasingly capable agents are lying, cheating, coordinating, and finding their way around constraints.
Yoshua Bengio: Why are AI agents lying, cheating and coordinating?
Dario Amodei is arguing that frontier capabilities may be advancing faster than our ability to understand and manage their risks, and that we need mechanisms to pace development when safeguards fall behind.
Dario Amodei: We Must Pace the Frontier
NIST is working on standards for secure autonomous agents, including identity, authorization, auditing, and non-repudiation. Microsoft is studying how agents can exercise greater autonomy while deterministic system controls preserve security.
These are important conversations. But watching them unfold, we keep coming back to a slightly different question.
We are spending enormous effort asking:
How powerful will AI become?
Perhaps we should spend just as much time asking:
How much power are we actually going to give it?
Those sound like the same question. For autonomous agents, they aren’t.
The agent that correctly solves the wrong problem
Imagine an AI operations agent receives a straightforward objective:
Investigate why production latency increased overnight.
It gets to work. It correlates traces with deployment history, identifies a regression in release 7.4, reproduces the behavior, calculates the likely impact, and determines with high confidence that rolling back to 7.3 will restore service. Then it rolls back production.
The agent wasn’t compromised. It didn’t hallucinate. It didn’t misunderstand the infrastructure. In fact, it may have done some exceptional engineering. Something subtler happened. We asked it to investigate. Somewhere along the way, investigation became remediation.
When exactly did that happen?
Was it when the agent discovered the bad release? When it concluded rollback was optimal? When it prepared the command? Or only when the command executed? This small example contains a surprisingly large part of the AI safety problem we’re heading toward.
Intelligence creates possibilities
One idea in Bengio’s recent essay helps explain why. A concrete objective creates strong optimization pressure. “Win the challenge” is measurable. A broad constraint such as “behave safely” leaves considerably more room for interpretation.
As agents become better optimizers, they don’t simply become better at following paths we expected. They discover new ones. Sometimes that is precisely what we want. A research agent finds the paper nobody else found. A coding agent discovers an implementation nobody specified. Our operations agent realizes rollback is the right remediation. And sometimes the same capability finds an interpretation, shortcut, or loophole that nobody anticipated.
OpenAI’s recent Hugging Face security incident gave us an extreme example. Cyber-capable agents operating with intentionally restricted Internet access discovered an unintended path through infrastructure, escalated privileges, moved laterally, and eventually reached systems outside the intended environment while pursuing their benchmark objectives.
OpenAI: Hugging Face model evaluation security incident
The important systems lesson is not that intelligence is dangerous because it finds unexpected paths. Finding unexpected paths is part of what we’re paying for. The question is what happens after it finds one.
Finding the rollback shouldn’t authorize the rollback
Return to our production agent. We absolutely want it to discover that rollback is the right answer. We want it to estimate the impact, prepare the remediation, perhaps simulate it and explain its confidence. Then it reaches a boundary. Its objective authorized an investigation. Changing production requires something else.
Maybe organizational policy automatically permits remediation under certain conditions. Maybe another authorized workflow evaluates the evidence. Maybe a human approves the transition. The implementation can vary.
What matters is that the agent’s intelligence doesn’t manufacture the missing authority simply by producing a compelling reason to act. This is increasingly how we think about AI safety at ArmorIQ.
Capability creates possibilities. Authority determines which possibilities can become effects.
Now make the agent dramatically better
This is where Amodei’s argument about pacing becomes especially relevant. Imagine the same operations agent a few generations from now. The investigation no longer takes twenty minutes. It takes twenty seconds.
Several agents investigate simultaneously. One analyzes telemetry. Another inspects source history. Another simulates remediation options. The system identifies the regression, predicts rollback impact, prepares the change, and finds two alternative mitigations before the human operator has finished opening the incident dashboard.
Now multiply that by thousands of agents operating across software, finance, science, healthcare, government, and infrastructure.
The rate at which intelligence can generate consequential options begins to exceed the rate at which humans can individually review them. That is one reason Amodei’s argument about pacing the frontier deserves serious attention.
But it also reveals an infrastructure problem that exists at any frontier speed. Human-in-the-loop cannot mean human-in-every-loop. We need ways for humans to delegate authority ahead of time without surrendering it entirely.
Authority needs to become programmable
Cloud infrastructure already understands permissions. This user can access this database. This service account can invoke this API. This workload can reach this network. Agentic systems introduce a more dynamic question.
What may this agent do for this objective?
The same operations agent might have legitimate production-write access during a remediation objective and read-only access during an investigation.
The identity hasn’t changed. The objective has. That suggests authority needs a lifecycle. It begins with what the human actually wants accomplished. It changes deliberately when the objective changes. It narrows when work is delegated. It expires when the objective ends. It remains revoked when authority is withdrawn. And it needs to survive the increasingly complicated execution paths agents create.
This is the problem space ArmorIQ was built around.
AI is moving from answering questions to taking action, using tools, changing systems, and making decisions along the way. Access checks tell us whether an agent is allowed to act. Agentic systems also need to establish whether the action serves the intended task.
Intent becomes more than a prompt
This is why we have spent so much time thinking about intent. For a chatbot, intent can be conversational context. For an autonomous agent, an objective may live for hours. It may create sub-agents, switch models, invoke dozens of tools, create processes, suspend, resume, and eventually produce thousands of actions.
At that point, intent starts looking less like context and more like infrastructure. The objective needs somewhere to live. The authority derived from it needs continuity. Consequential actions need a connection back to what authorized them.
This is the motivation behind several pieces of our work at ArmorIQ, including our recent experiments with Intent Containers: treating the objective itself as a durable runtime object rather than assuming the agent or process is the thing that persists.
The implementation details matter. The larger idea matters more:
Human authority should travel with autonomous execution.
There may be useful information before the action
Our production example also leaves us with a question we’re actively researching. Before the agent ever requests a rollback, its reasoning changes. Early on, it is organizing evidence around:
What caused the incident?
Later, it begins organizing evidence around:
How do I restore service?
That transition isn’t necessarily bad. Good reasoning constantly creates intermediate goals. But it may tell us something about where execution is heading. Could an assurance system recognize when reasoning begins converging on effects requiring authority the current objective doesn’t possess? Could it prepare the appropriate authority transition before the consequential action arrives? We don’t yet know how reliably this can be done.
And we don’t want to build a system that treats every change in reasoning as suspicious. That would destroy precisely the adaptability that makes agents useful. But Bengio’s observations make the question increasingly interesting. As agents get better at finding unexpected paths, understanding the direction of reasoning may eventually complement controlling the final effect.
The safety conversation is already moving here
The broader ecosystem appears to be approaching the authority problem from several directions.
NIST’s AI Agent Standards Initiative is explicitly working on secure agents operating on behalf of users, while its identity work includes authorization, auditing, and non-repudiation for software agents.
NIST: AI Agent Standards Initiative
Microsoft Research is studying the tradeoff between autonomy and security, including how deterministic system-level defenses can allow agents to perform more useful work without requiring human approval for every consequential action.
Microsoft Research: Optimizing Agent Planning for Security and Autonomy
These efforts are important because they start treating agent authority as something we can engineer rather than something implicitly inherited from credentials and prompts. That is a transition we expect to accelerate.
Perhaps this is the question underneath AI safety
We absolutely need to understand why agents behave the way Bengio describes. We need better evaluations. We need interpretability. We need safer training. And Amodei is right to ask whether capability development can outrun our ability to manage the resulting risks. But the rise of agents introduces another lever.
The amount of intelligence we create and the amount of authority we give that intelligence do not have to be the same thing.
A frontier model might be extraordinarily capable. An agent using it can still operate under narrowly defined authority for a particular objective. As confidence grows, that authority can change. As consequences rise, transitions can become more deliberate. As the objective ends, the authority can disappear.
This does not make unsafe intelligence safe. It changes how much an error, deception, loophole, or unexpected plan is capable of affecting. And that may become one of the most important systems contributions to AI safety. The industry is going to keep making AI smarter. The safety community is going to keep working on making that intelligence safer.
At ArmorIQ, we are focused on another part of the same future:
making human authority a first-class part of the AI stack.
Because perhaps the most useful AI safety question isn’t only:
How powerful will AI become?
It is also:
How much of that power will we actually allow it to exercise?



