ARMORIQ

Execution State Can Travel Backward. Authority Cannot.

Part 3 of our Intent Container series: suspend and resume exposes a security problem that process identity, workload identity, and sandbox isolation cannot solve on their own.

Sep 2, 202610 min read
Execution State Can Travel Backward. Authority Cannot.// Cover

In Part 1 of this series, we argued that autonomous AI needs a new runtime abstraction: the Intent Container, a boundary around the objective rather than simply around an agent or process.

In Part 2, we followed that objective down the stack. PAP bounds the authority available to an objective as it is refined. IAP cryptographically commits the accepted plan and proves its execution lineage. intentd maintains the current Intent Container state and carries the resulting assurance state across runtime boundaries. In our first reference implementation, KAP enforces that authority in the guest kernel against real execution. PAP’s refinement model is explicitly designed to keep authority bounded as plans become more concrete, while IAP supports re-anchoring, delegation, and revocation as authenticated changes to the committed lineage.

Thanks for reading ArmorIQ - Intent is the New Perimeter! Subscribe for free to receive new posts and support my work.

That gets us from an objective to a syscall. Then the agent goes to sleep. And things get considerably more interesting.

Agents are not always running

Long-running agents spend a surprising amount of their lives doing nothing. An agent may be waiting for a user, another agent, an external event, a long-running job, or a scheduled continuation. Keeping a complete execution environment consuming resources during that time makes little sense. This is one reason suspend and resume is such an important primitive for emerging agent infrastructure.

In the first intentd reference implementation, Google’s Agent Substrate is the runtime adapter and an Actor runs inside a Kata microVM. Cloud Hypervisor can pause that VM and checkpoint its state. When work becomes available again, the environment can be restored and the Actor continues from where it stopped.

From an agent infrastructure perspective, this is exactly what we want. From an authority perspective, however, checkpointing creates a subtle problem. A snapshot captures the past. Authorization lives in the present.

Imagine an agent goes to sleep authorized

Consider our running example. An enterprise agent is working on:

Analyze quarterly revenue and prepare a board-ready report using internal financial data.

The objective has been approved. IAP has committed its execution lineage. intentd has created the Intent Container. The Actor is running inside a KAP-enabled guest with authority epoch 1. The kernel allows the operations associated with that objective. Then the Actor reaches a point where it needs to wait. We suspend it.

Cloud Hypervisor checkpoints the guest. Importantly, this is not simply a copy of the application’s files. The snapshot captures guest memory, which means it also captures kernel state. For KAP, this initially looks like a gift. The Actor’s intent group survives. Its lineage root survives. Its authority epoch survives. When the Actor resumes, we do not need to reconstruct the entire security context from scratch. But suppose something changes while it sleeps.

The user revokes the objective. The control plane moves the authority from epoch 1 to epoch 2. Now restore the old snapshot. The guest kernel wakes up exactly where it was. Including at epoch 1.

💥

We have just restored revoked authority.

Nothing in the snapshot is necessarily invalid

This is what makes the problem interesting. There is nothing obviously malicious about the restored environment. It may be exactly the snapshot the infrastructure created earlier. The Actor identity is legitimate. The workload identity is legitimate. The VM is legitimate. The snapshot may pass every integrity check we have. Even its KAP state is authentic. It is simply old.

That distinction matters because integrity and authorization answer different questions. Snapshot integrity tells us:

Is this the execution state we previously saved?

What we actually need to know is:

Is this execution state still authorized to operate now?

The answer can change while the Actor is asleep. That makes authority fundamentally different from execution state.

Why identity cannot solve this

It is tempting to treat this as another identity problem. Before restoring the Actor, authenticate it again. Confirm its workload identity. Check the VM measurement. Validate the snapshot.

All useful. None answers the question. The same Actor can be authorized at 10:00 AM and revoked at 10:15 AM. Its identity did not change. The objective did. This is exactly why we have been arguing that the objective needs to become a first-class runtime object.

Authority does not belong simply to:

user

↓

agent

↓

process

↓

VM

For autonomous execution, we need another lineage:

objective

↓

committed intent

↓

delegated authority

↓

Actor

↓

process

The Actor’s identity tells us what woke up. The objective lineage tells us whether it should still be allowed to act.

Authority needs a clock of its own

Our current design handles this with an authority epoch. Think of the epoch as the current version of the objective’s executable authority. When the objective is first committed:

Objective: ic-001

Authority Epoch: 1

Status: Active

The Actor is bound to epoch 1. If authority changes or is revoked, the control plane advances the epoch:

Objective: ic-001

Authority Epoch: 2

Status: Revoked

Now consider restoring a snapshot containing:

Objective: ic-001

Authority Epoch: 1

The snapshot may be perfectly authentic. But it is stale. And stale authority must fail closed. That is a much simpler enforcement problem than asking the restored agent whether it still believes it has permission to continue.

Restore becomes an authorization boundary

This leads to an important design decision: restore is an authorization event. Before a restored guest becomes active, the restore path compares the authority epoch captured in the snapshot with the current epoch for the objective. A match allows restoration to continue. A mismatch because authority changed or the objective was revoked fails closed. The check belongs to the restore boundary itself, so it applies whether Resume arrives through intentd, kubectl, or Agent Substrate.

At the restore boundary, conceptually:

Restore snapshot

↓

Snapshot says epoch 1

↓

Current authority says epoch 2

↓

DENY

The snapshot cannot vote itself back into power. Its saved authority is evidence from the past, not permission to execute now.

The execution state can travel backward. Authority cannot.

This gives us an invariant we have come to like:

The execution state can travel backward in time. Authority cannot.

Checkpointing is fundamentally a form of time travel.

Files, memory, processes, and kernel structures can all return to an earlier state. That is precisely what makes checkpoint and restore useful. But security state cannot blindly follow the same semantics.

Suppose an agent’s financial-data access was revoked after the checkpoint.

Or a delegated agent was removed.

Or a budget was exhausted.

Or an MCP server was removed from the objective’s authority.

Or the user stopped the objective entirely.

Restoring an old execution environment should not undo any of those decisions. The execution can return to yesterday. The authorization has to remain today.

This is bigger than revocation

Once we started looking at suspend and resume through the objective rather than the VM, we realized the same issue applies to much more than a binary revoked/not-revoked state.

Imagine an agent is checkpointed with a $5 budget and has spent $1. While it sleeps, sibling agents belonging to the same objective consume another $3.50. When the Actor wakes up, its local snapshot cannot reasonably believe the objective still has $4 available. Or suppose an agent was originally allowed to use three MCP interfaces. While it sleeps, an administrator removes one because of a security incident. Restoring the Actor should not restore that interface.

The same applies to delegation. A child Actor may have been validly delegated authority when the snapshot was created, but that delegation may have subsequently expired or been revoked. IAP’s trust-update model already treats re-anchoring, delegation, and revocation as changes to the committed trust lineage rather than properties frozen permanently at initial authorization. IAP_Paper.pdfPDF

The general principle becomes: A snapshot can preserve the execution state. It cannot be the source of truth for current authority. That source of truth belongs to the objective.

This is where the Intent Container earns its name

Without an Intent Container, this problem becomes awkwardly distributed. The agent runtime knows about the Actor. The VM manager knows about the snapshot. The kernel knows about processes. The identity system knows about the workload. The policy system knows that something was revoked. Someone has to connect those facts.

intentd gives us the object around which they can be connected.

The Intent Container exists independently of any particular Actor snapshot. It maintains the current objective identity, lineage, authority epoch, delegation state, and eventually other objective-level state such as budgets and data constraints.

Actors can come and go. Snapshots can be created and restored. Processes can die and restart. The objective remains the reference point. That is why we increasingly think the Intent Container is not merely a convenient grouping abstraction. It is the right place to establish continuity across agent lifecycles.

The kernel still has to enforce the answer

Of course, knowing that a snapshot is stale is useful only if the workload cannot bypass the decision. In our current reference implementation, the KAP-enabled guest records its authority epoch to durable shared storage. On restore, after the volumes return but before the VM starts, the host compares that saved epoch with current Intent Container authority. If they differ, the restore is refused before the Actor becomes active. Because the check runs in the restore path itself, a direct wake through kubectl or Agent Substrate reaches the same boundary. Other runtimes and policy enforcement points need equivalent integration.

Our current conversation with the Agent Substrate team has clarified the portable requirement. What we first presented through #1223, #1224, and #1284 as separate requests is really one need at two lifecycle moments: establish authority before an Actor first becomes active, and reconcile it before a restored Actor becomes reachable. Short-lived authority narrows the revocation window, but it cannot detect rollback when a snapshot also restores the guest’s clock. We are not asking Substrate to adopt intentd or KAP semantics, only for stable pre-activation extension points that Kata, gVisor, or another runtime can implement at its own enforcement boundary.

intentd owns objective continuity.

KAP owns kernel enforcement continuity in this reference backend. The Intent Container says which authority is current. The kernel can enforce only the authority installed in the guest. The restore lifecycle reconnects those two facts before resumed execution proceeds.

Next: putting the whole Intent Container on screen

The first three posts in this series have moved progressively downward. Part 1 introduced the objective as the runtime boundary. Part 2 followed that objective from PAP and IAP through intentd and our Google Agent Substrate, Kata, and KAP reference path until it became a kernel-enforced syscall decision.This post introduced time into the equation: what happens when execution disappears and later comes back after authority has changed.

Next, we want to stop drawing the architecture and show the full intentd lifecycle, including the restore boundary that keeps saved execution state from carrying stale authority forward:

Objective → Actor → bind → allow → deny → suspend → resume → revoke → stale snapshot refused at restore.

Along with the demo, we’ll show the architecture behind it and start sharing more of the implementation details and integration points.

This is the moment our demo will spw

We will start an Actor at authority epoch 1, perform an allowed operation, checkpoint it, and restore it normally with the same enforcement state. We then revoke the objective. When the pre-revocation snapshot is restored, the host reads the saved epoch from durable shared storage, compares it with current authority after the volumes return but before the VM starts, and refuses the stale restore. Its authority belongs to the past. The restore is refused before the Actor becomes active, even when the wake does not originate in intentd.

For now, we take leave with one rule that we think agent infrastructure will increasingly need: Agents can sleep. Agents can move. Agents can resume from the past. Their authority must always come from the present.

Thanks for reading ArmorIQ - Intent is the New Perimeter! Subscribe for free to receive new posts and support my work.

Onboarding open

Ready to control what your AI agents actually do?

Join the teams shipping safer, compliant AI agent deployments. White-glove onboarding for the first 50 design partners.

Read Docs →
Live Intent Assurance↗