Durable Execution for AI Agents
An AI agent may spend minutes, hours, or days retrieving documents, calling APIs, waiting for approval, and continuing after new information arrives. During that time, processes crash, deployments restart, networks fail, and services time out. If progress exists only in memory, or a saved conversation lacks the state needed for recovery, the workflow may have to start over. Starting over can waste earlier model calls and discard completed work. Durable execution stores enough progress to resume from a known point, provided the state is stored durably and compatible execution code remains available.
What a Durable Workflow Records
A durable workflow may preserve its original request, identifier, completed steps, results, retry schedules, tool calls, approval decisions, error information, and other state needed for recovery. Application-specific approval details and version metadata may need to be recorded explicitly. Platforms can use checkpoints, ordered event histories for replay, or both. Temporal records workflow events such as activity completions and timers, which a worker can replay after a crash. Microsoft’s Durable Task programming model and Durable Task Scheduler provide state management, checkpointing, and distributed coordination for long-running agent workflows.
This is different from saving a conversation transcript. Chat history may show what the model said without confirming whether a tool completed an approval was granted, or an external action remains unresolved.
Keeping External Work Out of Repayable Logic
Replay-based systems may run workflow code again while rebuilding state. Calls to the system clock, unrecorded randomness, model APIs, or external services can return different results during replay. Framework APIs for time, timers, and deterministic randomness can be safe inside orchestration logic, while activities or tasks handle model calls, database operations, and API requests. Once their outcomes are durably recorded, recovery can reuse them.
This separation matters for agents because model outputs can vary. Repeating an unrecorded request could produce a different plan or tool selection, while persisting the output preserves the earlier decision. A call that finishes before its result is recorded may still run again. Checkpoint frequency also requires judgment: frequent writes increase storage and coordination overhead, while larger work units or delayed persistence can increase the work repeated after a crash.
Durability Alone Does Not Prevent Duplicate Actions
A checkpoint may show that an agent requested an action without confirming whether the external system completed it. An agent may submit a purchase order, and the purchasing service may create it before the connection fails. If confirmation never reaches the workflow, the operation may still appear incomplete when execution resumes. Repeating it could create a duplicate. Durable execution cannot resolve that uncertainty alone.
Consequential tools should use service-enforced idempotency keys, stable business identifiers backed by uniqueness constraints, or status checks and reconciliation. A unique identifier alone does not prevent duplication. The workflow should distinguish confirmed success, confirmed failure, and unknown outcome. An uncertain request may be retried with the same key and parameters while the service's guarantees remain valid. Otherwise, the workflow may need to inspect the external system or escalate before repeating the side effect. Retries also need limits, backoff, deadlines, and escalation rules because a durable workflow can keep retrying for a long time.
Waiting for People and External Events
Durable execution lets a workflow pause without keeping a worker process occupied. An agent may wait for spending approval, code review, or more information from a customer, then resume after the expected event is delivered and processed.
LangGraph checkpointers save graph state and support human participation and recovery, but surviving process loss requires persistent storage rather than an in-memory checkpointer. Durability settings affect whether the latest progress has been saved when a crash occurs. In LangGraph, resuming an interrupt restarts the containing node, so code before the interrupt may run again. Side effects in that code must be safe to repeat or moved to a suitable separate step.
Approval records should capture the reviewer, action, parameters, decision, and expiration conditions. Before execution, the workflow should verify the reviewer’s authority, match the response to the correct task and action, and confirm that approval and authorization remain valid. Event handlers should tolerate duplicate delivery where the underlying system allows it. Long pauses also create versioning concerns because code, prompts, schemas, policies, models, and tools may change or disappear. Teams need compatible changes, versioned workers, or tested migration paths so older executions can resume with appropriate logic.
Designing for Recovery
Durable execution helps agents recover from infrastructure failures using recorded progress, but the record alone is not enough. Reliable recovery depends on careful state design, replay-safe orchestration, durable model and tool results, and clear handling of uncertain outcomes. Teams should test crashes at important boundaries and plan for workflows that outlive a deployment. With those controls in place, long-running agents can resume from saved progress while safely retrying or reconciling work that was still in flight.
