State Management and Checkpointing
An agent run that needs twenty tool calls and dies on the nineteenth from a network glitch is expensive if it has to restart from scratch — and more expensive still if that happens often. State management and checkpointing are the two mechanisms that prevent exactly this: the agent keeps track of its intermediate state, and writes it durably to storage at defined points (a checkpoint), so a run picks up from the last safe point after a crash instead of from zero. Taken to its logical conclusion, this principle is called durable execution — execution that survives crashes, timeouts and server restarts as if they never happened.
Why state has to outlive a single step
An agent is rarely a single model call. The typical shape is a chain of plan, call a tool, read the result, decide the next step — spanning seconds to hours, and across multiple processes entirely in multi-agent systems. Any of these steps can fail: an API stops responding, a process gets redeployed, or a user steps in via human-in-the-loop and pauses the run for hours.
Without persisted state, every one of these cases means the entire run starts over — including every LLM and tool call already paid for. For short chat replies that’s a nuisance. For long agent pipelines with many expensive intermediate steps (research, code generation, GPU jobs), it’s a real cost problem — and for processes that trigger real-world actions (orders, payments, approvals), it’s also a consistency problem: did the step actually run, or not?
Checkpointing: a snapshot after every step
A checkpointer is the storage component that writes a run’s state to a database after every so-called super-step — in LangGraph, that means a full snapshot of the graph state, referenced by a thread ID. If the process crashes or gets paused, the next start loads the same thread and reconstructs exactly the state before the interruption — no guessing which step ran last.
Common checkpointer backends differ by use case:
- In-memory (e.g. MemorySaver). Development only — state disappears the moment the process ends.
- SQLite. Persists across restarts, fine for tests and small deployments.
- Postgres / distributed databases. The production standard: survives crashes, scales horizontally, and lets multiple workers share the same thread store.
Checkpoints also bring a side effect that’s often the real reason to use them: time travel. Because every intermediate state is stored individually, a run can be resumed not just from the end but from any earlier step — useful for debugging, and for continuing exactly from the point where a human-in-the-loop correction happened.
Durable execution: more than a save point
A plain checkpoint is, by itself, only a save point — someone (code or a human) still has to notice that a run was interrupted and explicitly trigger the restart from that point. Durable-execution runtimes like Temporal go a step further: the actual workflow logic runs inside a function the platform itself supervises, and every model or tool call executes as a separate “activity” that gets retried automatically on failure, without the surrounding workflow code having to program that itself. The foundation for this is an event history — a complete, append-only log of every event in a run, from which the state can be fully replayed whenever needed.
The promise behind it: the workflow code executes “exactly once and to completion” — whether it runs for seconds or weeks, regardless of hardware or network failures in between. That guarantee is stronger than plain checkpointing, because resumption happens automatically instead of needing a manual trigger.
In practice: LangGraph, Temporal and the new wave
Demand for durable execution in AI agents pushed several platforms into the mainstream during 2026:
- LangGraph checkpointers are the default path for agent frameworks: built in, thread-based, and well suited for human-in-the-loop and time travel within a single graph.
- Temporal offers a generic durable-execution model built on workflows and activities, and has shipped official integrations for popular agent SDKs since 2026 — aimed at cases where reliability needs to extend beyond a single framework’s boundaries.
- Cloud-native variants such as AWS Durable Functions, Cloudflare Workflows (GA since 2025) and the Vercel Workflow DevKit bring the same principle as a managed platform building block, with no infrastructure of your own to run.
The performance cost is modest: durable-execution engines typically add single-digit milliseconds per activity call — for tool calls that already take several seconds (GPU jobs, for instance), that overhead is a rounding error. It becomes noticeable only when a run (re)starts and the event history has to be replayed once.
When the effort pays off
State management and checkpointing are worth the effort once any of these apply:
- The run consists of many costly steps (LLM calls, search APIs, GPU inference) — restarting from step one would be expensive.
- The run takes a long time — minutes to days — and has to survive process restarts, deployments or server maintenance.
- The run includes human-in-the-loop pauses, where a person may take hours or days to approve.
- The run triggers real, sometimes irreversible actions, where being able to prove what ran when — and whether it ran twice — is itself a requirement.
For short, cheap, single-shot tasks with no side effects, the overhead rarely pays off — a plain retry without persisted state is enough there.
Pitfalls in practice
- A checkpoint is not the same as automatic recovery. Plain checkpointers only save state — deciding when and how a run gets resumed is often still left to your own code. If you need real fault tolerance without manual intervention, you need a durable-execution runtime, not just checkpointing.
- Idempotency is mandatory, not optional. If a step gets repeated after a crash, it must not trigger duplicate side effects (a duplicate charge, a duplicate email). See Error Handling: Retry and Idempotency.
- Storage growth. Every checkpoint is an additional record. Without a cleanup routine for old, completed threads, the database grows without bound.
- Schema drift. If the shape of the state object changes between two deployments, old checkpoints become incompatible — a run that was paused across a deployment can no longer resume cleanly.
- Over-granular checkpointing. Not every tiny step needs its own save point — overly fine-grained checkpoints add storage and latency overhead without meaningfully improving reliability.
FAQ
Is checkpointing the same as agent memory? No, even though both can share the same technical foundation. Agent memory stores knowledge and context so an agent can pick up a topic later. Checkpointing stores the technical execution state of a running process so it can resume after an interruption — it’s about resumption, not recollection.
Do I only need this for multi-agent systems? No. Even a single agent with a long tool chain benefits from checkpointing. It becomes especially important in multi-agent orchestration, because multiple processes are coordinated and can each fail independently.
What’s the difference between checkpointing and a plain retry? A plain retry repeats one failed call without knowing the overall state of the run. Checkpointing saves the entire intermediate state of a multi-step process, so after an interruption it’s not just the single call but the whole run that resumes from the last safe point.
Do I have to build checkpointing myself? In most cases, no. Agent frameworks like LangGraph ship checkpointers out of the box; for stricter reliability needs, dedicated durable-execution platforms like Temporal or cloud-native workflow services are an option. Building it yourself only pays off for very specific requirements.
Does checkpointing add noticeable latency or cost? The running overhead sits in the single-digit-millisecond range per step and barely registers against typical tool calls that already take several seconds. The more noticeable cost is storage across many long runs — a cleanup routine for completed threads helps there.
Discover more
Bulk Content Engine: How Context and RAG Tags Make the Orchestrator Smarter
My Bulk Content Engine now pauses and resumes at any point, because the orchestrator maintains its own context — plus RAG tags per task.
EncyclopediaHuman-in-the-Loop — Approval Gates for AI Agents
What human-in-the-loop means in agent workflows: approval gates, intervention points before critical actions, and why they are mandatory for irreversible steps.
NewsExecution engine without an IDE: tickets from the dashboard to any repo
You write a ticket in the boostN dashboard and it runs on the right repository — no IDE needed. Multiple repos, in parallel, in seconds.