Design docs for production agents are lopsided. They describe triggers, plans, tool calls, retries and hand-offs in exhausting detail, then close with a sentence like “and we can always stop it later.” In practice, “later” is the moment your concurrency is saturated, your model bill is climbing, and the only tool anyone has reached for is kill -9 against a container that does not own the run.
This guide is the stop side of the problem: cancellation semantics, kill switches, timeouts and spend limits — a production checklist for the long-tail question of how to stop a runaway AI agent with kill switches, cancellation semantics and spend limits. We take each execution engine at its word: how Inngest, Temporal, Cloudflare Workflows and LangGraph actually cancel, pause, terminate, restart or delete a run, how Restate caps concurrent cost through flow control, and how OpenAI, LiteLLM and OpenRouter stop the money when the engine will not stop the work. Resume and replay appear here only as the reason a naive process kill is unsafe; the full argument lives in checkpointing is not durable execution. For why these runs run away in the first place, start with why multi-agent systems fail: the MAST taxonomy, then seven failure patterns that break production systems and AI agent cost optimization in 2026.
The missing stop path: why production agents need cancellation and spend guardrails
Production agents need cancellation and spend guardrails because the stop path is missing by default: Inngest documents that cancellation stops queued, sleeping, or active runs when remaining work is no longer needed, while an already-executing step still finishes (Inngest cancellation). OpenAI’s hard spend limit independently returns a 429 (OpenAI spend limits).
Three failure shapes make that gap dangerous. First, loops: a self-correcting agent keeps re-planning long after the goal is unreachable, and engine retries stretch the tail — Restate retries a run action with exponential backoff when a called API times out, with retry policy configurable in service configuration (Restate key concepts). Second, fan-out: concurrency multiplies spend per wall-clock second, which is exactly why Restate frames flow control as a cost ceiling for “AI agent calls that translate directly into model or API spend” (Restate key concepts). Third, ownership: your agent is not a local process you own, so killing the client container leaves the server-side execution running.
That last point is why the stop path must be designed rather than improvised. An engine-native cancellation only stops work the engine can still withhold, and it explicitly does not undo a call that already completed (Inngest cancellation). A spend cap only stops work after the meter has moved. Neither control substitutes for the other, which is the verdict this guide builds toward.
Engine cancellation semantics at a glance
Engine cancellation semantics differ per system, and Temporal’s own docs are explicit: a Workflow Execution closes as Cancelled only when it successfully handles a cancellation request, while Terminated is the abrupt stop and Timed Out is its own closed state (Temporal workflow execution).
Read down the Primitive column before you read across it. The five engines do not share a vocabulary for “stop,” and the differences are semantic, not cosmetic: one system gives you a cooperative cancellation, another gives you an abrupt termination with optional rollback, another gives you only a pause or a concurrency ceiling. Match the primitive to the incident you are actually trying to handle.
| Category | System | Primitive | Stop semantics from source | Key caveat | Source |
|---|---|---|---|---|---|
| Engine stop | Inngest | Cancellation via timeouts, cancelOn events, bulk cancellation |
Stops queued, sleeping, or active runs; an executing step finishes, then the run is marked canceled and later steps do not run | Cancellation does not undo a call that already completed | Inngest cancellation |
| Engine stop | Temporal | Cancelled / Terminated / TimedOut / Paused | Cancelled means the execution successfully handled a cancellation request; Terminated is the abrupt stop; Timed Out is a distinct closed state; Paused stops dispatch of new Workflow Tasks | Cancellation is cooperative — it succeeds only when the workflow handles the request | Temporal workflow execution |
| Engine stop | Cloudflare Workflows | pause()/resume(), terminate({ rollback: true }), restart(), delete() |
terminate() stops an instance and can run registered rollback handlers first; restart() cancels in-progress steps and erases intermediate state |
A stopped or terminated instance cannot be resumed; delete() stops execution without running rollback handlers |
Cloudflare Workflows |
| Engine stop | LangGraph | interrupt() |
Pauses graph execution at any point, saves state through the checkpointer, and waits indefinitely until resumed with Command(resume=...) |
It is a pause primitive, not a cancellation signal; static breakpoints are not recommended for human-in-the-loop | LangGraph interrupts |
| Engine stop | Restate | Concurrency limits (flow control) | Caps concurrent invocations for a scope, explicitly to put a ceiling on expensive concurrent work such as AI agent calls | Rate limits and priorities are planned next; the journal replays completed steps and retries with exponential backoff | Restate key concepts |
| Spend cap | OpenAI | Hard spend limit, plus spend alert | Affected API requests return a 429 error when the hard spend limit is reached |
A spend alert only sends a notification — API traffic continues | OpenAI spend limits |
| Spend cap | LiteLLM | Budgets: global, team, team-member, virtual-key, model, agent | Exceeding a budget returns an ExceededBudget error; a hard budget enforcement (fail closed) option is available |
Budgets require Postgres; a DB-less deployment fails open and requests keep being served past the limit | LiteLLM proxy users |
| Spend cap | OpenRouter | Credit limits vs rate limits | Credit limits return HTTP 402; rate limits return HTTP 429; the in-flight spending budget can reject a request before it reaches a provider |
error.metadata.limit_source names the credit source: in-flight budget, key limit, or credits |
OpenRouter limits |
Timeouts as the first-line stop signal
Timeouts are the first-line stop signal because Inngest cancels a run that waits too long for its first step or exceeds its execution window, with both windows set in function configuration (Inngest cancellation); Temporal records Timed Out as a distinct closed state (Temporal workflow execution).
A timeout is the only stop signal that requires no operator, no dashboard and no alert to be useful. It fires while nobody is watching, which makes it the correct default for the first line of defence: an agent that never reaches its first step, or never leaves it, should be dead before anyone opens a laptop. Inngest exposes exactly this as one of its three cancellation mechanisms — a run is canceled if it waits too long for its first step or exceeds its execution window, both configured in function configuration (Inngest cancellation).
Treat timeouts as a ladder. The engine-level window catches a wedged run; a workflow-level timeout catches a run that is progressing but unproductive; and the spend cap catches anything that slips past both. Keep the engine timeout visible in your function configuration rather than hard-coded in a wrapper script, because a wrapper dies with the process that started it (inngest/inngest). Timeouts also give you a clean status to audit: Temporal closes an execution as Timed Out rather than leaving it ambiguously open (Temporal workflow execution).
Cooperative cancellation vs termination vs delete
Cooperative cancellation, termination, and delete are three different stop strengths: Temporal’s Cancelled state means the execution successfully handled the request, Cloudflare’s delete() stops a running instance without running rollback handlers, and restart() immediately erases intermediate state and cancels in-progress steps (Temporal workflow execution, Cloudflare Workflows).
Cooperative cancellation is a request, not a command. Temporal closes an execution as Cancelled only when the execution “successfully handles a cancellation request,” and a workflow may only block on a Temporal SDK Awaitable — one of which is progress on cancellation of another workflow execution, alongside Signals, child workflows, activity executions and timers (Temporal workflow execution). If your workflow never awaits anything cancellation-aware, cancellation is a request that lands politely and does nothing until it does. That is why Terminated exists: it is the abrupt stop, and it is auditable as a closed state rather than a silent kill.
Inngest draws the same boundary from the other side: if a step is already executing, that step finishes, the run is then marked canceled, and later steps do not run (Inngest cancellation). The documentation warns plainly that cancellation does not undo a call that already completed — so any external effect (an email sent, a tool charged, a row written) needs its own compensating action, not a prayer.
Cloudflare Workflows gives you the full ladder of stop verbs, and the verb determines the side effects. terminate() stops an instance; pass { rollback: true } to run registered rollback handlers before terminating, or use the CLI form npx wrangler workflows instances terminate <WORKFLOW_NAME> <INSTANCE_ID> --rollback (Cloudflare Workflows). Once stopped or terminated, the instance cannot be resumed. restart() immediately cancels any in-progress steps, erases intermediate state and treats the workflow as if it were run for the first time. delete() stops a running instance without running rollback handlers. Choose deliberately: rollback, reset, or discard.
Pause is not cancel: LangGraph interrupts and Cloudflare pause/resume
Pause is not cancel in both engines: LangGraph’s interrupt() pauses execution at any point, saves state through the checkpointer, and waits indefinitely until resumed with Command(resume=...), while Cloudflare’s paused instance still exists under a paused status value (LangGraph interrupts, Cloudflare Workflows).
A pause is a hold, and holds cost money and hold locks. LangGraph documents interrupt() as a pause primitive: it halts graph execution at any point, persists state via the checkpointer, and waits indefinitely until you resume execution; the resume value becomes the return value of the interrupt() call (LangGraph interrupts). That is ideal for human review and approval gates, and useless as a kill switch — a paused run still occupies the state you were trying to reclaim. The project’s own guidance is that static breakpoints (interrupt_before=[...] / interrupt_after=[...] on compile() or per invocation) are not recommended for human-in-the-loop workflows; use interrupt() instead (langchain-ai/langgraph).
Two operational details matter when a pause is your gate. Typed resume values require a response_schema, which needs langgraph>=1.2.12; invalid input raises a Pydantic ValidationError, and you can retry with a corrected Command(resume=...) (LangGraph interrupts). Interrupts surface through stream.interrupts, so an operator can see a paused run — but visibility is not termination.
Cloudflare makes the same distinction in its status values: an instance sits at paused alongside queued, running, errored, terminated, complete, waiting, waitingForPause and unknown, with instance.pause() and instance.resume() as the pair of controls (Cloudflare Workflows). Note the asymmetry: pause and resume are reversible; terminate and delete are not. waiting and waitingForPause are the states that look like progress but are not, and they are where a runaway run hides when you are scanning a list for something to stop.
Bulk cancellation and operational stop controls
Bulk cancellation is the operational stop control, and Inngest scopes it by function and time range from the Dashboard or REST API, but it does not prevent new runs from being enqueued (Inngest cancellation), so pause the function or stop the event source too.
Bulk stop is where incident response usually fails, because it is a filter, not a fence. Inngest lets you stop many existing runs selected by function and time range from the Dashboard or REST API, but the documentation is explicit that bulk cancellation does not prevent new runs from being enqueued — if the trigger keeps producing runs, you must pause the function or stop the source of events as well (Inngest cancellation). A kill switch that clears a queue while the producer keeps filling it is a rate limiter you did not intend to build.
The complementary control is refusing work before it starts. Restate’s flow control caps concurrent invocations for a scope, described directly as a way to control cost by putting a ceiling on expensive concurrent work such as AI agent calls that translate directly into model or API spend (Restate key concepts); rate limits and priorities are listed as planned next. Combined with an idempotency key — which deduplicates duplicate requests automatically — a scope ceiling gives you a stop that is architectural rather than reactive (Restate key concepts).
Cloudflare’s batch control belongs in the same drawer: deleteBatch([...]) deletes up to 100 instances in one call, while delete() removes a single instance, and deleting a running instance stops its current execution without running rollback handlers (Cloudflare Workflows). Bulk operations need an owner, a time window, and a documented answer to the question “what stops the next one?”
Spend-side kill switches: OpenAI limits, tiers, hard limits, alerts
The spend-side kill switch is OpenAI’s hard spend limit, which makes affected API requests return a 429, while a spend alert only sends a notification and lets API traffic continue (OpenAI spend limits), separate from the approved monthly usage limit (OpenAI rate limits).
Start with the shape of the controls. OpenAI defines rate limits in RPM, RPD, TPM, TPD, IPM (images) and audio-minutes, set at the organization and project level — not the user level — varying by model, with long-context models carrying a separate limit (OpenAI rate limits). The company also sets an approved monthly usage limit for each organization, which is distinct from the spend limits you configure yourself. That distinction is the whole point: the usage limit is a platform policy, your spend limit is your circuit breaker.
OpenAI documents two configurable spend controls with deliberately different behaviour: a spend alert, which sends a notification while API traffic continues, and a hard spend limit, where affected API requests return a 429 error (OpenAI spend limits). An alert is a pager; a hard limit is a stop. Deploy both, but never describe an alert as a cap — it is not one, and the difference is measured in dollars at 3 a.m.
Usage tiers bound the ceiling you can even configure: Free $100/month; Tier 1 ($5 paid) $100/month; Tier 2 ($50 paid) $500/month; Tier 3 ($100 paid) $1,000/month; Tier 4 ($250 paid) $5,000/month; Tier 5 ($1,000 paid) $200,000/month (OpenAI rate limits). Note also that Batch API queue limits are calculated from total input tokens queued per model: pending batch tokens count against the queue limit, completed batches do not. An overnight batch is therefore a spend decision with a queue-shaped failure mode.
LiteLLM budgets and the Postgres fail-open caveat
LiteLLM budgets fail open without Postgres, because every budget is enforced against spend read from the database, so on a DB-less deployment the global check is skipped and requests keep being served past the limit, and key, team, and user budgets are unavailable without a database (LiteLLM proxy users).
Budget scope is the feature that makes a self-hosted proxy useful as a kill switch: global proxy budgets (litellm_settings.max_budget with a budget_duration such as 30d), team budgets, team-member budgets via max_budget_in_team, virtual-key budgets, model-specific budgets, and agent budgets that combine rate limits (tpm/rpm) with session-level caps on iterations and dollar budget (LiteLLM proxy users). Agent-level dollar caps are the closest thing in this stack to “this agent may not spend more than X,” and they sit below your engine’s cancellation in the stack.
Then read the caveat, because it is the kind of line that survives only until your first incident: every budget on the page is enforced against spend read from the database, so none of them cap anything on a DB-less deployment. Without a database, max_budget fails open, the global check is skipped, and requests keep being served past the limit (LiteLLM proxy users). A budget that silently does nothing is worse than no budget, because your runbook assumes it works.
When the budget does fire, exceeding it returns an ExceededBudget error, and the proxy offers a hard budget enforcement (fail closed) option plus a budget-reservation mechanism; temporary increases can be granted with temp_budget_increase and temp_budget_expiry (LiteLLM proxy users). Enable fail-closed enforcement for any key that reaches a production model, and treat a temporary increase as an incident artifact with an expiry — not as a quiet config edit. Pair this with the broader unit-economics work in AI agent cost optimization in 2026.
OpenRouter credit limits, in-flight budget, and 402 vs 429
OpenRouter separates spend from speed: credit limits — account balance, per-key credit limits, or the in-flight spending budget — return HTTP 402 before a request reaches a provider, while rate limits return HTTP 429 (OpenRouter limits), which is why a 402 needs budget and a 429 needs pacing.
Credit limits come from three places, and you should be able to name which one fired without guessing. The first is your account balance. The second is a per-key credit limit — an optional spending cap configured on an individual API key, exposed as limit, limit_reset and limit_remaining from GET /api/v1/key. The third is the in-flight spending budget: a cap on the estimated token cost of the paid requests you have running or recently completed, relative to your balance — a fraction of balance up to a fixed ceiling (OpenRouter limits).
The in-flight budget is the interesting one for runaway agents, because it rejects a request whose estimated cost does not fit before it reaches a provider. That is a pre-flight spend stop rather than a post-hoc bill, and its identity arrives in error.metadata.limit_source as openrouter_in_flight_budget, openrouter_key_limit or openrouter_credits (OpenRouter limits). Route those three values to three different alerts: one for “you are spending too fast,” one for “this key is capped,” one for “you are out of money.”
Rate limits are a separate axis. For free-model request IDs ending in :free, an account with under 10 credits purchased all-time gets 20 requests per minute and 50 requests per day; at least 10 credits purchased gives 20 requests per minute and 1,000 requests per day, while paid variants have no platform-level request cap (OpenRouter limits). Monitor with GET https://openrouter.ai/api/v1/key, checking limit_remaining, usage_daily and free_model_daily_requests.
Lifecycle observability: events, status values, and audit trails
Lifecycle observability is the audit trail for every stop: Temporal’s Event History records each state transition, Cloudflare exposes status values from queued through terminated, and Inngest emits an inngest/function.cancelled system event naming the function and run (Temporal workflow execution, Cloudflare Workflows, Inngest cancellation).
Instrument the stop, not just the start. Temporal records every state transition in Event History, which means cancellation, termination and timeout are auditable events rather than log lines someone grepped for (Temporal workflow execution). Inngest emits an inngest/function.cancelled system event per canceled run whose payload identifies the original function and run via function_id and run_id — attach a separate cleanup function to that event so releasing a run also releases what the run was holding. Inngest also notes a canceled run can be replayed from the Dashboard, which is an operational convenience, not a reason to skip designing the stop (Inngest cancellation).
Cloudflare’s status vocabulary should map one-to-one onto your dashboards: queued, running, paused, errored, terminated, complete, waiting, waitingForPause, unknown (Cloudflare Workflows). Every one of those states needs an owner and an expected next transition; unknown and waitingForPause are the two that hide incidents. Restate adds a journal in which every step of durable execution is tracked, so a replay skips completed steps — useful for diagnosis, and a reminder that the stop you issued is recorded somewhere whether or not you recorded it (Restate key concepts).
Ship the audit trail with the kill switch: who issued the stop, which primitive was used (cancel, terminate, restart, delete, budget, credit), which status value resulted, and what the spend meter read before and after. That record is what turns the next incident review into a fix instead of a debate — and it belongs alongside the failure taxonomies in seven failure patterns that break production systems and why multi-agent systems fail.
Frequently asked questions
1. Can I just kill -9 a Temporal or Inngest workflow?
No. kill -9 kills your client process, not the server-side run: Temporal’s execution continues and closes later as Cancelled, Terminated or Timed Out, and Inngest marks a run canceled only after the executing step finishes (Temporal workflow execution, Inngest cancellation). Stop through the engine’s cancellation API so the transition lands in Event History.
2. Does Cloudflare terminate() always run rollback handlers?
No — rollback is opt-in. You must pass { rollback: true } to terminate() to run registered rollback handlers before terminating, or use the CLI --rollback flag; delete() stops a running instance without running rollback handlers at all (Cloudflare Workflows). Choose the stop verb deliberately, because a stopped or terminated instance cannot be resumed.
3. Is a LangGraph interrupt() a cancellation signal?
No. interrupt() is a pause: it halts graph execution at any point, persists state through the checkpointer, and waits indefinitely until you resume with Command(resume=...), where the resume value becomes the call’s return value (LangGraph interrupts). Use it for human review, and use engine-level cancellation when the run must actually stop.
4. Does LiteLLM max_budget cap spend without Postgres?
No. Every budget is enforced against spend read from the database, so on a DB-less deployment max_budget fails open: the global check is skipped and requests keep being served past the limit (LiteLLM proxy users). Run the proxy with Postgres, and enable hard budget enforcement so an exceeded budget returns an ExceededBudget error.
5. What is the difference between OpenRouter 402 and 429?
402 means a credit limit stopped you: account balance, a per-key credit limit, or the in-flight spending budget, identified by error.metadata.limit_source (OpenRouter limits). 429 means a rate limit stopped you — request volume, including free-model caps of 20 requests per minute. Fix the first with budget, the second with pacing.
The Bottom Line
A production agent needs two independent circuit breakers, and neither one is a substitute for the other. The first is engine-native cancellation: an explicit, auditable stop verb — Inngest cancellation, Temporal Cancelled or Terminated, Cloudflare terminate({ rollback: true }) — issued through the engine, never through the process table, because cancellation does not undo a call that already completed (Inngest cancellation). The second is a spend-side cap that fails closed: an OpenAI hard spend limit returning 429 (OpenAI spend limits), a LiteLLM budget running against Postgres with hard enforcement (LiteLLM proxy users), or OpenRouter credit limits that reject with 402 before a request reaches a provider (OpenRouter limits). If you only build one, build the spend cap — it bounds the damage. If you only build the engine stop, you have bounded the work but not the money; if you only build the spend cap, a wedged run can still hold state and concurrency indefinitely. Deploy both, wire the stop events to an owner, and make timeout the default that fires when nobody is watching.
How this guide was built
Desk research only, measured 2026-10-06: no agent was deployed, no workflow was run, no spend limit was configured, and no account was created. Every fact, status value, error code and limit in this guide was read from vendors’ own primary documentation — Inngest cancellation, Temporal workflow execution, Cloudflare Workflows, LangGraph interrupts, Restate key concepts, OpenAI rate limits, OpenAI spend limits, LiteLLM proxy users and OpenRouter limits — plus repository metadata from the GitHub API on the same date: temporalio/temporal (23,502 stars, MIT, Go), inngest/inngest (5,917 stars, Go), restatedev/restate (4,523 stars, Rust) and langchain-ai/langgraph (42,784 stars, MIT, Python). Nothing here is a hands-on test result, and no version, price or endpoint behaviour has been inferred beyond what those pages state.
← Back to all posts


