Skip to main content
THE_COLUMN // AI

Agent Idempotency: How Infrastructure Teams Keep Retried Agent Actions From Double-Writing Into Systems of Record

Written by: iSimplifyMe·Created on: Sep 28, 2026·13 min read

How many times did your ticketing agent issue the same refund last quarter? If you run agents with write access to NetSuite, Salesforce, or ServiceNow, you may already know the answer — and it is rarely zero.

Agents inside systems of record now retry automatically after timeouts, throttling, and tool failures. That behavior is correct for availability, yet every retry of a write is a second chance to create a duplicate invoice, a duplicate case, or a second credit memo against the same customer.

This post covers the three controls that keep retried agent actions from double-writing: idempotency keys, dedupe ledgers, and replay-safe tool contracts. Our pieces on agent fallback design and agent tool design cover what happens when a tool fails; this one covers what happens when the tool succeeded and nobody heard about it.

Why Retried Agent Writes Are Harder Than Retried API Calls

A conventional integration retries the same HTTP request with the same body. An agent retries an intent, and the model may re-plan that intent into a different tool call with a different payload on the second attempt.

Consider a Claude-based billing agent on Bedrock that calls create_credit_memo for $4,200. The NetSuite call commits, the response takes 31 seconds, and the Lambda running the tool hits its 30-second timeout before returning.

The orchestrator sees a failure and hands control back to the model. The model reads the conversation, concludes the memo was never written, and calls the tool again — this time with a memo description it rephrased, which means any hash of the raw request body no longer matches.

The result is $8,400 in credits against a $4,200 dispute. Keep in mind that nothing in this sequence was a bug in the conventional sense — every component behaved exactly as designed.

Where Do Agent Retries Actually Come From?

A single agent write can be retried by five layers at once: the model's re-plan, the orchestrator, the SDK client, the queue or event bus, and the API gateway. Each layer retries without knowing the others did.

Most teams instrument one retry layer and assume it is the only one. In practice, the retry sources in a production agent stack include but are not limited to:

  • Model re-planning. After a tool error or timeout, the model decides on its own to try again. This is the only layer that can change the payload between attempts.
  • Orchestrator retry policy. Step Functions Retry blocks, LangGraph node retries, or a custom loop with exponential backoff. These typically replay the exact tool input.
  • SDK and HTTP client retries. The AWS SDK retries throttled and 5xx responses by default, and most vendor SDKs for Salesforce and Zendesk do the same underneath your code.
  • Queue and event redelivery. SQS standard queues deliver at least once, Lambda retries failed asynchronous invocations twice by default, and EventBridge can retry a target for up to 24 hours.
  • Human re-runs. An operator sees a stuck workflow in the console and presses "retry" — often the most dangerous replay of all, because it happens hours later.

Multiply these layers together and a single intent can fan out into far more attempts than any one retry policy suggests. This is why deduplication has to live at the write boundary, not inside any one retrying component.

What Is An Idempotency Key For An Agent Action?

An agent idempotency key is a deterministic ID built from the write's business intent — tenant, target record, operation, payload hash — so every retry of that intent maps to one key and one write.

The critical word is deterministic. A UUID generated at call time is useless for agents, because the model's second attempt generates a fresh UUID and the system treats it as a new write.

Instead, derive the key in the tool layer from structured fields the model cannot paraphrase. For the credit memo example, those fields would be:

  • Tenant and environment. The customer account and whether this is production or sandbox, so keys never collide across tenants.
  • Workflow run ID. The orchestrator's execution ID, which stays stable across model re-plans within one run.
  • Target record. The NetSuite customer internal ID and the originating dispute ticket number.
  • Operation. A fixed verb such as credit_memo.create, not the tool's display name.
  • Normalized payload hash. A SHA-256 over the amount, currency, and line items — explicitly excluding free-text fields like memo descriptions.

That last exclusion is what defeats the paraphrase problem. The model can rewrite the description a dozen ways, and the key still resolves to the same value because the description never enters the hash.

Note that the run ID belongs in the key only when the business rule is "once per run." If the rule is "once per dispute ever," drop the run ID so a human re-run tomorrow also dedupes — the key's scope should be a written policy decision, not an accident of implementation.

How Does A Dedupe Ledger Work?

A dedupe ledger is a table keyed on the idempotency key that records each write's state — pending, committed, or failed — plus its result, so a retry returns the stored outcome instead of executing again.

DynamoDB is the common choice on AWS because a conditional PutItem with attribute_not_exists(pk) gives you an atomic claim on the key. Redis with SET NX works for short windows, and Postgres works if you already hold a transaction open against the same database.

The execution sequence for every mutating tool call follows the same steps:

  1. Claim. Conditionally write the key with status pending, a lease expiry, and the payload hash. If the write fails because the key exists, skip to step four.
  2. Execute. Call the downstream system, passing the idempotency key or an external ID wherever the target API accepts one.
  3. Commit. Update the ledger entry to committed and store the downstream record ID and response body.
  4. Replay. On a duplicate, read the entry — return the stored result if committed, wait or reconcile if pending, and allow a fresh attempt only if it is marked failed with a retryable error.

AWS Lambda Powertools ships an idempotency utility that implements most of this pattern against DynamoDB for Python, TypeScript, Java, and .NET. It is a reasonable starting point; however, you will still need to own key derivation, because the utility hashes whatever payload you hand it.

The Pending-State Problem

The hardest case is a process that dies between steps two and three. The downstream write committed, but the ledger still says pending, and the lease eventually expires.

A naive implementation treats an expired lease as permission to retry, which reintroduces the exact duplicate the ledger was built to stop. The safer rule is to reconcile before re-executing — query the system of record by external ID, correlation field, or idempotency header to learn whether the first attempt landed.

This is why the ledger alone is not sufficient. It needs a system of record that can answer the question "did this already happen?"

Replay-Safe Tool Contracts

A replay-safe tool contract requires an idempotency key as input, returns the same response for the same key, and rejects a reused key carrying a different payload with an explicit conflict error.

The contract is the promise a tool makes to the orchestrator and the model. If the contract is vague, every caller has to guess, and guessing is how duplicates ship.

Here's how the rules of a replay-safe contract break down:

  • Key is required, not optional. The tool schema rejects mutating calls without a key. The tool layer computes it, so the model never has to supply one.
  • Same key, same payload returns the original result. The response is byte-for-byte what the first successful call returned, including the downstream record ID, so the model can proceed as if nothing unusual happened.
  • Same key, different payload is a hard conflict. Return a 409-equivalent with a clear message. Silently accepting the new payload hides a real change in intent; silently returning the old result hides a bug.
  • Responses tell the model what happened. Include a field such as replayed: true so the model does not narrate a second refund to the customer when only one exists.
  • Timeouts are reported as unknown, not failed. A tool that cannot confirm the outcome should say so, which pushes the orchestrator toward reconciliation rather than blind retry.

That final rule matters more than it looks. Most double-writes in agent systems start with a timeout that the tool reported as a failure, and the model reasonably believed it.

All of these rules belong in the tool registry alongside the schema. When a new tool ships without them, it should fail review the same way a tool without IAM scoping would under your agent identity and access policy.

What Systems Of Record Give You Natively

You do not have to build every guarantee yourself. Many systems of record expose a native deduplication hook, and a replay-safe tool should use it as the second line of defense behind the ledger.

The native mechanisms we lean on most often include:

SystemNative mechanismWhat it coversWhat it misses
SalesforceUpsert on an External ID fieldDuplicate record creation for the same external keySide effects such as Flows and outbound emails triggered on insert
NetSuiteexternalId on transactionsDuplicate transaction creationUpdates that re-apply amounts to existing records
HubSpotBatch upsert by unique idPropertyDuplicate contacts and objectsTimeline events and associations created alongside
ServiceNowcorrelation_id on task recordsLookups to reconcile before creating an incidentEnforcement — uniqueness is not guaranteed without a rule you add
StripeIdempotency-Key headerReplayed POSTs return the original responseKeys are pruned after 24 hours or more, so late replays write again
SQS FIFOMessage deduplication IDDuplicate sends within a five-minute windowAnything outside five minutes or upstream of the queue

Read the last column carefully. Every native mechanism has a boundary — a time window, a record type, or a side effect it does not cover — and your ledger has to span those gaps.

How Long Should Idempotency Keys Live?

Retain keys at least as long as the longest retry window in the chain. EventBridge can retry for up to 24 hours, so a ledger TTL under 24 hours leaves a gap where a late retry writes twice.

Map every retry source from the list above and write down its maximum window. Then add the longest plausible human re-run interval — for most finance and support workflows, that means days, not minutes.

We typically set DynamoDB TTL between 7 and 30 days for financial writes and 72 hours for ticketing writes. At DynamoDB on-demand pricing, a ledger holding a few million small items costs single-digit dollars per month, which is trivial next to a single duplicated credit memo.

Where Idempotency Keys Still Fail

Keys and ledgers solve most duplicate writes, yet two failure modes survive a correct implementation. Both come from getting the key's scope wrong.

Keys That Are Too Narrow

If the payload hash includes a timestamp, a model-generated summary, or a floating-point amount the model recomputed, each retry produces a new key. The ledger then faithfully records two distinct writes, both of which it believes are legitimate.

The fix is normalization. Round amounts to the currency's minor unit, sort line items, strip free text, and exclude anything the model authors in prose.

Keys That Are Too Broad

The opposite mistake collapses two genuinely separate actions into one. If a customer legitimately disputes two $4,200 charges on the same day, a key built from customer ID and amount alone will silently drop the second credit.

Accordingly, every key should include the identifier of the originating business event — the dispute ID, the ticket number, the order line. That is the field that distinguishes "the same request again" from "a new request that looks identical."

Side Effects Outside The Write

Finally, remember that the write itself is rarely the only effect. A Salesforce upsert can dedupe the record while a Flow on insert still sends the customer two confirmation emails, one per attempt.

Push side effects behind the same ledger, or trigger them from the committed ledger event rather than from the downstream insert. Our agent orchestration guide covers how to sequence those downstream steps so each one is individually keyed.

Compensations need keys too: when a duplicate slips through and the agent issues a reversal, key the reversal from the original key plus the operation name. Otherwise a retried rollback can reverse twice, and you have traded one incident for another.

How Do You Know Idempotency Is Working?

Inject failures after the downstream commit but before the acknowledgment, then confirm the retry returns the ledger's stored result and the system of record shows exactly one write.

Unit tests on the happy path prove very little here. The failure you care about only appears when the network, the timeout, and the model's re-plan all line up, so you have to manufacture that alignment deliberately.

A validation suite for replay safety should cover at least these cases:

  • Post-commit timeout. Force the tool to time out after the downstream API returns success. Expect one record and a replayed: true response on retry.
  • Paraphrased re-plan. Replay the same intent with a rewritten free-text field. Expect the same key and no second write.
  • Payload conflict. Reuse a key with a changed amount. Expect an explicit conflict error that the model surfaces rather than retries.
  • Expired lease. Kill the process mid-write, let the lease lapse, and retry. Expect reconciliation against the system of record, not re-execution.
  • Late redelivery. Replay an EventBridge event 20 hours later. Expect the ledger to still hold the key.

Run the suite in shadow mode against a sandbox before any tool gains production write access. After that, it belongs in the release gate described in our agent release management post, so a model-version change cannot quietly break the paraphrase case.

What To Watch In Production

In production, the ledger itself becomes one of your best signals. Track replay rate per tool, conflict rate per tool, and the count of pending entries older than their lease.

A replay rate that climbs from 0.5% to 6% after a deploy usually means a timeout budget shrank or a downstream API slowed down. A nonzero conflict rate almost always means a key-derivation bug, and it deserves a page, not a dashboard tile.

Emit the idempotency key on every span so a single trace in your agent observability stack shows every attempt of one intent side by side. Store it in your agent audit trail as well, so an auditor asking why a customer received one credit instead of two can see the replay and the ledger decision that suppressed it.

This is the same discipline we argue for in reliability engineering for regulated AI — retries are a reliability feature only when every retried write is provably safe to repeat. Without that proof, retry policy is just a faster way to corrupt the ledger of record.

Frequently Asked Questions

Should the language model generate the idempotency key?

No — models produce different text on retries, so a model-generated key changes between attempts. Derive the key in the tool layer from structured fields — tenant, record ID, operation, normalized payload hash — never from free text.

Is exactly-once delivery from SQS FIFO enough to prevent duplicate agent writes?

No — SQS FIFO deduplicates within a five-minute window and only at the queue boundary. A retry that arrives later, or originates in the orchestrator or the model, bypasses it entirely, so the ledger still sits in front of the write.

What happens if the process crashes between the downstream write and the ledger update?

The ledger entry stays pending. On retry, the tool queries the system of record by external ID or correlation field to reconcile, then marks the entry committed rather than issuing a second write.

Do read-only agent tools need idempotency keys?

Not for correctness, because reads do not mutate state. They still benefit from a correlation ID so traces link each read to the write it informed, which matters when reconstructing an incident.

How should compensating actions such as reversals be keyed?

Give each compensation its own key derived from the original key plus the operation name. That keeps a retried rollback from reversing twice and ties it to the write it undoes in the audit trail.

Scoping Replay Safety Before Your Next Agent Gets Write Access

Idempotency is cheapest to add before an agent touches a system of record and most expensive to retrofit after the first duplicated invoice reaches a customer. The ledger, the key policy, and the contract are a few days of engineering; the incident review is weeks.

If you're about to give an agent write access to NetSuite, Salesforce, ServiceNow, or Zendesk and want a second set of eyes on retry safety, the team at iSimplifyMe builds and operates production agent systems across ERP, CRM, and ticketing environments every week.

Reach out for a working session — we'll map every retry layer in your stack, define the key scope for each mutating tool, and leave you with a replay test suite you can run before launch. You can also start with our overview of AI agent operations.

Ready to Grow?

Let's build something extraordinary together.

Start a Project
Apex Architecture

Every site we build runs on Apex — sub-500ms, AI-native, zero maintenance.

Explore Apex Architecture

Stay Ahead of the Curve

AI strategies, case studies & industry insights — delivered monthly.

⌘ K