
Idempotency in Tool Calls: the Bug Nobody Sees Until the AI Agent Double-Charges Someone
An AI agent calls the charge_customer tool. The HTTP call takes longer than expected, the timeout fires, and the model itself (not a script, the model) decides to try again. Or the agent is midway through a long task, the conversation history goes through context compaction, and when it resumes reasoning it "forgets" it already completed that step and redoes it. Or the user simply asked "try again, I think it hung" without knowing the first call had already succeeded.
In any of these three cases, the customer gets charged twice. This isn't a hypothesis from an academic paper. It is a class of failure that shows up as soon as an agent stops answering questions and starts operating systems, and it is usually discovered after the first incident.
Why This Is Different From Traditional API Retries
Reliability engineering has always known how to handle retries: exponential backoff, circuit breakers, at-least-once delivery with consumer-side deduplication. What changes with AI agents is who decides to call the tool again.
In a traditional system, the client doing the retry is deterministic code. It knows exactly why it's trying again (timeout, 503, dropped connection) and it was usually written with some notion of idempotency in mind, or at least with a predictable retry policy.
In an AI agent, the "client" is a probabilistic reasoning process. It can call the same tool again for reasons that have nothing to do with network failure:
- The model's own uncertainty: the tool's response wasn't clear in context, and the model "prefers" trying again over admitting it doesn't know whether it worked.
- Context compaction: long agent sessions often summarize their history to fit the context window. If the summary doesn't precisely preserve "step X already succeeded", the agent can rebuild its plan from scratch and repeat steps that already had a real side effect.
- A new turn picking up the same goal: the user comes back in a different session and asks the agent to "continue" or "make sure it happened", and the agent, without perfect visibility into what already ran, redoes the work to be safe.
None of these scenarios is covered by axios-retry or a circuit breaker. The problem isn't in the transport layer; it's in how the agent decides to act.
The Pattern: An Idempotency Key at the Tool Boundary
The technique isn't new. It's the same pattern payment APIs like Stripe's have used for years: every side-effecting operation takes an idempotency key, and the backend guarantees the same key never produces the effect twice.
What changes with agents is where the key comes from. It has to be deterministic relative to the intent: the same on the original call and on any repetition of it. A random key per call defeats deduplication, because every new attempt goes through it as if it were a different charge. In order of preference:
- A natural key from the domain. If the tool charges an invoice, the invoice is the identity of the charge:
charge:${invoiceId}. The model doesn't have to remember anything, because the key comes from the request itself. - A key injected by the orchestrator. For actions with no natural identity, like sending an email, whatever runs the agent can derive the key from the task id and the plan step and pass it to the tool, out of the model's reach.
- A key chosen by the model, with instructions in the tool description. It's the last resort, and it fails exactly in the scenario that motivated the pattern: after context compaction, the model may no longer remember which key it used.
The first option also changes the tool's design. Instead of "charge $X from customer Y", it becomes "charge invoice Z", and the amount comes from the backend, not from the model:
// tool schema exposed to the agent (JSON Schema)
const chargeCustomerTool = {
name: 'charge_customer',
description:
"Charges one of the customer's invoices. Charging the same invoice again returns " +
'the result of the first charge without charging again.',
input_schema: {
type: 'object',
properties: {
invoiceId: {
type: 'string',
description: 'The invoice to charge. It is the identity of the charge.',
},
},
required: ['invoiceId'],
},
};
On the backend, the detail that breaks most implementations is the order. Checking the key, charging and only then storing it looks right, but two concurrent calls pass the check together and charge twice; the UNIQUE constraint only complains on the INSERT, after the money is already gone. The key has to be reserved before the side effect:
// charges.service.ts (NestJS + Drizzle)
async chargeInvoice(invoiceId: string): Promise<ChargeResult> {
const key = `charge:${invoiceId}`;
// 1. Reserve. The UNIQUE on idempotency_keys.key guarantees a single winner,
// even when two calls arrive at the same time.
const [reserved] = await this.db
.insert(idempotencyKeys)
.values({ key, status: 'in_progress' })
.onConflictDoNothing({ target: idempotencyKeys.key })
.returning();
if (!reserved) {
const existing = await this.db.query.idempotencyKeys.findFirst({
where: eq(idempotencyKeys.key, key),
});
// Already done: return the stored result without charging again.
if (existing?.status === 'done') return existing.result;
// Another call is in the middle of the charge: the agent gets "in progress".
throw new ConflictException('This charge is already in progress.');
}
const invoice = await this.invoices.findById(invoiceId);
// 2. The same key goes to the provider (on Stripe, the Idempotency-Key header).
// If the process dies between the charge and step 3, the next attempt doesn't charge again.
const result = await this.paymentProvider.charge(invoice.customerId, invoice.amountCents, {
idempotencyKey: key,
});
// 3. Complete the reservation with the result.
await this.db
.update(idempotencyKeys)
.set({ status: 'done', result })
.where(eq(idempotencyKeys.key, key));
return result;
}
Two protections work together here. The reservation in the database blocks a second charge while the first one is in progress and after it. The key passed to the provider covers the gap between charging and storing the result. The provider, though, only keeps the key for a while (Stripe may discard it after 24 hours), so the database reservation remains the long-term protection.
One decision is left out of the code above on purpose: a reservation stuck in in_progress because the process died midway or the call to the provider timed out. It has to expire, or the invoice can never be charged again. Since the provider also received the key, whoever takes over an expired reservation can repeat the charge safely, as long as the provider still holds the key (on Stripe, up to 24 hours after the first call); past that window, the same key may create a new charge. That is why the reservation has to expire well within that window. A definitive failure is a different case: since there is no try/catch, an exception from charge (a declined card, for instance) also leaves the reservation in in_progress, and the agent keeps getting "in progress" for a charge that has already failed. That failure has to be caught and recorded, and the next attempt needs a different key (one carrying the attempt number, for instance), because Stripe returns the same error for the same key.
Not Every Tool Needs This, but More Do Than You'd Think
The practical rule: if calling the tool twice with the same intent could leave the world in a visibly different state than calling it once, it needs idempotency. That covers much more than payments:
send_email: duplicating a transactional email is embarrassing at best; for a "your password was changed" or "refund processed" email, it can cause real panic or confusion.create_support_ticket: two identical tickets for the same issue clutter the queue and make the human team waste time figuring out it's a duplicate.post_comment,create_orderor anyINSERTtriggered by a tool: without a natural key to deduplicate against, the agent can create two rows for a single user intent.
Read-only tools (get_order_status, search_products) need none of this: calling them twice doesn't change the state of the world. The question only needs to be asked for side-effecting tools, and it needs to be asked when the tool is designed, not discovered after an incident.
Why This Matters More Now Than Two Years Ago
Context compaction, the technique that lets an agent keep a long session going without blowing past the model's window, is exactly the kind of mechanism that can erase the precise record of "this step already ran". Agent sessions are getting longer, more autonomous and less often watched by a human confirming each action before it happens.
That changes a premise of reliability engineering: you can no longer assume the caller of an API is well-behaved code with a sensible retry policy. It might be a model that decides to try again for reasons no retry policy anticipated, and the only defense that works regardless of the reason is making the action itself safe to run more than once.
Conclusion
Idempotency has always been good API practice. With AI agents executing real actions (charging, sending, creating, deleting), it stops being good practice and becomes a design requirement for any side-effecting tool, because the caller's retry pattern is no longer something you control or can predict.
Before exposing a tool to an agent, ask: what happens if it gets called twice with the same intent? If the answer bothers you, the idempotency key isn't optional. And if you get to choose, make it a key the model doesn't have to remember.
Comments
Loading comments...