The first id in a payment should be yours, not the processor’s.
Most teams get this backwards. The code calls the processor, waits, gets back a ch_xxx or a txn_xxx, and writes that down as the id of the payment. The processor’s id becomes the primary key of money in your system.
That works until it doesn’t. And the way it stops working is always the same: a request times out, someone retries, and now you have two charges and one order. Or you have one charge and no record of it at all.
There are two mental models here.
The false one: a retry is safe because the network failed, so nothing happened. The processor gives me an id, so I have a handle on the payment.
The real one: a retry is a second attempt at moving real money, and you have no idea whether the first one landed. You need a handle on the payment before you make the call, because the call is exactly the part that might not come back.
So: generate your own payment id first. Treat it as the source of truth. Everything the processor tells you later — ids, statuses, webhooks — gets attached to your id, not the other way around.
Timeouts are how you get doubles
A timeout is not a failure. A timeout is the absence of information.
Walk through it. Your service posts a charge. Thirty seconds pass. Your HTTP client gives up and raises. What do you know?
You know your request left. You don’t know if it arrived. You don’t know if the processor authorized the card. You don’t know if the response is sitting in a buffer somewhere, or if it got written to the processor’s ledger a millisecond after your client stopped listening.
All four of these are consistent with the same exception object:
- Request never arrived. Nothing happened.
- Request arrived, processor failed. Nothing happened.
- Request arrived, processor charged the card, response lost. Money moved.
- Request arrived, processor charged the card, response slow. Money moved, and the response may still show up.
Your code sees one thing. Reality has four branches, and two of them took the customer’s money.
Now the retry. Something retries — your own retry wrapper, a queue redelivery, a user hitting the button again, an ops person re-running a failed job. Without an idempotency key, the processor has no way to know this is the same payment. It’s a fresh, valid, well-formed charge request. So it charges again.
Nobody wrote a bug. Every layer did what it was told. The customer got billed twice.
The processor’s id arrives too late
Here’s the structural problem with using the processor’s id as your identifier: you only get it on a successful response. The cases where you most need an identifier are exactly the cases where no response comes back.
You need to be able to say “payment 7f3a” while the call is in flight. You need it so you can:
- Write a row before you call, so a crash mid-call leaves a trace
- Log it, so support can find the attempt
- Match the webhook that arrives before your response does (this happens more than people expect)
- Decide, on retry, whether you’re repeating work or starting new work
An id that only exists after success is useless for handling failure. That’s the whole job.
And there’s a second reason, less dramatic but it compounds: provider ids are the provider’s. Switch processors, add a second one for redundancy, route by currency or region — and suddenly your payments table has two id formats, two shapes of null, and reconciliation queries with CASE statements in them. Your own id is stable across all of that. The provider id becomes just another column.
Idempotency keys are not optional
Every serious payment API supports an idempotency key. You pass it in a header or a field, and the processor promises: same key, same outcome. If it already processed that key, it returns the original result instead of doing the work again.
Send your payment id as the key.
That’s it. That’s the whole mechanism. One line of plumbing turns “retry might double-charge” into “retry is safe.”
But a few things matter about how you do it.
Generate the key before the call, not inside the retry loop. If each attempt makes a new key, you have no idempotency — you have a counter of charges. The key belongs to the payment, not the attempt.
Persist it before you call. Write your row, commit, then call the processor. If your process dies mid-call, the row is still there and a reconciliation job can go ask the processor what happened to that key. If you only hold the key in memory, a crash erases your only link to real money.
Keep the key scoped to one intent. Same key means same outcome, so if you reuse a key for a different amount, you’ll either get an error or get the old amount back. Both are confusing at 2am. One payment, one key, forever.
Know the retention window. Processors expire idempotency keys — often 24 hours, sometimes less. A retry a week later is not deduplicated. If your retries can span days, you need your own state check before you call, not just the processor’s.
Webhooks and retries need the same thing
Your own id does double duty. It protects the outbound call, and it’s also what makes inbound events safe to handle.
Webhooks get delivered more than once. That’s by design — the sender retries until it gets a 200, and it will happily send you the same event twice because your ack got lost. If your handler does work per delivery instead of per payment, you’ll mark things settled twice, send two receipts, or credit an account twice.
The fix is the same shape: every event resolves to your payment id, and the handler is written as “move this payment to this state if it isn’t already there” rather than “do this thing now.” Idempotent writes, not idempotent-ish writes.
Two things we’ve written about before and won’t rehash: a webhook arriving is not the same as money settling, and a refund is its own money movement that needs its own id. Both are true and both depend on this post. If you don’t have your own id, you have nothing to hang either one on.
What good looks like
It’s not much code. A table and a state machine.
create table payments (
id uuid primary key, -- yours. generated before any call.
idempotency_key text not null unique, -- usually just id, as text
provider text not null,
provider_id text, -- nullable. arrives later, or never.
amount_cents bigint not null,
currency char(3) not null,
state text not null, -- requested | submitted | settled | failed
attempts int not null default 0,
created_at timestamptz not null default now(),
updated_at timestamptz not null default now()
);
The lifecycle:
- requested — you wrote the row. No call has gone out. This is your “I intend to move money” record.
- submitted — you called the processor. You do not know the outcome. This is the honest state for a timeout, and most systems are missing it. They go straight from nothing to settled-or-failed, which means a timeout has nowhere to live.
- settled — confirmed by the processor, with
provider_idfilled in. - failed — confirmed failure. Not “we didn’t hear back.” Confirmed.
Two rules make it work:
Only the processor’s confirmation moves you out of submitted. Not a timeout, not a guess, not a 500. Something in submitted for too long is a reconciliation job’s problem — go ask the processor about that idempotency key and find out what actually happened. Don’t let a timeout write failed; that’s how you lose money you already took.
Transitions are idempotent. settled → settled is a no-op, not a second receipt. Write the handler so replaying every event you’ve ever received produces the same end state. That property is what lets you replay events at all, and one day you will want to.
That’s it. One id, four states, one nullable provider column.
Boundary
Futurify builds software. We’re not a payment processor and we don’t hold funds. The money moves through your processor, under your merchant account, on your rails.
What we do is the part around it: the ids, the states, the retries, the reconciliation jobs, the thing that tells you at 9am which payments are stuck in submitted and need a human. That’s Build. If you want someone watching those jobs and fixing them when they break, that’s Run.
Next step
Open your codebase and answer these. They take about ten minutes and they’re more informative than any audit.
- Is there a row in your database before the processor call goes out? If the first write happens after the response, a crash or a timeout leaves you with money moved and no record of it.
- Do you send an idempotency key, and is it the same across retries? Grep for your retry wrapper. If the key is generated inside it, you don’t have idempotency.
- What state does a timeout write? If the answer is
failed, you’re marking real charges as failures. If the answer is “nothing, it throws,” you’re losing them entirely.
If any of those answers made you uncomfortable, that’s normal — most systems grow in this order and the gap only shows up under load. Email us at hello@futurify.io and we’ll talk through it.