← Back to Writing

A Payment Timeout Is Not a Decline

I used to think of a payment call as having two useful outcomes: it worked, or it failed.

That model breaks as soon as the payment provider runs in another process.

In CommerceCore, the application sends an authorization request to a separate Payment Service over gRPC. The provider can commit the authorization to its own PostgreSQL database and then, before the response reaches CommerceCore, the RPC can hit its deadline.

From CommerceCore's point of view, the call timed out. From the provider's point of view, the customer has already been authorized. Those are both true.

That small gap between them ended up shaping most of the payment design.

The dangerous version

The tempting implementation looks roughly like this:

authorize payment

if success:
    payment = AUTHORIZED
else:
    payment = FAILED

A timeout goes through the else. That is wrong.

A timeout tells me that I did not receive a result before my deadline. It does not tell me that the provider rejected the payment. The request may have died before reaching the provider, failed there, or succeeded before the response was lost.

CommerceCore therefore has a third outcome: UNKNOWN. It is a real business state: an authorization may have happened, and the application lacks enough information to make a destructive decision.

Why FAILED would be unsafe

CommerceCore connects payment state to inventory reservations and order state. A proven failure permits cancellation and eventual inventory release. Misclassifying an ambiguous timeout as FAILED could release stock even though the provider authorized the payment. The customer is then charged for an order the application believes it cancelled.

The opposite shortcut is unsafe too. Marking every timeout AUTHORIZED would advance an order on unproven payment. The system has to preserve uncertainty until it gets better evidence.

Retrying creates another problem

Sending the request again is safe only if repetition cannot create another authorization. CommerceCore gives every authorization a stable provider request ID derived from its payment identity. The Payment Service persists that identity with the result.

same logical payment
+
same amount
=
same provider result

A repeated provider request with a different amount is a conflict, not a retry. The decision lives in durable provider state, so restarting the Payment Service does not erase it. Process-local deduplication solves almost none of the failures I care about; the interesting retry often follows a crash.

Reconciliation should ask, not act

Suppose CommerceCore stores UNKNOWN while the provider stores AUTHORIZED. My reconciliation path uses a provider lookup operation. It does not call authorization again.

reconciliation asks what happened
reconciliation does not make it happen again

Reconciliation recovers knowledge. If it repeats the original side effect, a mechanism intended to resolve ambiguity can produce a new ambiguous side effect. Once CommerceCore retrieves provider truth, it transitions based on evidence rather than what the original RPC appeared to do.

At-least-once shows up again

A provider can deliver a webhook more than once. CommerceCore stores event receipts and makes the business transition idempotent instead of pretending the transport is exactly-once.

The transactional outbox uses the same philosophy. A business fact and the intent to publish it commit in one local PostgreSQL transaction. Kafka publication happens afterward and is at-least-once. An event can reach a consumer and then be retried before the publisher records completion, so consumers tolerate duplicate delivery.

Authorization after inventory expiry

Another race makes the state machine less tidy:

inventory reserved
payment authorization starts
reservation expires
inventory becomes available again
payment provider says AUTHORIZED

Automatically confirming is unsafe because the stock may now belong to someone else. Pretending the payment never happened is also wrong. CommerceCore moves the case to REQUIRES_REVIEW and avoids automatically mutating inventory.

Sometimes correct software preserves an uncomfortable fact rather than forcing reality into a clean success/failure enum.

What PostgreSQL owns

Inside CommerceCore, PostgreSQL provides a strong local boundary. Checkout can atomically create the order, attach reservations, persist the idempotency mapping, and write the outbox intent.

That transaction cannot also own an independently operated Payment Service and Kafka without introducing a distributed transaction protocol. CommerceCore's database owns CommerceCore state. The Payment Service database owns provider state. Kafka distributes already-committed facts. Idempotency, durable intermediate states, and reconciliation cover the gaps.

The test that made it concrete

One failure case runs the Payment Service over a real localhost TCP gRPC connection. The provider commits an authorization, but CommerceCore's deadline expires before it receives success.

Payment Service: AUTHORIZED
CommerceCore:    UNKNOWN

Reconciliation queries the provider's persisted state and resolves the ambiguity without issuing another authorization. The successful path proves the API works. The timeout path proves the state model means something.

What changed in how I think about remote calls

I no longer treat an RPC error as a business result. A deadline exceeded means the transport stopped waiting; it says nothing definitive about whether the remote side committed.

Once I treated that distinction seriously, UNKNOWN, durable request identity, lookup-based reconciliation, at-least-once consumers, and REQUIRES_REVIEW stopped looking like incidental complexity. The payment API did not get simpler. Its behavior got more honest.

Related project: CommerceCore case study →