Skip to content

Available now

Retries and idempotency ​

A publish can time out after the bus has accepted it. The bus makes the retry safe, as long as you resend the same message.

How the bus recognises a retry ​

The bus remembers every accepted message_id, and a hash of its content, for at least 7 days.

You sendThe bus answers
A message_id it hasn't seenChecks and publishes it as usual
A message_id it has seen and published, with the same content200 {"accepted":true,"duplicate":true,"message_id":"…"}. Nothing is published again
A message_id it has seen but never published, with the same contentSee A message recorded but never published
A message_id it has seen, with any other content409 message_id.reused

"Same content" is compared as JSON, ignoring key order and whitespace, and includes everything in the message: the envelope timestamp too. A message rebuilt for a retry, with a new timestamp, is a different message under a reused id.

After 7 days a message_id is forgotten, and a resend is checked as a new message. For a snapshot that usually means a refusal as stale, since its sequence_number is no longer the latest.

A message recorded but never published ​

The bus records a message_id just before it publishes the message, and marks it published just after. If publishing fails, it removes the record again and answers 503 bus.publish_failed, so your retry is checked and published as usual. Rarely, a record is left without a publication: the bus stops between the two steps (its function ends just after recording), or publishing fails and removing the record fails as well. A retry of such a message is never answered 200 duplicate. It gets:

Your retry arrivesThe bus answers
Less than 30 seconds after the message was recorded503 message.in_flight. Nothing new is accepted: the first attempt may still be publishing it. Retry with backoff
30 seconds or more afterThe bus publishes it and answers 202, as for a first attempt. tolerated then lists only SOM warnings, not the snapshot-order warnings of the first attempt
30 seconds or more after, for a story.context whose story has accepted a newer snapshot since409 message.superseded. The old snapshot is not published, so consumers never receive an older snapshot after a newer one. The newer snapshot is complete: don't resend the old one
  • Keep retrying for at least 30 seconds. A client that gives up sooner can leave such a message unpublished. With the backoff below, capped at 30 s, a handful of attempts is enough.
  • Only one retry publishes it. Two retries of the same message at the same moment can't both publish it: one gets 202, the other 503 message.in_flight or 200 duplicate.
  • A newer snapshot waits for the old one. While an old snapshot is being republished (usually a fraction of a second; at most 30 seconds if the bus stops again meanwhile), a newer snapshot of the same story can get 409 commit.contention. Nothing was recorded for it: send it again after a moment.

One edge case remains. If the bus stops just after publishing a message and before marking it published, a retry 30 seconds or more later publishes it again. When that retry comes within 5 minutes of the first publication, the bus's topic drops it as a duplicate of the same message_id; later than that, consumers can receive it twice. For a snapshot this happens only while it is still the story's latest, so the order consumers see is never broken. Consumers de-duplicate on message_id in any case (below).

Retry safely ​

  • Build once, keep it. Keep the exact message (or its serialised JSON) until you have a final answer, and resend exactly that.
  • Retry only what can succeed: timeouts, network errors, 429 and 5xx. A 401 is worth one retry, with a new token. Nothing else: see Handle the bus's answers.
  • Back off exponentially, with jitter: for example 0.5 s, 1 s, 2 s, 4 s…, capped at 30 s, and give up after a bounded number of attempts. Then keep the message in your own outbox for an operator or a later resend, unchanged.
  • Retry per story, in order. While one of a story's messages is being retried, hold that story's next messages. Other stories carry on.
  • A refusal is final for that message. After a 409 on a snapshot, don't resend it: build a new snapshot from current state, with a new message_id.

What consumers see ​

Delivery to consumers is at least once. A retried publish answered 200 duplicate is never published again, and a republished one only in the edge case above. A consumer can also see a message_id again after its own failure, a replay or a redrive. Consumers de-duplicate on message_id. See Order, duplicates and snapshots.

Test it ​

ProvokeHowExpect
A duplicateSend the same message twice200, duplicate: true
A lost answerDrop the connection just after sending (a client timeout of a few milliseconds), then retry202 or 200 duplicate, perhaps after a 503 message.in_flight: never two deliveries
A reused idChange one field and resend with the same message_id409 message_id.reused
401, then recoverySend an expired or altered tokenYour client refreshes once and succeeds

429 and 503 are hard to provoke on the Sandbox. Test those branches against a stub in your own tests.

Checkpoint ​

  • [ ] My client retries only timeouts, 429 and 5xx, always with the same message_id and body, with backoff.
  • [ ] A lost answer never leads to a second message about the same change.
  • [ ] Messages that exhaust their retries are kept, not dropped.

Next: Prove your error paths.

SOM is an open standard maintained by the SOM working group. This service is not endorsed by it.