Appearance
Retries and idempotency
A publish can time out after the bus has accepted it. The bus makes the retry safe, as long as you resend the same message.
How the bus recognises a retry
The bus remembers every accepted message_id, and a hash of its content, for at least 7 days.
| You send | The bus answers |
|---|---|
A message_id it hasn't seen | Checks and publishes it as usual |
A message_id it has seen and published, with the same content | 200 {"accepted":true,"duplicate":true,"message_id":"…"}. Nothing is published again |
A message_id it has seen but never published, with the same content | See A message recorded but never published |
A message_id it has seen, with any other content | 409 message_id.reused |
"Same content" is compared as JSON, ignoring key order and whitespace, and includes everything in the message: the envelope timestamp too. A message rebuilt for a retry, with a new timestamp, is a different message under a reused id.
After 7 days a message_id is forgotten, and a resend is checked as a new message. For a snapshot that usually means a refusal as stale, since its sequence_number is no longer the latest.
A message recorded but never published
The bus records a message_id just before it publishes the message, and marks it published just after. If publishing fails, it removes the record again and answers 503 bus.publish_failed, so your retry is checked and published as usual. Rarely, a record is left without a publication: the bus stops between the two steps (its function ends just after recording), or publishing fails and removing the record fails as well. A retry of such a message is never answered 200 duplicate. It gets:
| Your retry arrives | The bus answers |
|---|---|
| Less than 30 seconds after the message was recorded | 503 message.in_flight. Nothing new is accepted: the first attempt may still be publishing it. Retry with backoff |
| 30 seconds or more after | The bus publishes it and answers 202, as for a first attempt. tolerated then lists only SOM warnings, not the snapshot-order warnings of the first attempt |
30 seconds or more after, for a story.context whose story has accepted a newer snapshot since | 409 message.superseded. The old snapshot is not published, so consumers never receive an older snapshot after a newer one. The newer snapshot is complete: don't resend the old one |
- Keep retrying for at least 30 seconds. A client that gives up sooner can leave such a message unpublished. With the backoff below, capped at 30 s, a handful of attempts is enough.
- Only one retry publishes it. Two retries of the same message at the same moment can't both publish it: one gets
202, the other503message.in_flightor200duplicate. - A newer snapshot waits for the old one. While an old snapshot is being republished (usually a fraction of a second; at most 30 seconds if the bus stops again meanwhile), a newer snapshot of the same story can get
409commit.contention. Nothing was recorded for it: send it again after a moment.
One edge case remains. If the bus stops just after publishing a message and before marking it published, a retry 30 seconds or more later publishes it again. When that retry comes within 5 minutes of the first publication, the bus's topic drops it as a duplicate of the same message_id; later than that, consumers can receive it twice. For a snapshot this happens only while it is still the story's latest, so the order consumers see is never broken. Consumers de-duplicate on message_id in any case (below).
Retry safely
- Build once, keep it. Keep the exact message (or its serialised JSON) until you have a final answer, and resend exactly that.
- Retry only what can succeed: timeouts, network errors,
429and5xx. A401is worth one retry, with a new token. Nothing else: see Handle the bus's answers. - Back off exponentially, with jitter: for example 0.5 s, 1 s, 2 s, 4 s…, capped at 30 s, and give up after a bounded number of attempts. Then keep the message in your own outbox for an operator or a later resend, unchanged.
- Retry per story, in order. While one of a story's messages is being retried, hold that story's next messages. Other stories carry on.
- A refusal is final for that message. After a
409on a snapshot, don't resend it: build a new snapshot from current state, with a newmessage_id.
What consumers see
Delivery to consumers is at least once. A retried publish answered 200 duplicate is never published again, and a republished one only in the edge case above. A consumer can also see a message_id again after its own failure, a replay or a redrive. Consumers de-duplicate on message_id. See Order, duplicates and snapshots.
Test it
| Provoke | How | Expect |
|---|---|---|
| A duplicate | Send the same message twice | 200, duplicate: true |
| A lost answer | Drop the connection just after sending (a client timeout of a few milliseconds), then retry | 202 or 200 duplicate, perhaps after a 503 message.in_flight: never two deliveries |
| A reused id | Change one field and resend with the same message_id | 409 message_id.reused |
401, then recovery | Send an expired or altered token | Your client refreshes once and succeeds |
429 and 503 are hard to provoke on the Sandbox. Test those branches against a stub in your own tests.
Checkpoint
- [ ] My client retries only timeouts,
429and5xx, always with the samemessage_idand body, with backoff. - [ ] A lost answer never leads to a second message about the same change.
- [ ] Messages that exhaust their retries are kept, not dropped.
Next: Prove your error paths.