Skip to main content
Surfais keeps trying until your receiver accepts an event or the retry schedule runs out. This page covers what your receiver has to do to stay correct, and how to recover when something was missed.

At-least-once delivery

Delivery is at-least-once per endpoint. Your receiver can see the same event more than once:
  • A failed attempt is retried. Every retry carries the same X-Surfais-Event-Id and the same body bytes. Only X-Surfais-Timestamp and X-Surfais-Signature change, because each attempt is signed when it is sent.
  • A duplicate can follow a successful attempt too. If Surfais cannot record your 2xx, the delivery is sent again.
  • A replay sends the same event again on purpose.
Events can also arrive out of order. Order them by data.run_date, then by occurred_at, not by arrival. So dedupe on X-Surfais-Event-Id. It always equals id in the body. Keep a table keyed on the event id, and let one insert-or-ignore statement decide:
The same pattern in Python:
Answer a duplicate with a 2xx. Any other status tells Surfais the attempt failed, and it is retried.
Insert the id in the same transaction that stores the event for processing. Then a crash between the two steps cannot leave you with an id marked as seen and no work queued.

Retry schedule

There are 9 attempts at most: the first one and 8 retries, spread over just under 8 hours. When attempt 9 fails, the delivery becomes dead and is not retried again. You can still replay it. The first 2xx ends the schedule and the delivery becomes succeeded.

What counts as a failure

Only a 2xx status is a success. Each of these is a failed attempt, and is retried:
  • a 4xx or 5xx status
  • any 3xx status: Surfais never follows a redirect with a signed body, so point the endpoint URL at its final location
  • no answer within the attempt timeout
  • a connection error
  • the endpoint URL no longer passing the registration rules, for example because its hostname now resolves to a private address. The URL is checked again before every attempt.
The attempt timeout is 10 seconds. It covers the whole attempt, from resolving your hostname to reading your response, so answer first and process afterwards. Failed attempts also count towards the endpoint’s circuit breaker. Enough of them in a row, with no success in between, make the endpoint failing.

The delivery log

GET /webhook-endpoints/{endpointId}/deliveries lists an endpoint’s deliveries, newest first. Add ?status= to narrow it to pending, delivering, succeeded or dead. The list is paginated. Keep the same status while you page: a continuation that names a different one is 400 invalid_cursor.
Each row is one event for one endpoint.

Values of last_error

Match on the prefix. The text after it carries detail that can change. Other values can appear: last_error is diagnostic text for people, so treat a value you do not recognise as an unclassified failure.

Replay

POST /webhook-deliveries/{deliveryId}/replay queues a delivery for an immediate new attempt. It works on a delivery that is succeeded, dead or pending, and needs the write scope. The 202 response:
  • The replayed attempt carries the same event id and the same body, with a fresh timestamp and signature.
  • attempts is not reset. A replayed dead delivery gets one more attempt, and becomes dead again if that attempt fails.
  • A delivery that is being attempted right now answers 409 delivery_in_flight. Wait for it to settle, then replay.
  • A delivery whose organisation is no longer linked to your partner account answers 409 link_revoked.
  • An unknown id, or another partner’s, answers 404 not_found.
A replay has the same event id as the original, so your dedupe table treats it as a duplicate. To reprocess an event you already handled, remove its id from your table before you replay it.
This example finds every dead delivery on an endpoint and replays it. It stops paging on has_more: false, waits out a 429 rate_limited, and reports a 409 instead of failing on it.
Run the Python example as python replay_dead.py <endpointId> and the Node.js example as node replay_dead.mjs <endpointId>.

Revocation

Your access to an organisation’s events depends on its link to your partner account. The link is checked when a delivery is queued, again when it is taken off the queue, and once more immediately before the request is sent. When a link is revoked, the organisation is deleted, or your partner account is suspended:
  • no new events are queued for that organisation, from that moment
  • deliveries still waiting in the queue are never sent. They become dead with last_error set to link_revoked
  • a replay of any of its deliveries answers 409 link_revoked
If the link is granted again, or the organisation is restored, new events flow and replay works again. The only attempt a revocation cannot stop is one already in flight when it takes effect.

Timing

Events are produced when a scan finishes, so they follow each organisation’s scan schedule.
  • A day with no scan has no events. That is by design.
  • A scan you request with POST /orgs/{orgId}/brands/{brandId}/scan-requests produces its events when it completes, like any other scan. See Run a scan.
  • Each scan event type is produced at most once per own brand per scan date. A second scan on the same date does not send you a second event of a type you were already sent for that brand and date; an event type that did not fire the first time can still fire. The second scan’s results do replace the earlier ones for that date, so the numbers on GET /orgs/{orgId}/brands/summary can change with no event. After an on-demand scan on a day that was already scanned, poll instead. See Request a scan.

Retention

Events and their deliveries are kept for a limited time and then removed. After that a delivery can no longer be listed or replayed. Do not use the delivery log as your archive: store what you need when you receive it.