> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfais.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Delivery, retries and replay

> How webhooks are delivered: at-least-once delivery, the retry schedule, what counts as a failure, the delivery log, replay and revocation.

Surfais keeps trying until your receiver accepts an event or the retry schedule runs out. This page covers what your receiver has to do to stay correct, and how to recover when something was missed.

## At-least-once delivery

Delivery is **at-least-once** per endpoint. Your receiver can see the same event more than once:

* A failed attempt is retried. Every retry carries the **same** `X-Surfais-Event-Id` and the **same body bytes**. Only `X-Surfais-Timestamp` and `X-Surfais-Signature` change, because each attempt is signed when it is sent.
* A duplicate can follow a **successful** attempt too. If Surfais cannot record your `2xx`, the delivery is sent again.
* A [replay](#replay) sends the same event again on purpose.

Events can also arrive out of order. Order them by `data.run_date`, then by `occurred_at`, not by arrival.

So dedupe on `X-Surfais-Event-Id`. It always equals `id` in the body. Keep a table keyed on the event id, and let one insert-or-ignore statement decide:

```sql theme={null}
CREATE TABLE surfais_events (
  event_id    text PRIMARY KEY,
  received_at timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP
);

-- One row inserted: first time you have seen this event. Process it.
-- No row inserted: a duplicate. Answer 2xx and skip it.
INSERT INTO surfais_events (event_id) VALUES (:event_id)
ON CONFLICT (event_id) DO NOTHING;
```

The same pattern in Python:

```python theme={null}
import sqlite3

db = sqlite3.connect("webhooks.db")
db.execute(
    "CREATE TABLE IF NOT EXISTS surfais_events ("
    "event_id TEXT PRIMARY KEY, "
    "received_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP)"
)


def first_time_seen(event_id):
    """Insert-or-ignore. True the first time an id arrives, False for every duplicate."""
    with db:
        cursor = db.execute(
            "INSERT OR IGNORE INTO surfais_events (event_id) VALUES (?)",
            (event_id,),
        )
    return cursor.rowcount == 1
```

Answer a duplicate with a `2xx`. Any other status tells Surfais the attempt failed, and it is retried.

<Tip>
  Insert the id in the same transaction that stores the event for processing. Then a crash between the two steps cannot leave you with an id marked as seen and no work queued.
</Tip>

## Retry schedule

| Attempt | When it is sent                   |
| ------- | --------------------------------- |
| 1       | When the event is produced        |
| 2       | 30 seconds after attempt 1 failed |
| 3       | 2 minutes after attempt 2 failed  |
| 4       | 5 minutes after attempt 3 failed  |
| 5       | 15 minutes after attempt 4 failed |
| 6       | 30 minutes after attempt 5 failed |
| 7       | 1 hour after attempt 6 failed     |
| 8       | 2 hours after attempt 7 failed    |
| 9       | 4 hours after attempt 8 failed    |

There are 9 attempts at most: the first one and 8 retries, spread over just under 8 hours. When attempt 9 fails, the delivery becomes `dead` and is not retried again. You can still [replay](#replay) it.

The first `2xx` ends the schedule and the delivery becomes `succeeded`.

## What counts as a failure

Only a `2xx` status is a success. Each of these is a failed attempt, and is retried:

* a `4xx` or `5xx` status
* **any `3xx` status**: Surfais never follows a redirect with a signed body, so point the endpoint URL at its final location
* no answer within the attempt timeout
* a connection error
* the endpoint URL no longer passing the registration rules, for example because its hostname now resolves to a private address. The URL is checked again before every attempt.

The attempt timeout is 10 seconds. It covers the whole attempt, from resolving your hostname to reading your response, so [answer first and process afterwards](/api/webhooks/verify-signatures#respond-fast).

Failed attempts also count towards the endpoint's circuit breaker. Enough of them in a row, with no success in between, make the endpoint [`failing`](/api/webhooks/overview#endpoint-states).

## The delivery log

`GET /webhook-endpoints/{endpointId}/deliveries` lists an endpoint's deliveries, newest first. Add `?status=` to narrow it to `pending`, `delivering`, `succeeded` or `dead`. The list is [paginated](/api/pagination). Keep the same `status` while you page: a continuation that names a different one is `400 invalid_cursor`.

```json theme={null}
{
  "data": [
    {
      "id": "a8c6e4d2-1f3b-4a79-9c5e-7d2b0f6a4e13",
      "event_id": "evt_01M319GV4Y8T5V6W7Y8ZQR4X2M",
      "event_type": "scan.completed",
      "endpoint_id": "e5f1a9c3-2d6b-4f80-a7c4-1b9e3d5f7a20",
      "subscription_id": "2e7c9b14-5a3d-4f68-9b02-8d1e6f4a3c57",
      "status": "pending",
      "attempts": 3,
      "next_attempt_at": "2026-09-21T06:20:21Z",
      "last_status_code": 503,
      "last_duration_ms": 412,
      "last_error": "http_503",
      "last_response_snippet": "upstream unavailable",
      "created_at": "2026-09-21T06:12:44Z",
      "updated_at": "2026-09-21T06:15:21Z"
    }
  ],
  "pagination": { "next_cursor": null, "has_more": false }
}
```

Each row is one event for one endpoint.

| Field                   | Description                                                                                                                                                                                  |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `event_id`              | The event's `id`, as sent in the body and in `X-Surfais-Event-Id`.                                                                                                                           |
| `event_type`            | One of the four event types, `webhook.test` included.                                                                                                                                        |
| `subscription_id`       | The subscription that matched. `null` for a test ping, and `null` once that subscription has been deleted.                                                                                   |
| `status`                | `pending`: waiting, due at `next_attempt_at`. `delivering`: an attempt is in flight. `succeeded`: your receiver answered `2xx`. `dead`: the schedule ran out, or the delivery was cancelled. |
| `attempts`              | Attempts made so far.                                                                                                                                                                        |
| `next_attempt_at`       | When the next attempt is due. Only meaningful while `status` is `pending`.                                                                                                                   |
| `last_status_code`      | The HTTP status of the last attempt. `null` if no response arrived.                                                                                                                          |
| `last_duration_ms`      | How long the last attempt took, in milliseconds. Can be `null` when no request was sent.                                                                                                     |
| `last_error`            | Why the last attempt failed. `null` after a success.                                                                                                                                         |
| `last_response_snippet` | The first bytes of your last response body. Put a short, useful message in your error responses and you can read it here.                                                                    |

### Values of `last_error`

| `last_error` starts with       | What happened                                                                                                                                                                                                                                                                       |
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `http_`                        | Your receiver answered with a `4xx` or `5xx` status. The status follows, for example `http_503`.                                                                                                                                                                                    |
| `redirect_refused`             | Your receiver answered with a `3xx`. The redirect was not followed.                                                                                                                                                                                                                 |
| `timeout`                      | Your receiver did not finish answering within the attempt timeout.                                                                                                                                                                                                                  |
| `request_failed`               | The connection failed before a response arrived.                                                                                                                                                                                                                                    |
| `url_refused`                  | The endpoint URL failed its check before sending, for example because it no longer resolves to a public address. Nothing was sent.                                                                                                                                                  |
| `url_check_timeout`            | The check of the endpoint URL ran out of time, typically because the DNS lookup for its hostname did not answer in time. Nothing was sent.                                                                                                                                          |
| `endpoint_not_active`          | The endpoint was `disabled`, or not yet verified, so the delivery was held. Nothing was sent, and the attempt does not count towards the circuit breaker. It still follows the retry schedule and becomes `dead` when that runs out. Replay it once the endpoint is `active` again. |
| `url_check_failed`             | The check of the endpoint URL could not be completed. Detail follows. Nothing was sent; the delivery is retried on the schedule.                                                                                                                                                    |
| `scan_not_completed`           | The scan that produced the event did not complete, so the event was withdrawn. Nothing was sent, and the delivery is `dead`.                                                                                                                                                        |
| `orphaned_after_final_attempt` | The last attempt was interrupted on our side. The delivery was closed as `dead` instead of being sent again. Replay it.                                                                                                                                                             |
| `link_revoked`                 | The organisation is no longer linked to your partner account. See [Revocation](#revocation). Final: the delivery is `dead` and is not retried.                                                                                                                                      |

Match on the prefix. The text after it carries detail that can change. Other values can appear: `last_error` is diagnostic text for people, so treat a value you do not recognise as an unclassified failure.

## Replay

`POST /webhook-deliveries/{deliveryId}/replay` queues a delivery for an immediate new attempt. It works on a delivery that is `succeeded`, `dead` or `pending`, and needs the `write` scope.

The `202` response:

```json theme={null}
{
  "data": {
    "id": "a8c6e4d2-1f3b-4a79-9c5e-7d2b0f6a4e13",
    "status": "pending",
    "next_attempt_at": "2026-09-21T15:03:10Z",
    "attempts": 9
  }
}
```

* The replayed attempt carries the **same event id and the same body**, with a fresh timestamp and signature.
* `attempts` is **not reset**. A replayed `dead` delivery gets one more attempt, and becomes `dead` again if that attempt fails.
* A delivery that is being attempted right now answers `409 delivery_in_flight`. Wait for it to settle, then replay.
* A delivery whose organisation is no longer linked to your partner account answers `409 link_revoked`.
* An unknown id, or another partner's, answers `404 not_found`.

<Warning>
  A replay has the same event id as the original, so your dedupe table treats it as a duplicate. To reprocess an event you already handled, remove its id from your table before you replay it.
</Warning>

This example finds every `dead` delivery on an endpoint and replays it. It stops paging on `has_more: false`, waits out a `429 rate_limited`, and reports a `409` instead of failing on it.

<CodeGroup>
  ```bash cURL theme={null}
  # List the dead deliveries on an endpoint
  curl "https://api.surfais.com/v1/webhook-endpoints/$ENDPOINT_ID/deliveries?status=dead" \
    -H "Authorization: Bearer $SURFAIS_API_KEY"

  # Replay one of them
  curl -X POST "https://api.surfais.com/v1/webhook-deliveries/$DELIVERY_ID/replay" \
    -H "Authorization: Bearer $SURFAIS_API_KEY"
  ```

  ```python Python theme={null}
  import os
  import sys
  import time

  import requests

  BASE_URL = "https://api.surfais.com/v1"
  HEADERS = {"Authorization": f"Bearer {os.environ['SURFAIS_API_KEY']}"}


  def call(method, path, params=None):
      """Send one request. Wait and try again on 429 rate_limited."""
      while True:
          response = requests.request(method, BASE_URL + path, headers=HEADERS, params=params, timeout=30)
          if response.status_code == 429 and response.json()["error"]["code"] == "rate_limited":
              time.sleep(int(response.headers.get("Retry-After", "1")))
              continue
          return response


  def dead_deliveries(endpoint_id):
      params = {"status": "dead"}
      while True:
          response = call("GET", f"/webhook-endpoints/{endpoint_id}/deliveries", params)
          response.raise_for_status()
          page = response.json()
          yield from page["data"]
          if not page["pagination"]["has_more"]:
              return
          # Continue with the cursor alone. It carries the status filter.
          params = {"cursor": page["pagination"]["next_cursor"]}


  def replay(delivery_id):
      response = call("POST", f"/webhook-deliveries/{delivery_id}/replay")
      if response.status_code == 202:
          return "replayed"
      if response.status_code == 409:
          return response.json()["error"]["code"]  # delivery_in_flight or link_revoked
      response.raise_for_status()
      return f"unexpected status {response.status_code}"


  if __name__ == "__main__":
      # Collect first, then replay: a replayed delivery leaves the dead list.
      for delivery in list(dead_deliveries(sys.argv[1])):
          print(delivery["id"], delivery["event_id"], delivery["last_error"], "->", replay(delivery["id"]))
  ```

  ```javascript Node.js theme={null}
  const BASE_URL = "https://api.surfais.com/v1";
  const HEADERS = { Authorization: `Bearer ${process.env.SURFAIS_API_KEY}` };

  // Send one request. Wait and try again on 429 rate_limited.
  async function call(method, path) {
    for (;;) {
      const response = await fetch(BASE_URL + path, { method, headers: HEADERS });
      const body = await response.json();
      if (response.status === 429 && body.error.code === "rate_limited") {
        const seconds = Number(response.headers.get("Retry-After") ?? "1");
        await new Promise((resolve) => setTimeout(resolve, seconds * 1000));
        continue;
      }
      return { status: response.status, body };
    }
  }

  async function deadDeliveries(endpointId) {
    const deliveries = [];
    let query = "?status=dead";
    for (;;) {
      const { status, body } = await call("GET", `/webhook-endpoints/${endpointId}/deliveries${query}`);
      if (status !== 200) throw new Error(`${status} ${body.error.code} (request ${body.request_id})`);
      deliveries.push(...body.data);
      if (!body.pagination.has_more) return deliveries;
      // Continue with the cursor alone. It carries the status filter.
      query = `?cursor=${encodeURIComponent(body.pagination.next_cursor)}`;
    }
  }

  async function replay(deliveryId) {
    const { status, body } = await call("POST", `/webhook-deliveries/${deliveryId}/replay`);
    if (status === 202) return "replayed";
    if (status === 409) return body.error.code; // delivery_in_flight or link_revoked
    throw new Error(`${status} ${body.error.code} (request ${body.request_id})`);
  }

  // Collect first, then replay: a replayed delivery leaves the dead list.
  for (const delivery of await deadDeliveries(process.argv[2])) {
    console.log(delivery.id, delivery.event_id, delivery.last_error, "->", await replay(delivery.id));
  }
  ```
</CodeGroup>

Run the Python example as `python replay_dead.py <endpointId>` and the Node.js example as `node replay_dead.mjs <endpointId>`.

## Revocation

Your access to an organisation's events depends on its link to your partner account. The link is checked when a delivery is queued, again when it is taken off the queue, and once more immediately before the request is sent.

When a link is revoked, the organisation is deleted, or your partner account is suspended:

* no new events are queued for that organisation, from that moment
* deliveries still waiting in the queue are **never sent**. They become `dead` with `last_error` set to `link_revoked`
* a replay of any of its deliveries answers `409 link_revoked`

If the link is granted again, or the organisation is restored, new events flow and replay works again. The only attempt a revocation cannot stop is one already in flight when it takes effect.

## Timing

Events are produced when a scan finishes, so they follow each organisation's [scan schedule](/scans/scan-schedule).

* A day with no scan has no events. That is by design.
* A scan you request with `POST /orgs/{orgId}/brands/{brandId}/scan-requests` produces its events when it completes, like any other scan. See [Run a scan](/api/guides/run-a-scan).
* Each scan event type is produced at most once per own brand per scan date. A second scan on the same date does not send you a second event of a type you were already sent for that brand and date; an event type that did not fire the first time can still fire. The second scan’s results do replace the earlier ones for that date, so the numbers on `GET /orgs/{orgId}/brands/summary` can change with no event. After an on-demand scan on a day that was already scanned, poll instead. See [Request a scan](/api/guides/run-a-scan).

## Retention

Events and their deliveries are kept for a limited time and then removed. After that a delivery can no longer be listed or replayed. Do not use the delivery log as your archive: store what you need when you receive it.
