> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfais.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate limits and quotas

> The two request allowances on every API key, the headers that report them, and how to back off on 429 and 503.

Every API key has two request allowances. Each response tells you where you stand, and a refusal tells you how long to wait.

## Two allowances per key

| Allowance           | Window                     | When it is used up   |
| ------------------- | -------------------------- | -------------------- |
| Requests per minute | One clock minute           | `429 rate_limited`   |
| Requests per month  | One calendar month, in UTC | `429 quota_exceeded` |

Both are set on the key when it is issued, so two keys can have different allowances. The monthly quota can be unlimited.

Read your per-minute allowance from the `X-RateLimit-Limit` header on any response. The monthly quota is agreed when the key is issued and is not reported in a header. If you do not know yours, [ask support](/api/support).

The per-minute window is a fixed clock minute, not a rolling one. The allowance refills when the minute ends, at the time in `X-RateLimit-Reset`.

## The rate limit headers

Every response to a request the API counted carries three headers, on a success and on an error alike:

| Header                  | Value                                                           |
| ----------------------- | --------------------------------------------------------------- |
| `X-RateLimit-Limit`     | Requests this key may make per minute.                          |
| `X-RateLimit-Remaining` | Requests left in the current minute. Never below `0`.           |
| `X-RateLimit-Reset`     | Unix time, in seconds, at which the current minute window ends. |

All three describe the per-minute window. None of them describes the monthly quota.

Responses sent before the request is counted do not carry them: `401`, `405` and `503 api_unavailable`. Do not expect them on `503 rate_limit_unavailable` either: it is usually sent before the request could be counted. Treat them as optional on any `5xx`.

## When you go over

| Response             | Meaning                              | `Retry-After`                                                             |
| -------------------- | ------------------------------------ | ------------------------------------------------------------------------- |
| `429 rate_limited`   | The per-minute allowance is used up. | Seconds until the current minute ends.                                    |
| `429 quota_exceeded` | The monthly quota is used up.        | Seconds until the calendar month ends, in UTC. That can be days or weeks. |

```json theme={null}
{
  "error": {
    "code": "rate_limited",
    "message": "Per-minute rate limit exceeded."
  },
  "request_id": "0b8f6c2e-5d1a-4f3b-9c7e-2a4d6e8f0b13"
}
```

Both carry `Retry-After`, in whole seconds. Wait that long before you send the request again.

A request refused by the per-minute limit does not count towards the monthly quota.

On `429 quota_exceeded` the `X-RateLimit-*` headers still describe the minute window. Use `Retry-After`, not `X-RateLimit-Reset`, to know when the quota comes back. Do not sleep through it. Stop the job and schedule it for the new month, or ask for a higher quota.

## What counts

Every authenticated request is charged to the per-minute allowance. That includes requests the API goes on to refuse: a `400` for a bad parameter, a `403` for a missing scope, a `404` for an id it cannot find, a `413` for an oversized body. A replayed [idempotent request](/api/idempotency) counts too.

A request that the per-minute limit lets through is then charged to the monthly quota. There is one exception: `403 tier_not_entitled`, answered when the key's organisation has lost API access, is charged to the minute only.

A `401` is not charged to any key. Failed authentication is throttled by client address instead. After too many `401` responses from one address in a minute, that address gets `429 rate_limited` for the rest of the minute, even with a valid key. The `X-RateLimit-*` headers on that response describe the address's allowance of failed attempts. See [Authentication](/api/authentication).

## When the API cannot count, it refuses

The API never serves a request it could not meter. If it cannot check your key or count the request, it refuses the request instead. Every `503` carries `Retry-After`.

| Response                     | Meaning                                                                             | `X-RateLimit-*` headers                                    |
| ---------------------------- | ----------------------------------------------------------------------------------- | ---------------------------------------------------------- |
| `503 rate_limit_unavailable` | The API could not check your key or count the request.                              | Do not expect them. The request may not have been counted. |
| `503 store_unavailable`      | The request was let through and counted, but the data store did not answer in time. | Present.                                                   |
| `503 api_unavailable`        | The API is switched off. Every request gets this answer.                            | Absent.                                                    |

<Warning>
  Do not write a client that requires the `X-RateLimit-*` headers on every response. A `503` can arrive without them. Read them when they are present.
</Warning>

## Back off and retry

This helper sends one request and retries it on `429` and `503`. It waits for `Retry-After`. If the header is missing, it falls back to exponential backoff with jitter. It gives up after a fixed number of attempts, and it raises instead of sleeping when the wait is long, which is what a spent monthly quota looks like.

Retrying a read is always safe. Before you retry a write, give it an [`Idempotency-Key`](/api/idempotency).

<CodeGroup>
  ```python Python theme={null}
  import os
  import random
  import time

  import requests

  BASE_URL = "https://api.surfais.com/v1"
  API_KEY = os.environ["SURFAIS_API_KEY"]
  MAX_ATTEMPTS = 5
  MAX_WAIT_SECONDS = 120


  class RetryLater(Exception):
      """The API asked for a wait too long to sleep through, such as a spent monthly quota."""

      def __init__(self, code, retry_after, request_id):
          super().__init__(f"{code}: retry in {retry_after} s (request {request_id})")
          self.code = code
          self.retry_after = retry_after
          self.request_id = request_id


  def request_with_backoff(method, path, **kwargs):
      """Send one request. On 429 or 503, wait and try again, up to MAX_ATTEMPTS times."""
      headers = {"Authorization": f"Bearer {API_KEY}", **kwargs.pop("headers", {})}
      for attempt in range(MAX_ATTEMPTS):
          response = requests.request(
              method, f"{BASE_URL}{path}", headers=headers, timeout=30, **kwargs
          )
          if response.status_code not in (429, 503) or attempt == MAX_ATTEMPTS - 1:
              return response

          retry_after = response.headers.get("Retry-After")
          if retry_after is None:
              # No header: exponential backoff with full jitter.
              wait = random.uniform(0, min(MAX_WAIT_SECONDS, 2**attempt))
          else:
              wait = int(retry_after)
              if wait > MAX_WAIT_SECONDS:
                  body = response.json()
                  raise RetryLater(body["error"]["code"], wait, body["request_id"])
              # Many clients are released at the same instant. Spread them out a little.
              wait += random.uniform(0, 1)
          time.sleep(wait)


  if __name__ == "__main__":
      response = request_with_backoff("GET", "/orgs")
      response.raise_for_status()
      print("Requests left this minute:", response.headers.get("X-RateLimit-Remaining"))
      for org in response.json()["data"]:
          print(org["id"], org["name"])
  ```

  ```javascript Node.js theme={null}
  // Node 18+. Save as backoff.mjs and run: node backoff.mjs
  const BASE_URL = "https://api.surfais.com/v1";
  const API_KEY = process.env.SURFAIS_API_KEY;
  const MAX_ATTEMPTS = 5;
  const MAX_WAIT_SECONDS = 120;

  const sleep = (seconds) => new Promise((resolve) => setTimeout(resolve, seconds * 1000));

  // The API asked for a wait too long to sleep through, such as a spent monthly quota.
  class RetryLater extends Error {
    constructor(code, retryAfter, requestId) {
      super(`${code}: retry in ${retryAfter} s (request ${requestId})`);
      this.code = code;
      this.retryAfter = retryAfter;
      this.requestId = requestId;
    }
  }

  // Send one request. On 429 or 503, wait and try again, up to MAX_ATTEMPTS times.
  async function requestWithBackoff(method, path, options = {}) {
    const headers = { Authorization: `Bearer ${API_KEY}`, ...options.headers };
    for (let attempt = 0; ; attempt++) {
      const response = await fetch(`${BASE_URL}${path}`, { ...options, method, headers });
      const retryable = response.status === 429 || response.status === 503;
      if (!retryable || attempt === MAX_ATTEMPTS - 1) return response;

      const retryAfter = response.headers.get("Retry-After");
      let wait;
      if (retryAfter === null) {
        // No header: exponential backoff with full jitter.
        wait = Math.random() * Math.min(MAX_WAIT_SECONDS, 2 ** attempt);
      } else {
        wait = Number(retryAfter);
        if (wait > MAX_WAIT_SECONDS) {
          const body = await response.json();
          throw new RetryLater(body.error.code, wait, body.request_id);
        }
        // Many clients are released at the same instant. Spread them out a little.
        wait += Math.random();
      }
      await sleep(wait);
    }
  }

  const response = await requestWithBackoff("GET", "/orgs");
  if (!response.ok) throw new Error(`Request failed with ${response.status}`);
  console.log("Requests left this minute:", response.headers.get("X-RateLimit-Remaining"));
  for (const org of (await response.json()).data) {
    console.log(org.id, org.name);
  }
  ```
</CodeGroup>

## Staying under your limits

* **Use the rollup.** `GET /orgs/{orgId}/brands/summary` returns one row per own brand, with its latest score, the change since the previous one, sentiment, share of voice and scan freshness. That is one request instead of one per brand.
* **Poll the change token, not the data.** `changed_at` on each summary row, and `last_scan_changed_at` on a brand, move whenever something changed for that brand. Read results again only when the value has moved. See [Data model](/api/data-model).
* **Ask for full pages.** A `limit` of 200, or 1000 on the mentions endpoint, fetches the same data in fewer requests. While a scan is running, smaller pages are the better choice for results and mentions. See [Pagination and date windows](/api/pagination#reading-results-while-a-scan-is-running).
* **Let webhooks tell you.** Platform partners can subscribe to `scan.completed` and stop polling for new results. See [Webhooks](/api/webhooks/overview).
