> ## Documentation Index
> Fetch the complete documentation index at: https://docs.surfais.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Export results in bulk

> Copy an own brand's result history into your own systems, keep the copy in sync, and work with the API's paging rules.

This guide copies results out of Surfais in two stages: a one-off backfill of a brand's history, then an incremental sync that re-reads only the brands that changed. Read [Data model](/api/data-model) first if result units, mentions and the change token are new to you.

## Choose the endpoint

| Endpoint                                       | One row is                                                                                                       | Largest page | Use it for                                                         |
| ---------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- | ------------ | ------------------------------------------------------------------ |
| `GET /orgs/{orgId}/brands/{brandId}/mentions`  | One result unit with this own brand's mention only. It covers every prompt of the brand, inactive ones included. | 1,000        | Bulk history and warehouse loads. This is the bulk path.           |
| `GET /orgs/{orgId}/prompts/{promptId}/results` | One result unit of one prompt, with every tracked brand's mention and the answer's citations.                    | 200          | Competitor-by-competitor detail, and the citations of each answer. |
| `GET /orgs/{orgId}/brands/{brandId}/sources`   | One cited domain, aggregated over the window.                                                                    | 200          | The sites the AI platforms cite for the brand.                     |

All three serve finalised results only, and mentions and sources take own brands only. Mentions and results come newest date first. Without `limit`, a page holds 50 rows.

`GET …/mentions` returns one row per result unit, whether or not the brand was mentioned, so you can compute presence rates from it:

```json Response theme={null}
{
  "data": [
    {
      "prompt_id": "6a2c8e14-7b3d-4f59-a1c6-0d9e8f7b5a23",
      "run_date": "2026-09-22",
      "platform": "chatgpt",
      "country": "GB",
      "finalized_at": "2026-09-22T09:41:07Z",
      "mentioned": true,
      "position": 2,
      "sentiment": 78
    },
    {
      "prompt_id": "6a2c8e14-7b3d-4f59-a1c6-0d9e8f7b5a23",
      "run_date": "2026-09-22",
      "platform": "gemini",
      "country": "GB",
      "finalized_at": "2026-09-22T09:43:52Z",
      "mentioned": false,
      "position": null,
      "sentiment": null
    }
  ],
  "pagination": {
    "next_cursor": "eyJrIjoibWVudGlvbnNA…",
    "has_more": true
  }
}
```

## Backfill a brand's history

The script below walks `GET …/mentions` for one own brand and one date window, and writes each row as a line of JSON. It handles every failure the walk can meet:

* **`429` and `503`**: it waits for `Retry-After` seconds, then repeats the request. It stops instead of sleeping for more than five minutes, which is what an exhausted monthly quota asks for.
* **A void walk**: `400 invalid_cursor` with `details.reason` set to `stale`, or `409 scan_in_progress`. Both mean the brand's data changed under the walk. The script waits until the brand's change token stops moving, then **starts again from the first page**.
* **Partial output**: it writes to a temporary file and renames it when the walk completes, so the output file is never partial.
* **The minute allowance**: when `X-RateLimit-Remaining` reaches `0`, it sleeps until `X-RateLimit-Reset`.

<CodeGroup>
  ```python Python theme={null}
  """Backfill one own brand's mention rows into a newline-delimited JSON file.

  Usage: python backfill_mentions.py <org_id> <brand_id> <from> <to>
  Dates are YYYY-MM-DD (UTC), inclusive.
  """
  import json
  import os
  import sys
  import time

  import requests

  API = "https://api.surfais.com/v1"
  PAGE_SIZE = 1000        # the largest page /mentions serves
  MAX_RETRY_AFTER = 300   # never sleep longer than this on a 429 or 503
  MAX_WALKS = 5           # give up after this many restarted walks
  SETTLE_SECONDS = 60     # how long the change token must hold still

  session = requests.Session()
  session.headers["Authorization"] = f"Bearer {os.environ['SURFAIS_API_KEY']}"


  class RestartWalk(Exception):
      """The brand's data changed under the walk. Start again from the first page."""


  def get(path, params=None):
      """GET one JSON body. Waits out 429 and 503, and raises RestartWalk when the walk is void."""
      while True:
          resp = session.get(f"{API}{path}", params=params, timeout=60)
          if resp.status_code == 200:
              pause_when_minute_is_spent(resp)
              return resp.json()
          try:
              body = resp.json()
              error, request_id = body["error"], body["request_id"]
          except (ValueError, KeyError):
              error, request_id = {"code": "unknown", "message": resp.text[:200]}, None
          wait = int(resp.headers.get("Retry-After", "0"))
          if resp.status_code in (429, 503) and 0 < wait <= MAX_RETRY_AFTER:
              time.sleep(wait)
              continue
          if resp.status_code == 409 and error["code"] == "scan_in_progress":
              raise RestartWalk("scan_in_progress")  # the settle wait outlasts its Retry-After
          details = error.get("details")
          stale = isinstance(details, dict) and details.get("reason") == "stale"
          if resp.status_code == 400 and error["code"] == "invalid_cursor" and stale:
              raise RestartWalk("stale cursor")
          sys.exit(f"{resp.status_code} {error['code']}: {error['message']} (request {request_id})")


  def pause_when_minute_is_spent(resp):
      """Sleep to the end of the minute window rather than run into a 429."""
      if resp.headers.get("X-RateLimit-Remaining") == "0":
          reset = int(resp.headers.get("X-RateLimit-Reset", "0"))
          time.sleep(max(0.0, reset - time.time()) + 1)


  def walk_mentions(org_id, brand_id, date_from, date_to):
      """Yield every mention row in the window, newest date first."""
      params = {"from": date_from, "to": date_to, "limit": PAGE_SIZE}
      while True:
          page = get(f"/orgs/{org_id}/brands/{brand_id}/mentions", params)
          yield from page["data"]
          if not page["pagination"]["has_more"]:
              return
          # The cursor carries the window and the filters. Send it alone.
          params = {"cursor": page["pagination"]["next_cursor"], "limit": PAGE_SIZE}


  def wait_until_settled(org_id, brand_id):
      """Block until the brand's change token stops moving."""
      path = f"/orgs/{org_id}/brands/{brand_id}"
      token = get(path)["data"]["last_scan_changed_at"]
      while True:
          time.sleep(SETTLE_SECONDS)
          latest = get(path)["data"]["last_scan_changed_at"]
          if latest == token:
              return
          token = latest


  def main():
      org_id, brand_id, date_from, date_to = sys.argv[1:5]
      out_path = f"mentions-{brand_id}.ndjson"
      for walk in range(1, MAX_WALKS + 1):
          try:
              rows = 0
              # Write to a temporary file, so a restarted walk never leaves half a result.
              with open(f"{out_path}.part", "w", encoding="utf-8") as out:
                  for row in walk_mentions(org_id, brand_id, date_from, date_to):
                      out.write(json.dumps(row) + "\n")
                      rows += 1
              os.replace(f"{out_path}.part", out_path)
              print(f"wrote {rows} rows to {out_path}")
              return
          except RestartWalk as reason:
              print(f"walk {walk} abandoned ({reason}); waiting for the brand to settle", file=sys.stderr)
              wait_until_settled(org_id, brand_id)
      sys.exit("the brand's data kept changing; run the backfill again once its scan has finished")


  if __name__ == "__main__":
      main()
  ```

  ```javascript Node.js theme={null}
  // backfill-mentions.mjs — Node.js 18 or later
  // Usage: node backfill-mentions.mjs <orgId> <brandId> <from> <to>
  // Dates are YYYY-MM-DD (UTC), inclusive.
  import { once } from "node:events";
  import { createWriteStream } from "node:fs";
  import { rename } from "node:fs/promises";
  import { setTimeout as sleep } from "node:timers/promises";

  const API = "https://api.surfais.com/v1";
  const PAGE_SIZE = 1000; // the largest page /mentions serves
  const MAX_RETRY_AFTER = 300; // never sleep longer than this on a 429 or 503
  const MAX_WALKS = 5; // give up after this many restarted walks
  const SETTLE_SECONDS = 60; // how long the change token must hold still

  const headers = { Authorization: `Bearer ${process.env.SURFAIS_API_KEY}` };

  // The brand's data changed under the walk. Start again from the first page.
  class RestartWalk extends Error {}

  // GET one JSON body. Waits out 429 and 503, and throws RestartWalk when the walk is void.
  async function get(path, params = {}) {
    const url = `${API}${path}?${new URLSearchParams(params)}`;
    for (;;) {
      const res = await fetch(url, { headers, signal: AbortSignal.timeout(60_000) });
      const text = await res.text();
      let body;
      try {
        body = JSON.parse(text);
      } catch {
        body = { error: { code: "unknown", message: text.slice(0, 200) }, request_id: null };
      }
      if (res.status === 200) {
        await pauseWhenMinuteIsSpent(res);
        return body;
      }
      const wait = Number(res.headers.get("retry-after") ?? 0);
      if ((res.status === 429 || res.status === 503) && wait > 0 && wait <= MAX_RETRY_AFTER) {
        await sleep(wait * 1000);
        continue;
      }
      if (res.status === 409 && body.error.code === "scan_in_progress") {
        throw new RestartWalk("scan_in_progress"); // the settle wait outlasts its Retry-After
      }
      const stale = body.error.details?.reason === "stale";
      if (res.status === 400 && body.error.code === "invalid_cursor" && stale) {
        throw new RestartWalk("stale cursor");
      }
      throw new Error(`${res.status} ${body.error.code}: ${body.error.message} (request ${body.request_id})`);
    }
  }

  // Sleep to the end of the minute window rather than run into a 429.
  async function pauseWhenMinuteIsSpent(res) {
    if (res.headers.get("x-ratelimit-remaining") === "0") {
      const reset = Number(res.headers.get("x-ratelimit-reset") ?? 0);
      await sleep(Math.max(0, reset * 1000 - Date.now()) + 1000);
    }
  }

  // Yield every mention row in the window, newest date first.
  async function* walkMentions(orgId, brandId, from, to) {
    let params = { from, to, limit: PAGE_SIZE };
    for (;;) {
      const page = await get(`/orgs/${orgId}/brands/${brandId}/mentions`, params);
      yield* page.data;
      if (!page.pagination.has_more) return;
      // The cursor carries the window and the filters. Send it alone.
      params = { cursor: page.pagination.next_cursor, limit: PAGE_SIZE };
    }
  }

  // Resolve once the brand's change token stops moving.
  async function waitUntilSettled(orgId, brandId) {
    const path = `/orgs/${orgId}/brands/${brandId}`;
    let token = (await get(path)).data.last_scan_changed_at;
    for (;;) {
      await sleep(SETTLE_SECONDS * 1000);
      const latest = (await get(path)).data.last_scan_changed_at;
      if (latest === token) return;
      token = latest;
    }
  }

  async function main() {
    const [orgId, brandId, from, to] = process.argv.slice(2);
    const outPath = `mentions-${brandId}.ndjson`;
    for (let walk = 1; walk <= MAX_WALKS; walk++) {
      // Write to a temporary file, so a restarted walk never leaves half a result.
      const out = createWriteStream(`${outPath}.part`, { encoding: "utf8" });
      try {
        let rows = 0;
        for await (const row of walkMentions(orgId, brandId, from, to)) {
          if (!out.write(`${JSON.stringify(row)}\n`)) await once(out, "drain");
          rows++;
        }
        out.end();
        await once(out, "finish");
        await rename(`${outPath}.part`, outPath);
        console.log(`wrote ${rows} rows to ${outPath}`);
        return;
      } catch (err) {
        out.destroy();
        if (!(err instanceof RestartWalk)) throw err;
        console.error(`walk ${walk} abandoned (${err.message}); waiting for the brand to settle`);
        await waitUntilSettled(orgId, brandId);
      }
    }
    throw new Error("the brand's data kept changing; run the backfill again once its scan has finished");
  }

  main().catch((err) => {
    console.error(err.message);
    process.exit(1);
  });
  ```
</CodeGroup>

Run it with the organisation id, the brand id and an inclusive date window:

```bash theme={null}
python backfill_mentions.py "$SURFAIS_ORG_ID" "$SURFAIS_BRAND_ID" 2026-08-19 2026-09-22
```

### Why a walk can become void

Every page of a walk belongs to one state of the brand's data. If the brand's [change token](/api/data-model#freshness-two-different-fields) moves while you page, the API refuses the next page instead of serving rows from two states. While a scan of the brand runs, the token moves with every result, so expect refusals then. A large first page can be refused the same way, with `409 scan_in_progress` and a `Retry-After`.

The remedy is always the same: discard the rows from the abandoned walk and start again from the first page. [Reading results while a scan is running](/api/pagination#reading-results-while-a-scan-is-running) has the full rules.

## Keep the copy in sync

After the backfill, do not walk every brand on every run. Let the change token tell you which brands to read.

<Steps>
  <Step title="Store a token and a date for each own brand">
    Keep the last `changed_at` you synced, and the UTC date you synced on.
  </Step>

  <Step title="List the summaries">
    Call `GET /orgs/{orgId}/brands/summary` with `limit=200`, the largest page. One call covers up to 200 own brands. Follow `next_cursor` if `has_more` is `true`.
  </Step>

  <Step title="Skip the brands that did not change">
    Compare each row's `changed_at` with the stored value, for equality only. Skip the brand when they match. Skip it too when `changed_at` is `null`: nothing has been recorded for the brand yet.
  </Step>

  <Step title="Re-pull a short trailing window for the rest">
    Walk `GET …/mentions` from the day before your last sync to today, with `walk_mentions` from the backfill script (`walkMentions` in Node.js). Routine scans only add or replace units on the newest dates, so a short window is enough for them.
  </Step>

  <Step title="Upsert on the unit's natural key">
    The key is `prompt_id`, `platform`, `country` and `run_date`. A second scan on the same day replaces a unit, so an insert-only load would store it twice.
  </Step>

  <Step title="Store the token you read in step 2">
    Store the `changed_at` from the summary, not one you read after the pull. If more data arrived while you were pulling, the stored token is already out of date, and the next run pulls the brand again.
  </Step>
</Steps>

The token tells you that something changed, not what. It also moves for changes that leave results alone, such as an edited brand name, so some pulls find nothing new.

<Note>
  A trailing window does not see every change. A unit can disappear: deleting a prompt in the dashboard removes its results. An older unit can also change: when a brand alias is added or removed, Surfais recounts the brand's recent results. If your copy must match exactly, re-run the backfill over a longer window from time to time, and replace that window in your copy rather than upserting into it.
</Note>

## Windows and filters

* `from` and `to` are inclusive UTC dates, written `YYYY-MM-DD`. `to` defaults to today, and `from` to 30 days before `to`.
* A window can span at most 366 days on mentions and results, and at most 90 days on sources. A wider one answers `400 validation_error` with `details[].code` set to `window_too_wide`. For a longer history, run one walk per consecutive window.
* Mentions and results accept `platform` and `country` filters. Sources accepts neither.
* **The first page fixes the window and the filters.** Continue with `cursor` and `limit` alone, as the script does. Naming a different value on a later page answers `400 invalid_cursor`, with `details.reason` set to `window_changed` or `filter_changed`. To change either, start a new walk.

[Pagination and date windows](/api/pagination) covers cursors in full.

## Truncated brand lists on result units

`GET …/results` lists each tracked brand's mention in the unit's `brands` array:

```json Response theme={null}
{
  "data": [
    {
      "run_date": "2026-09-22",
      "platform": "perplexity",
      "country": "GB",
      "finalized_at": "2026-09-22T09:41:07Z",
      "brands": [
        {
          "brand_id": "a41d9c6e-0b2f-4f3a-8c75-9e1d2b3a4c5d",
          "mentioned": true,
          "position": 1,
          "sentiment": null
        },
        {
          "brand_id": "c2a4e8f0-6b1d-4c3a-8e5f-9d7b2a1c4e60",
          "mentioned": true,
          "position": 2,
          "sentiment": 78
        }
      ],
      "brands_total": 2,
      "brands_truncated": false,
      "citations": [
        {
          "url": "https://travel-weekly.example/best-boutique-hotels",
          "domain": "travel-weekly.example",
          "cited": true
        }
      ]
    }
  ],
  "pagination": {
    "next_cursor": null,
    "has_more": false
  }
}
```

A unit lists at most 250 brands. When it held more, `brands_truncated` is `true` and `brands_total` says how many there were. Mentioned brands are kept first, then brands by position, so a truncated list is the top of a ranking. It is not a complete list of mentions: a unit that mentions more brands than the cap loses some of them.

Within the array, entries are ordered by `brand_id`. The cap never affects `GET …/mentions`, which returns one brand's mention per unit.

## Be a good citizen

* **Export when the data is quiet.** Start bulk walks once the brand's change token has settled. Scheduled scans start early in the UTC day, so a daily export is better placed later. See [Scan schedule & cadence](/scans/scan-schedule).
* **Use smaller pages while a scan runs.** A smaller first page is less likely to meet `409 scan_in_progress`.
* **Watch `X-RateLimit-Remaining`.** Your limits are set on your key, so read them from the response headers and pause before you reach zero. See [Rate limits and quotas](/api/rate-limits).
* **Do not sleep on a monthly quota.** `Retry-After` on `429 quota_exceeded` counts to the end of the month. Stop the job and raise it with your team instead.
* **Ask the summary first.** One `GET /orgs/{orgId}/brands/summary` page tells you which brands changed. Read the per-brand endpoints only for those.
