---
title: "AI crawler tracking — Tracwell docs"
description: "Track AI page fetches from your server without another browser script."
canonical: "https://tracwell.app/docs/ai-crawlers"
markdown: "https://tracwell.app/docs/ai-crawlers.md"
---

# AI crawler tracking

See when supported AI crawlers request your pages. Keep your existing browser tracking installation.

## Two different kinds of AI traffic

AI referrals in Overview and Acquisition show people arriving from identifiable AI sources. They work with your existing tracker. AI crawlers report automated page requests, including clients that never execute browser JavaScript.

## One-time server setup

Generate a server key in Settings → Website → Server keys for an active website. Store it as TRACWELL\_SERVER\_KEY in your hosting environment. Never use a public browser environment variable. Crawler tracking works in Private and Product modes and requires active plan access.

Install the server helper, then choose your framework. It accepts a standard web Request and adds nothing to your browser bundle. Merge the example into your existing request handling; keep authentication, redirects, and routing intact.

Install

```bash
pnpm add tracwell@0.4.0
```

### Next.js

Add tracking to your existing proxy. Next.js 15 and earlier use middleware.ts and export middleware instead. Preserve existing authentication, redirects, and matchers.

proxy.ts

```ts
import { NextResponse, type NextRequest, type NextFetchEvent } from "next/server";
import { trackAiCrawler } from "tracwell/server";

export function proxy(request: NextRequest, event: NextFetchEvent) {
  event.waitUntil(trackAiCrawler(request, {
    serverKey: process.env.TRACWELL_SERVER_KEY!,
  }).then(result => {
    if (result.status === "failed") console.warn("Crawler tracking:", result.code);
  }));
  return NextResponse.next();
}

export const config = {
  matcher: ["/((?!api|_next/static|_next/image|favicon.ico).*)"],
};
```

[Next.js integration reference](https://nextjs.org/docs/app/api-reference/file-conventions/proxy)

### Cloudflare Workers

This example serves a site through an ASSETS binding. In an existing Worker, keep your current response handler in place of env.ASSETS.fetch. Configure TRACWELL\_SERVER\_KEY as a runtime Worker secret, not only a build variable. Requests must reach the Worker; for Workers Static Assets, configure assets.run\_worker\_first for the routes you want to track.

src/index.ts

```ts
import { trackAiCrawler } from "tracwell/server";

interface Env {
  ASSETS: Fetcher;
  TRACWELL_SERVER_KEY: string;
}

export default {
  async fetch(request, env, ctx) {
    const response = await env.ASSETS.fetch(request);
    ctx.waitUntil(trackAiCrawler(request, {
      serverKey: env.TRACWELL_SERVER_KEY,
      statusCode: response.status,
    }).then(result => {
      if (result.status === "failed") console.warn("Crawler tracking:", result.code);
    }));
    return response;
  },
} satisfies ExportedHandler<Env>;
```

[Cloudflare Workers integration reference](https://developers.cloudflare.com/workers/static-assets/routing/worker-script/)

### Astro

For Astro 5+ with a server adapter and on-demand rendering. Add the tracking call around your existing next() call, or compose middleware with Astro’s sequence helper. Prerendered pages need tracking at the hosting edge.

src/middleware.ts

```ts
import { defineMiddleware } from "astro:middleware";
import { getSecret } from "astro:env/server";
import { trackAiCrawler } from "tracwell/server";

export const onRequest = defineMiddleware(async (context, next) => {
  const response = await next();
  if (!context.isPrerendered) {
    const result = await trackAiCrawler(context.request, {
      serverKey: getSecret("TRACWELL_SERVER_KEY")!,
      statusCode: response.status,
    });
    if (result.status === "failed") console.warn("Crawler tracking:", result.code);
  }
  return response;
});
```

[Astro integration reference](https://docs.astro.build/en/guides/middleware/)

### SvelteKit

For a server deployment using private runtime environment variables. Keep your existing handle logic, or compose it with sequence from @sveltejs/kit/hooks. For Cloudflare bindings, read the key from event.platform.env instead. Prerendered pages and static assets bypass this hook.

src/hooks.server.ts

```ts
import { building } from "$app/environment";
import { env } from "$env/dynamic/private";
import type { Handle } from "@sveltejs/kit";
import { trackAiCrawler } from "tracwell/server";

export const handle: Handle = async ({ event, resolve }) => {
  const response = await resolve(event);
  if (!building) {
    const result = await trackAiCrawler(event.request, {
      serverKey: env.TRACWELL_SERVER_KEY!,
      statusCode: response.status,
    });
    if (result.status === "failed") console.warn("Crawler tracking:", result.code);
  }
  return response;
};
```

[SvelteKit integration reference](https://svelte.dev/docs/kit/hooks)

### Nuxt

For Nuxt 3/4 on Nitro. In nuxt.config.ts, add runtimeConfig: { tracwellServerKey: '' } and set the private runtime variable NUXT\_TRACWELL\_SERVER\_KEY to your server key. Nuxt uses that name to override runtime config. This server middleware continues to your existing routes; it does not return a response or read the request body.

server/middleware/tracwell.ts

```ts
import { defineEventHandler, getHeader, getMethod, getRequestURL } from "h3";
import { trackAiCrawler } from "tracwell/server";

export default defineEventHandler(async (event) => {
  const method = getMethod(event);
  if (method !== "GET" && method !== "HEAD") return;
  const config = useRuntimeConfig(event);
  const request = new Request(getRequestURL(event), {
    method,
    headers: { "user-agent": getHeader(event, "user-agent") ?? "" },
  });
  const result = await trackAiCrawler(request, {
    serverKey: config.tracwellServerKey,
  });
  if (result.status === "failed") console.warn("Crawler tracking:", result.code);
});
```

[Nuxt integration reference](https://nuxt.com/docs/4.x/directory-structure/server)

### Express

For Express 5 on Node.js 22+. Register this on your existing app before static files and page routes. Set SITE\_ORIGIN to your canonical origin, such as https\://example.com, so proxy headers do not determine the tracked hostname. This middleware leaves existing routes and request bodies intact.

server.ts — before your page routes

```ts
import { trackAiCrawler } from "tracwell/server";

// app is your existing Express application.
app.use(async (req, res, next) => {
  if (req.method !== "GET" && req.method !== "HEAD") return next();
  try {
    const url = new URL(process.env.SITE_ORIGIN!);
    const incoming = new URL(req.originalUrl, url);
    url.pathname = incoming.pathname;
    const result = await trackAiCrawler(new Request(url, {
      method: req.method,
      headers: { "user-agent": req.get("user-agent") ?? "" },
    }), { serverKey: process.env.TRACWELL_SERVER_KEY! });
    if (result.status === "failed") console.warn("Crawler tracking:", result.code);
  } catch {
    console.warn("Crawler tracking: request could not be prepared");
  }
  next();
});
```

[Express integration reference](https://expressjs.com/en/guide/using-middleware.html)

### Direct HTTP

For other languages or backends, send the same JSON from your server. This shell example requires curl and jq. REQUEST\_URL and USER\_AGENT must come from the incoming page request; EVENT\_ID is a generated UUID and TIMESTAMP is the event time in UTC with milliseconds. Strip URL queries and fragments before sending. Only send supported crawler GET/HEAD page requests. Keep the payload unchanged on retries.

Server-side request template

```bash
# Set TRACWELL_SERVER_KEY from your server's secret store.
# Supply EVENT_ID, TIMESTAMP, REQUEST_URL, USER_AGENT per incoming request.
payload=$(jq -n \
  --arg id "$EVENT_ID" \
  --arg timestamp "$TIMESTAMP" \
  --arg url "$REQUEST_URL" \
  --arg agent "$USER_AGENT" \
  '{schema_version: 1, event_id: $id, timestamp: $timestamp,
    url: $url, user_agent: $agent}')

curl --silent --show-error --include --max-time 10 \
  https://collect.tracwell.app/v1/server/crawlers \
  --header "Authorization: Bearer $TRACWELL_SERVER_KEY" \
  --header 'Content-Type: application/json' \
  --data "$payload"

# Expect HTTP 202 and a JSON receipt with:
# receipt_version: 1, batch_id: EVENT_ID, accepted_events: 1
# Retry network failures, 408, 429 and 5xx with the same payload.
# Do not follow redirects or retry other 4xx responses.
```

Await delivery, as shown in the server examples, or pass the promise to your host’s supported waitUntil to avoid delaying the page response. Always inspect failed results. Do not start an unawaited promise in a serverless handler. Pass statusCode only when you have the final response.

Only requests reaching the handler can be tracked. Static exports, prerendered pages, and responses served by a CDN before your handler require tracking at the hosting edge. Keep your existing cache behavior; installing middleware does not guarantee visibility into every cached request.

No extra browser script is needed. The server helper ignores ordinary visitors, common static assets, and API routes. Keep robots.txt, sitemaps, and Markdown pages included. There is no historical backfill.

## Supported crawlers

Expanded coverage requires tracwell 0.4.0 or newer. Upgrade the package and redeploy your server integration; no additional browser script or server key is needed. Older SDKs filter out the newly supported bots before sending.

| Crawler                 | Provider     | Category       |
| ----------------------- | ------------ | -------------- |
| `ChatGPT-User`          | OpenAI       | AI fetches     |
| `OAI-SearchBot`         | OpenAI       | Indexing       |
| `GPTBot`                | OpenAI       | Training       |
| `Claude-User`           | Anthropic    | AI fetches     |
| `Claude-SearchBot`      | Anthropic    | Indexing       |
| `ClaudeBot`             | Anthropic    | Training       |
| `Perplexity-User`       | Perplexity   | AI fetches     |
| `PerplexityBot`         | Perplexity   | Indexing       |
| `MistralAI-User`        | Mistral      | AI fetches     |
| `MistralAI-Index`       | Mistral      | Indexing       |
| `MistralAI-Training`    | Mistral      | Training       |
| `Amzn-User`             | Amazon       | AI fetches     |
| `Amzn-SearchBot`        | Amazon       | Indexing       |
| `Amazonbot`             | Amazon       | Other crawlers |
| `meta-externalfetcher`  | Meta         | AI fetches     |
| `meta-externalagent`    | Meta         | Other crawlers |
| `FacebookBot`           | Meta         | Other crawlers |
| `DuckAssistBot`         | DuckDuckGo   | AI fetches     |
| `Applebot`              | Apple        | Other crawlers |
| `Bytespider`            | ByteDance    | Other crawlers |
| `CCBot`                 | Common Crawl | Other crawlers |
| `Google-CloudVertexBot` | Google       | Other crawlers |
| `AI2Bot`                | Ai2          | Training       |

Google-Extended and Applebot-Extended control data usage; they are not separate crawling user agents. Ordinary Googlebot and Bingbot traffic is not classified as AI traffic.

## What the numbers mean

Overview → AI crawlers separates AI fetches, indexing, training, and other crawler requests. Other crawlers includes mixed-purpose and dataset crawlers whose individual requests do not establish training use. Identity is based on self-declared user agents without IP verification. Requests are not counts of people, prompts, citations, or confirmed training use.

Tracwell strips URL queries and fragments and does not store IPs, raw user agents, cookies, or authorization headers. Paths may still contain sensitive information; only instrument appropriate public routes. Crawler requests stay out of human visitor and conversion reports. Persisted requests count toward your existing event allowance and use the same retention and deletion policy.

## Delivery and direct integrations

The helper returns ignored, accepted, or failed. Accepted means the collection Queue accepted the request, not that database persistence is complete. It retries transient failures at most twice, preserving event identity. Use your runtime’s waitUntil and handle failed results; retries do not survive process termination.

Other backends can POST JSON to https\://collect.tracwell.app/v1/server/crawlers with Authorization: Bearer YOUR\_SERVER\_KEY. Send schema\_version: 1, a UUID event\_id, an ISO timestamp with milliseconds, the requested url, the original user\_agent, and an optional numeric status\_code. Reuse the entire payload on retries. The URL must match an allowed origin for the key’s website. Unknown crawlers and malformed requests are rejected; a matching 202 receipt confirms acceptance.
