AI crawler tracking

See when supported AI crawlers request your pages. Keep your existing browser tracking installation.

Two different kinds of AI traffic

AI referrals in Overview and Acquisition show people arriving from identifiable AI sources. They work with your existing tracker. AI crawlers report automated page requests, including clients that never execute browser JavaScript.

One-time server setup

Generate a server key in Settings → Website → Server keys for an active website. Store it as TRACWELL_SERVER_KEY in your hosting environment. Never use a public browser environment variable. Crawler tracking works in Private and Product modes and requires active plan access.

Install the server helper, then choose your framework. It accepts a standard web Request and adds nothing to your browser bundle. Merge the example into your existing request handling; keep authentication, redirects, and routing intact.

Install
pnpm add tracwell@0.4.0

Next.js

Add tracking to your existing proxy. Next.js 15 and earlier use middleware.ts and export middleware instead. Preserve existing authentication, redirects, and matchers.

proxy.ts
import { NextResponse, type NextRequest, type NextFetchEvent } from "next/server";
import { trackAiCrawler } from "tracwell/server";

export function proxy(request: NextRequest, event: NextFetchEvent) {
  event.waitUntil(trackAiCrawler(request, {
    serverKey: process.env.TRACWELL_SERVER_KEY!,
  }).then(result => {
    if (result.status === "failed") console.warn("Crawler tracking:", result.code);
  }));
  return NextResponse.next();
}

export const config = {
  matcher: ["/((?!api|_next/static|_next/image|favicon.ico).*)"],
};

Next.js integration reference

Await delivery, as shown in the server examples, or pass the promise to your host’s supported waitUntil to avoid delaying the page response. Always inspect failed results. Do not start an unawaited promise in a serverless handler. Pass statusCode only when you have the final response.

Only requests reaching the handler can be tracked. Static exports, prerendered pages, and responses served by a CDN before your handler require tracking at the hosting edge. Keep your existing cache behavior; installing middleware does not guarantee visibility into every cached request.

No extra browser script is needed. The server helper ignores ordinary visitors, common static assets, and API routes. Keep robots.txt, sitemaps, and Markdown pages included. There is no historical backfill.

Supported crawlers

Expanded coverage requires tracwell 0.4.0 or newer. Upgrade the package and redeploy your server integration; no additional browser script or server key is needed. Older SDKs filter out the newly supported bots before sending.

CrawlerProviderCategory
ChatGPT-UserOpenAIAI fetches
OAI-SearchBotOpenAIIndexing
GPTBotOpenAITraining
Claude-UserAnthropicAI fetches
Claude-SearchBotAnthropicIndexing
ClaudeBotAnthropicTraining
Perplexity-UserPerplexityAI fetches
PerplexityBotPerplexityIndexing
MistralAI-UserMistralAI fetches
MistralAI-IndexMistralIndexing
MistralAI-TrainingMistralTraining
Amzn-UserAmazonAI fetches
Amzn-SearchBotAmazonIndexing
AmazonbotAmazonOther crawlers
meta-externalfetcherMetaAI fetches
meta-externalagentMetaOther crawlers
FacebookBotMetaOther crawlers
DuckAssistBotDuckDuckGoAI fetches
ApplebotAppleOther crawlers
BytespiderByteDanceOther crawlers
CCBotCommon CrawlOther crawlers
Google-CloudVertexBotGoogleOther crawlers
AI2BotAi2Training

Google-Extended and Applebot-Extended control data usage; they are not separate crawling user agents. Ordinary Googlebot and Bingbot traffic is not classified as AI traffic.

What the numbers mean

Overview → AI crawlers separates AI fetches, indexing, training, and other crawler requests. Other crawlers includes mixed-purpose and dataset crawlers whose individual requests do not establish training use. Identity is based on self-declared user agents without IP verification. Requests are not counts of people, prompts, citations, or confirmed training use.

Tracwell strips URL queries and fragments and does not store IPs, raw user agents, cookies, or authorization headers. Paths may still contain sensitive information; only instrument appropriate public routes. Crawler requests stay out of human visitor and conversion reports. Persisted requests count toward your existing event allowance and use the same retention and deletion policy.

Delivery and direct integrations

The helper returns ignored, accepted, or failed. Accepted means the collection Queue accepted the request, not that database persistence is complete. It retries transient failures at most twice, preserving event identity. Use your runtime’s waitUntil and handle failed results; retries do not survive process termination.

Other backends can POST JSON to https://collect.tracwell.app/v1/server/crawlers with Authorization: Bearer YOUR_SERVER_KEY. Send schema_version: 1, a UUID event_id, an ISO timestamp with milliseconds, the requested url, the original user_agent, and an optional numeric status_code. Reuse the entire payload on retries. The URL must match an allowed origin for the key’s website. Unknown crawlers and malformed requests are rejected; a matching 202 receipt confirms acceptance.