Bots
How bot traffic is detected, what it's kept out of, and how to read the AI and All bots views.
What this page shows
Every request that reaches Trace is checked for signs of a bot before it's stored. Bot requests are still recorded, but only here: they never count as visitors, pageviews, sessions, goal conversions, alerts or "online now" anywhere else in the dashboard.
Only bots that run your tracking script
Trace sees a request only when a page ran its script. Crawlers that download your HTML without executing JavaScript — which includes most AI training crawlers — never reach Trace, so a low number here doesn't mean nobody is crawling you. Your server or CDN logs are the place to see those.
How detection works
Two independent layers run on every request. Both always run, so a request caught by both is counted under each layer's card.
1. UA pattern
The User-Agent is matched against a curated list of known bots, each tagged with its name, the company or project behind it (the operator) and its purpose. The list covers AI crawlers and agents, search engines, social link-preview fetchers, SEO crawlers, uptime and performance monitors, security scanners, HTTP libraries and command-line tools, and headless browsers. A User-Agent that isn't on the list but carries the generic marks of a bot — a SomethingBot/1.0 product token, a +https://… contact URL, or the word crawler or spider — is also caught, as an unidentified bot with no purpose.
HTTP libraries (curl, python-requests, Node's fetch, Go, Guzzle, …) get one exemption: if the request carries the screen and viewport size the tracking script always sends, it was built by a real page and relayed by a first-party proxy that didn't forward the visitor's User-Agent — so it's counted as a human.
2. Header heuristics
Deliberately narrow rules for clients that pretend to be a browser. Each one is something a real browser cannot produce on a tracking request:
| Rule | Fires when |
|---|---|
| No User-Agent | The request has no User-Agent at all and its payload wasn't built by a page (no screen or viewport size). |
| Missing browser headers | The User-Agent claims desktop Chrome, Edge or Opera 80+ or Firefox 90+, the request carries an Origin header, yet it has neither Accept-Language nor any Sec-Fetch-* header. Those browsers send both on every tracking request. |
| Page-load headers | Sec-Fetch-Mode: navigate or Sec-Fetch-Dest: document. The tracking script only sends background requests, so this is a replayed page load. |
Mobile browsers and Safari are never judged by these rules. Requiring Origin keeps proxies out: a proxy that strips headers usually forwards the User-Agent but not Origin. Requiring both Accept-Language and Sec-Fetch to be missing covers plain-HTTP setups (where browsers don't send Sec-Fetch) and privacy tools that trim Accept-Language. A request caught only by this layer is shown as an unidentified Scripted client.
Purposes
| Purpose | What it is |
|---|---|
| AI training | Collects pages to train a model (GPTBot, ClaudeBot, CCBot, Bytespider, …). Sends no readers back. |
| AI answer engine | Indexes pages so an assistant can cite them (OAI-SearchBot, PerplexityBot, Amazonbot, …). |
| AI agent | An assistant opening a page because a person just asked it to (ChatGPT-User, Claude-User, …). |
| Search engine | Googlebot, Bingbot, Applebot, DuckDuckBot, Yandex, Baidu and other search and ad crawlers. |
| Link preview | Chat and social apps fetching a link to build its preview card (Slack, WhatsApp, LinkedIn, X, …). |
| SEO crawler | Backlink and site-audit crawlers (Ahrefs, Semrush, Majestic, Moz, Screaming Frog, …). |
| Monitoring | Uptime, synthetic and performance checks (UptimeRobot, Pingdom, Checkly, Lighthouse, …). |
| Security scanner | Vulnerability and internet-wide scanners (Censys, Nuclei, WPScan, Qualys, …). |
| Scripted client | HTTP libraries and command-line tools, and anything caught only by header heuristics. |
| Headless browser | Headless Chrome, Puppeteer, Playwright, Selenium and similar automation. |
| Unclassified | Caught by a generic bot pattern, so its purpose isn't known. |
Headless browsers are usually stopped earlier
The tracking script itself refuses to send anything from an automated browser it can recognise (WebDriver, Puppeteer, Playwright, …). The server-side checks catch what gets past it, such as requests posted to the collect endpoint directly.
The two lenses
The switch at the top of the page picks what the page answers. Everything below it follows the lens.
AI & agents
Which AI systems read the site, what they read, and what they sent back.
| Metric | Meaning |
|---|---|
| AI requests | Requests from bots with an AI purpose (training, answer engine or agent). |
| Share of traffic | AI requests as a share of every request: all bot requests plus human pageviews and events. |
| Training / Answer-engine crawls / Agent fetches | AI requests split by purpose. |
| AI-referred visitors | Real people who arrived from an AI assistant (the AI Assistants channel). They're humans, not bots, and count in every report. |
| AI requests chart | Requests over time, by bot or by purpose. |
| AI operators | Per company: requests by purpose, the visitors its assistant referred, and requests per referral — how much reading it does for each visit it sends back. |
All bots
Every bot, and how it was caught.
| Metric | Meaning |
|---|---|
| Bot requests | Every bot request in the current scope. |
| Share of traffic | Bot requests as a share of every request (all bots plus human pageviews and events). |
| UA pattern / Header heuristics | Requests each detection layer caught. Click a card to scope the whole page to that layer; click again to clear it. |
| Named bots | How many distinct identified bots were seen. |
| Pages hit | How many distinct paths bots requested. |
| Bot requests chart | Requests over time, by purpose or by bot. |
Breakdowns and controls
- Bots (bot, operator, purpose), Pages (path, hostname), Referrers, Devices (browser, OS, device) and Locations (country, city) list the top ten; the expand button lists up to 200.
- Clicking a page, hostname, referrer, browser, OS, device, country or city row adds it as a dashboard filter — the same filter the rest of the dashboard uses.
- Clicking a purpose row scopes the page to that purpose. Scopes show as chips next to the lens switch; ✕ clears one. Switching lens clears both.
- The date range, granularity and filters work as on every page. Filters that describe human visits (goals, properties, entry/exit page, visit count) don't apply to bots and are ignored here.
- The lens and scope live in the URL (
?lens=all&layer=header_heuristics&purpose=seo), so a scoped view can be bookmarked or shared.
False positives
The rules are tuned to miss a bot rather than hide a real visitor, but a few setups can still trip them. If traffic you know is human shows up here:
- First-party proxy (a custom
data-api-url): make it forward the visitor'sUser-Agent,Accept-Language,OriginandSec-Fetch-*headers (andX-Forwarded-Forfor location). A proxy that forwards everything, like a framework rewrite, is never a problem. - Your own monitoring or tests (Lighthouse, Checkly, a Playwright suite) are counted as bots on purpose — that keeps them out of your visitor numbers.
- An in-app browser or kiosk whose User-Agent contains a bot-like token: filter the Bots page by its browser or page to confirm, then contact support with the User-Agent so the list can be corrected.
There is no per-site allowlist yet; detection changes apply to new requests only and never rewrite stored data.
FAQ
Why is my AI crawler count so low?
Most AI training crawlers fetch HTML and never run JavaScript, so the tracking script never sees them. What shows here are the ones that do run it — mostly agents and answer engines rendering pages.
Are bots counted in my visitors or usage?
Not in visitors or any other report — bot requests are stored with a bot flag and excluded. They do count toward your workspace's monthly event usage, since they are still received and stored.
Why is Googlebot here and not in my visitors?
Googlebot renders pages with a real browser engine, so it runs the script. Trace recognises it as a search-engine bot and keeps it out of your visitor numbers.
Does the main dashboard's AI traffic card show all of this?
No. It shows only AI training and answer-engine/agent crawlers; search engines, previews, monitors and scripts appear only on this page. You can hide that card in Settings → Dashboard.