How to see if GPTBot, ClaudeBot and PerplexityBot crawl your site
Three ways to see AI bot traffic (server logs, Cloudflare's AI Crawl Control, host log drains) and how to verify real bots against published IP ranges.
LLaunchScaler·Published ·8 min read
To see whether GPTBot, ClaudeBot, PerplexityBot and the other AI crawlers visit your site, search your server access logs for their user-agent tokens, or open Cloudflare's AI Crawl Control if your domain is proxied through Cloudflare, or send your host's traffic logs to a log drain and filter there. Then verify the hits against the IP ranges each operator publishes, because anyone can put "GPTBot" in a user agent.
A crawl tells you access works. It does not tell you the page was worth citing. Below are the three ways to read AI bot traffic, how to tell real bots from spoofed ones, and what a crawl with no citation means, checked in September 2026.
Which user agents should you look for?
Look for the tokens each operator documents. OpenAI and Anthropic each run three: a search crawler, a training crawler and a fetcher that visits pages when a user asks. Perplexity runs two, for search and for user requests. Search your logs for these names exactly as written.
Operator
Token in the user agent
What the operator says it does
Published IP list
OpenAI
OAI-SearchBot
Surfaces websites in ChatGPT's search features
openai.com/searchbot.json
OpenAI
GPTBot
Crawls content that may be used to train its foundation models
Questions, answered
What people ask about this
01
How do I know if ChatGPT crawls my website?
Search your server access logs for OpenAI's user-agent tokens: OAI-SearchBot for ChatGPT search, GPTBot for training, and ChatGPT-User for fetches a user triggers. Then confirm the request IPs fall inside the ranges OpenAI publishes, because user agents are easy to fake.
Visits pages for user actions in ChatGPT and Custom GPTs
openai.com/chatgpt-user.json
Anthropic
ClaudeBot
Collects content that could contribute to model training
claude.com/crawling/bots.json
Anthropic
Claude-SearchBot
Indexes content to improve search results
claude.com/crawling/bots.json
Anthropic
Claude-User
Fetches pages when a Claude user asks a question
claude.com/crawling/bots.json
Perplexity
PerplexityBot
Surfaces and links websites in Perplexity search results
perplexity.com/perplexitybot.json
Perplexity
Perplexity-User
Visits pages to answer a user's question
perplexity.com/perplexity-user.json
Google is the exception. Google-Extended "doesn't have a separate HTTP request user agent string," so Gemini training and grounding permissions do not show up as their own bot. You see Googlebot, and verify it by DNS as described below. The full list of AI user agents is in the AI crawler user agents list.
How do you find AI bots in your server logs?
Search the access log for the user-agent tokens, then count hits per bot, per path and per status code. On a server running nginx or Apache with the standard combined log format, the user agent is the last quoted field of each line, and one grep finds every AI crawler request.
Look at the status column first. Vercel measured ChatGPT's crawler spending 34.82% of its fetches on 404 pages, so a long run of 404s and 301s in this output would match what Vercel saw, and they are the first thing to fix. Why that matters is covered in how AI crawlers handle 404s and redirects.
How do you tell real AI bots from fake ones?
Check the request's IP address against the operator's published ranges. A user-agent string is just text any script can send; the IP range is what the operator vouches for. OpenAI, Anthropic and Perplexity each publish JSON files listing their crawlers' IP prefixes, and Google publishes both IP lists and a DNS method.
The files share one format: a creationTime and a list of prefixes, each with an ipv4Prefix such as "104.210.140.128/28". This short Python script tests IPs from your log against OpenAI's search crawler list:
import ipaddress, json, sys, urllib.request
data = json.load(urllib.request.urlopen("https://openai.com/searchbot.json"))
nets = [ipaddress.ip_network(p["ipv4Prefix"]) for p in data["prefixes"] if "ipv4Prefix" in p]
for line in sys.stdin:
ip = ipaddress.ip_address(line.strip())
print(ip, "verified" if any(ip in n for n in nets) else "NOT in published range")
Pipe in the IPs from the previous command: grep "OAI-SearchBot" access.log | awk '{print $1}' | sort -u | python3 verify.py. Swap the URL for gptbot.json, chatgpt-user.json, claude.com/crawling/bots.json or perplexity.com/perplexitybot.json to check the others.
For Googlebot, Google documents a DNS check: run host on the IP, confirm the name ends in googlebot.com, google.com or googleusercontent.com, then run host on that name and confirm it returns the original IP. Google also publishes IP range files for its crawlers.
Anthropic adds one caution for anyone tempted to block by IP instead: blocking its IP addresses "may not work correctly or persistently guarantee an opt-out, as doing so impedes our ability to read your robots.txt file." Control access in robots.txt; use IP lists to verify. How to write those rules is in robots.txt for AI bots.
How do you see AI crawlers in Cloudflare?
Open AI Crawl Control for your domain in the Cloudflare dashboard. Cloudflare says it is "Available on all plans" and works automatically once your domain proxies traffic through Cloudflare. The Overview tab shows AI crawler activity, and filters narrow it by date range, crawler, operator, hostname or path.
In the Cloudflare dashboard, select the domain and go to AI Crawl Control.
Review the Overview tab for a snapshot of crawler activity.
Open the Crawlers tab to see each AI crawler, with an Action column where you can choose Block.
Open the Metrics tab for breakdowns by date range, crawler, operator, status code, hostname or path. On free plans it shows only the past 24 hours.
Use Track robots.txt to see which crawlers are violating your directives.
On all plans Cloudflare detects AI crawlers by user agent string; Enterprise plans with Bot Management add detection through Bot Management detection IDs and configurable time ranges. If Cloudflare is blocking crawlers you want, that is covered in Cloudflare blocking AI crawlers.
Cloudflare Radar's AI Insights page is different: it shows AI bot traffic across Cloudflare's whole network, not yours. It charts request trends for the most active AI bots, the status codes they receive, a crawl-to-refer ratio by platform and the AI user agents named in robots.txt files of the top 10,000 domains. Use it for context, and your own dashboard for your site.
How do you see AI crawlers on Vercel or Netlify?
Send your traffic logs to an external service with a log drain, then filter by user agent there. Both hosts offer drains on paid plans: Vercel on Pro and Enterprise, Netlify on Enterprise. The drain streams request logs to a destination such as Datadog or Axiom, or to your own HTTP endpoint, where you search the user-agent field for the tokens above.
Vercel: drains forward "Runtime, build, and static logs" to a custom HTTP endpoint or a native integration. Through the REST API, a log drain uses the log schema at version v1.
Netlify: a Team Owner goes to the site's Analytics & metrics, then Log Drains, and selects Enable a log drain. Choose site traffic logs. Netlify offers an option to exclude personally identifiable information, which drops client_ip and user_agent from traffic logs; leave it off for this job, or you will have nothing to filter or verify.
How soon should a change show up in your logs?
Within about a day for robots.txt changes, longer for everything else. OpenAI says "it can take ~24 hours from a site's robots.txt update for our systems to adjust," and Perplexity says it "may take up to 24 hours for our systems to reflect changes." After you allow a bot, expect its user agent to reappear in your logs from the next day onward.
For other fixes, watch the mix rather than the count. When you repair broken internal links and clean your sitemap, the share of AI crawler requests that return 200 should rise over the following weeks, and the share of 404s and 301s should fall. When you server-render a page, the same bot's requests to it will not change, but the HTML it receives will.
A monthly check is enough for most sites: run the three log commands above, verify the IPs, compare the status mix with last month, and note any bot that has stopped visiting.
What does it mean if bots crawl your pages but AI answers never cite them?
It usually means a page problem, not a bot problem. A crawl proves the bot reached you. Whether the page gets cited depends on what the bot received and whether the page answers a question well. Check the status code in your logs, the raw HTML, and the page's answer to the question you hope to be cited for.
Work through it in order:
Status code: if the bot got 404, 301 or 403 for the page, it never saw your content.
Raw HTML: fetch the page with curl and confirm the main text is there. Vercel found that OpenAI's, Anthropic's and Perplexity's crawlers do not render JavaScript.
The right bot: GPTBot hits mean training crawls. For ChatGPT search, look for OAI-SearchBot; for Claude search, Claude-SearchBot.
The answer: the page should answer a specific question near the top, in text that stands alone.
Logs show who came; they do not show what the bot could use. Run the free scan on LaunchScaler to check that side: it takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. Its AI visibility checks test whether robots.txt blocks OAI-SearchBot, Claude-SearchBot or PerplexityBot, whether a firewall answers them with a 403 or a challenge, whether ChatGPT-User and Claude-User get a real page, and whether your main content is in the HTML those crawlers read.
No. GA4 collects data through a JavaScript tag, and AI crawlers such as GPTBot do not run JavaScript, according to Vercel's measurements, so their visits do not appear there. GA4 shows the human visits AI assistants send you, in its AI Assistant channel. Crawls themselves are in server, CDN or host logs.
03
How do I see AI crawlers in Cloudflare?
Open AI Crawl Control in the Cloudflare dashboard for your domain. The Overview tab shows AI crawler activity with filters by crawler, operator, hostname and path, and the Metrics tab breaks it down by status code. On free plans the Metrics tab covers the past 24 hours.
04
How do I verify a request really came from GPTBot or ClaudeBot?
Match the request's IP address against the operator's published list: openai.com/gptbot.json, openai.com/searchbot.json and openai.com/chatgpt-user.json for OpenAI, claude.com/crawling/bots.json for Anthropic, and perplexity.com/perplexitybot.json for Perplexity.
05
My site is crawled by AI bots but never cited. Why?
A crawl proves access, not usefulness. Check what the bot actually received: the status code, whether the main text was in the HTML, and whether the page answers a question directly. Those page problems are the usual reason a crawled page is not cited.
ClaudeBot collects training data, Claude-SearchBot indexes pages for Claude's search and Claude-User fetches pages users ask for. Each obeys its own group.
Perplexity cited content about 3x fresher than Google in one study. Put an accurate date in three places, visible, schema and sitemap, and never fake it.