AI crawler user agents: the list for robots.txt and logs
Every major AI crawler user agent and robots.txt token, with its operator, purpose, JavaScript rendering and published IP list, checked September 2026.
LLaunchScaler·Published ·8 min read
The AI crawler user agents that matter for robots.txt and log analysis are OpenAI's GPTBot, OAI-SearchBot and ChatGPT-User, Anthropic's ClaudeBot, Claude-SearchBot and Claude-User, Perplexity's PerplexityBot and Perplexity-User, and Common Crawl's CCBot, plus two robots.txt tokens with no crawler behind them, Google-Extended and Applebot-Extended. The table below gives each one's operator, purpose, whether it renders JavaScript and where its IP ranges are published, checked in September 2026.
Operators add agents and change version numbers, so treat this as a dated snapshot and confirm against each operator's page before you ship a rule.
Which AI crawler user agents should you know?
Sort them by purpose first, because that decides whether you block them. Training crawlers collect content for models, search crawlers index pages for an assistant's answers, and user-triggered agents fetch a page because a person asked. Each operator's documentation states the purpose; the rendering column comes from Vercel's measurement of AI crawler traffic.
Token
Operator
Purpose
Follows robots.txt
Renders JavaScript
Published IP list
GPTBot
OpenAI
Training
Yes
No (Vercel)
openai.com/gptbot.json
Questions, answered
What people ask about this
01
What are the main AI crawler user agents?
OpenAI runs GPTBot, OAI-SearchBot and ChatGPT-User; Anthropic runs ClaudeBot, Claude-SearchBot and Claude-User; Perplexity runs PerplexityBot and Perplexity-User; Common Crawl runs CCBot. Google-Extended and Applebot-Extended are robots.txt tokens with no crawler of their own.
Search; "not used to crawl content for AI foundation models"
Yes
No (Vercel)
www.perplexity.com/perplexitybot.json
Perplexity-User
Perplexity
User-triggered fetch
"Generally ignores robots.txt rules"
Not published
www.perplexity.com/perplexity-user.json
CCBot
Common Crawl
Crawl for an open dataset
Blocking via robots.txt documented
Not published
index.commoncrawl.org/ccbot.json
Google-Extended
Google
Token: Gemini training and Gemini app grounding
Read from robots.txt
Crawling is done by Googlebot, which renders
Googlebot's lists apply
Applebot-Extended
Apple
Token: Apple foundation model training
Read from robots.txt
Crawling is done by Applebot, which renders
Applebot's list applies
"Not published" means the operator's documentation, and Vercel's study, do not say. Vercel's research found that "none of the major AI crawlers currently render JavaScript," naming OpenAI's three agents, ClaudeBot and PerplexityBot among others, while Gemini uses Googlebot's rendering and Applebot renders through a browser-based crawler. The study did not cover Claude-SearchBot, Claude-User or Perplexity-User, and their operators do not document rendering; plan on them reading raw HTML too.
Cloudflare's managed robots.txt also disallows Amazonbot, Bytespider and meta-externalagent. Those operators' own documentation was not checked for this list, so they are named here only so you recognise them in a log.
What are the full user-agent strings?
The robots.txt token is a substring of the full user-agent header, which is what appears in your logs. OpenAI, Perplexity, Apple and Common Crawl publish the full strings; Anthropic's help article names only the tokens. OpenAI notes its version numbers may change, so match on the token rather than the whole string.
Agent
User-agent string as published
OAI-SearchBot
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot
GPTBot
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
ChatGPT-User
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
OAI-AdsBot
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot
PerplexityBot
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Perplexity-User
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
CCBot
CCBot/2.0 (https://commoncrawl.org/faq/)
Applebot
Contains Applebot/ followed by a version and +http://www.apple.com/go/applebot
ClaudeBot, Claude-SearchBot, Claude-User
Not published in full; search logs for the token
OpenAI adds one detail worth knowing when you read logs: when OAI-SearchBot or GPTBot fetch robots.txt, the string may carry an extra robots.txt; marker, so those requests are easy to tell apart from page requests.
Which AI crawlers are tokens, not crawlers?
Google-Extended and Applebot-Extended never make a request. Google says Google-Extended "doesn't have a separate HTTP request user agent string" and that crawling "is done with existing Google user agent strings." Apple says Applebot-Extended "does not crawl webpages" and "is only used to determine how to use the data crawled by the Applebot user agent."
That has two practical effects. You cannot see them in logs, so robots.txt is the only record of what you told them. And blocking them does not reduce crawling, because Googlebot and Applebot still fetch your pages for Search, Spotlight, Siri and Safari. The guide to whether you should block AI crawlers explains what each token does and does not switch off.
How do you verify an AI crawler by IP address?
Match the request's source IP against the list its operator publishes, because anyone can send a user-agent string. OpenAI, Anthropic and Perplexity publish JSON lists; Google, Apple and Common Crawl publish lists and also support reverse DNS checks. A request whose address is not on the operator's list did not come from that operator.
JSON lists for common crawlers, special-case crawlers and user-triggered fetchers
Hostname ends in googlebot.com, google.com or googleusercontent.com, confirmed by forward lookup
Apple
Applebot's CIDR list, linked from Apple's Applebot page
Hostname in applebot.apple.com
Common Crawl
index.commoncrawl.org/ccbot.json
Hostname in crawl.commoncrawl.org
The JSON files from OpenAI, Anthropic and Perplexity share one shape: a creationTime and a prefixes array of ipv4Prefix entries. This checks one address against one list:
python3 - <<'EOF'
import ipaddress, json, urllib.request
ip = ipaddress.ip_address("203.0.113.7") # address from your log
url = "https://openai.com/searchbot.json" # the list for the claimed bot
data = json.load(urllib.request.urlopen(url))
nets = [ipaddress.ip_network(p["ipv4Prefix"]) for p in data["prefixes"] if "ipv4Prefix" in p]
print("verified" if any(ip in n for n in nets) else "not on the list")
EOF
For reverse DNS, Google's documented method is two lookups: host 66.249.66.1 to get the hostname, then host on that hostname to confirm it resolves back to the same address. Apple documents the same pattern for *.applebot.apple.com, and Common Crawl gives the example of an address resolving to a hostname under crawl.commoncrawl.org. Anthropic warns against using IP blocking as your opt-out, since it "impedes our ability to read your robots.txt file"; use its list to verify, and robots.txt to set policy.
Should rules match the user agent or the IP address?
In robots.txt you can only use the token. In a firewall, match both: the token tells you which bot claims to be calling, and the published address list tells you whether the claim is true. The operators' guidance lines up on this, with one caution from Anthropic.
OpenAI recommends "allowing OAI-SearchBot in your site's robots.txt file and allowing requests from our published IP ranges."
Perplexity's WAF instructions combine a user-agent condition with an IP condition, and set the action to allow.
Anthropic says blocking its IP addresses "may not work correctly or persistently guarantee an opt-out," because it stops them reading your robots.txt, so use robots.txt to opt out and the address list to verify.
A rule that matches the user agent alone lets anyone who copies the string through your allow rule. A rule that matches the address alone breaks the day the operator adds a range you have not loaded.
How do you find AI crawlers in your server logs?
Search your access log for each token, then count status codes per token. A crawler that gets 200 is reading your pages; one that gets 403, 429 or a challenge page is being stopped before robots.txt matters. This one command covers every token in the table:
Then look at the codes for the ones you want to allow, for example grep "OAI-SearchBot" access.log | awk '{print $9}' | sort | uniq -c on a standard combined log format, where the status is the ninth field. Verify a sample of addresses against the published lists before trusting the counts. The guide to checking whether AI bots crawl your site covers log setups on common hosts and CDNs.
How often do these lists change?
Often enough that a rule written once will drift. Every IP list above carries a creationTime, and in September 2026 they ranged from February 2025 to that same month. Version numbers inside the strings change too, which is why operators tell you to match the token rather than the full header.
IP list
creationTime in September 2026
openai.com/chatgpt-user.json
2026-09-25
openai.com/gptbot.json
2026-09-22
claude.com/crawling/bots.json
2026-08-18
index.commoncrawl.org/ccbot.json
2026-08-11
openai.com/searchbot.json
2026-01-02
www.perplexity.com/perplexity-user.json
2025-10-17
www.perplexity.com/perplexitybot.json
2025-02-07
Two habits keep you current. Pull the lists on a schedule rather than pasting addresses into a firewall rule by hand; Perplexity recommends automating this in its own WAF instructions. And follow the operators' change notices: OpenAI's crawler page offers an RSS feed for updates, and Anthropic's help article has a form to be notified of substantial changes. Common Crawl's file notes that forward-confirmed reverse DNS is also recommended for verifying its addresses.
How do you use these tokens in robots.txt?
Use the token exactly as the operator spells it; matching is case-insensitive but spelling-exact under RFC 9309, and a misspelled token silently falls back to your User-agent: * group. Group training tokens under Disallow: / if you opt out of training, and give search and user agents an allow group with your private paths repeated inside it.
Knowing the tokens is half the job; the other half is what your site returns to each of them. Run the free scan on LaunchScaler: it takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. Its AI visibility checks read your robots.txt for OAI-SearchBot, Claude-SearchBot and PerplexityBot, fail a wildcard rule that catches them, request your pages as the retrieval bots and as ChatGPT-User and Claude-User to catch a 403, a challenge or an empty client-rendered shell, and check whether your main content is in the raw HTML.
OpenAI documents three: GPTBot for training, OAI-SearchBot for ChatGPT search and ChatGPT-User for fetches a user triggers. For example, OAI-SearchBot's string ends in 'compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot', and OpenAI notes version numbers may change.
03
How do I verify that a request really came from an AI crawler?
Match the source IP against the operator's published list, such as openai.com/searchbot.json, www.perplexity.com/perplexitybot.json or claude.com/crawling/bots.json. Google, Apple and Common Crawl also support reverse DNS checks. The user-agent string alone can be faked.
04
Do AI crawlers execute JavaScript?
Mostly no. Vercel's crawler study found OpenAI's, Anthropic's and Perplexity's crawlers do not render JavaScript, while Gemini uses Googlebot's rendering and Applebot renders with a browser-based crawler.
05
Why don't I see Google-Extended in my server logs?
Because it never visits. Google says Google-Extended has no separate user-agent string; Googlebot does the crawling and the token is only read from robots.txt as a control. Apple says the same of Applebot-Extended.
Vercel measured ChatGPT's crawler spending 34.82% of fetches on 404s and 14.36% on redirects, against 8.22% and 1.49% for Googlebot. How to cut the waste.
When ChatGPT names competitors and not you, list who wins each prompt, open the sources it cited, and get onto those pages. A step-by-step gap analysis.