OAI-SearchBot vs GPTBot vs ChatGPT-User: which OpenAI crawler to allow
GPTBot collects training data, OAI-SearchBot feeds ChatGPT search and ChatGPT-User fetches pages for users. Block GPTBot, allow the other two.
LLaunchScaler·Published ·8 min read
GPTBot collects content that may be used to train OpenAI's models, OAI-SearchBot crawls pages so they can appear in ChatGPT's search answers, and ChatGPT-User fetches a page when a ChatGPT user's request needs it. Most SaaS sites that want to be cited but not trained on should disallow GPTBot and allow OAI-SearchBot and ChatGPT-User; OpenAI says the settings are independent, so blocking training does not remove you from ChatGPT search.
The quotes and settings below come from OpenAI's crawler documentation, checked in September 2026. OpenAI updates that page when it adds agents, so check it again before you ship a change.
What is the difference between OAI-SearchBot, GPTBot and ChatGPT-User?
They are three user agents with three jobs. OAI-SearchBot finds pages for ChatGPT search, GPTBot gathers training data, and ChatGPT-User acts on a user's request in ChatGPT and Custom GPTs. Each has its own robots.txt token and its own published IP list, and blocking one has no effect on the others.
OAI-SearchBot
GPTBot
ChatGPT-User
OpenAI's description
"used to surface websites in search results in ChatGPT's search features"
"used to crawl content that may be used in training our generative AI foundation models"
Used "for certain user actions in ChatGPT and Custom GPTs"
How it crawls
Automatically
Automatically
Only when a user's request triggers it; "not used for crawling the web in an automatic fashion"
Questions, answered
What people ask about this
01
What is the difference between OAI-SearchBot and ChatGPT-User?
OAI-SearchBot crawls automatically to surface websites in ChatGPT's search features, and it follows robots.txt. ChatGPT-User visits a page only when a ChatGPT user's request calls for it, and OpenAI says robots.txt rules may not apply to it because a user started the action.
OpenAI documents a fourth agent, OAI-AdsBot, which visits only landing pages submitted as ads on ChatGPT to check them against its policies. OpenAI says its data "is not used to train generative AI foundation models." If you do not advertise on ChatGPT you will not see it.
Does blocking GPTBot remove you from ChatGPT?
No. OpenAI states that "each setting is independent of the others," and gives this exact case as its example: a webmaster "can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training." Blocking GPTBot changes what may be trained on, not what ChatGPT search can show.
OpenAI adds one efficiency detail: "If your site has allowed both bots, we may use the results from just one crawl for both use cases to avoid duplicative crawling." So allowing both does not double the load on your server.
The opt-out tells OpenAI your content "should not be used in training generative AI foundation models." OpenAI's documentation does not say whether it reaches content collected before you added the rule, so treat it as a setting for the future. Whether to block it is a rights decision rather than a visibility one; the guide on whether to block AI crawlers sets out the trade-off bot by bot.
What happens if you block OAI-SearchBot?
Your pages stop appearing as sources in ChatGPT's search answers. OpenAI's wording: "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links." For a product that wants ChatGPT to recommend it, this is the one OpenAI agent you cannot afford to block.
OpenAI recommends the opposite of blocking: "To help ensure your site appears in search results, we recommend allowing OAI-SearchBot in your site's robots.txt file and allowing requests from our published IP ranges." The second half of that sentence matters as much as the first. A firewall or bot-protection rule that challenges OAI-SearchBot's addresses blocks it just as effectively as a Disallow line, and it will not show up when you read robots.txt.
Changes are not instant. OpenAI says "it can take ~24 hours from a site's robots.txt update for our systems to adjust." If you unblock OAI-SearchBot today, give it a day before testing ChatGPT search. The guide to ranking in ChatGPT covers what happens after access is fixed.
Does ChatGPT-User follow robots.txt?
Not reliably, by OpenAI's own account. ChatGPT-User acts when "users ask ChatGPT or a CustomGPT a question" and it "may visit a web page." OpenAI says: "Because these actions are initiated by a user, robots.txt rules may not apply." It also says ChatGPT-User "is not used to determine whether content may appear in Search."
That has two consequences. A ChatGPT-User Disallow line is a request, not a guarantee; if you truly need to keep it out of a path, the enforcement point is your server or firewall, matched against openai.com/chatgpt-user.json. And blocking it does not take you out of ChatGPT search, because search inclusion is governed by OAI-SearchBot.
For most product sites the right answer is to let it in. When someone pastes your pricing page into ChatGPT and asks "does this plan include SSO?", ChatGPT-User is the agent that reads the page to answer. A block or a challenge page means ChatGPT answers without your page, or tells the user it could not access it.
What robots.txt should a SaaS site use for OpenAI's bots?
Disallow GPTBot, allow OAI-SearchBot and ChatGPT-User, and repeat your private paths in the allowed group. That keeps your content out of training, keeps you eligible for ChatGPT search, and lets ChatGPT read pages users send it.
# OpenAI: model training
User-agent: GPTBot
Disallow: /
# OpenAI: ChatGPT search and user-requested fetches
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
Allow: /
Disallow: /app/
Disallow: /api/
Disallow: /account/
Three rules from the robots.txt standard, RFC 9309, explain the layout. A crawler follows the group that names it and ignores the User-agent: * group, so private paths you disallow for everyone must be repeated in the named group. Several User-agent lines can share one group of rules, which is why OAI-SearchBot and ChatGPT-User sit together. And matching on the token is case-insensitive but spelling-exact: OAI-Searchbot matches, OAI-Search-Bot does not, and an unmatched token silently falls back to the * group.
If you want OpenAI's agents fully open, you do not need an OpenAI section at all, as long as your * group does not disallow the pages you want cited. The robots.txt guide for AI bots has a complete file covering Anthropic, Perplexity, Google and Apple as well.
How do you verify a request really comes from OpenAI?
Match the IP address, not the user-agent. Anyone can send GPTBot/1.4 in a header, but only OpenAI's servers use the addresses in its published lists. Each JSON file holds a prefixes array of CIDR blocks under ipv4Prefix, so checking an address from your logs takes a few lines.
python3 - <<'EOF'
import ipaddress, json, urllib.request
ip = ipaddress.ip_address("203.0.113.7") # the address from your log
data = json.load(urllib.request.urlopen("https://openai.com/searchbot.json"))
nets = [ipaddress.ip_network(p["ipv4Prefix"]) for p in data["prefixes"] if "ipv4Prefix" in p]
print("OpenAI OAI-SearchBot" if any(ip in n for n in nets) else "not in OAI-SearchBot's ranges")
EOF
Swap searchbot.json for gptbot.json or chatgpt-user.json to check the other agents. For firewall rules, combine both conditions: user-agent contains OAI-SearchBot and source IP is in the published list. OpenAI also notes that when it fetches robots.txt it may add a robots.txt marker to the user-agent string, such as OAI-SearchBot/1.4; robots.txt;, so log lines for your robots file are easy to separate from page requests.
Refresh the lists regularly. The files carry a creationTime field, and in September 2026 the GPTBot and ChatGPT-User files had been regenerated that same month, while the OAI-SearchBot file dated from January 2026.
How do you check what your site tells OpenAI's bots today?
Read your robots.txt as each bot would, then request a page with each user-agent and look at the status code. A 200 with your real content is what you want. A 403, a 429 or a page containing a challenge means a firewall is answering before your site does.
Fetch the live file: curl -s https://yoursite.com/robots.txt. Find the group each token matches, falling back to * if it has none.
Request a page as each agent: curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" https://yoursite.com/pricing. Repeat with the GPTBot and ChatGPT-User strings.
Check the page body, not only the code. Search it for a sentence from your page; if it is missing, the content is rendered by JavaScript, which Vercel's crawler study found OpenAI's agents do not execute.
Search your access logs for OAI-SearchBot, GPTBot and ChatGPT-User and verify a few addresses against the published lists.
A curl test from your laptop comes from your IP, not OpenAI's, so a firewall rule that allows verified OpenAI addresses may still block your test. That is a correct rule, not a failure. The list of AI crawler user agents has the strings and IP lists for every other operator.
Check every AI crawler rule in one pass
Reading robots.txt by hand works until a deploy, a CDN setting or a plugin changes it without telling you. Run the free scan on LaunchScaler: it takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. Its AI visibility checks read robots.txt for OAI-SearchBot, Claude-SearchBot and PerplexityBot, fail a wildcard "block all AI" rule that catches them, pass a training-only block such as GPTBot as a legitimate choice, request your pages as ChatGPT-User to confirm a real page comes back rather than a challenge, and check whether your main content is in the raw HTML these agents read.
Does blocking GPTBot remove my site from ChatGPT?
No. OpenAI says each setting is independent: you can disallow GPTBot to keep your content out of model training and still allow OAI-SearchBot so your pages appear in ChatGPT search results.
03
What happens if I block OAI-SearchBot?
OpenAI says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. It can take about 24 hours for a robots.txt change to take effect.
04
Does ChatGPT-User respect robots.txt?
Not reliably. OpenAI's documentation says that because ChatGPT-User's actions are initiated by a user, robots.txt rules may not apply, and that ChatGPT-User is not used to decide whether content appears in Search.
05
How do I verify a request really came from OpenAI?
Check the IP address against the list OpenAI publishes for that agent: openai.com/gptbot.json, openai.com/searchbot.json and openai.com/chatgpt-user.json. A user-agent string can be copied by anyone; the published IP ranges cannot.
Search Console now lets you pull a site out of AI Overviews and AI Mode together. Page-level options, what each costs, and why Google-Extended isn't one.
Allow PerplexityBot in robots.txt and in your firewall to appear as a Perplexity source. Perplexity-User fetches for users and mostly skips robots.txt.