PerplexityBot and Perplexity-User: robots.txt rules to stay citable
Allow PerplexityBot in robots.txt and in your firewall to appear as a Perplexity source. Perplexity-User fetches for users and mostly skips robots.txt.
LLaunchScaler·Published ·8 min read
To stay citable in Perplexity, let PerplexityBot crawl your public pages: it is the crawler that surfaces and links websites in Perplexity's search results, and Perplexity says it is not used to train AI models. Allow it in robots.txt and in your firewall, because Perplexity recommends both, and leave Perplexity-User alone, since it fetches pages for users and generally ignores robots.txt anyway.
The quotes below come from Perplexity's crawler documentation, checked in September 2026.
What is PerplexityBot?
PerplexityBot is Perplexity's search crawler. Perplexity says it is "designed to surface and link websites in search results on Perplexity" and that "it is not used to crawl content for AI foundation models." It is the agent that decides whether your pages are in the pool Perplexity can cite.
Its full user-agent string is:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Perplexity's own recommendation is short: "To ensure your site appears in search results, we recommend allowing PerplexityBot in your site's robots.txt file and permitting requests from our published IP ranges." Both halves matter. The robots.txt rule is a permission; the IP ranges are what your firewall needs to let the permission through.
Because Perplexity says PerplexityBot is not used for model training, blocking it does not protect your content from training. It only removes you from Perplexity's results. That makes it a different decision from blocking GPTBot or ClaudeBot, which are training crawlers.
What is Perplexity-User, and does it follow robots.txt?
Perplexity-User visits a page when a user's question needs it. Perplexity says it "supports user actions within Perplexity," may "visit a web page to help provide an accurate answer and include a link to the page in its response," and is not used for crawling or for training. On robots.txt, the documentation is direct: it "generally ignores robots.txt rules."
The reason Perplexity gives is that "a user requested the fetch." OpenAI takes the same position for ChatGPT-User. The practical effect is that a Disallow line is not a dependable control. If you need to keep it out of a path, that has to happen at the server or firewall, matched against its published address list.
Questions, answered
What people ask about this
01
What is PerplexityBot?
PerplexityBot is Perplexity's crawler for surfacing and linking websites in Perplexity's search results. Perplexity says it is not used to crawl content for AI foundation models, and recommends allowing it in robots.txt and allowing its published IP ranges.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
For a product site, blocking it rarely makes sense. When a buyer asks Perplexity about your pricing, Perplexity-User is what reads your pricing page and links it in the answer.
What robots.txt rules keep you citable in Perplexity?
Give PerplexityBot and Perplexity-User a group that allows your public pages and repeats your private-path rules. Under RFC 9309, a crawler that finds a group naming it ignores the User-agent: * group, so any path you want kept private must be listed again inside the Perplexity group.
If your * group already allows everything public, you do not need a Perplexity group at all. You need one when a broader rule would catch PerplexityBot, which is how sites end up blocking Perplexity without meaning to.
What is in robots.txt
What PerplexityBot does
Fix
User-agent: * then Disallow: /, no Perplexity group
Obeys the * group and crawls nothing
Add a User-agent: PerplexityBot group that allows public paths
A "block AI bots" list that includes PerplexityBot
Crawls nothing; you leave Perplexity's results
Remove PerplexityBot from the list; it is not a training crawler
User-agent: Perplexity-Bot
Matches nothing and falls back to *
Use the token exactly: PerplexityBot
User-agent: PerplexityBot with only Allow: /
Crawls everything, including paths the * group hid
Repeat your private Disallow lines in the group
Matching on the token is case-insensitive, so perplexitybot works, but the spelling must be exact. After any change, Perplexity says "it may take up to 24 hours for our systems to reflect changes." The robots.txt guide for AI bots has a complete file covering every operator.
Which rules fit which goal?
Decide what you want from Perplexity first, then write the smallest set of rules that does it. Because Perplexity says neither agent is used for model training, there is no "block training, keep citations" split to make here, unlike with OpenAI or Anthropic. The choice is between being a source and not being one.
Your goal
robots.txt
Firewall
Result
Be cited in Perplexity answers
Allow PerplexityBot, or make sure no rule catches it
Allow PerplexityBot's and Perplexity-User's published ranges
Pages can be indexed and fetched for users
Keep one area private, such as /app/
Disallow: /app/ inside the Perplexity group
Require a login on that area
Public pages stay citable; the private area is not crawled
Stay out of Perplexity search
User-agent: PerplexityBot then Disallow: /
Nothing needed for the crawler
No search indexing; users can still send Perplexity-User to a page
Stay out of Perplexity completely
Disallow both tokens
Block both published IP lists
No indexing and no user fetches
The last row is the only one that needs the firewall to do real work, because robots.txt alone does not reliably stop Perplexity-User. Anything that must never be fetched by any agent belongs behind authentication, not behind a Disallow line: robots.txt rules, in the words of RFC 9309, "are not a form of access authorization."
How do you verify PerplexityBot by IP?
Check the source address against Perplexity's published lists: https://www.perplexity.com/perplexitybot.json for PerplexityBot and https://www.perplexity.com/perplexity-user.json for Perplexity-User. Each file holds a prefixes array of ipv4Prefix entries. A request that says PerplexityBot from any other address did not come from Perplexity.
python3 - <<'EOF'
import ipaddress, json, urllib.request
ip = ipaddress.ip_address("203.0.113.7") # the address from your log
data = json.load(urllib.request.urlopen("https://www.perplexity.com/perplexitybot.json"))
nets = [ipaddress.ip_network(p["ipv4Prefix"]) for p in data["prefixes"] if "ipv4Prefix" in p]
print("PerplexityBot" if any(ip in n for n in nets) else "not PerplexityBot")
EOF
Perplexity advises combining both checks in firewall rules: user-agent contains PerplexityBot or Perplexity-User, and source IP is in the published ranges. Matching the user-agent alone lets anyone who copies the string through; matching the IP alone breaks when Perplexity adds addresses. Perplexity also recommends automating updates to the lists, since the ranges change.
Why can a firewall hide you from Perplexity when robots.txt is open?
A firewall answers the request before your site does, so robots.txt never gets a say. If a WAF rule, a bot-protection mode or a CDN setting returns a 403 or a challenge page to PerplexityBot, Perplexity gets no content from you, even though your robots.txt allows it.
On Cloudflare, three settings can cause it. Bot Fight Mode issues challenges to traffic matching known bot patterns, and Cloudflare's documentation says it "cannot be customized, adjusted, or reconfigured via WAF custom rules," so a skip rule will not exempt PerplexityBot from it. The AI bot policies under Security Settings let you block the Search, Agent and Training behaviours, and Cloudflare defines Search as crawlers that "collect or index your content to answer questions about it later," which is PerplexityBot's job. And from September 15, 2026, new Cloudflare domains block Agent traffic on pages that display ads by default, and Cloudflare's Agent category includes "chat fetch bots" of the kind Perplexity-User is.
Test what Perplexity's agents actually receive:
Request a page as PerplexityBot and print the status: curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)" https://yoursite.com/.
Repeat with the Perplexity-User string.
Check headers with curl -sI for cf-mitigated: challenge, which Cloudflare sets on every challenge page.
Read your firewall or security event log for requests from Perplexity's published addresses and the action applied.
Your laptop is not on Perplexity's address list, so a correctly written rule that allows only verified Perplexity addresses may still block your curl test. The firewall log is the real answer. Perplexity's documentation walks through the rule for Cloudflare and AWS WAF, and the guide to Cloudflare blocking AI crawlers covers every Cloudflare setting involved.
Does PerplexityBot render JavaScript?
No, by Vercel's measurement. Vercel's study of AI crawler traffic listed PerplexityBot among the crawlers that do not render JavaScript and counted 24.4 million PerplexityBot fetches across its network in one month. A page whose content appears only after scripts run reaches Perplexity as an empty shell.
Search the raw HTML for a sentence from your page: curl -s https://yoursite.com/pricing | grep -c "per seat". A 0 means Perplexity cannot quote it. Server rendering, static generation or prerendering fixes it. Run the same test for your <title>, your meta description and any JSON-LD block: if a script injects them after load, a crawler that does not execute JavaScript never sees them either, and Perplexity has to describe your page from whatever text is left.
How much does Perplexity send back?
More than other AI platforms per page crawled. Cloudflare's 2025 Radar review found Perplexity had the lowest crawl-to-refer ratios of the major AI platforms, with peaks generally below 400 crawls per referral and below 200 from September 2025 onwards. Allowing PerplexityBot costs you crawl load and returns visits at a better rate than the other platforms Cloudflare measured. To see those visits yourself, filter your analytics for referrals from perplexity.ai and compare them month by month with the PerplexityBot requests in your server logs.
A robots.txt line is easy to read; a firewall rule that challenges one user agent is not. Run the free scan on LaunchScaler: it takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. Its AI visibility checks flag robots.txt rules that block PerplexityBot, OAI-SearchBot or Claude-SearchBot, fail a wildcard block that catches them, request your page with a retrieval bot's user agent to catch a 403 or a challenge that a browser never sees, and check whether your main content is in the raw HTML Perplexity reads.
Only if you do not want to appear as a source in Perplexity answers. Perplexity says PerplexityBot is not used for model training, so blocking it does not protect your content from training; it only removes you from Perplexity's search results.
03
Does Perplexity-User respect robots.txt?
Mostly not. Perplexity's documentation says that since a user requested the fetch, Perplexity-User generally ignores robots.txt rules. It visits a page when a user's question needs it and may link that page in the answer.
04
How long does a robots.txt change take to reach Perplexity?
Perplexity says each setting works independently and it may take up to 24 hours for its systems to reflect changes.
05
Why is Perplexity not citing my site when robots.txt allows PerplexityBot?
A firewall or bot-protection rule may be challenging it. Firewall rules answer a request before robots.txt matters, so a 403 or a challenge page hides your site even with an open robots.txt. Test with PerplexityBot's user agent and check your firewall events.
AI Mode splits a question into many searches and cites pages that answer each part. Be indexed, answer every sub-question, and track it apart from AIO.
Reddit is 46.7% of the citations among Perplexity's top ten sources. Why AI answers lean on Reddit, and how to take part without breaking subreddit rules.
A copy-paste robots.txt that disallows AI training crawlers and allows AI search and user agents, one commented group per bot, plus the traps to avoid.