robots.txt example for a SaaS site: what to allow, what to block
A commented robots.txt for a SaaS site: marketing pages open, /app/ and /api/ closed, a Sitemap line, and how Google treats a missing or failing file.
LLaunchScaler·Published ·9 min read
A robots.txt for a typical SaaS site leaves the marketing pages open, blocks the logged-in app and the API, and names the sitemap with a full URL. The whole file can be five lines: User-agent: *, Allow: /, Disallow: /app/, Disallow: /api/, and Sitemap: https://example.com/sitemap.xml.
Below is that file with every line explained, the optional rules for search and filter URLs, what Google does when the file is missing or failing, and the mistakes that block a whole site.
What does a robots.txt example for a SaaS site look like?
A SaaS robots.txt usually needs one group for all crawlers, two or three Disallow lines for the product and API paths, and a Sitemap line. Everything you do not disallow is crawlable by default, so the file lists exceptions, not pages. This is a complete, commented version for https://example.com:
# robots.txt for https://example.com
# One group for every crawler that has no group of its own.
User-agent: *
# Optional: crawling is allowed by default. Kept here so the intent is explicit.
Allow: /
# The logged-in product. Nothing behind the login is useful to a crawler.
Disallow: /app/
# API routes return JSON for your own front end, not pages for searchers.
Disallow: /api/
# Internal search results: every query creates a new URL.
Disallow: /search?
# Sort and filter parameters that reorder the same list.
Disallow: /*?*sort=
Disallow: /*?*filter=
# Absolute URL, including https and the host.
Sitemap: https://example.com/sitemap.xml
Lines starting with # are comments and are ignored. Field names such as User-agent and disallow are case-insensitive, but the paths are not: Disallow: /App/ does not block .
Questions, answered
What people ask about this
01
What should a basic robots.txt file contain?
A User-agent: * line, the Disallow rules for paths crawlers should skip (for a SaaS, usually /app/ and /api/), and a Sitemap: line with the full https URL of your sitemap. Everything not disallowed is crawlable by default.
If your site runs on Next.js, you can generate the same file from app/robots.ts, which returns an object with rules (userAgent, allow, disallow) and a sitemap URL, and Next.js serves it at /robots.txt. A static app/robots.txt works too.
What should a SaaS block, and what should it never block?
Block URLs that are infinite, private or useless to a searcher: the logged-in app, the API, internal search results and parameter URLs that only sort or filter a list. Never block the CSS, JavaScript and image files your public pages need to render, because Google renders pages with them and indexes what it sees.
Path or pattern
Block it?
Why
/app/, /dashboard/
Yes
Behind a login, so crawlers only get the login redirect.
/api/
Yes
JSON endpoints, not pages.
/search? (internal search results)
Yes
Every query string is a new URL, and an empty results page is one of Google's own examples of a soft 404.
?sort=, ?filter=, ?view= parameters
Yes, if the unfiltered page exists
Google's faceted navigation guide calls these infinite URL spaces and recommends disallowing them.
/_next/, /static/, /assets/, *.css, *.js
No
Google needs them to render the page.
/pricing, /features/, /blog/, /docs/
No
These are the pages you want found.
/login, /signup/thank-you
Optional
Harmless to crawl on a small site. To keep them out of results, use noindex instead, which requires crawling.
Google's own list of URLs to keep crawlers away from covers faceted navigation and session identifiers, infinite spaces, shopping cart and infinite scrolling pages, and pages that perform an action such as "sign up" or "buy now". Its robots.txt introduction draws the line on resources: you can block unimportant scripts or styles, but "if the absence of these resources make the page harder for Google's crawler to understand the page, don't block them."
Watch the prefix matching. Rules match from the start of the path, so Disallow: /app (no trailing slash) also blocks /apple-pay-guide and /app-integrations. Disallow: /app/ blocks only the folder. To block the bare /app URL as well, add Disallow: /app$, where $ marks the end of the URL.
How does Google decide which rule applies?
Google picks one group per crawler, then the most specific matching rule inside it. A crawler follows the group with the most specific user agent that matches its name and ignores all others, including User-agent: *. Within that group, the longest matching path wins, and on a tie the least restrictive rule wins.
Two consequences matter for a SaaS file:
A named group replaces the * group for that crawler. If you add User-agent: Googlebot with one Disallow, Googlebot stops reading your * rules entirely, so repeat /app/ and /api/ inside it.
Allow can carve an exception out of a block. With Disallow: /docs/ and Allow: /docs/public/, the longer Allow path wins for /docs/public/setup.
Google's spec gives the tie case directly: with allow: /folder and disallow: /folder, the URL /folder/page is allowed, because Google uses the least restrictive rule. The * wildcard matches any run of characters and works in Allow and Disallow paths, but not in the Sitemap line.
Where does robots.txt go, and what status must it return?
The file must be named robots.txt, sit at the root of the host, and apply only to that protocol, host and port. https://example.com/robots.txt does not cover https://docs.example.com/, https://www.example.com/ or http://example.com/. Each host you serve needs its own file, and it should answer with a 200.
For a SaaS with a marketing site, a docs subdomain and an app subdomain, that means three files:
https://example.com/robots.txt: the file above.
https://docs.example.com/robots.txt: usually User-agent: * with nothing disallowed, plus its own Sitemap line.
https://app.example.com/robots.txt: User-agent: * and Disallow: / if nothing on that host should be crawled.
The file must be UTF-8 plain text, and Google ignores anything past 500 KiB. It caches the file for up to 24 hours, so an edit is not read instantly. After an urgent change, open the robots.txt report in Search Console (Settings), select the file, and choose Request a recrawl.
One SaaS-specific trap: a single-page app that serves index.html for every unknown path will answer /robots.txt with your HTML shell and a 200. Google then tries to parse the HTML as rules and ignores everything it cannot read, so your Disallow lines effectively do not exist. Check the response with curl -i https://example.com/robots.txt and confirm the body is the plain-text file.
What happens if robots.txt is missing or returns an error?
A missing file is fine: Google treats every 4xx response except 429 as if no robots.txt existed and crawls without restrictions. A 5xx is not fine. Google stops crawling the whole site for the first 12 hours while retrying, then relies on the last good copy for up to 30 days.
robots.txt response
What Google does
200 with valid rules
Applies the rules as written.
3xx redirect
Follows at least five hops, then treats the file as a 404.
404, 410 or another 4xx (except 429)
Assumes there are no crawl restrictions.
5xx, 429, timeout or DNS failure
Stops crawling the site for 12 hours while retrying, then uses the last good version for 30 days. With no cached copy, it assumes no restrictions.
Still failing after 30 days
Behaves as if there is no file if the site is otherwise available; stops crawling if the site has availability problems.
This asymmetry catches teams that put robots.txt behind the same middleware as the app. If an auth check, a rate limiter or a crashed edge function answers /robots.txt with a 500, a 503 or a 429, Google treats your whole site as off limits for half a day at a time. Serve the file statically, outside any auth or rate limiting.
Why does 'Disallow: /' block the whole site?
Disallow: / matches every path, because every URL path starts with /. Under User-agent: * it tells every crawler without its own group to crawl nothing on that host. It is correct on a staging host and a site-wide outage in search when it ships to production.
It usually arrives one of two ways: a robots.txt committed for staging that gets deployed everywhere, or a framework or host setting that serves a blocking file on preview domains and gets copied into the production config. Nothing warns you when it happens. Google picks up the new file within about a day (it caches robots.txt for up to 24 hours), stops reading your pages, and URLs that stay indexed because other sites link to them show up with no description. The noindex shipped to production guide covers catching this class of deploy mistake before it costs traffic.
Protect staging with HTTP authentication instead. Google's own robots.txt introduction says the file cannot enforce crawler behavior and recommends password protection for anything private, and a login-protected staging site never needs a blocking robots.txt at all.
Is robots.txt the right tool to keep a page out of Google?
No. robots.txt controls crawling, not indexing. Google says a disallowed URL can still be indexed if other pages link to it, showing the URL without a description. To keep a crawlable page out of results, allow crawling and send noindex in a meta tag or the X-Robots-Tag header.
The two tools also fight each other. A page blocked in robots.txt is never fetched, so Google never sees its noindex. The X-Robots-Tag vs robots.txt comparison shows which to use for pages, PDFs and whole folders, and the blocked by robots.txt guide covers the Search Console status you get when a page you want indexed is disallowed.
What about AI crawlers like GPTBot and OAI-SearchBot?
AI crawlers read the same file, and each one follows the most specific group that matches its name. That means a group for GPTBot or ClaudeBot replaces your * rules for that bot, so it must repeat /app/ and /api/. The companies also split their bots by job: OpenAI uses OAI-SearchBot for ChatGPT search and GPTBot for training, Anthropic uses Claude-SearchBot for search and ClaudeBot for training, and Perplexity says PerplexityBot is not used to crawl content for AI foundation models.
Which ones to allow is a separate decision with its own trade-offs, and a blanket block written for training bots often catches the retrieval bots that put you in ChatGPT, Claude and Perplexity answers. The robots.txt file for AI bots gives the per-bot groups and the reasoning for each.
How do you test robots.txt before and after you deploy?
Test in this order after any change:
Open https://example.com/robots.txt in a private window and confirm you see plain text, not your app's HTML.
Run curl -I https://example.com/robots.txt and confirm the status is 200, from every host you serve.
In Search Console, open Settings, then the robots.txt report, and check the fetched version and any parse errors.
Inspect one marketing URL and one /app/ URL in URL Inspection; "Crawl allowed?" should read Yes for the first and No for the second.
To check the file against your live pages in one pass, run the free scan on LaunchScaler with just your address and no account. It runs 156 checks across 6 of its 7 categories, and its robots checks flag a Disallow rule that blocks a page you want ranked, a noindex hidden behind a robots.txt block that Google can never read, a missing sitemap or a robots.txt with no Sitemap line, and rules that block the AI retrieval bots (OAI-SearchBot, Claude-SearchBot and PerplexityBot) along with the training bots.
Under User-agent: * it tells every crawler without its own group not to crawl any URL on that host. It is the right file for a staging site and the wrong one for production, where it stops Google reading every page.
03
What happens if my robots.txt returns a 404?
Google treats every 4xx response except 429 as if no robots.txt existed, so it crawls the whole site with no restrictions. That is fine if you have nothing to block.
04
What happens if my robots.txt returns a 500 error?
Google stops crawling the site for the first 12 hours while it retries the file, then uses the last good copy for up to 30 days. A robots.txt that keeps failing can halt crawling of a site that is otherwise working.
05
Should I block CSS and JavaScript in robots.txt?
No. Google renders pages with their CSS and JavaScript, and its documentation says not to block resources whose absence makes a page harder to understand. Blocking /_next/, /static/ or /assets/ can leave Google indexing a broken layout.
06
Does robots.txt keep a page out of Google?
No. It only stops crawling, and Google can still index a blocked URL if other pages link to it, without a description. To keep a page out of results, allow crawling and add a noindex meta tag or X-Robots-Tag header.
Server error (5xx) means Googlebot got a 500-level status or a timeout. Find when it happened in Crawl Stats, match it to your logs, then fix the cause.
Couldn't fetch can be transient, but a sitemap that stays unread is usually blocked, redirected, not XML or over 50,000 URLs. Check each cause with curl.
Sitemap lastmod uses W3C Datetime, such as 2026-09-28T09:00:00+00:00. Google trusts it only when it is consistently accurate, so set it from real edits.