Time to First Byte is slow: server, CDN and cache fixes
TTFB is Good at 0.8 s or less and poor above 1.8 s. Walk the request path (redirects, DNS, TLS, CDN, server, database) and fix the stage that's slow.
LLaunchScaler·Published ·8 min read
To reduce Time to First Byte, measure where the time goes along the request path (redirects, DNS, the TCP and TLS connection, the CDN, the server and the database) and fix the stage that is slow. The fixes that usually matter most are serving HTML from a CDN cache where the content allows it, removing redirects, putting server functions next to their database, and cutting sequential database queries per request. Good TTFB is 0.8 seconds or less at the 75th percentile; poor is above 1.8 seconds.
TTFB is not a Core Web Vital, but every other loading metric waits for it. web.dev puts it plainly: "nothing can happen on the frontend until the backend delivers that first byte." A slow server makes a Good Largest Contentful Paint hard to reach whatever you do on the page.
What is TTFB, and what counts as slow?
TTFB is "the time between starting navigating to a page and when the first byte of a response begins to arrive." web.dev lists its parts as redirect time, service worker startup time (if any), DNS lookup, connection and TLS negotiation, and the request up to the first byte of the response. Good is 0.8 seconds or less; poor is greater than 1.8 seconds.
TTFB at the 75th percentile
Rating (PageSpeed Insights)
800 ms or less
Good
Over 800 ms up to 1,800 ms
Needs improvement
Over 1,800 ms
Poor
web.dev calls the thresholds "a rough guide." A server-rendered page can afford a slightly higher TTFB than a client-rendered one, because its first bytes already contain the content. The real target is a TTFB small enough that LCP can still finish within 2.5 seconds; improving Largest Contentful Paint shows how TTFB fits into the LCP budget as its first sub-part, about 40% of the total.
Questions, answered
What people ask about this
01
What is TTFB?
Time to First Byte is the time from starting to navigate to a page until the first byte of the response arrives. It is the sum of redirect time, service worker startup, DNS lookup, connection and TLS negotiation, and the request itself up to the first byte.
Use field data to know whether TTFB is a problem, then break a single request into stages to find where. PageSpeed Insights shows TTFB from real Chrome users. web.dev warns that lab tools often test "the final URL," missing redirects, and that Lighthouse's server response time audit "excludes DNS lookup and redirect times."
For one request, curl's timing variables split the time by stage:
The DNS, TCP, TLS and first-byte values are each measured from the start of the request, so subtract the previous one to get a stage's own time; time_redirect is the total spent on redirects before the final request. time_starttransfer is the time "until the first byte is received," which "includes the time the server needed to calculate the result." If most of the time sits between time_appconnect and time_starttransfer, the server is slow; if it sits before time_appconnect, the connection is.
To see inside the server, add a Server-Timing header. web.dev's example is Server-Timing: db;desc="Database";dur=121.3, ssr;desc="Server-side Rendering";dur=212.2, which Chrome DevTools shows in the Network panel and which real-user tools can read from the Navigation Timing API. It turns "the server is slow" into "the database takes 121 ms and rendering takes 212 ms."
Stage (curl timing)
Slow when
Typical fix
Redirects (time_redirect)
Links point at URLs that redirect
Link the final URL; add HSTS
DNS (time_namelookup)
A slow or distant DNS provider
Use a fast DNS provider or your CDN's DNS
TCP and TLS (time_connect, time_appconnect)
Server far from the visitor, old TLS
Serve from a CDN edge; use TLS 1.3
Server work (to time_starttransfer)
Uncached HTML, slow queries, cold starts
Cache at the edge; fix queries; move the function next to the database
How do redirects add to TTFB?
Every redirect is a full extra request before the real one starts, and web.dev names redirects as "one common contributor to a high TTFB." Same-origin redirects are the ones you control: links missing https://, which default to http:// and redirect, or links whose trailing slash does not match your canonical form.
Fix them in this order:
Find internal links that return 301 or 302 and point them at the final URL.
Give advertisers, newsletters and social profiles the final URL, so their own tracking redirect is the only extra hop.
Send Strict-Transport-Security. web.dev notes HSTS makes the browser go straight to HTTPS on later visits, and adding your domain to the HSTS preload list covers the first visit too.
Collapse chains such as http://example.com to https://example.com to https://www.example.com/ into one hop.
How do you cache HTML at the CDN?
Cache the HTML itself at the edge wherever the page is the same for every visitor. A CDN serves from "edge servers" close to the visitor, and web.dev notes that without cacheable content, "requiring a trip all the way back to the origin can negate much of the value of a CDN." Marketing pages, docs and blog posts are usually cacheable; logged-in dashboards are not.
The mechanics are Cache-Control headers the CDN honours:
s-maxage sets how long the response stays fresh in a shared cache such as a CDN, here 5 minutes, and stale-while-revalidate lets the cache reuse the stale copy while it revalidates, as MDN describes both. web.dev notes "even a short caching time can result in noticeable performance gains for busy sites," because only the first visitor in each window waits for the origin. Frameworks such as Next.js can prerender pages at build time or revalidate them on a schedule, which gives you the same edge-cached HTML without setting headers by hand.
Watch for two things that defeat the cache. Unique query parameters, such as analytics tags, "may look like different content to the CDN," so the cached copy is never used; configure the CDN to ignore them in the cache key. And personalisation in the HTML (a user's name in the header) makes the whole page uncacheable; render that part on the client or from a separate request.
web.dev also warns that caching "can mask a slow backend." Keep a way to bypass the cache, such as a header or parameter your CDN respects, so you can still measure the origin's real time.
Do serverless cold starts and function regions affect TTFB?
Yes, both do. A cold start adds the time to boot a new function instance to the request, and a function running far from its database adds the network distance to every query. On Vercel, functions run by default in Washington, D.C. (iad1), and Vercel recommends running them "close to your data source."
The region problem multiplies with the number of queries. If your database is in Europe and your function runs in iad1, every query crosses the Atlantic and back, so a page that makes five sequential queries pays that round trip five times before sending a byte. Set the function region to match the database, in the dashboard's Function Regions setting or with "regions" in vercel.json:
{
"regions": ["fra1"]
}
For cold starts, Vercel's Fluid compute lists "automatic bytecode optimization" and "function pre-warming on production deployments" to reduce them, with bytecode caching applying to Node.js 20 and later in production only. Beyond the platform, keep server bundles small, avoid heavy work at module load, and prerender pages that do not need a server at request time at all.
How do database round trips slow the first byte?
Every query that must finish before the HTML can start adds its full round trip plus its execution time. Sequential queries add up, so a page that fetches the user, then the account, then the plan, then the usage, waits four times. The fix is fewer, parallel and cached queries, measured with Server-Timing so you can see each one.
Run independent queries at the same time instead of one after another, for example with Promise.all in Node.js.
Replace per-row queries in a loop (the N+1 pattern) with one query that fetches all rows.
Cache results that many requests share, such as plan limits or navigation data, in memory or at the edge.
Add indexes for the queries the page runs, and check their plans with your database's EXPLAIN.
Stream the page. web.dev notes newer versions of React can stream server-rendered markup "as it's being rendered," so the <head> and page shell leave the server while slower data is still loading.
Do HTTP/2, HTTP/3 and 103 Early Hints help?
They help the connection and what happens after the first byte, not a slow backend. web.dev notes CDNs serve "modern protocols such as HTTP/2 or HTTP/3," that HTTP/3 avoids TCP's head-of-line blocking by running over UDP, and that TLS 1.3 "is designed to keep TLS negotiation as short as possible."
103 Early Hints is for pages where the server needs time to build the HTML. The server sends an early 103 response with Link headers so the browser can start fetching critical CSS or connecting to a font origin while it waits:
Chrome's documentation says Early Hints "are not useful if your server can send a 200 (or other final responses) right away," and web.dev adds that static sites "probably won't" benefit. It also changes the number you measure: web.dev notes the 103 response "counts as the 'first bytes'," so TTFB can look fast while the server is still slow. Measure the real server time with Server-Timing alongside it.
Does a slow TTFB affect crawling and indexing?
It can. Google's crawl budget documentation says that if a site's "response times (including latency and Time-to-First Byte) remain stable or improve, the limit goes up," and that if the site slows down, "the limit goes down and Google crawls less." A slow server means fewer pages fetched per visit.
That matters most for sites with many URLs waiting to be crawled. If Search Console shows pages as discovered but not crawled, server speed is one of the two sides of the problem; Discovered - currently not indexed covers the crawl capacity check. For the Core Web Vitals side, what a failed Core Web Vitals assessment means covers how TTFB feeds into the verdict through LCP.
Check TTFB and your server's consistency
To see your TTFB from real users and how your server behaves under repeated requests, run the free scan on LaunchScaler with the URL, no account needed. Its checks read TTFB from Chrome field data at the 75th percentile, measure median and slow-end response times with any timeouts, report whether the site serves HTTP/2 or HTTP/3, flag long redirect chains and the HTTP to HTTPS redirect, and check text compression and static-asset caching headers.
0.8 seconds or less at the 75th percentile is Good, and more than 1.8 seconds is poor. web.dev calls these a rough guide, because TTFB is not a Core Web Vital; what matters is that it leaves room for a Good LCP.
03
How do I test TTFB?
For real users, read TTFB in PageSpeed Insights field data or the web-vitals library. For a single request, run curl with a write-out format that prints time_namelookup, time_connect, time_appconnect and time_starttransfer, which break the time into DNS, TCP, TLS and first byte.
04
What causes a slow TTFB?
The usual causes are redirects before the final URL, a server far from the visitor with no CDN cache, HTML that cannot be cached at the edge, slow backend work such as several sequential database queries, and serverless cold starts or a function running far from its database.
05
Does TTFB affect SEO?
Indirectly. It is not a Core Web Vital, but it is part of every LCP, and Google's crawl budget documentation says that when response times, including TTFB, get longer, Google's crawl capacity limit goes down and it crawls less.
The six security headers every site should send, the value for each, the Mozilla Observatory penalty for a missing one, and how to set them on any host.