Warmup Cache Request: A Practical Guide to Cache Warming
Deploys have a nasty habit. You push new code and purge the cache, and the very next visitor pays for all of it. Their page loads a second or two slower than normal, and nobody on your team ever sees it happen. A warmup cache request fixes that by making the first visit yourself before a real person arrives. This guide covers what it is, how to send one, and the part most articles skip: how to prove it worked.
What Is a Warmup Cache Request?
A warmup cache request is an automated HTTP request sent to a URL so your caching layers build and store the response ahead of time. A deploy script, a scheduled job, or a CI step sends it. A human doesn’t.
The request follows the same path a visitor would: CDN edge, reverse proxy, application, database. Every layer that misses stores a copy on the way back. The next real visitor gets that stored copy and skips the origin entirely.
You’ll also hear cache warming, preloading, or priming. Same idea, different labels. The warmer is just a script playing the role of visitor number one.
Cold Cache vs Warm Cache: What Actually Changes
On a cold cache, the edge finds nothing and forwards the request to your origin. The origin runs its logic, queries the database, renders the template, and only then responds. Whoever triggered that chain waits through every step.
On a warm cache, the edge already holds a copy and answers right away. No query, no render.
Here’s where people overstate things. A cold cache raises Time to First Byte, but TTFB isn’t a Core Web Vitals metric. Google’s web.dev guidance treats it as a metric that comes before FCP and LCP, with a rough guide of 0.8 seconds or less for most sites. So warming helps your vitals indirectly by giving LCP a head start. Any article that lists TTFB as vital is stretching the facts.
When to Send One
Four moments matter most. Right after a deploy, because releases often clear caches. Right after a manual purge or invalidation. Before a planned spike, like a product drop or a newsletter send. And on a schedule that lands shortly before your TTL expires, so popular pages never go cold on their own.
How to Send a Warmup Cache Request
Start small. Pull your top URLs from analytics: homepage, main category pages, best sellers, and key landing pages. Warm those first and ignore the long tail.
The simplest version is a shell loop:
bash
while read -r url; do
curl -s -o /dev/null \
-w "%{http_code} %{time_starttransfer}s %{url_effective}\n" "$url"
sleep 0.5
done < urls.txt
That sleep matters. Fire hundreds of requests at once, and you’ve built your own traffic spike, which can hit the origin harder than real visitors would. Keep it to a couple of requests per second unless you’ve tested higher.
Run the script as a post-deploy step in CI/CD, so it happens every time without anyone remembering. Sites with a sitemap can parse it instead of maintaining a list by hand. JavaScript-heavy pages are a different case: a headless browser like Playwright also pulls the fonts, scripts, and API calls that plain curl skips.
Two safety rules. Never warm logged-in, cart, or account pages, since a cached personal page can end up in front of the wrong person. And set a clear user agent so you can filter warmup traffic out of analytics and allow it through your firewall.
Why Warmups Quietly Fail: Cache Keys
Here’s the mistake I’d bet on. Your script reports 200 on every URL, yet the hit ratio barely moves. Usually the cache key doesn’t match.
A cache stores each response under a key built from the URL and often from query strings, cookies, and headers named in Vary. Warm shoes, and a visitor arriving at /shoes?utm_source=email may get a separate entry, unless your CDN ignores tracking parameters. Compression works the same way: a script that doesn’t send Accept-Encoding can warm a variant real browsers never request.
Geography matters too. Most CDN edges keep their own cache, so a script running from one country only warms nodes near it. If your audience spans continents, run warmups from several regions, or use your CDN’s tiered caching where it’s offered.
How to Prove the Warmup Worked
Running the script proves nothing on its own. Check three things.
Cache headers. Request a page and read the response headers:
bash
curl -sI https://example.com/ | grep -iE "cache|age"
Cloudflare returns CF-Cache-Status, where HIT is the goal. Many setups add X-Cache or X-Cache-Status, and the standard Age header shows how long a copy has been stored. A MISS right after warming usually means your Cache-Control headers are blocking storage.
TTFB. Compare before and after using curl’s time_starttransfer or a synthetic test from several regions. [YOUR DATA: add your own before and after numbers here.]
Hit ratio. Your CDN dashboard shows cache hits against origin fetches. Watch it for the first few minutes after a deploy. A sharp drop in origin requests is the clearest sign the warmup landed. I won’t give you a “healthy” percentage as gospel, because sensible targets depend on how much of your content is cacheable.
The Other Meaning: App Engine Warmup Requests
Search this phrase and you’ll sometimes land on platform docs instead of CDN guides. Google App Engine has its own feature: with warmup requests enabled, it issues GET requests to /_ah/warmup, and you can add handlers there for tasks like pre-caching data. Google is upfront that warmup requests are not guaranteed to be called. A first instance or a steep traffic ramp can produce a loading request instead. Treat it as best effort, and don’t make it your only defense.
When You Can Skip It
Small site, light traffic, few deploys? Real visitors rebuild a small cache within minutes, and a warmer may cost more engineering time than it saves. Warming earns its place on busy sites, frequent releases, and revenue pages where a slow first load costs sales.
It also can’t repair a slow query. If a cache miss takes four seconds, warming just hides that until the next miss.