The cheapest request is the one you never make

Every response your application sends has a cost measured in latency, CPU, database queries, and money. Caching is the discipline of paying that cost once and reusing the answer. But "add caching" is not a single decision — a modern web stack has at least four distinct layers where answers can be stored, each with different latency, capacity, freshness guarantees, and failure modes.

Teams get in trouble in two opposite ways: they cache nothing and melt under load, or they cache everything and ship stale prices, ghost inventory, and deleted content that refuses to die. The fix is not more or less caching — it is putting the right data at the right layer with an explicit freshness contract.

This article walks through the four layers from the outside in: the browser, the CDN, the application, and the database. For each one we will look at what belongs there, the headers or code that control it, and how to invalidate it when the world changes.

Layer 1: The browser cache

The closest cache to the user is the one built into their browser. It is free, it is instant, and it is controlled entirely by the Cache-Control header your server sends.

http
Cache-Control: public, max-age=3600, stale-while-revalidate=86400

The directives that matter in practice:

The classic pattern is immutable static assets plus revalidated HTML. Build tools like Next.js fingerprint your JS and CSS bundles (app-a1b2c3.js), so you can cache them for a year:

http
Cache-Control: public, max-age=31536000, immutable

When the content changes, the filename changes, and the old cache entry simply becomes unreachable. The HTML document, which references those fingerprinted URLs, gets no-cache so users always pick up the latest references. This pairing gives you year-long caching with instant deploys — the best of both worlds.

Layer 2: The CDN edge

A CDN puts a cache in dozens of points of presence worldwide, cutting a 200 ms transatlantic round trip to 10 ms. The same Cache-Control headers apply, with one addition: s-maxage overrides max-age for shared caches only.

http
Cache-Control: public, max-age=60, s-maxage=3600, stale-while-revalidate=86400

This says: browsers may keep it for a minute, but the CDN may keep it for an hour and serve stale for up to a day while revalidating. That asymmetry is powerful — the CDN absorbs your traffic spikes, and a single revalidation from the edge refreshes what thousands of users see.

In a Next.js app deployed on a platform with edge caching, you express this in code:

ts
// app/api/products/route.ts
export async function GET() {
  const products = await db.query('SELECT * FROM products WHERE active = true');
  return Response.json(products, {
    headers: {
      'Cache-Control': 'public, s-maxage=300, stale-while-revalidate=3600',
    },
  });
}

The hard part at this layer is purging. Time-based expiry (s-maxage) is your baseline, but some changes cannot wait five minutes — a price fix, a takedown, a correction. Every serious CDN offers tag-based or path-based invalidation APIs. Next.js formalizes this with revalidateTag and revalidatePath, which purge the platform cache when your data mutates:

ts
import { revalidateTag } from 'next/cache';

export async function POST(req: Request) {
  const product = await updateProduct(await req.json());
  revalidateTag('products'); // purges every cached fetch tagged 'products'
  return Response.json(product);
}

Design rule: cache by tag, not by URL list. When a product changes, you do not want to enumerate every page that shows it — you want to invalidate the products tag and let the cache figure out the rest.

Layer 3: The application cache

Below the CDN sits your server, and this is where you cache things that are expensive to compute but not worth a network hop to the browser's or CDN's store: rendered fragments, aggregated query results, third-party API responses, session data. The canonical tool is Redis.

The dominant pattern is cache-aside (lazy loading):

ts
import { redis } from './redis';

async function getProductById(id: string) {
  const cacheKey = `product:${id}`;

  const cached = await redis.get(cacheKey);
  if (cached) return JSON.parse(cached);

  const product = await db.query('SELECT * FROM products WHERE id = $1', [id]);
  if (product) {
    // Cache for 10 minutes, with a small random jitter (see stampede section)
    await redis.set(cacheKey, JSON.stringify(product), 'EX', 600 + Math.floor(Math.random() * 60));
  }
  return product;
}

Three failure modes deserve explicit handling:

Cache stampede (thundering herd). A hot key expires, and ten thousand requests all miss simultaneously, all hitting the database at once. Mitigations: add random jitter to TTLs so keys do not expire in lockstep (shown above), or use request coalescing — the first miss acquires a lock and recomputes while everyone else waits for or reuses the stale value.

Cache penetration. Requests for IDs that do not exist always miss, and an attacker can exploit this to hammer your database. Cache negative results (null with a short TTL) or put a Bloom filter in front.

Stale-after-write. Cache-aside does not know when the database changes. The fix is a write-through or write-invalidate step in your mutation path:

ts
async function updateProduct(id: string, data: ProductInput) {
  const product = await db.query('UPDATE products SET ... WHERE id = $1 RETURNING *', [id]);
  await redis.del(`product:${id}`); // invalidate; next read repopulates
  return product;
}

Delete, don't update. Deleting on write is idempotent and race-resistant; updating on write can interleave with a concurrent read and leave the cache holding the older value.

Layer 4: The database

Your database is already caching, whether you configured it or not. PostgreSQL keeps hot pages in its shared buffer pool, and the OS page cache sits underneath that. You tune this layer with configuration (shared_buffers, typically around 25% of RAM as a starting point) and, more importantly, with indexes — the right index is effectively a cache of "where does this row live" that turns a table scan into a direct lookup.

Measure before adding another layer on top. In PostgreSQL, a quick buffer hit ratio check tells you whether your working set fits in memory:

sql
SELECT
  sum(heap_blks_hit)::numeric / NULLIF(sum(heap_blks_hit) + sum(heap_blks_read), 0) AS buffer_hit_ratio
FROM pg_statio_user_tables;

A healthy OLTP workload should show 0.99 or higher. If it is far below, the fix is usually more memory or better indexes — not a Redis layer in front of queries you have not examined.

Choosing a layer: a decision framework

When you identify something worth caching, ask three questions:

  1. Is it the same for every user? Yes → it can go to the CDN. No (per-user data) → browser private or application cache keyed by user.
  2. How stale can it be? Seconds to minutes → time-based expiry anywhere. Zero tolerance → short TTL plus tag-based purging, or do not cache at all.
  3. How expensive is a miss? Cheap query → browser/CDN caching is enough. N+1 aggregation or third-party API call → application cache with cache-aside.

The general principle: cache as far out (close to the user) as the freshness requirements allow. Every layer you move outward is an order of magnitude cheaper and faster, but harder to invalidate. Public, slow-changing content belongs at the edge; private, fast-changing data belongs close to the database or uncached.

Invalidation: the part everyone gets wrong

Phil Karlton's old joke — "there are only two hard things in computer science: cache invalidation and naming things" — survives because invalidation is where caching systems fail. A workable strategy combines three mechanisms:

And test the invalidation path, not just the caching path. Caching bugs that serve stale data are far more damaging than the performance problems you were solving.

Conclusion

Caching is not one knob but four: immutable fingerprinted assets and revalidated HTML in the browser, s-maxage and tag-based purging at the CDN, cache-aside with jitter and delete-on-write in Redis, and buffer tuning plus proper indexes in the database. The rules of thumb are simple to state: cache as close to the user as freshness allows, put a TTL on everything, invalidate by tag on writes, and measure the database before wrapping it in more infrastructure. Do those four things and you get the speed of caching without the 2 a.m. "why is the old price still showing" incident.