Cache hit ratio: what 80 % vs 95 % means for your database

Short answer: the database sees the misses, not the hits. Going from 80 % to 95 % hits cuts its read load from 20 % to 5 % of the traffic: 4× less. Going from 95 % to 99 % cuts it another 5×.

The flip side: a database sized for 95 % hits cannot survive the cache going cold. After a restart every read is a miss until the cache warms up again, and that is when it falls over.

The arithmetic

database reads = (1 − hit ratio) × reads
For 10,000 reads per second in front of the cache. Writes go to the database either way.
Hit ratioMissesDatabase reads/svs 95 %
0 % (cold cache)100 %10,00020×
50 %50 %5,00010×
80 %20 %2,0004×
90 %10 %1,0002×
95 %5 %5001×
99 %1 %1000.2×

Two consequences. First, a few points of hit ratio are worth more than they look: 95 → 90 % sounds like a small drop and doubles the database's read load. Second, a high hit ratio is a dependency: the higher it is, the bigger the gap between normal load and cold-cache load. At 98 % hits the database gets 50× its usual reads when the cache is empty.

The hit ratio you get depends on how popular your keys are, not only on cache size. Real traffic is skewed: a few products, links or users get most of the reads, so a cache holding a small share of the keys can answer most requests. Measure it (hits ÷ lookups, per cache) rather than assume it.

A simulated example: the cache restart

Stackrig's Cache restart template is a product catalog API: 1,500 requests/s, each reading one of 500,000 products through Redis (cache-aside, popularity skewed like real shops), plus a few writes. Redis answers 98 % of the reads, so the Postgres behind it only sees 30 reads/s and was sized small (2 vCPU). Then Redis is flushed, as in a restart or failover. These are Stackrig model results for that design, not measurements:

Stackrig model results for the Cache restart template (live model, AWS prices). The warm-up follows key popularity: hot products are back within a second, the long tail takes minutes.
ScenarioPostgres loadErrors
Warm cache (98 % hits)23 %0; p99 28 ms
Redis flushed, first 10 s (hit ratio 31 %)157 %38 % of requests
Until the cache is warm again–errors for ~130 s, ~29,000 failed requests
Fix: request coalescing on Redis–still ~29,000 failed
Fix: one read replica (+$240/month)–0

Coalescing barely helps here, and that is the lesson: it merges concurrent misses for the same key, but after a flush the misses are for thousands of different products. Headroom helps, and it costs money every month. In the smaller Web app with database design the same arithmetic runs the other way: a Redis cache in front of an overloaded Postgres takes it from 132 % to 36 % load at 3× traffic (also a model result; see the live demo on the start page).

Fixes that work

  1. Warm up before taking traffic. Pre-load the hottest keys (from yesterday's access log or the old node) before a new cache node joins, and restart cache nodes one at a time, never all at once.
  2. Request coalescing. Let one request load a missing key while others wait for it (a lock or "single flight"). It stops the herd on a hot key; as the example shows, it does not help when the misses are spread over many keys.
  3. TTL jitter. Give keys slightly random expiry times, so keys cached at the same moment don't all expire at the same moment and refill together.
  4. Serve stale while refreshing. Return the expired value and refresh it in the background, where slightly old data is acceptable.
  5. Headroom or shedding. Size the database for more than the warm load (a replica, a bigger instance), or limit the rate of misses that reach it and fail fast for the rest. Decide which one you pay for.

See it before it happens

In Stackrig a cache has a hit ratio, a warm-up that follows key popularity after a flush, and optional request coalescing. Flush it in the simulation, watch the database turn red, and compare the fixes with their cost. How far the model can be trusted is published on How accurate is Stackrig?

Open “Cache restart” in the playground Watch the live demo Get early access

More