Slide the hit rate. Crank the TTL. Watch the database breathe — or panic.
What problem are we solving?
A database query at 50ms is fine — until you multiply it by 1,000 requests per second. That's 50 seconds of DB work every wall-clock second, and your users see it as latency. A cache sits in front of the slow thing and answers the questions it's already answered. The tradeoff: every cached answer is potentially stale, and a cache miss is more expensive than no cache at all.
Live simulation
updates every 1s
App
Cache80%
DB40/s
Hits this tick
0
Misses this tick
0
DB reqs this tick
0
Cache hit ≈ 5ms · DB read ≈ 50ms.
Controls
200 req/s
80%
60s
Cache available
Serving reads from cache
Live metrics
Cache hits00 this tick
Cache misses00 this tick
DB requests0
Avg latency0.0 ms
Stale-data risk20%TTL 60s
Worked example
At 200 req/s with 80% hit: DB sees 40 req/s. Without the cache, DB would see all 200 req/s.
What just happened?
80% of requests are served from cache at 5ms each. Only 40 req/s reach the DB. This is the payoff: cache absorbs the majority of load.
Try this
Turn the cache off
Your cache cluster just went down. Toggle the cache off and watch what happens to DB load and average latency. Then bring it back up.
If your DB can handle 500 req/s and traffic is 1000 req/s with an 80% hit rate, what happens to the DB when the cache fails? Where does the request go?
Key takeaway
A cache trades consistency for capacity. Every hit is a DB query you didn't have to run — but every entry has a TTL, and after the TTL you may serve stale data. The three dials — hit rate, TTL, and cache availability — together determine whether your cache is a force multiplier or a liability.