There's a special kind of education that only comes from watching your server melt under real traffic. No amount of architecture astronautics prepares you for the moment your database falls over at 2 AM because someone linked you from a popular forum. Here's what serving millions of requests a month actually teaches you.

Lesson 1: Cache everything, then cache more

The single highest-leverage thing you can do for a read-heavy workload is cache aggressively. In our case, over 60% of requests never touched the origin server, they were served from edge cache. That one decision did more for our infrastructure bill and our latency than every other optimization combined.

The key insight: most web traffic is repetitive. The same pages, the same API responses, requested over and over. A CDN with sensible cache headers turns your origin server from a workhorse into a rarely-consulted oracle. Set your TTLs thoughtfully, even 60 seconds of caching on a hot endpoint can collapse thousands of requests into one.

A CDN with sensible cache headers turns your origin server from a workhorse into a rarely-consulted oracle.

Lesson 2: Your database is fine; your queries aren't

Enjoying this story?

Get the five most important stories in tech, every morning. Free.

When things get slow, the instinct is to blame the database. Resist it. In nearly every incident we've investigated, the database was perfectly capable, the queries were the problem. A missing index here, an N+1 query there, a full-table scan hiding inside an ORM call.

Before you reach for read replicas, connection poolers, or a different database entirely, run EXPLAIN on your slow queries. The fix is usually an index that takes thirty seconds to add. We once cut a p99 latency spike by 90% with a single composite index. The database wasn't the bottleneck. Our understanding of it was.

Lesson 3: Boring technology wins

There's enormous pressure in our industry to adopt the new thing, the new database, the new framework, the new deployment paradigm. At scale, novelty is a liability. Every unfamiliar component is a 2 AM page you haven't learned how to debug yet.

Our stack is deliberately unfashionable: a relational database, a boring web framework, static files on a CDN. None of it will impress anyone at a conference. All of it has decades of operational knowledge behind it, which means when something breaks, the answer is a search query away instead of a research project.

Every unfamiliar component is a 2 AM page you haven't learned how to debug yet.

Lesson 4: Measure before you optimize

Server infrastructure
Five million requests a second takes serious hardware. (Photo: Unsplash)

We spent a week once optimizing an endpoint that turned out to account for 0.3% of total request time. The real bottleneck was a serialization step we'd never profiled. Profiling first would have saved the week.

This sounds obvious. It isn't, in practice. Optimization feels productive; measurement feels passive. But unmeasured optimization is just expensive guessing. Instrument everything, look at the flame graphs, and let the data tell you where the time goes. It's almost never where you think.

Lesson 5: Failure is a feature to design for

Everything fails eventually: disks, networks, DNS, your cloud provider's control plane. The question isn't whether you'll have an outage, but whether your system degrades gracefully or falls over completely.

Our approach: timeouts on everything, circuit breakers on external calls, and cached fallbacks wherever possible. When a downstream service dies, users should see slightly stale data, not an error page. Designing for failure feels pessimistic. It's actually the most optimistic thing you can do, because it's what lets you sleep.

The meta-lesson

Infrastructure at scale is mostly about restraint. Cache instead of compute. Index instead of scale. Boring instead of novel. Measure instead of guess. The systems that survive are rarely the cleverest: they're the ones with the fewest ways to fail.

Five million requests a month sounds like a lot until you're serving it. Then it just feels like Tuesday, as long as you built for the boring case, not the impressive one.

Lesson 6: Automate the toil, not the judgment

Every infrastructure has toil: certificate renewals, log rotation, dependency updates, backup verification. The temptation is to automate everything. The better rule is to automate everything boring and keep humans in the loop for everything consequential.

Deployments are the classic example. Fully automated deploys feel modern until a bad deploy propagates globally in ninety seconds with nobody watching. We deploy automatically to staging, require a human glance for production, and keep one-click rollbacks within arm's reach. It's slightly slower. It's enormously safer.

The heuristic: automate the things you'd do identically every time, and keep humans where judgment matters, assessing risk, interpreting novel failures, deciding when the runbook doesn't apply. Toil is what burns teams out; judgment is what makes them valuable. Don't automate away the wrong one.