Preparing Your E-commerce Site for Black Friday Traffic: A DevOps Checklist

Preparing Your E-commerce Site for Black Friday Traffic: A DevOps Checklist

Every November the pattern repeats. Traffic builds through the week, spikes hard on the Friday, then keeps testing you well into Monday. When a shop slows down at the wrong moment, customers don't wait — they go elsewhere, and plenty never come back. The encouraging part is that most Black Friday outages are entirely predictable, and predictable problems can be fixed in advance. Here's a checklist covering the three areas that cause the most pain: load testing, caching and scaling.

Plan capacity from a number, not a feeling

Start with last year's peak hour: orders per hour, sessions per hour, and requests per second at the edge. Multiply by the growth you're expecting, then add real headroom — sizing to your forecast exactly means any surprise becomes an incident.

Break that figure down by route. Product and category pages might account for most requests, but basket and checkout endpoints cost the most: they write to the database, hold sessions open and call payment providers. Write the numbers down and agree on them. "We'll see how it goes" is not a capacity plan.

Load test with realistic journeys

A test that hammers the homepage with thousands of concurrent connections tells you almost nothing about how the shop behaves when people are actually buying. Build scripts that follow real journeys: land on a category page from a paid ad, filter, open three products, add to basket, apply a discount code, check out as a guest. Include think time between steps, a share of logged-in users with existing baskets, and the mobile traffic mix, which is usually the majority.

Then test the failure paths. What happens when the payment provider takes six seconds to respond? Does the request queue back up, or does it fail cleanly? A slow dependency is far more dangerous than a dead one.

Run the test against an environment that mirrors production — same instance sizes, same cache configuration, same CDN settings — and ramp up gradually. Watch for the point where latency starts curving upwards; that's your real ceiling, and it arrives well before errors appear. One caution: aggressive tests can trip fraud checks or a provider's rate limits, so warn suppliers and use sandboxes where you can.

Caching is the cheapest performance you will ever buy

Most sale traffic is anonymous and reads the same catalogue, which makes it ideal to serve from cache. Work through the layers:

  • CDN edge: fingerprinted filenames, long cache lifetimes, and a purge process you have actually tested.
  • Full-page cache: category and product pages for guests. Exclude anything personalised — baskets, wishlists, stock messages.
  • Micro-caching: a lifetime of a few seconds at the edge absorbs enormous spikes on pages that change often.
  • Object cache: Redis or similar for sessions, query results and fragments. Check the hit ratio, not just that the service is running.
  • Opcode cache: enabled, correctly sized, with the hot files preloaded.

Two things go wrong repeatedly. The first is a cache stampede: a popular key expires and every request stampedes the origin at once. Use locks or request coalescing so only one request rebuilds the value. The second is stale stock. Decide in advance how much staleness customers will tolerate on a product page, and make sure inventory changes purge the right keys.

Scale out before you scale up

Assume any individual instance can disappear without warning, and design so that it doesn't matter. The application tier should be stateless: sessions in Redis or a database, uploads in object storage, no scheduled tasks pinned to one machine.

Autoscaling needs the right signal. If the app is I/O bound — waiting on the database, the cache, a payment API — CPU will stay low while response times climb. Scale on request rate, queue depth or latency instead, and pre-scale ahead of the sale rather than reacting to it.

The database is usually the first thing to break. Put connection pooling in front of it, send catalogue reads to replicas, and keep basket and checkout on the primary — remembering that replica lag is visible to anyone who adds an item and immediately reloads. Move anything unnecessary for the response into a queue: confirmation emails, ERP sync, analytics events.

Finally, check your limits: cloud quotas, database connections, CDN plan, email throughput. Quota increases take time to approve, and nobody wants to be filing a support ticket on Friday afternoon.

Protect the checkout path

Everything else can break. Search can crawl, recommendations can vanish, reviews can disappear. The route from basket to payment must stay up.

Rate-limit anything expensive — search, login, discount code validation — and add bot management, because a large share of sale traffic is scripts hitting inventory endpoints. If demand exceeds what you can serve, a virtual waiting room is a legitimate tool: an honest queue beats a timeout every time.

Decide now which features to switch off under load: recommendations, recently viewed, reviews, loyalty balances. Wire them to feature flags so the decision takes seconds rather than a deployment. And test how the shop behaves when the payment provider is slow — a clear "confirming your payment" state is much better than a spinner that never ends.

Watch the right things, and write it down

Latency percentiles by route (p95 and p99 matter more than averages), error rates, origin request rate, cache hit ratio, slow database queries, queue depth. Alert on the thresholds that came out of your load test, and make sure alerts reach someone who can act on them at 8am on the day.

Then write a one-page runbook: who is on call, who talks to the business, how to roll back, how to disable a feature, who contacts the payment provider and the hosting company. Freeze deploys from Thursday. A last-minute change is a poor trade against a quiet, well-understood codebase.

A workable run-up

  1. Week one: build the capacity model, write load test scripts, notify third parties.
  2. Week two: run the full test, fix the top three bottlenecks, review every caching layer.
  3. Days before: pre-scale, raise quotas, verify dashboards and alerts, circulate the runbook, start the deploy freeze.

When the peak arrives, resist the urge to make changes. Watch the graphs, keep the team small and calm, and let the quiet work you did earlier do its job.

Photo: Mediamodifier / Pixabay