Autoscaling for Traffic Spikes: What Actually Scales, and What Shouldn't
A scheduled spike peaks in seconds, while a new server takes minutes. Learn what scales under a burst, why pre-scaling beats reactive rules, how to run SSR and PHP-FPM tiers, and how to load test it.
Author
Anichur Rahaman
3 weeks ago11 min read3 views
Twenty requests per second, flat: that is the traffic of an online store at 11:59 p.m. on the night a flash sale opens, as its platform lead watches the dashboard. The autoscaling rule is armed: add a server when CPU stays above 70% for five minutes. (This is an illustrative scene, not a real shop.)
At midnight the sale opens and about a thousand shoppers act within the same minute. Load jumps to 440 requests per second. At 12:01 the dashboard is still green, because the five-minute average has barely moved. At 12:03 checkout is timing out. The new server joins at 12:06, after the shoppers have left for a competitor.
Autoscaling reacts to load, and a scheduled spike arrives before the reaction does. This article covers what really scales under a sharp burst, what should not be asked to scale, and how to find out before the event instead of during it.
The numbers below come from my own tests on one Laravel and Next.js setup. They show the method, not a universal benchmark, so run your own.
This is part 1 of the five-part series "Engineering for High Volume". It covers the request path, from the edge to the application. Queues, the data tier, observability and safe deploys follow in parts 2 to 5.
A spike is not a busy day
Two shapes of traffic need two different answers. A gradual rise gives any control loop time to notice and respond. A step does not. The 20-to-440 jump in the opening scene is a step, and a rule that averages CPU over five minutes cannot see it in time.
The good news is that a scheduled spike is predictable. You know the start time, you usually know roughly how many people are coming, and you can often replay the request mix from the last event. That predictability is your biggest tool, and most of this article is about using it.
The request path, edge to app
Before deciding what to scale, draw the path a request takes. In the setups I run, it looks like this: CDN and WAF at the edge, a managed load balancer, two backend web nodes, one or more server-side rendering (SSR) frontend nodes, a worker node for background jobs, and a data tier behind them.
The request path from edge to data. Each layer scales in a different way and at a different speed.
Each box has its own failure mode and its own scaling lever. Treating them as one "server" is how a team ends up adding RAM to the one tier that was never the problem.
The edge and the load balancer
Put a CDN in front first. Static files, images and any page that is the same for every guest should never reach your servers during a spike. The WAF and bot rules at the same layer also keep scrapers from eating capacity you reserved for buyers.
Behind the CDN, use a managed load balancer with health checks and connection draining. Health checks pull a sick node out automatically. Draining lets a node finish its in-flight requests before it leaves, which is what makes deploys and scale-in invisible to users.
Two details are worth checking. First, the balancing algorithm. Many cloud load balancers default to round robin, which assumes every request costs the same. Checkout and search do not. A "least outstanding requests" or "least connections" mode sends the next request to the node with the least work in progress. On AWS, for instance, round robin is the default for Application Load Balancers and least outstanding requests is a per-target-group setting. Second, the load balancer is itself a service that scales. AWS documents that an Application Load Balancer can roughly double its capacity in five minutes, and offers reserved capacity for events that more than double traffic in under five minutes (see the capacity reservation guide). Other providers have similar limits, so ask yours.
In my load tests the load balancer was never the bottleneck: CPU stayed at or below 8% with zero errors at the highest rate. That is typical. Look further down the path first.
Stateless web nodes: the price of admission
You can only put two web nodes behind a load balancer if either node can serve any request. That property is called being stateless, and it is cheap to get right and painful to retrofit. Use this checklist before you add a second node.
Sessions in Redis or Valkey, not in files on the node's disk.
Application cache in Redis or Valkey, shared by every node.
Uploads and generated files in object storage (S3-compatible), never on local disk.
One scheduler. Cron-style jobs must run on exactly one node, or every task runs twice.
Identical builds. Every node runs the same image, with configuration from the environment.
Aligned upload limits at every layer. Each proxy has its own body limit (nginx client_max_body_size, PHP post_max_size and upload_max_filesize). In one setup a 1 MB default on the host proxy silently rejected uploads until all layers matched.
A self-hosted platform such as StoreConsole keeps sessions and cache in Redis for this reason, so nodes stay interchangeable.
The SSR tier: one Node process is one core
This tier hides the sharpest trap. A Next.js server renders pages in a Node.js process, and Node runs your JavaScript on a single thread. One process can use about one core for rendering, however large the machine is.
I replayed the real request mix of a previous peak with k6 against a single frontend process. It saturated at roughly 220 to 250 requests per second. Past that point the p95 latency jumped from about 260 ms to about 3 seconds, and p99 reached 6.4 seconds. During all of this the host still had more than two idle cores. The peak minute of the real event needed about 440 requests per second, so one process would have failed at the exact moment it mattered.
The fix was plain: run several identical frontend containers from the same image (three, in my case) behind nginx with least_conn, which passes each request to the server with the fewest active connections. Then I re-ran the test at the target rate. One more detail applies to Next.js specifically: by default each instance keeps its own local cache, so with several instances you should point the framework at a shared cache store, as its documentation describes.
The arithmetic
Capacity planning here is simple multiplication. Take the measured limit per process, multiply by the number of processes, and compare it with the peak you expect, with headroom.
Item
Value (one setup)
Note
Peak minute needed
~440 req/s
From replaying the last real event
One SSR process, measured
220-250 req/s
p95 rose from ~260 ms to ~3 s beyond this
Three SSR processes
~660-750 req/s
About 1.5 times the peak as headroom
Load balancer CPU
8% or less
Not the bottleneck
PHP-FPM: size the pool on purpose
The backend web tier has the opposite problem. PHP-FPM runs a pool of worker processes, and each handles one request at a time. If the pool is too small, requests wait in a queue. If it is too large, the node runs out of memory and starts swapping or killing processes.
A rule of thumb that works: the number of busy workers you need is roughly requests per second multiplied by the average time per request. An illustrative example: 200 requests per second at 150 ms each needs about 30 busy workers, so a pool of 60 leaves room for slow requests. Then check memory: if one worker uses about 100 MB, 60 workers need about 6 GB, so the node needs more than that.
The settings I used in one setup were pm.max_children=60, start_servers=30, min_spare_servers=20, max_spare_servers=30 and max_requests=1000, with keepalive between nginx and FPM (keepalive 16 and fastcgi_keep_conn on). Starting with 30 servers matters: the pool is already warm when the burst arrives instead of forking under pressure.
Two more habits pay off. Turn on OPcache in production with validate_timestamps=0 and restart the workers on deploy. And consider fastcgi_next_upstream error timeout http_500 http_503 with fastcgi_next_upstream_tries 2, which turned transient FPM hiccups into silent retries. This is only safe for idempotent requests. Retrying a payment POST is not a fix, it is a bug.
Why reactive autoscaling arrives late
Now the core idea. Reactive autoscaling watches a metric, waits for a threshold, launches a machine, waits for it to boot, join the load balancer and warm up its caches. In practice that takes one to three minutes at best. Defaults can make it slower: AWS EC2 Auto Scaling, for example, applies a 300-second default cooldown, and basic monitoring publishes instance metrics only every five minutes unless you pay for one-minute detailed monitoring.
A scheduled spike peaks in seconds. So the new machine arrives after the people have already left, or worse, after they have already seen errors. A virtual machine also cannot grow its RAM without a reboot, so "scale up in place" does not help either.
Reactive scaling follows the load. Pre-scaling stands ready before it. Illustrative timeline.
Pre-scale for events you know about
For a scheduled event, raise the minimum node count before it starts, then lower it afterwards. It is cheaper than it sounds: an extra node for a few hours costs very little compared with a failed sale. It is also more reliable than any rule, because it does not depend on a metric crossing a line in time.
Keep reactive rules as a safety net for the unexpected, not as the plan. And scale what is fast to scale. Processes inside a container start in seconds, machines take minutes, which is why I scale queue workers by process rather than by machine (part 2 of this series).
Approach
Reaction time
Best for
Pre-scale (raise the minimum)
Ready before the event
Scheduled spikes
Reactive VM autoscaling
1-3 minutes or more
Gradual growth, unplanned traffic
Process-level scaling
Seconds
Queue workers inside a container
Which lever to pull, and when: three questions decide between pre-scaling, reactive scaling, process-level scaling and edge protection.
Load test it like the real event
None of the numbers above come from guessing. They come from tests that look like the event. A test that hammers only the home page proves very little. Use this sequence.
Capture the real request mix from the last event's access logs: which URLs, in what proportions, with what think time.
Replay it with k6 (or a similar tool) at the rate you expect, then at 1.5 times that rate.
Set abort thresholds in the script, such as p95 above 3 seconds, or any 5xx or timeout, so a failing test stops instead of hurting a shared system.
Add a guard script that watches CPU on the nodes and stops the run if limits are hit.
Compare your metrics with the cloud provider's. In my tests they agreed within a few percent, which told me the dashboards were trustworthy.
Change one thing at a time, and re-run at the target rate after every fix.
Do not upgrade hardware before a test proves where the limit is. In my case, the machines were fine. The single-threaded process was the limit, and the fix cost nothing but configuration.
The same midnight, prepared
Run the opening scene again with the preparation done. At 11:00 p.m. the minimum goes from two web nodes to three, and the SSR tier runs three processes behind least_conn. The cache is warm and the FPM pool already holds 30 workers.
At midnight, 440 requests per second arrive. The measured ceiling is about 660, so p95 stays near its normal level and the autoscaling rule never fires, because nothing needs it. At 12:30 a.m. the minimum drops back. The event cost a few hours of one extra node.
A scheduled spike is a step, not a ramp. Reactive autoscaling, which takes minutes, reacts after it has passed.
For known events, pre-scale: raise the minimum node count before the start and lower it afterwards.
Keep web nodes stateless: sessions and cache in Redis, files in object storage, one scheduler.
One Node.js SSR process uses about one core. Run several identical processes behind least_conn and test at the target rate.
Size the PHP-FPM pool from requests per second, request time and memory per worker, and pre-warm it.
Load test with the real request mix, abort thresholds and a guard script, and fix the measured bottleneck before buying hardware.
Anichur Rahaman is a software architect and the creator of StoreConsole. He designs commerce and ERP systems for growing businesses, with a focus on event-driven architecture, data integrity and self-hosted operations.