Part 2: Two Load Balancers, Two Pools, and Why Adding Servers Was Not Enough
How we rebuilt the application layer of an exam platform: two load balancers, two autoscale pools, a Next.js frontend with two containers per node, a tag-based backend, a CPU metric that lags 5 to 8 minutes, and a rollout with 219 probes and 0 errors.
One hour after we switched DNS to the new frontend, the old server was still receiving about 36% of the traffic. We had lowered the TTL to 300 seconds. We had tested the new setup through a hosts-file entry. And still more than a third of the visitors walked straight past the new front door, because their resolvers remembered the old address.
That number is why the old server stayed alive as our way back. It also sums up this stage: every step looked simple on paper, and every step hid a measured surprise.
In Part 1 we found that our first bottlenecks were a rate limiter and a flood of Valkey connections, not a shortage of servers. Only after that did adding servers become the right tool, and even then it needed an architecture. This article covers it: two load balancers, two autoscale pools, a rebuilt frontend, a backend where any node can be replaced, and a deploy method that does not drop requests. The first version of that method did.
This is part 2 of the five-part case study "From One Server to Exam-Day Ready". NovaCommerce is a fictional name; the architecture, numbers and mistakes are real.
Why more servers was not the first answer
In Part 1 the bottleneck moved three times. First a rate limiter that counted students by IP address, so a whole school shared one bucket and got HTTP 429. Then Valkey, where a load test returned 2,854 application errors because every request opened a fresh TLS connection to a server that could not accept them fast enough. Only the third, plain backend CPU, is the kind that more servers solve: one 4 vCPU / 8 GB node served about 65 to 70 requests per second of the real API mix.
Adding nodes to the first two would have made them worse. Every new node would run the same limiter code and open its own connections to the same Valkey. We would have paid for more servers and seen the same errors, only faster. So this article is about the third bottleneck, done properly.
Two load balancers, not one
Our first sketch used a single load balancer for everything. One is cheaper than two, and there is less to look after. Status: Considered, then Rejected.
A DigitalOcean load balancer cannot route by host name or by URL path, and each one targets exactly one tag, which means one group of servers. Frontend and backend nodes are different groups, with different sizes, health checks and scaling rules. It is a reception desk that hands every visitor the same list of rooms.
DNS already split the website and the API into separate domains, so the split came for free: one load balancer per tier, each in front of its own pool. Status: Implemented. The two together cost $48 a month.
No warm pool, so we kept the cooks in the kitchen
Early on we wrote down an attractive requirement: a warm standby that can take traffic within 3 to 4 seconds. It does not exist here. DigitalOcean autoscale pools have no warm pool, and a new droplet needs minutes to be created, booted, checked and admitted by the load balancer. A taxi already idling at your door is a different product from a taxi you have to call.
So we did two boring things instead.
Spare capacity that is already serving. Both pools have a minimum of 2 nodes, and both sit inside the load balancer all the time. Losing one leaves the other already taking traffic.
Pre-scaling. About 45 minutes before a big exam we raise the pool minimum, so the extra nodes are booted, checked and serving before the first student clicks Start. An extra backend node costs about $0.08 per hour.
We also considered Kubernetes (DOKS) for seconds-level scaling. It would mean much more to run, so we left it for later. Status: Considered, not implemented.
The frontend: eight CPUs and one cashier
The old frontend was one 8 vCPU / 16 GB droplet running one Next.js container, with TLS by certbot on the droplet and no load balancer. If it died, the site was down. It also wasted money quietly: a Node.js process renders on roughly one core, so most of those eight CPUs did nothing. A supermarket with eight checkout counters and one cashier.
The new design runs several small, identical nodes instead of one large one. Figure 1 shows the whole application layer; we walk through it from the left.
Two tiers, each with its own load balancer and pool. The worker is the one fixed server outside both pools.
One frontend node
Each node runs nginx in front of two identical Next.js containers, one per vCPU. The settings that matter:
Container ports are bound to localhost only, so only nginx can reach a container.
nginx balances between the two containers with least_conn: each request goes to the container with the fewest active connections.
Each container has a 1.5 GB memory cap, restart: always and rotating logs.
nginx takes the real client IP from the load balancer, so per-student limits keep working instead of seeing one address for everybody.
A micro-cache serves static files, images and the home page without waking Next.js.
/lb-health passes only when Next.js answers, so a node with a dead app leaves the rotation even though nginx is alive.
HTTPS ends at the frontend load balancer, and the cloud firewall lets a frontend node accept port 80 only from that load balancer.
Switching DNS, with a way back
We built the new frontend next to the old server and tested it through a hosts-file entry on our own machines, so the real domain pointed at the new load balancer for us and nobody else. We compared 12 real pages with the old server, lowered the DNS TTL to 300 seconds and switched.
The old server stayed for rollback, because pointing DNS back would have been the whole undo. We needed it longer than expected: after one hour it still received about 36% of the traffic from cached DNS. A TTL is a request, not an order. Destroying the server early would have sent about a third of the visitors to an address that no longer answers.
The load test, and a cheaper node
Two frontend nodes served about 356 requests per second with a p95 of 0.05 seconds and 0 errors, at 55 to 60% CPU. The pool first used dedicated-CPU droplets. That test showed plenty of headroom, so we moved to shared AMD droplets with 2 vCPU / 4 GB, about $28 a month each. A measurement, not a guess, made the cheaper node safe.
Before
After
Servers
1 droplet, 8 vCPU / 16 GB, 1 Next.js container
Pool of 2 to 10 nodes, 2 vCPU / 4 GB shared AMD, 2 containers each
If one server dies
The site is down
The other node keeps serving
Measured with 2 nodes
Rendering limited to about one core
About 356 requests/s, p95 0.05 s, 0 errors, CPU 55 to 60%
The backend: from droplet IDs to a tag
Before, the backend load balancer pointed at two droplets by ID. A third server could never join by itself, because the load balancer only knew two names. We changed the target from IDs to a tag. It is the difference between a guest list and a staff badge: anyone wearing the badge gets in.
The change was one API update that kept every other setting of the load balancer, with 0 errors during the switch. From then on, any node the pool creates carries the tag, joins the load balancer on its own, and falls under the same tag-based firewall and database access rules. HTTPS runs end to end: the load balancer re-encrypts to each backend over the private network. Every node boots from one golden snapshot.
Frontend pool
Backend pool
Node
2 vCPU / 4 GB shared AMD, about $28 a month
4 vCPU / 8 GB, about $56 a month
Minimum / maximum
2 / 10
2 / 10
Scale-out rule
70% average CPU, cooldown 5 minutes
70% average CPU (later 55%), cooldown 5 minutes
Health check
/lb-health passes when Next.js answers
/lb-health answered by nginx alone
Stateless nodes, and the one server that is not
The web nodes were already stateless, and we checked again before trusting a pool with them. Sessions, cache and queues live in managed Valkey, logs go to stderr, and uploads, including students' written-answer PDFs, go to Spaces object storage. On 7 October we verified it: the backend nodes wrote 0 files to disk in 24 hours. That is why any node can be deleted at any time.
Each backend node runs only web (nginx) and app (php-fpm). Queues, the scheduler, the WebSocket server and OMR are switched off on web nodes, so extra servers never run a job twice. They all live on the single fixed worker outside both pools, a known single point of failure that Part 4 returns to.
Health checks and draining
On the backend, /lb-health is answered by nginx alone and never boots the framework, so a health check cannot be slowed down by PHP or the database. It also gives us a drain switch. To take a node out of service, we make /lb-health return 503. The load balancer stops sending new requests to it within about 30 seconds, and requests already running finish normally.
The autoscale test and the five-minute lie
On 5 October, between 21:17 and 21:33, we tested the autoscaling itself. k6 replayed the real exam-peak request mix, GET requests only so nothing was written: 300 to 450 requests per second, then about 750 requests per second for 6 minutes.
What we watched
Result
Backend pool
2 to 3 nodes at 21:29, when the pool average reached 75%
Frontend pool
2 to 3 nodes at 21:30, at 71%
Errors and timeouts
0 server errors, 0 timeouts; new nodes joined the load balancers by themselves
Overall p95
132 to 153 ms
API p95
320 to 705 ms, while the backend sat near 90% before the third node joined
HTTP 429
About 4%: all the load came from one test IP, so the per-IP limits did their job
Look at the API p95, 320 to 705 ms. That is the price of waiting. The backend sat near 90% CPU until the third node joined, because DigitalOcean's CPU metric lags real load by 5 to 8 minutes. In the earlier backend test, the first extra node appeared about 7 minutes after CPU hit 99%. The autoscaler is not broken. It is reading old news.
The lag works in the other direction too. When the load stopped, the metric still looked high, and the backend briefly scaled to 4 nodes. It is a shower tap: you turn it further because nothing has changed yet, and then the hot water arrives all at once.
The load is a step, the metric is a slow curve, and the autoscaler follows the curve. Schematic of the effect, not a recording.
The general argument is in Autoscaling for Traffic Spikes; here it has our own numbers behind it. For a scheduled exam we raise the minimum about 45 minutes ahead and lower it only afterwards. Autoscaling stays on as the safety net for the unplanned. Status of the test: Tested.
Deploying by replacing, never by editing
Pool nodes are disposable, so nobody edits one by hand: it would vanish at the next scale-in, and the next new node would boot from the snapshot, not from the edit. Our flow:
Change one node and test it.
Take a snapshot of it.
Point the pool template at the snapshot.
DigitalOcean creates the new nodes, then deletes the old ones.
We keep the last 3 good images for rollback and never roll a template during exam hours. We pass every setting on every pool update, so nothing silently resets. The account's droplet limit of 25 also had to be raised before both pools could reach their maximum during a rollout, when old and new nodes exist together.
The 503 blip, and the guarded rollout
The first backend rollout was a plain template change, and it produced a short burst of 503 errors. DigitalOcean deleted the old droplets while the load balancer was still routing requests to them. It is like closing the old dining room while the host is still seating guests there.
The fix was to stop relying on the provider's order of events and run the rollout as a guarded sequence. Figure 2 shows it.
The old nodes are drained before they are deleted, so the load balancer never sends a request to a node that is about to disappear.
Change only the image on the pool template.
Wait until every new node answers /lb-health.
Wait about 40 seconds for the load balancer to admit them.
Drain the old nodes first: their /lb-health returns 503.
DigitalOcean removes the old nodes after its cooldown, with nothing left to serve.
The later full rollout of both pools, on 7 October, ran with an uptime probe every 2 seconds on real pages. The result: 219 probes and 0 errors. Status: Tested, then fixed.
Swap and hardening on every node
Memory was the quiet risk on the frontend. Those nodes had no swap, and each container may use up to 1.5 GB of a 4 GB node. On 7 October we added a 2 GB swap file with swappiness 10, so the kernel prefers RAM, and baked it into the frontend image. Backend nodes already had 4 GB of swap. The same set of changes covered hardening on every node.
Area
What is in place
Access
Key-only SSH and Fail2Ban on all nodes
Network
Cloud firewalls by tag: frontend nodes accept port 80 only from their load balancer; backend nodes accept 80 and 443 only from theirs
Processes
Containers run their request-serving processes as non-root users
Secrets
Secret env files with mode 600
Database
The app user has only SELECT, INSERT, UPDATE and DELETE, with no DDL
Firewalls follow the tag, so a node the pool creates in the middle of the night gets its rules the moment it exists. Nobody has to remember anything. Status: Implemented on 7 October.
What we learned
Adding servers is not architecture. The limiter and Valkey needed a code change and a data-tier fix; more nodes would only have copied the problem.
No warm pool? Build the warmth yourself. Spare capacity already serving, plus a raised minimum about 45 minutes before a big exam.
Pre-scale for scheduled exams. The CPU metric lags by 5 to 8 minutes, so reactive autoscaling is the safety net, not the plan.
Make nodes disposable. Stateless web nodes, a tag instead of IDs, replacement by image instead of editing.
Drain before you delete. A guarded rollout is a procedure, not a hope: 219 probes, 0 errors.
Keep the old thing until its traffic has left. A TTL is a request; an hour later the old server still had 36%.
One more rule, learned on a real exam evening and told in a later part: scale-in does not drain by itself.
Where each piece stands
Piece
Status
Two load balancers, two autoscale pools, golden-image deploys, swap and hardening
Implemented
Autoscale test at about 750 requests/s; first rollout and its 503 blip
Tested, then fixed
One load balancer for both tiers
Rejected
Warm standby in seconds; Kubernetes (DOKS)
Considered; DOKS for later
Static assets from the CDN; Fluent Bit on frontend nodes, whose nginx logs vanish with the node
Planned
The application layer can now grow and be replaced without anyone noticing. The next question is whether a student's answer survives: the database. In Part 3 we move MySQL in a planned 26-minute window, and we meet the legacy admin panel that kept writing to the wrong database.