Part 1: What High Availability Really Meant for NovaCommerce
A fictional exam platform with real numbers: one frontend server, two fixed backends, one database node and two powered-off standbys. Part 1 defines high availability and shows how the bottleneck kept moving.
On 25 September, during exams, students began to see HTTP 429, "too many requests". The cause was not a shortage of servers. It was a rule that counted the wrong thing: a rate limiter was counting by IP address, and a school puts hundreds of students behind one address. To the limiter, a whole building looked like one very impatient person.
Eleven days later, at 02:28 on 6 October, a backend load test returned 2,854 application errors. Every one said the same thing: a connection to Valkey had timed out. A different layer, a different cause, a different fix. After that came plain CPU, which is the one problem that more servers really do solve.
That sequence is the real story of this project: the bottleneck kept moving. Adding servers is not architecture. It is one move among several, and the last one.
Between late September and 7 October 2026, we took a Laravel and Next.js platform on DigitalOcean from one frontend server and two fixed backends to a setup ready for a scheduled exam peak. This first part covers the starting point: the business problem, what we inherited, why it was not enough, what "high availability" meant for us, and what the first tests taught us.
Before we start: NovaCommerce is a fictional name, used for this case study. I left out anything that would identify the company, its servers or its students.
This is part 1 of the five-part case study "From One Server to Exam-Day Ready". NovaCommerce is a fictional name; the architecture, numbers and mistakes are real.
A platform whose traffic arrives on a schedule
NovaCommerce is an online platform that sells courses and runs timed online exams for students. People sign up, pay, take live exams, see their results and use AI study help.
Most web traffic is a hill: it rises in the morning and falls at night. Exam traffic is a cliff. Think of the doors of a concert hall. Nothing happens for hours, then the doors open and everyone arrives together. When a scheduled exam starts, thousands of students hit the platform in the same minute.
The good news is that the cliff has a timetable. We know when it comes. I describe that advantage in general terms in the guide to autoscaling for traffic spikes. This series is the other half: a real platform, with real numbers.
The business gave us four requirements, in plain language.
Exam spikes are scheduled and sharp. Thousands of students arrive in the same minute.
Submissions must never be lost. A slow page is annoying. A lost exam submission is a student's work gone.
Deploys must be invisible. The site has to stay up while we ship a release.
Cost must stay sane. Capacity nobody needs counts as a defect too.
Every technical choice in this series traces back to one of those four lines.
The starting point: what we inherited
Here is the setup as we found it before 5 October 2026, in one DigitalOcean region and one private network.
Two single points of failure (the frontend and the database), a fixed list of two backends, and two standby servers that were switched off but still billed.
The frontend: one server, eight CPUs, one cashier
The frontend was a single droplet with 8 vCPU and 16 GB of memory, running one Next.js container. The domain pointed straight at it, and TLS came from certbot on the droplet itself. There was no load balancer. If that server died, the site was down.
It also wasted money. One Node.js process renders on roughly one core, so most of the eight vCPUs sat idle: a shop with eight tills and one cashier.
The backend: two servers a load balancer could not grow
A DigitalOcean load balancer fronted two fixed backend droplets, each with 4 vCPU and 8 GB. The application is Laravel (PHP-FPM) behind nginx in Docker, as two containers per node: web for nginx and app for php-fpm. Queues (Horizon), the scheduler and Reverb, the WebSocket server for live updates, ran on a worker instead.
The load balancer targeted each backend by its own fixed ID, not by tag. That is the difference between a guest list with names on it and a rule that says "anyone wearing a staff badge". With names, a new server can never walk in on its own. Someone has to edit the list.
There was good news here, and it mattered more than anything else. The backend web nodes were already stateless. Sessions, cache and queues lived in managed Valkey, and logs went to stderr. A server you can delete without losing anything is a server you can multiply. Without this property the rest of the series could not have happened.
The database and the switched-off standbys
The database was one Managed MySQL "Advanced" node with 8 vCPU and 32 GB: a single node, with no standby.
Then there were two "standby" droplets, kept as a backup. They were powered off. A powered-off droplet still costs money, and bringing one up takes minutes. It is a spare tyre in a locked garage across town: it exists, you pay for it, and it does not help while you are stuck on the road. That is not high availability.
Two smaller things sat in the corner. The load balancer's TLS certificates were manual uploads with an expiry date, so renewal was a recurring task someone had to remember. And the account's droplet limit was 25, which would matter once two pools could grow and replace nodes (part 2).
Why this was not enough
Part of the setup
How it was built
What happens under stress or failure
Frontend
One 8 vCPU / 16 GB droplet, one Next.js container, no load balancer
Site down if it dies; most cores idle, because one Node process uses about one core
Backend
Two fixed 4 vCPU / 8 GB droplets behind a load balancer, targeted by ID
No automatic growth; a new server can never join by itself
Database
One Managed MySQL node, 8 vCPU / 32 GB, no standby
No second node ready to take over if the node fails
Standbys
Two droplets, powered off
Billed while off; minutes to power on
Look at what is missing from that table: a bug. Nothing was broken. The setup did what it was built to do. The trouble was a mismatch between a scheduled cliff and a setup that could not lose a server, could not add one by itself, and kept its spare capacity switched off. You cannot patch a mismatch. You have to change the shape of the system, and that is what architecture means.
What "high availability" really meant here
"High availability" is one of those phrases everyone nods at and nobody defines. Before we changed anything, we wrote down what it meant for this platform. It came down to four lines.
No single server whose loss takes the site down. The database survives the loss of a node. Deploys and scale-in are invisible to students. Capacity is ready before an exam starts, not five minutes after.
I like a definition you can test with a question. Can we pull the plug on any one server and keep serving? Can a database node fail without the exam stopping? Can we ship a release without a student noticing? Is the capacity already there when the first student clicks "start"?
Two business rules sit on top of it: submissions are never lost, and cost stays sane. The figure maps each requirement to the layer that must deliver it and to the part of the series where that layer appears.
Every requirement has an owner: one layer that must deliver it, and one part of this series that shows how.
The last line of the definition, "not five minutes after", is the one that hurt. In our tests, autoscaling added the first extra node about seven minutes after CPU reached 99%. Seven minutes is a long time when an exam starts in one. Capacity has to be there before the cliff, not after it.
The first tests: the bottleneck keeps moving
With the definition written, we did what the method demands: measure first. Between late September and 6 October we watched production and ran load tests, and found three problems in a row. They look unrelated, and that is the point.
Three bottlenecks in a row, each with a different kind of fix. Only the third is solved by adding servers.
1. The limiter: a code problem (25 September)
Measured on production: students were getting HTTP 429 during exams. The global API rate limiter was keyed on IP, because the default auth guard was empty for token-based students, so the limiter could not tell who a student was. A school or a mobile carrier puts hundreds of students behind one NAT address, and the whole building shared one bucket.
Status: Implemented. We keyed the API limiter per student or instructor account, and by IP only for guests. More servers would not have helped; the rule was saying no, not the hardware. A close cousin of this bug came back on 7 October in the route-level throttles, and that story is in part 4.
2. Valkey connections: a data-tier sizing problem (6 October)
At 02:28 on 6 October, a backend load test returned 2,854 application 500 errors. Every one was RedisException: Operation timed out while connecting, with a connect timeout of 5 seconds. Valkey (4 GB, primary plus standby) could not accept connections fast enough under load.
How the app used it added to the pressure. Each request opened a fresh TLS connection to Valkey (about 5.4 ms of CPU per request) and another to MySQL (about 3.3 ms). Picture a shop where every customer is buzzed in and checked at the door, every single time. Under a crowd, the door becomes the queue.
Status: Implemented, then Tested. We resized Valkey in place to 8 GB with two nodes (primary and standby), data kept and host unchanged. Then we retested with a gradual ramp from 25 to 500 requests per second, starting on two backend nodes while the pool scaled out under the load: 0 application errors and 0 backend 5xx. There were 143 errors at the load balancer, but only while the nodes were CPU-saturated. That is the expected "out of capacity" signal, not a bug. The data tier gets part 3, and Valkey returns in part 4. For the general theory, see the data tier under load.
3. Backend CPU: a capacity problem
With the first two fixed, what remained was plain arithmetic. One 4 vCPU / 8 GB backend node serves about 65 to 70 requests per second of the real API mix at full CPU, at about 59 ms of CPU per request. Every later capacity decision is built on that number.
This is the only one of the three bottlenecks that more servers actually solve, and even here the timing is the trap: the first extra node arrived about seven minutes after CPU hit 99%. How we made capacity arrive early is the subject of part 2.
So the order was a limiter (code), then Valkey connections (data-tier sizing), then backend CPU (capacity). If we had started by adding servers, the limiter would still have said 429 and Valkey would still have timed out.
What the old setup cost
Before talking about the new shape, it helps to know what the old one cost. These are monthly DigitalOcean list prices for the old web tier.
Item
Cost
What it bought
Old web tier: frontend, two backends, worker, one load balancer, two standbys
about $416 a month
A running site with several single points of failure
Two powered-off standby droplets
$112 a month
No availability: billed while off, minutes to power on
One extra backend node, for comparison
about $0.08 an hour
Capacity that actually serves requests
The $112, a little over a quarter of the web tier, is the number that stayed with me. It bought two servers that could not serve a single request. Pre-scaling an extra backend node costs about $0.08 an hour, so raising capacity for an evening costs cents. Paying all month for capacity that is switched off is the more expensive way to feel safe.
One honest note: the $416 covers only the web tier, not the databases, Valkey or storage, so it cannot be compared with a full bill. In part 5 I compare like with like.
The method, and the plan for the series
The pattern we repeated all through the project is the one you just saw: measure, find the bottleneck, understand the cause, change the architecture, test again, and measure again. Adding servers is not architecture. Architecture is choosing which layer changes, and why.
Here is how the five parts fit together.
Part
Layer
What you will see
1 (this part)
Starting point
The business problem, the old setup, the definition of high availability, the first bottlenecks and the old cost
The real exam evening of 7 October, load and soak testing, and the final architecture
We also considered ideas we did not build, and I label every one. One load balancer for both tiers was Considered, then Rejected. A "warm standby ready in seconds" was Considered, but the autoscale pools have no warm pool, so we used spare capacity already serving plus pre-scaling. Kubernetes was Considered for later and is not implemented.
Not everything is finished, and I will say so. The single worker server is a known single point of failure; fixing it is Planned, not done (part 4). A formal multi-hour soak test is also Planned (part 5).
What we wanted to achieve
The goal as a checklist, with the part where each line is built or tested.
Frontend and backend tiers behind load balancers, with at least two nodes each (part 2).
New servers that join and leave on their own, with no hand-edited nodes (part 2).
Capacity raised before an exam, with autoscaling as the safety net (parts 2 and 5).
A database with a standby that costs less per month than one big node (part 3).
Web nodes with no local state, so submissions survive any one node dying (part 4).
Deploys and scale-in that students cannot see (parts 2 and 4).
A capacity model built from real exam data (part 5).
A bill we can explain line by line (part 5).
What we learned
Write "high availability" down in terms of failures and events before buying anything. A definition you can test with a question is better than a slogan.
A powered-off standby is a cost, not availability. Ours was $112 a month for no real protection.
The bottleneck moves. Fixing the first one reveals the next, so measure again after every change.
Different bottlenecks need different fixes: code (the limiter), data-tier sizing (Valkey) and capacity (backend CPU). Only the last is solved by more servers.
Stateless nodes are the price of admission. Without them, the later steps would not have been safe.
Next, in part 2, we build the application layer: two load balancers, two autoscale pools, pre-scaling, and a rollout method learned from a short 503 blip.