One self-hosted console to run your entire business — commerce, ERP, HRM, CRM & manufacturing

Part 5: Measure, Change, Measure Again: Load Tests, a Real Exam Night and the Final Architecture

Seven measurements, one real exam evening and one scale-in mistake: how load tests, a capacity model and an honest soak plan shaped the final NovaCommerce architecture, its cost, and a Solutions Architect's checklist. Part 5 closes the series.

Author

Anichur Rahaman

8 hours ago18 min read1 views
Part 5: Measure, Change, Measure Again: Load Tests, a Real Exam Night and the Final Architecture

On the evening of 7 October, 661 students started an exam on NovaCommerce, at a peak of 86 starts a minute, and 439 more sat a second exam after it. The backend peaked at about 69 requests per second and reached 47% CPU on a two-minute average. The database used less than a quarter of its CPU. No background job failed.

Then a pool minimum was lowered in the middle of the exam. A serving node vanished without draining, and about 360 requests failed in two minutes.

Both halves of that evening were test results: the calm half came from a short run of tests in the days before, the ugly half from something no test had covered yet. This part walks the journey in order (baseline, load test, bottleneck, optimization, retest, scaling, soak, final result), then shows the final architecture, its cost, and the checklist I now use in every review.

This is part 5 of the five-part case study "From One Server to Exam-Day Ready". NovaCommerce is a fictional name; the architecture, numbers and mistakes are real.

The loop we kept repeating

Every change in this series came out of one loop: measure, find the bottleneck, understand the cause, change the design, test again, measure again. A change made before the cause is understood is a guess, and a guess that happens to be a server costs money every month. Think of a garden hose with a kink: if the water is weak, a bigger tap does nothing, so you walk the hose until you find the kink. Ours moved three times, and each time it was a different kind of problem.

Timeline of seven measurements from the 25 September baseline to the 7 October exams, plus the planned load test and soak, with what each one found and what it changed
Seven measurements, each with a number behind it, and one dashed card for the test that is still planned.

Baseline and load test: the bottleneck moved three times

Baseline: a status code, not a CPU graph

The first thing we measured on production was not CPU. On 25 September, students were getting HTTP 429 (too many requests) during exams. The API rate limiter keyed on the IP address, because the default auth guard was empty for token-based students. A school or a mobile carrier puts hundreds of students behind one NAT IP, so a whole building shared one bucket. A 429 is the application saying no, not a server out of breath, so no extra node could have helped. Implemented: the limiter now keys on the student or instructor account, and on the IP only for guests.

Load test: 2,854 errors, one sentence

At 02:28 on 6 October, a backend load test returned 2,854 application 500s, every one RedisException: Operation timed out while connecting (5-second connect timeout). The cause: Valkey (4 GB, primary and standby) could not accept connections fast enough, and each request opened a fresh TLS connection to it, about 5.4 ms of CPU per request against about 3.3 ms for MySQL. More web nodes would only have opened more connections to the same Valkey.

Implemented: Valkey resized in place to 8 GB on two nodes (primary and standby), data kept, host unchanged. It costs more money, and it removed a real failure.

Retest: a gradual climb from 25 to 500 requests per second, starting on two backend nodes while the pool scaled out, gave 0 application errors and 0 backend 5xx. There were 143 errors at the load balancer, but only while the nodes were saturated on CPU. That is the expected "out of capacity" signal, not a bug, and it showed us where each node runs out.

How much one node can do

The backend tests also gave us the number every later decision leans on. At full CPU, one 4 vCPU / 8 GB backend node served about 65 to 70 requests per second of the real API mix, about 59 ms of CPU per request. The figures agree: 4 vCPU is 4,000 ms of CPU per second, and 4,000 divided by 59 is about 68. Autoscale added the first extra node about 7 minutes after CPU hit 99%.

That makes three bottlenecks in order: a rule (the limiter), the data tier (Valkey connections), and only then plain backend CPU. Each needed a different fix, and only the last is solved by more servers.

Scaling tests: do the new servers arrive, and when?

Frontend: two nodes, 356 requests per second

With the frontend rebuilt as a pool (nginx in front of two identical Next.js containers per node), two nodes served about 356 requests per second, p95 0.05 s, 0 errors, at 55 to 60% CPU, and we compared 12 real pages against the old server. The headroom paid for itself: the pool moved from dedicated-CPU droplets to shared AMD 2 vCPU / 4 GB droplets. Implemented.

The 16-minute autoscale test

On 5 October, from 21:17 to 21:33, k6 replayed the real exam-peak request mix (GET only, nothing written): 300 to 450 requests per second, then about 750 for six minutes.

  • The backend scaled from 2 to 3 nodes at 21:29, when the pool average reached 75%. The frontend scaled from 2 to 3 at 21:30, at 71%. New nodes joined the load balancer by themselves.
  • 0 server errors, 0 timeouts, overall p95 of 132 to 153 ms. API p95 was 320 to 705 ms while the backend sat near 90% waiting for the third node.
  • About 4% of requests got 429, because all the load came from one test IP: per-IP limits doing their job.

The lesson was in the timing. DigitalOcean's CPU metric lags real load by 5 to 8 minutes, so scaling starts late and then overshoots: the backend briefly scaled to 4 after the load had stopped. The first extra backend node arrived twelve minutes after the test began, and an exam does not wait twelve minutes. So for scheduled exams we pre-scale, raising the pool minimum about 45 minutes before a big one, and keep reactive autoscaling as the safety net. The generic version of this argument is in Autoscaling for Traffic Spikes.

The rollout probe

Tested, then fixed: our first backend template rollout produced a short 503 blip, because DigitalOcean deleted the old droplets while the load balancer was still routing to them. The guarded rollout changes only the image, waits until the new nodes answer /lb-health, waits about 40 seconds for the balancer to admit them, and drains the old nodes first. The full rollout of both pools on 7 October ran under an uptime probe every 2 seconds on real pages: 219 probes, 0 errors.

Where scaling helped, and where it did not

This is the table I wish I had seen before starting. For every problem the tests or production exposed, it asks one question: would more servers have fixed it?

Problem foundMore servers?What fixed itStatus
Backend CPU saturated at 65 to 70 requests/s per nodeYesAutoscale pool (2 to 10 nodes) plus pre-scalingImplemented
Whole schools getting 429 (25 Sep)NoLimiter keyed per student account, IP only for guestsImplemented
Route throttles still counted by IP: 74% of calls to one dashboard endpoint returned 429 in an exam (7 Oct)NoThrottles counted per student per route; 0 such 429s afterwardsImplemented
2,854 Valkey connection timeouts (6 Oct)NoValkey resized in place to 8 GB x 2Implemented
Merit cards (a 0.34-second job) waiting 10 to 115 minutes behind AI jobs of about 10 seconds each (6 Oct)NoSeparate queue lane for the AI jobs; merit cards ready in about 1 minuteImplemented
Backend sees one frontend IP for every studentNoPass the real student IP in a signed headerPlanned

Of six problems, one was a capacity problem. The other five were two rules, a connection limit, a queue layout and a missing header. The data-tier side of this pattern is covered in The Data Tier Under Load.

Soak testing: what we ran and what we did not

A load test asks how much a system can take. A soak test asks how long it can take it: an engine at full revs for a minute proves little about three hours, because the slow problems show up late (memory that creeps up, connections that leak, disks that fill, queues that drift). "We soak tested it" is easy to say and hard to defend, so I will be exact: we have not run a formal soak test yet.

EvidenceWhat it coversStatus
Short sustained runs: the 25 to 500 requests/s climb, the 16-minute autoscale test at up to about 750 requests/s, a 2-second uptime probe through a rolloutMinutes of sustained load: 0 application errors in the climb, 0 server errors in the autoscale testTested
The 7 Oct exam evening: about two hours, two exams back to backSteady CPU, 0 failed jobs, one scale-in mistakeObserved in production
A multi-hour soak together with the next large load test, in an agreed maintenance windowMemory growth, connection leaks, queue driftPlanned

The exam evening is the closest thing we have, but it is an observation, not a controlled test. One related check does hold: backend nodes wrote 0 files to disk in 24 hours, so disks are not a soak risk there. Frontend nodes still keep about 440 MB of nginx logs a day locally, which is why Fluent Bit on the frontend is on the Planned list.

The real exam: 7 October as the production proof

Nothing replaces real students. Two exams ran back to back, and our once-a-minute digest watched them: students started and submitted, requests and 5xx per node, pool CPU, worker load and failed jobs.

MeasureResult
Exam A (20 minutes)661 started, 638 submitted (96.5%); starts peaked at 86 per minute
Exam B (25 minutes)439 started, 425 submitted (96.8%)
FrontendUp to 1,741 unique visitors in exam A, peak 403 in one minute; about 732,000 requests from 6 to 8 pm; CPU up to 24%
BackendPeak about 69 requests/s (per-minute logs; 63 on the provider's two-minute average); CPU up to 47% on a two-minute average; one node at 86% for a minute
MySQLCPU max 24.5% (average 13.4%), at most 5 running queries, 0 lock waits
Valkey12.5% memory
WorkerCPU 46%, 0 failed jobs

Three things stand out.

  1. The bottleneck is backend CPU per node, and the database had about 6 times headroom. The model below puts the MySQL limit near 450 requests/s, against 69 on the night.
  2. The test numbers predicted the evening. 63 requests per second at 59 ms of CPU each is about 3.7 CPU-seconds per second. Two 4 vCPU nodes have 8 vCPU, so the average should be about 46%. The dashboard showed 47%.
  3. The average hides the node that hurts. Traffic split 65/35 between the two backend nodes, because long-lived keep-alive connections from the frontend proxies pin traffic. The dashboard said 47% while one node touched 86% for a minute.

The scale-in mistake

In the middle of an exam, someone lowered the backend pool minimum. DigitalOcean removed a serving node without draining it, and about 360 requests failed in two minutes. The earlier tests had taught us that scale-out is slow and late. The exam taught the other half: scale-in is quick and does not say goodbye. Implemented as a runbook rule: raise the minimum before an exam, lower it only after, and never trust autoscale scale-in to drain a node.

The capacity model: what it says, and where it stops

After the exam I turned the real data into a small model, so the next exam can be planned with arithmetic instead of fear. It has four measured inputs: about 15 requests per student when an exam opens, then about 1.5 a minute (0.025 a second); about 30 requests/s of baseline traffic; 50 requests/s as the safe load for one backend node (75% CPU); and a 30% margin for uneven splitting.

In plain words: take the nodes you will have, multiply by 50, divide by 1.3 for the margin, and subtract the 30 of baseline traffic. What remains is what students may use at the moment the last one joins. Each student costs their 15 opening requests spread over the join window, plus 0.025 for being present.

Take 4 nodes. 4 x 50 / 1.3 is about 154, and minus 30 leaves about 124. If students join over 10 minutes, each costs 15 / 600 + 0.025 = 0.05 requests/s, so 124 / 0.05 gives roughly 2,500 students; the table rounds down to 2,400. If they all join within 2 minutes, each costs 15 / 120 + 0.025 = 0.15, which gives roughly 800.

Backend nodesJoining over ~10 minutesAll joining within ~2 minutesSafe backend load (derived)
2about 900about 300about 77 requests/s
4about 2,400about 800about 154 requests/s
6about 4,000about 1,300about 231 requests/s
8 to 105,500 or more1,800 to 2,300about 308 to 385 requests/s
Bar chart of students supported by 2, 4, 6 and 8 to 10 backend nodes for a ten-minute join and a two-minute join, next to a line chart of safe backend requests per second against the MySQL limit near 450
Backend nodes are the limit all the way to 10 nodes; the MySQL line sits above even the 10-node ceiling.

The real evening fits the first row: a peak of 86 starts a minute against the roughly 90 a minute that row allows. Where the model stops:

  • It comes from one multiple-choice exam. Written exams with PDF uploads load the worker much more and must be measured separately.
  • Rows above 3 nodes are extrapolated; the most we have seen in a test is 3 nodes, briefly 4. Planned: a full load test in a maintenance window.
  • Nodes only count once they exist. With a metric that lags 5 to 8 minutes and a first extra node about 7 minutes after CPU hit 99%, the table is a pre-scaling guide, not a promise of reactive scaling.
  • It covers the backend request path only. The worker, Valkey and the AI provider's rate limit have their own ceilings.

The final architecture, layer by layer

Solid boxes run today; the dashed strip is planned.

Final production architecture: users, frontend load balancer and pool, backend load balancer and pool, MySQL primary and standby, Valkey, PostgreSQL with pgvector, a worker, Spaces with CDN, OpenSearch fed by Fluent Bit, the ops console, and a dashed strip of planned items
Every box exists because a test or a real exam asked for it; the dashed strip is what we have not done yet.
  1. Edge. Two load balancers, one per tier. Rejected: one balancer for both, because a DigitalOcean balancer cannot route by host or path.
  2. Frontend pool. 2 to 10 nodes, nginx in front of two Next.js containers (one per vCPU), scaling at 70% CPU.
  3. Backend pool. 2 to 10 nodes from one golden snapshot, only nginx and php-fpm, CPU target 55% (it began at 70%). No jobs on web nodes, so extra nodes never run one twice.
  4. Data. MySQL Standard 8.4 (4 vCPU / 16 GB, primary and standby, reads on the standby, sticky = true); Valkey 8 GB on two nodes; PostgreSQL with pgvector (2 vCPU / 4 GB) kept apart, so vector queries never compete with exam writes.
  5. Worker. One fixed 4 vCPU / 8 GB droplet for the queue lanes (including a separate AI lane), the scheduler, WebSockets and all SMS, because the SMS gateway whitelists one IP. A known single point of failure.
  6. Files, logs, watching. Spaces behind a CDN; backend logs through Fluent Bit into OpenSearch; our ops console samples every tier read-only and runs the guarded deploys.

Before and after

AreaBefore (until 5 Oct)Now
FrontendOne 8 vCPU / 16 GB droplet, one Next.js container, no load balancerLoad-balanced pool of 2 to 10 nodes, two containers per node
BackendTwo fixed droplets, targeted by ID, so new servers could never joinBalancer targets a tag; pool of 2 to 10 from one snapshot, pre-scaled before exams
MySQLOne Advanced node, 8 vCPU / 32 GB, no standbyStandard 4 vCPU / 16 GB, primary and standby
"Standby" serversTwo powered-off droplets, $112 a month, minutes to power onRemoved; real standbys inside MySQL and Valkey (now 8 GB, up from 4)

What the money buys

Item (monthly, DigitalOcean list prices)Cost
Before: whole web tier (16 GB frontend, 2 backend, worker, 1 load balancer, 2 powered-off standbys)about $416
Now, web tier: 2 backend ($56 each), 2 frontend ($28 each), worker, two load balancers$112 + $56 + $56 + $48
Now, data and logs: MySQL pair, Valkey 8 GB x 2, PostgreSQL, OpenSearchabout $389 + $240 + $60 + $20
Now: snapshots, Spacesabout $5 + $5 and usage
Now: total production, backend pool at 2 nodesabout $990

The totals are not like for like: $416 was the web tier only, and $990 is the whole production stack. On the same list prices the web tier alone is now about $272. The other $709 is managed MySQL, Valkey, PostgreSQL and OpenSearch, and that is what the extra money buys: a database that survives a node failure, a Valkey with room and a standby, vector search that stays out of the way of exam writes, and logs that outlive a deleted node. Pre-scaling is the cheap part: an extra backend node costs about $0.08 an hour. The decisions behind the numbers are in the checklist below.

Planned, and not done

  • Planned: a full load test and a formal multi-hour soak in a maintenance window, before the next large exam.
  • Planned: Next.js static assets from the CDN, Fluent Bit on the frontend nodes, and a signed real-IP header from frontend to backend so every per-IP rule and log is exact.
  • Planned: worker high availability: a reserved IP whitelisted with the SMS gateway, a small standby worker, and an exam-only burst worker started before big exams.
  • Considered: Kubernetes (DOKS) for seconds-level scaling, which is much more to run. A warm standby ready in 3 to 4 seconds was also considered, but DigitalOcean pools have no warm pool, so we pre-scale and keep spare capacity serving inside the balancer.

A Solutions Architect's checklist

This is the checklist I now use in architecture reviews. Every item is something this project did or paid for.

Scalability

  • Web nodes are stateless: sessions, cache and queues in Valkey, uploads in Spaces, 0 files written to disk in 24 hours.
  • Each tier scales by its own unit: two Next.js containers per frontend node, php-fpm workers per backend node.
  • Known load is pre-scaled; reactive scaling is only the safety net.

Availability

  • No single node can take the site down: two balancers, pool minimums of 2, standbys for MySQL and Valkey.
  • The health check is answered by nginx alone, and a node is drained (503 on /lb-health) before it is removed.
  • The old system stays until traffic has really left: an hour after the DNS switch it still got about 36% of requests.

Performance

  • Know CPU per request (59 ms), not only requests per second.
  • Watch per-request setup cost: a fresh TLS connection cost about 5.4 ms of CPU for Valkey and 3.3 ms for MySQL.
  • Reads are shared between the standby and the primary, writes go to the primary, and a student still sees their own answer.

Reliability

  • Slow jobs get their own queue lane, so a 0.34-second job never waits behind a 10-second one.
  • Calls to outside providers are paced and retried (8 a minute shared, backoff from 1 to 15 minutes, up to 12 hours), so students never see an error.
  • Rate limits count per student, not per IP, wherever many students share one address.

Observability

  • A once-a-minute digest during exams: started and submitted, requests and 5xx per node, pool CPU, worker load, failed jobs.
  • Look at each node, not only the pool average (47% against 86%).
  • Logs outlive the node: Fluent Bit into OpenSearch for the backend; the frontend is Planned.

Failure recovery

  • Managed database backups and failover, golden snapshots, and the last 3 good images kept for rollback.
  • Every cutover keeps a way back: the old database frozen for 48 hours, the old frontend kept until DNS drained.
  • Single points of failure are written down. The worker is one: if it dies, SMS, the scheduler and exam processing stop while submissions wait safely in Valkey.

Cost

  • Measure headroom before choosing a size: the frontend moved to shared CPU after the test, and a Standard MySQL with a standby replaced one big Advanced node (cheaper and highly available).
  • Remove spend that buys nothing: two powered-off standbys cost $112 a month and gave no availability.
  • Pre-scale instead of over-provisioning, and spend where it removes a real failure, like the bigger Valkey.

Maintainability

  • Never hand-edit pool nodes: change one, test it, snapshot it, point the pool template at the snapshot.
  • Pass every setting on a pool update, so nothing silently resets.
  • New guards ship with tests that fail without the fix, as the per-student throttles did.

Future growth

  • Write the connection math down: 10 nodes x 80 php-fpm workers is 800, under MySQL's 1,601.
  • Know the next ceiling: MySQL near 450 requests/s, about 6 times the night's peak.
  • Measure written exams on their own, and check account limits early: the droplet limit of 25 had to be raised before both pools could reach their maximum during a rollout.

What we learned, and where this leaves us

  1. The first bottleneck is rarely the one you feared. We worried about servers; the first three problems were a rate-limit rule, a connection limit and a queue layout.
  2. Fix the cause, not the symptom. Only one problem in six was a capacity problem.
  3. Check production against the model. 63 requests/s at 59 ms predicted about 46% CPU, and we saw 47%.
  4. Averages hide the node that hurts, and scale-in is abrupt. 47% on average, 86% on one node; raise the minimum before an exam and lower it only after.
  5. Say what you have not tested. We have load tests and one real evening; a formal soak is Planned.

When this series began, the platform was one frontend droplet, two backend droplets pinned by ID and one big database node. The exam evening of 7 October ran on a different system, and none of the difference came from adding servers for their own sake. It came from the same loop, again and again, including the times the measurement told us we were wrong.

Measure. Find the bottleneck. Understand it. Change one thing. Test. Measure again.

That loop is the architecture; the pools, the standbys and the diagrams are what it leaves behind. If you keep one thing from these five parts, keep the loop, and use it before your next exam night, sale or launch.

Thank you for reading all five parts. To go back, start with part 1, the starting point and high availability, then part 2, scaling the application layer, part 3, the database and part 4, Valkey, workers and logging. This was the last part of the series.

About the Author

Anichur Rahaman

Continue Reading