One self-hosted console to run your entire business — commerce, ERP, HRM, CRM & manufacturing

Part 4: Valkey, Queues and Logs, the Quiet Parts That Break First

Part 4 of the NovaCommerce case study: 2,854 Valkey timeouts, one worker that is still a single point of failure, a queue that delayed merit cards by 115 minutes, an AI provider that said no, and a rate limiter that counted the building.

Author

Anichur Rahaman

1 day ago13 min read2 views
Part 4: Valkey, Queues and Logs, the Quiet Parts That Break First

A backend load test at 02:28 on 6 October came back with 2,854 application errors. Every one was the same sentence, thrown while connecting: RedisException: Operation timed out. The web servers were healthy and the database was healthy. The part that said no was the quiet service holding sessions, cache and queues.

The same week produced a second number: 115. That is how many minutes the slowest merit card waited to be recomputed after an exam. The worker's CPU was not the cause, and neither were WebSockets. The job was simply standing in the wrong line.

Part 3 covered the databases. This part covers the pieces that tie a stateless fleet together: Valkey, the one fixed worker, the queues, the rate limits, the logs and the signals we watch. They look dull on a diagram, and they gave us many of our real surprises. Each section below is a failure, a cause and a change.

This is part 4 of the five-part case study "From One Server to Exam-Day Ready". NovaCommerce is a fictional name; the architecture, numbers and mistakes are real.

A fleet with no memory needs somewhere to remember

A node that can be deleted at any minute cannot own anything. In this platform the owners are a short list: managed MySQL for application data, managed PostgreSQL with pgvector for AI embeddings, Spaces object storage for files, managed OpenSearch for logs, and managed Valkey for everything small, fast and shared: sessions, cache and queues.

Think of supermarket cashiers. The till drawer and the stock list sit in a back office. A cashier can go home mid-shift and the next one takes the lane, because nothing was in the cashier's pocket.

Layered diagram: frontend pool, backend pool and the fixed worker above a private network line, and five stores below it: Valkey, MySQL, PostgreSQL with pgvector, Spaces and OpenSearch
Nodes on top hold nothing. Everything that must survive a node being deleted lives in one of the five stores below.

We checked this instead of assuming it. On 7 October the backend nodes wrote 0 files to local disk in 24 hours. Uploads, including students' written-answer PDFs, go straight to Spaces; sessions and cache are in Valkey; logs go to stdout. Implemented and verified. That is why the autoscale pools may delete any node at any time.

StateWhere it livesWhat it buys us
SessionsValkey (primary + standby)Any backend node can serve any student
CacheValkeyOne cache shared by every node
Queued jobsValkeySubmissions wait safely if the worker is down
Uploads, answer PDFs, mediaSpaces with CDNNo file ever sits on a node's disk
Logsstdout, shipped to OpenSearchThey outlive the node (backend done, frontend planned)
Application dataManaged MySQLCovered in Part 3

One Valkey, on purpose

Cache, sessions and queues share one managed Valkey 8 with a standby node. Its eviction policy is noeviction: when memory is full, Valkey refuses new writes with an error instead of quietly deleting old keys. For a pure cache that sounds backwards. But this instance also holds sessions and queued jobs, and we prefer a loud error to a job vanishing without a trace. We are far from that edge: on the exam evening of 7 October it used about 12.5% of its memory.

2,854 timeouts: when the shared memory could not answer

Back to 02:28. The test was hitting a Valkey of 4 GB with a primary and a standby. Memory was not the problem. Under load it could not accept new connections fast enough, and every request that waited longer than the 5-second connect timeout became a 500.

Part of the pressure was ours. Each request opened a fresh TLS connection to Valkey, which we measured at about 5.4 ms of CPU per request (about 3.3 ms for MySQL). At the 65 to 70 requests per second one backend node can serve, that handshake alone costs roughly a third of a core. That is back-of-the-envelope arithmetic, not a profile.

The fix was a resize, done in place: Valkey went from 4 GB to 8 GB, two nodes (primary plus standby), with the data kept and the host name unchanged. No application change, no cutover. Implemented.

Then we measured again. Tested: a gradual climb from 25 to 500 requests per second, starting on two backend nodes while the pool scaled out, gave 0 application errors and 0 backend 5xx. The load balancer showed 143 errors, but only while the nodes were CPU-saturated: the expected "out of capacity" signal, not a bug.

The useful lesson is that the bottleneck moved three times in a few days, and each move needed a different kind of fix.

OrderBottleneckKind of problemFix
1API rate limiter keyed on IPCode and configurationKey by account (Implemented)
2Valkey connectionsData-tier sizing4 GB to 8 GB, 2 nodes (Implemented)
3Backend CPU per nodePlain capacityMore nodes, pre-scaled (Implemented)

Only the last row is solved by adding servers. For the second, more servers would have opened even more connections into the same Valkey.

One worker, and the jobs only it can do

Everything around the worker scales. The worker itself does not. It is one fixed droplet (4 vCPU, 8 GB) and it carries four jobs:

  • Queues. Default, exam, three notification queues, OMR and enrollment reports, run by Horizon.
  • The scheduler. Timed tasks must run exactly once.
  • WebSockets. Reverb serves the live updates. Its port is open only to the backend nodes, by tag.
  • All SMS. The SMS gateway whitelists a single IP address, so SMS can only leave from this machine.

This is why queues, the scheduler, WebSockets and OMR are switched off on the web nodes: a new autoscaled node must never run a job twice. What our case adds to the generic rule is that SMS is tied to an address, not only to a process.

It is also a single point of failure, and we say so plainly. If this droplet dies, SMS, the scheduler and exam processing stop. Students keep submitting, and their submissions wait safely in Valkey, which is the reason queues live there. But nothing processes them until a worker is back.

On the exam evening of 7 October the worker peaked at 46% CPU with 0 failed jobs, so this is a risk we plan for, not a failure we have seen. The plan, with honest labels:

StepWhyStatus
Reserved IP, whitelisted with the SMS gatewayA replacement worker sends SMS from a known addressPlanned
Small standby workerSomething ready to take over queues and schedulerPlanned
Exam-only burst worker image, started before big examsExtra queue capacity when it is neededPlanned
One snapshot of the workerA rebuild starts from an image, not from memoryImplemented

The reserved IP comes first: without it, a standby worker would send SMS from an address the gateway does not accept.

The merit card that waited 115 minutes

After an exam, a "merit recompute" job refreshes each student's merit card. On 6 October those cards started appearing 10 to 115 minutes late.

What we ruled out

The obvious suspects were the worker's CPU and the WebSocket server. Neither was the cause.

What it was

The merit job is tiny: 0.34 seconds. It shared one queue, served by 2 workers, with an AI job that makes one language-model call of about 10 seconds per student. After an exam, about 2,000 of those AI jobs were enqueued, and every merit job sat behind them.

The arithmetic shows the scale. One AI job takes as long as about 29 merit jobs. Two thousand of them is roughly 20,000 seconds of work, and across 2 workers that is nearly three hours of line. A wait of 10 to 115 minutes fits that picture.

It is the express lane at a supermarket: a loaf of bread should not wait behind a full trolley, and we had built one till for everybody.

Before and after diagram: one shared queue where tiny merit jobs wait behind about 2,000 ten-second AI jobs, versus separate lanes where the AI lane is paced by a shared limiter of 8 per minute
Same jobs, different topology: on the right the merit job never meets an AI job, and the AI lane has its own pace.

The change

Implemented: a separate queue lane and its own container, just for AI narratives, with 2 replicas. Merit and position jobs never wait for a language model again, and merit cards are ready in about a minute. The AI jobs already waiting were moved to the new lane with one atomic script, so each job was in exactly one place at every moment.

The fix was topology, not horsepower. The generic rule of one queue per kind of work is in Queues and Workers at Scale; what surprised us was how harmless the slow job looked.

Then the AI provider said no

The new lane protected the merit cards. It did not make the AI jobs faster. The lane ran into the provider's own rate limit, about 7 to 10 calls a minute, and once the provider account simply ran out of credit.

Both are the same failure: an outside service says "not now". A retry loop that hammers it only burns budget and fills the failed-jobs table, so the lane was redesigned to be patient. Implemented and live on the worker since 6 October:

MechanismSettingWhat it does
Shared rate limiter8 per minute across all workersKeeps calls inside what the provider allows
Release, not failExponential backoff, 1 to 15 minutes, with jitterA refused job goes back to the queue; jitter stops them returning together
Retry windowUp to 12 hoursThe job keeps trying long after the rush
Budget refundMonthly AI budget counterA refused call is not charged
FallbackTemplate textThe student sees a generic narrative meanwhile

Students never see an error. The template text is there when they open the page, and the real narrative replaces it when its turn comes.

It was tested for real within a day. The provider account ran out of credit, about 3,800 narratives piled up as delayed jobs, and none of them failed. The next afternoon credit was added, and the first new narrative was written within minutes, with no restart and no manual replay.

The arithmetic shows why the window is hours: 2,000 narratives at 8 a minute is about four hours of work. A window of a few minutes would drop most of them.

Sharing the limiter matters. Two replicas that each pace themselves at 4 a minute work until someone adds a third. One shared counter keeps the total honest however many workers exist.

Rate limits: count the student, not the building

A rate limiter is only as good as the thing it counts. We got this wrong twice, at two layers, and both times the symptom was students receiving HTTP 429 in the middle of an exam.

WhenWhat was countedWhat went wrongFix
25 SepGlobal API limiter, keyed on IP (empty default guard for token-based students)A school or mobile carrier puts hundreds of students behind one NAT IP: one shared bucketKey by student or instructor account, IP only for guests (Implemented)
7 OctRoute throttles such as "30 per minute", still by client IPThe backend saw one frontend proxy IP for everyone: 74% of calls to one dashboard endpoint got 429Count per student per route, guests by IP; 0 such 429s afterwards (Implemented)

Every browser API call goes through the frontend proxy, so a per-IP rule treats the whole exam hall as one visitor.

Both fixes came with tests that fail without the fix; the second was deployed one node at a time with zero downtime. Planned: pass the real student IP from the frontend to the backend in a signed header, so every per-IP rule and every log line is exact.

The stale trusted proxy

A smaller risk sat nearby. The backend still trusted the IP of the old frontend server, the one we had deleted. If the cloud ever handed that address to someone else, they could fake student IPs. We removed it on 7 October with the drain-one-node method, again with zero downtime. Implemented. A trust list is a list of promises; delete the ones whose owner is gone.

Logs that outlive the node

With autoscaling, the machine you want to inspect is often gone. So logs must leave the node as they are written. Our containers write only to stdout and stderr, and Fluent Bit ships the backend container logs to managed OpenSearch, the ELK-style stack: one searchable place that is kept after the node is gone. Implemented for the backend nodes. The generic method is in Observability for High-Volume Systems.

It paid for itself in the capacity work. DigitalOcean's two-minute average put the 7 October backend peak at 63 requests per second; the per-minute logs said about 69. Averages hide peaks, and a capacity model needs the peak.

There is a gap. Frontend nodes still keep their nginx logs on local disk, about 440 MB a day, and lose them when a node is removed. Planned: Fluent Bit on the frontend nodes too. Until then, the frontend is the one layer where a scale-in erases the evidence.

What we watch, and what we do about it

Our ops console is read-only: it samples and never changes anything. It records CPU per node and per container, pool CPU, MySQL threads running, connections and lock waits, Valkey memory, clients and operations per second, and queue health. DigitalOcean's own monitoring adds requests per second and response classes at the load balancers.

During an exam we do not stare at all of that. We read a digest once a minute: students started and submitted, requests and 5xx per node, pool CPU, worker load and failed jobs. Few numbers, each tied to a decision.

Left-to-right loop: signals from the ops console, load balancer metrics and OpenSearch logs, a once-a-minute digest, then actions such as drain a node, roll back the image and pre-scale, then re-measure
Signals feed a digest, the digest triggers one of a few known moves, and every incident returns as a runbook rule.

Acting on a signal

The loop has a small set of moves, described in Part 2:

  • Drain a node. Make /lb-health return 503, and the load balancer stops sending new requests within about 30 seconds.
  • Roll back the image. We keep the last three good images, and we never roll a template during exam hours.
  • Pre-scale. Raise the pool minimum about 45 minutes before a big exam. Reactive scaling is only the safety net, because DigitalOcean's CPU metric lags real load by 5 to 8 minutes.

In the full rollout of both pools on 7 October, an uptime probe hit real pages every 2 seconds: 219 probes, 0 errors. That is the loop closing: change, watch, measure.

The scale-in mistake

During the real exam on 7 October, someone lowered the backend pool minimum. DigitalOcean removed a serving node without draining it, and about 360 requests failed in 2 minutes.

It was a human slip and a platform behaviour at once: autoscale scale-in does not drain. The rule is now in the runbook (Implemented as a runbook rule): raise the minimum before an exam, and lower it only after. We know its limit: a runbook rule depends on someone remembering it in the middle of an exam. It is a guard, not a lock.

What we learned

  • Stateless nodes are only as safe as the shared stores behind them. Size Valkey for connections, not just memory.
  • One queue per kind of work. A 0.34-second job must never wait behind a 10-second one.
  • When a provider limits you, pace the calls, back off with jitter, retry for hours, refund the budget and show a fallback.
  • Check what every rate limiter counts. The student is the unit, not the building and not the proxy.
  • Ship logs off the node as they are written, and close the frontend gap before it costs an incident.
  • Write down your single points of failure with a status next to each. Ours is one worker, with planned fixes and no pretending.
  • Autoscale scale-in does not drain. Raise the minimum before an exam.

In Part 5, the last of the series, we put it together: the real exam evening of 7 October with production numbers, the capacity model, the monthly cost, and the load and soak test we still owe ourselves.

About the Author

Anichur Rahaman

Continue Reading