One self-hosted console to run your entire business — commerce, ERP, HRM, CRM & manufacturing

Queues and Workers at Scale: From One Serial Worker to an Autoscaled Process Pool

A burst exposes background jobs first. Follow one setup from a single serial worker, where only 20-26% of jobs finished on time, to separate queues and a Horizon process pool that scales in seconds.

Author

Anichur Rahaman

2 weeks ago12 min read2 views
Queues and Workers at Scale: From One Serial Worker to an Autoscaled Process Pool

The on-call engineer of an exam platform has one minute to spare at 8:59 on the morning of a test. At 9:00 sharp, 1,000 students press Submit within the same minute. The web servers answer every click in milliseconds and the dashboards stay green. By 9:03 the support inbox is filling up: students saw "Submitted" on screen, but no confirmation has arrived, and the grading team sees only a handful of answers. This is an illustrative scene, but I prepared a real platform for exactly this kind of spike.

The cause was not the web tier. A single background worker was processing the submissions one at a time. Queues turn a burst into a line, so the quality of your workers decides whether that line takes three seconds or sixteen minutes.

This article follows the path I took from one serial worker to a pool of processes that grows and shrinks on its own, plus the job-design habits that make the pool safe to run.

This is part 2 of the series "Engineering for High Volume". Part 1 covers the request path from edge to app: Autoscaling for Traffic Spikes: What Actually Scales.

Why background work fails first

A web request has a human waiting, so we notice when it is slow. A queued job has nobody waiting at that moment, so the slowness hides until the backlog is huge. Queues also turn a burst into a line: if 1,000 jobs arrive in one minute and each takes one second, a single worker needs more than 16 minutes to finish.

That was exactly my starting point. One worker, started with a command like queue:work --queue=retakes,default, handled 1,000 burst jobs one at a time. Worse, lower-priority work shared the same line as live submissions, so a batch of unimportant jobs could sit in front of something a customer was waiting on. In the replay of that peak, only 20 to 26 percent of the jobs finished on time. These are numbers from one setup, not a universal benchmark, but the shape of the failure is common.

The database was fine. The managed instance was nowhere near its limits. The bottleneck was a design decision: one process, one line, no priorities.

Step 1: separate the workers from the web

The first change was physical. Web and workers had shared an 8 GB machine, and when the workers got busy they starved PHP-FPM of CPU and memory. The site slowed down because of the very jobs it had queued.

Workers went onto their own node. Now a busy queue can use every core it wants without touching a single page load, and a web problem cannot stop the queue from draining. It also lets you size each machine for its own job: web nodes for many short requests, the worker node for memory-hungry jobs.

Before and after diagram: one shared web and worker node with a single queue, versus separate web and worker nodes with one queue per priority
Before: one machine and one line. After: web and workers apart, one queue per kind of work.

Step 2: one queue per priority and type

The second change was to stop mixing everything in one line. I created separate queues by what the work means to the customer:

  • Submissions: the live, customer-facing work that must finish within seconds.
  • Notifications: emails, SMS and push messages. Important, but a delay of a minute is tolerable.
  • Image processing: heavy, memory-hungry, never urgent.
  • Reports: slow exports and summaries that can wait until the spike is over.

Separate queues let you give each one its own number of workers, its own timeout and its own retry rules. A report job that runs for four minutes can no longer block a payment confirmation. The rule of thumb: if two kinds of work have different urgency, different memory needs or different failure behaviour, they belong in different queues.

Listing queues in priority order on a single worker (first queue always wins) is better than one mixed line, but it is still one process. Dedicated processes per queue are what remove the pre-emption problem for good.

Step 3: how many workers do you need?

Add workers and the line clears faster, up to a point. In my tests, about 12 worker processes cleared 1,000 jobs in roughly 3 seconds when the jobs were light, and in roughly 10 seconds when they were heavy. Past about 20 workers, adding more did not help: they simply queued up behind each other on database writes.

Worker processesWhat I observed (one setup)
1Jobs run one at a time; only 20–26% on time
About 121,000 light jobs in ~3 s, heavy jobs in ~10 s
Beyond ~20No gain; workers wait on database writes

The numbers hold together with simple arithmetic. These per-job times are what the results imply, so treat them as illustrative: a light job of about 0.03 seconds makes 1,000 jobs equal 30 seconds of work, and 12 workers share that into roughly 2.5 seconds. A heavy job of about 0.12 seconds makes 120 seconds of work, and 12 workers finish it in about 10 seconds. One worker would need the full 120 seconds, and by then the live submissions had long since missed their window.

This is the lesson behind the capacity worksheet at the end: the right number is not "as many as possible". It is the point where the next worker stops making the line shorter. Find it with a load test, not a guess, and do not upgrade the database before that test shows the database is actually saturated. I cover the data tier in part 3.

Scale processes, not machines

A fixed pool of 12 is wasteful on a quiet Tuesday and perhaps too small on a bigger day. Autoscaling is the answer, but the layer you scale matters. Starting a new virtual machine takes one to three minutes, and a scheduled spike peaks in seconds. As I explained in part 1, for known events you pre-scale. For queues there is a better option: scale the processes inside a machine that is already running.

Laravel Horizon does this with its balance setting. According to the official Laravel documentation, Horizon offers three strategies: simple splits incoming jobs evenly across worker processes, auto adjusts the number of processes per queue based on the current workload of each queue, and false leaves balancing off. With auto, a handful of options decide how fast it reacts.

SettingWhat it controlsValue I used for the submit queue
balanceStrategy: auto, simple or falseauto
minProcessesProcesses kept per queue even when idle1
maxProcessesUpper limit of processes Horizon may scale to10
balanceMaxShiftHow many processes can be added or removed in one adjustment3
balanceCooldownSeconds to wait between adjustments3

Horizon's documented defaults are gentler: one process per adjustment, every three seconds. I raised the shift to three because a burst should not take half a minute to ramp up. The result: when the line grew, Horizon forked more worker processes within seconds; when it emptied, it reaped them. No new machine, no cluster, no extra cost.

Illustrative chart of queue depth rising during a burst while the number of worker processes scales up and then back down
An illustrative burst: workers follow the queue depth up in seconds and back down when the line is empty.

Size the memory for the maximum

Process-level autoscaling has one trap. The container must be able to hold the maximum, not the average. If each job peaks at about 128 MB and you allow 10 processes, that is roughly 1.3 GB, so I set the container to about 2 GB. If the container is sized for the quiet state, the first real burst ends in out-of-memory kills, and the kernel removes workers in the middle of jobs.

Measure memory per job, not per worker, because one heavy job type can use up the whole budget.

Slim the heavy job before you buy hardware

One job in my setup peaked at 288 MB. It fetched its own images over HTTP and read files from local disk, holding all of it in memory at once. Streaming the files from object storage instead brought the peak down to about 128 MB. That change alone saved a whole machine size tier.

Before adding RAM or a bigger node, ask three questions about every heavy job:

  1. Does it load a whole file into memory when it could stream it?
  2. Does it fetch data over the network that it could receive in its payload, or read from a nearby store?
  3. Can it be split into many small jobs instead of one large one?

Smaller jobs also fit Horizon's balancing better, because it can move capacity in small steps.

Compare the autoscaling options

Horizon is not the only way to scale workers. Here is how the four common options compare for a Laravel-style queue.

OptionWhat scalesReaction timeCost and effort
Static replicas (Compose or Swarm)Nothing; you set a fixed countManualSimple, but you pay for peak capacity all the time
Kubernetes with HPAPods, on CPU or custom metricsMinutesNeeds a cluster and a metrics pipeline
Kubernetes with KEDAPods, driven by queue lengthMinutesNeeds a cluster; scales on the real signal
Horizon auto-balanceProcesses inside a containerSecondsFree, no cluster, bounded by one machine

KEDA is a good project. According to its documentation, it checks each trigger, such as the length of a Redis list, every polling interval (30 seconds by default), and scaling from one replica up to many is then handled by the Kubernetes Horizontal Pod Autoscaler. Adding pods also means scheduling and starting containers. For a scheduled spike that peaks in seconds, that is still too slow, and a cluster is a real cost for a small team.

My rule: use Horizon auto-balance for the fast reaction, and pre-scale the machine or the minimum before an event you know is coming. Consider KEDA when one node is no longer enough for the work you run every day.

Design jobs so retries are safe

A pool of fast workers will expose every weakness in how your jobs are written. These habits kept mine safe:

  • Make jobs idempotent. A job can run twice: after a timeout, a deploy or a retry. Running it twice must not send two emails or charge twice. Use a unique key or a check of the current state at the top of the job.
  • Dispatch after the commit. If a job is queued inside a database transaction, a worker can pick it up before the data exists, or the transaction can roll back and leave a job for nothing. Dispatch only after the commit succeeds.
  • Encrypt payloads that carry personal data. Queued payloads sit in Redis in readable form unless you encrypt them. Pass identifiers where you can, and encrypt the rest.
  • Set timeouts and retries on purpose. A job's timeout should be shorter than the queue's retry window, with backoff between attempts, and a failed-jobs table you actually read.
Flowchart of one queued job: commit, dispatch, queue by type, worker takes the job, a check for already done, run, success check, retry with delay or failed jobs table
One job, every outcome: skipped as a duplicate, completed, retried after a delay, or parked in the failed table where a person can see it.

The scheduler is a singleton

Cron-style scheduled tasks must run on exactly one node. If you scale web or worker nodes and every one of them runs the scheduler, each scheduled report is created several times, every reminder goes out twice and every cleanup collides with itself.

Pick one node, document it, and keep it out of any autoscaling group. Background commands also need the same care as web requests. In one self-hosted setup, a scheduler that booted the whole framework for every small command, with no CLI OPcache and tight memory limits, caused hundreds of out-of-memory kills. Running the relays inside a long-lived process and turning on CLI OPcache fixed it.

Drain the queues on every deploy

A worker is a long-running process, so it keeps old code in memory until you restart it, and a hard restart kills whatever job it is running. The safe order is to stop taking new jobs, let the in-flight ones finish, then start the new workers.

  1. Run Horizon's terminate command (horizon:terminate) so workers finish their current job and then exit.
  2. Give the container a generous stop grace period, longer than your longest job, so the orchestrator does not kill it early.
  3. Start the new workers on the new release.
  4. Start the scheduler last, so no scheduled task fires against a half-deployed system.

I go through the full deploy sequence, including blue/green and rollbacks, in part 5.

A worker capacity worksheet

Use this short list before your next spike:

  1. Measure how many jobs the event produces in its peak minute, per queue.
  2. Measure the run time and the peak memory of each job type on a real worker.
  3. Divide jobs by run time to estimate the processes you need to clear the peak within your target.
  4. Set the maximum slightly above that, and set the container memory for max processes times peak job memory.
  5. Load test with the real mix and watch database write time: stop adding workers when it becomes the limit.
  6. Confirm that the scheduler runs once, jobs are idempotent and a deploy drains cleanly, and decide which queue-depth and wait-time numbers you will watch (see part 4).

Back to 9:00 on exam morning. With the same 1,000 submissions, the submissions queue has its own processes and nothing else stands in front of it. Horizon grows from one process to ten within about ten seconds, the burst clears in roughly the time a student takes to look back at the screen, and the confirmations go out before the first support email is written. The reports and image jobs wait, as they should.

The next part, The Data Tier Under Load, looks at what happens once the workers are fast enough to push on the database: connection pooling, Redis roles and a separate store for AI embeddings.

Key takeaways

  • Background jobs usually fail first in a burst, and the slowness stays hidden until the backlog is large.
  • Move workers off the web nodes and give each priority or type of work its own queue.
  • Scale worker processes inside a machine with Horizon auto-balance for second-level reaction, and size memory for the maximum.
  • Slim heavy jobs before buying hardware, and stop adding workers when database writes become the limit.
  • Make jobs idempotent, dispatch after commit, encrypt personal data in payloads, and run the scheduler on one node.
  • Drain the queues on every deploy with a generous grace period.

Anichur Rahaman is a software architect and the creator of StoreConsole. He designs commerce and ERP systems for growing businesses, with a focus on event-driven architecture, data integrity and self-hosted operations.

About the Author

Anichur Rahaman

Continue Reading