One self-hosted console to run your entire business — commerce, ERP, HRM, CRM & manufacturing

Scaling a Self-Hosted Commerce Platform: A Server Roadmap from 100 to 100,000 Orders a Day

A stage-by-stage guide to scaling commerce infrastructure: what to add at each stage, the metrics that tell you when, what breaks first, zero-downtime blue/green deploys and backups you can actually restore.

Author

Anichur Rahaman

2 months ago9 min read1 views
Scaling a Self-Hosted Commerce Platform: A Server Roadmap from 100 to 100,000 Orders a Day

Every self-hosted commerce platform starts the same way: one server, one database, one person who knows where everything is. That is the right way to start. It is cheap, simple and fast enough for the first thousand customers.

The trouble begins when growth arrives faster than the architecture changes. A flash sale doubles traffic for an afternoon. A courier integration retries every few seconds. A finance report scans two years of orders during the lunch rush. Suddenly checkout is slow, and nobody can say why.

This guide is the scaling roadmap I use when planning infrastructure for commerce and ERP systems. It moves through four stages, from about 100 orders a day to 100,000 and beyond. For each stage it explains what to add, what to measure, and — most importantly — what breaks first, so you change the architecture when a number tells you to, not when a customer does.

Four stages of server architecture growth: all in one box, split by role, scale out, and a fleet
Four stages of growth. Each column lists what to add and the first thing that usually breaks.

First principle: measure before you scale

Scaling is expensive and every new component is something else that can fail. Before adding servers, make sure you can see these five numbers on a dashboard:

  • p95 response time for the home page, a product page, cart and checkout. Averages hide the slow requests that lose sales.
  • Database load: active connections, slow queries (anything over 100 ms) and cache hit ratio.
  • Queue depth and age: how many jobs are waiting, and how old the oldest one is.
  • Memory and swap on every host, plus out-of-memory kills in the kernel log.
  • Error rate: 5xx responses per minute and failed jobs per hour.

If you cannot see these, the first scaling project is observability. A server you cannot see is a server you will over-buy.

Stage 1 — Everything in one box (up to ~1,000 orders a day)

A single virtual server with 4 vCPUs, 8 GB of RAM and fast SSD storage runs a surprising amount of business. Containers keep it tidy: web (PHP-FPM behind nginx), PostgreSQL, Redis, one queue worker and one scheduler.

What to get right at this stage

  • OPcache on, with preload. PHP compiles each file once and keeps it in memory. This alone often cuts response times by half.
  • No work in provider boot. Anything that runs on every request (reading settings, building menus) must be cached. Ten milliseconds per request becomes a CPU core at scale.
  • Background jobs for slow work. Emails, courier bookings, image conversions and PDF invoices belong on the queue, never inside the checkout request.
  • Off-site backups, tested. A nightly database dump to object storage, and a monthly restore test. An untested backup is a hope, not a backup.

What breaks first

Memory. A heavy report, an import or an image batch competes with checkout for RAM. You see it as random slow pages and, eventually, the kernel killing a process. The scheduler is a common victim: give it its own memory limit (512 MB is a sensible floor) and run a single long-lived scheduler process rather than starting a new one every minute.

Stage 2 — Split by role (~1,000 to 10,000 orders a day)

The next step is not "more of the same server". It is separating roles so they stop competing with each other.

RoleWhy it movesTypical size
Database hostIts own disk I/O and memory for caching; no noisy neighbours.4–8 vCPU, 16–32 GB RAM, NVMe
Redis hostSessions, cache and queue stay fast under load.2 vCPU, 4–8 GB RAM
Web nodesCPU-bound PHP work, scaled independently.2–4 vCPU each
Queue workersSeparate pools per queue, so emails never delay payments.2 vCPU per pool

Cache the pages guests see

Most traffic to a store is anonymous: people browsing products and categories. Those pages can be cached for a few minutes as full HTML, at the application and at the CDN. Two rules prevent painful bugs:

  • The cache key must include the host name. If your store answers on two domains, a key without the host serves one domain's links and scripts to the other.
  • Never cache anything personal. Cart counts, prices for logged-in customers and account pages are rendered after the cached shell, or not cached at all.

What breaks first

Database connections and slow queries. Each PHP worker opens its own connection. Forty workers across two web nodes plus queue workers can exceed PostgreSQL's comfortable connection count. At the same time, one missing index turns a 5 ms query into a 900 ms one when the table grows past a million rows. Watch the slow query log weekly.

Stage 3 — Scale out (~10,000 to 50,000 orders a day)

Now you run several identical web nodes behind a load balancer. The system must be stateless at the web tier: sessions in Redis, uploads in object storage, nothing important on a web node's local disk.

Request path: CDN, gateway, web node with PHP-FPM, Redis, PgBouncer and PostgreSQL, with queue workers, scheduler, SSR and object storage off the request path
The request path. Everything in the bottom row must stay off it — a shopper should never wait for a queue job.

The components that matter at this stage

  • Connection pooling (PgBouncer). Hundreds of application connections share a few dozen real database connections. This is often the single biggest stability win.
  • A read replica for reports. Dashboards, exports and analytics read from a replica, so a heavy report never slows checkout.
  • Object storage for media and backups. S3 or a compatible service such as Cloudflare R2, served through the CDN. Your web nodes stop serving images entirely.
  • A separate SSR renderer. Server-side rendering makes pages fast and indexable, but it is a different workload from PHP. Scale it on its own.
  • Queue supervision. A dashboard (Laravel Horizon, for example) that shows throughput, failures and wait time per queue, with alerts.

What breaks first

Cache stampedes and hot rows. When a popular cache entry expires, hundreds of requests try to rebuild it at the same moment. Use locks or "stale-while-revalidate" so one request rebuilds while others serve the old value. Separately, a single row that everyone updates — a stock counter for a flash-sale product, a sequence number — becomes a queue of waiting transactions. Keep those updates short and atomic.

Stage 4 — A fleet (50,000+ orders a day)

At this size the technical patterns are well known. What changes is the discipline around them.

  • Partitioned queues, so payments, notifications, search indexing and AI jobs each have dedicated capacity.
  • Search and analytics offloaded to engines built for them, fed by events rather than nightly dumps.
  • Archive storage for history. Old orders, logs and audit trails move to cheaper storage while staying queryable.
  • Multi-region edge for static assets and cached pages, close to customers.
  • Runbooks and on-call. Every alert has a written first response. The worst outages I have seen were not caused by missing servers, but by nobody knowing what to do at 3 a.m.

Deploy without downtime: blue/green

A store that goes offline for every release teaches the team to release less often, which makes each release riskier. Blue/green deployment breaks that cycle.

Blue/green deployment in six steps: build once, start green, migrate with expand-only changes, warm and check, switch, keep blue for rollback
The new version boots, migrates and passes health checks beside the live one. Only then does traffic move.
  1. Build once. One image per commit, tested in CI, promoted unchanged from staging to production.
  2. Start green beside the live blue containers.
  3. Migrate with expand-only changes. Add columns and tables; never rename or drop in the same release. The old code must keep working against the new schema.
  4. Warm and check. Cache config and routes, hit a health endpoint, run a short smoke test.
  5. Switch the gateway upstream with a reload, not a restart, so no connection is dropped.
  6. Keep blue running for a quick rollback. Remove old columns in a later release, once nothing reads them.

Hard-won lesson: clearing caches or compiled files belonging to the live version while the new one is being prepared will break a few requests on every deploy. Prepare the new release in its own directory, and clear the application cache once, right after the switch.

Backups, RPO and RTO in plain words

Two numbers define your backup strategy, and the business — not IT — should choose them:

  • RPO (recovery point objective): how much data you can afford to lose. A nightly backup means up to 24 hours of orders.
  • RTO (recovery time objective): how long you can be offline while restoring.
StageSensible RPOSensible RTOHow
124 hours4 hoursNightly dump + files to object storage
21 hour1 hourHourly snapshots, scripted restore
3–4MinutesUnder 30 minutesContinuous WAL archiving, standby replica

Whatever you choose, schedule a restore test. In StoreConsole, the Backups module takes scheduled database and file snapshots, checks disk health, applies retention and restores in one click. The short tour below shows it.

Backups tour (0:46): scheduled snapshots, retention and one-click restore.

Five traps I see again and again

  1. Caching a 404 at the CDN. Upload an image after a page has already requested it, and the CDN may serve "not found" for hours. Write files first, link to them second.
  2. Tests that pass on one database and fail on another. SQLite in tests and PostgreSQL in production disagree on case-sensitive search, type comparisons and enum changes. Run at least one CI job on the production engine.
  3. Running heavy builds on the production host. Two asset builds at once can exhaust memory and take the store down. Build in CI, ship images.
  4. Background processes started with a plain shell "&". They die with the session. Use a supervisor or the container runtime.
  5. Debug tooling left on in production. Query listeners and profilers add work to every request, and a forgotten monitor can quietly consume a CPU for weeks.

A capacity checklist you can use today

  • p95 checkout time under 800 ms at your busiest hour.
  • Database CPU under 60% at peak; no query over 100 ms on the checkout path.
  • Oldest queued job under 60 seconds for payments and notifications.
  • At least 30% free memory on every host at peak; zero OOM kills in the last week.
  • Last restore test less than 30 days ago.
  • Every release deployable and reversible without downtime.

If you are planning hardware for StoreConsole specifically, the system requirements page lists recommended sizes per stage, and the documentation covers Docker, queues and the scheduler.

Key takeaways

  • Scale when a metric tells you to: p95 latency, database load, queue age, memory, error rate.
  • Stage 1 is one well-tuned box. Stage 2 splits roles. Stage 3 scales out with pooling, replicas and object storage. Stage 4 is about discipline.
  • Keep slow work off the request path; a shopper should never wait for a queue job.
  • Deploy blue/green with expand-only migrations, and keep the old version ready for rollback.
  • Let the business choose RPO and RTO, then test restores on a schedule.

Anichur Rahaman is a software architect and the creator of StoreConsole. He designs commerce and ERP systems for growing businesses, with a focus on event-driven architecture, data integrity and self-hosted operations.

About the Author

Anichur Rahaman