Securing and Shipping High-Volume Systems: A Hardened Edge and Zero-Downtime Deploys
A walk through the security layers from CDN and WAF to a private network, and a delivery pipeline of rolling and blue/green deploys with expand and contract migrations, ending with a go-live checklist for spike day.
Author
Anichur Rahaman
1 day ago13 min read
Seventeen minutes before the checkout complaint in part 4, at 9:44 on the same release morning, the on-call engineer of the ticket site faces a different problem. The sale opens at 10:00, and about 1,000 buyers will act in the same minute. A developer asks to ship a one-line fix for a typo in the order email. At 9:52 the dashboard shows 4,000 login attempts a minute from a handful of addresses, a scalper script warming up. (This is an illustrative scenario, not a real incident.)
Two decisions made weeks earlier decide how that morning goes: how much of the system the internet can reach, and whether a deploy needs the site to stop. If the answers are "everything" and "yes", the engineer has to refuse the fix and hope the script gives up.
This final article covers both: a hardened path from the internet to the data, and a delivery pipeline that never takes the site down. It draws on my own field notes from platforms prepared for scheduled spikes such as a flash sale, a ticket release or an exam start. They are one setup's lessons, not a universal recipe.
No single product secures a system. A firewall does not fix a leaked password, and a strong password does not help if the database is open to the internet. The practical approach is layers: each one assumes the layer in front of it will sometimes fail.
The OWASP Top 10 is a useful reality check here. The current edition, OWASP Top 10:2025, puts Broken Access Control first and Security Misconfiguration second. Both are mostly about how a system is wired, not about exotic exploits. That matches what I see: most incidents come from a port left open, a default left unchanged or a permission that was too wide.
Only the edge is public. Everything behind it lives on a private network and accepts traffic from named sources only.
The edge: CDN, WAF, bots and rate limits
Put a CDN with a web application firewall (WAF) in front of everything public. It absorbs volumetric floods, serves cached pages without touching your servers, and blocks the most common attack patterns before they reach your code. For a spike, it also protects your capacity: every request the edge answers is one your web nodes never see.
Three settings matter more than the rest:
Bot management. A flash sale attracts scalpers and scripts. Challenge suspicious clients on the checkout and login paths, not on the whole site, so real customers are not slowed down.
Rate limits at the edge. Limit login, password reset, search and checkout per client. Set the limit from real traffic: in a load test, replay the request mix of a previous peak and note the highest legitimate rate per client.
Rate limits in the application. The edge can be bypassed if someone finds the origin address. Keep a second limit in the app, per account and per action, so a single user cannot hammer an expensive endpoint.
Lock the origin so it only accepts traffic from the edge network. If your servers answer anyone on the internet directly, the WAF is a suggestion, not a control.
TLS, and where it ends
Encrypt everything in transit. Use TLS 1.3 where clients support it and keep TLS 1.2 as the floor. PCI DSS 4.0 requires strong cryptography for cardholder data on public networks and excludes SSL and early TLS versions, so TLS 1.0 and 1.1 should be switched off if you take cards.
The design question is where TLS terminates. Many setups terminate at the CDN, then again at a host reverse proxy, then pass plain HTTP to a container. That last hop is fine on a private host network, but only if you know it is private. If traffic crosses a network you do not control, encrypt that hop too.
A related trap is mismatched limits. In one platform I ran, two nginx layers sat in series: a host proxy that ended TLS and a container proxy in front of PHP-FPM. The host default body limit was 1 MB, so uploads failed silently until every layer agreed: client_max_body_size on both proxies, and post_max_size and upload_max_filesize in PHP. It is a reliability bug, but it also tempts people to raise a limit blindly. Decide the maximum once, set it in every layer, and write it down.
A private network: only the edge is public
The single most valuable rule is also the dullest: nothing is public except the edge and the load balancer. Web nodes, workers, the database, the cache and the search node sit on a private network. The managed database and cache accept connections only from that network, so a leaked password alone does not give an attacker a way in.
Add firewall rules that say who may talk to whom, for example:
Component
Who may connect
Public?
CDN / WAF
The internet
Yes
Load balancer
Edge network ranges only
Yes, restricted
Web and SSR nodes
Load balancer only
No
Workers and scheduler
Nobody inbound
No
Database and cache
Web nodes and workers, by private address
No
SSH and admin access
A jump host or VPN, keys only
No
Review this table whenever the topology changes. Most "we were open for a month" stories begin with a quick change that nobody wrote down.
The shared-alias incident: isolation inside the host
Network isolation is not only about the internet. On one host I found that several sibling projects shared a single Docker network, and each project named its PHP service app. The proxy was configured to send requests to app on port 9000, and that name resolved to several containers. Half the requests reached a different project's PHP-FPM.
Nothing crashed. Pages simply came from the wrong code, with the wrong configuration and sometimes the wrong database. That is a reliability problem and a data-exposure problem at once.
The fix is simple. Give every project its own network, use unique service names, and never expose a service to a network that only needs to reach something else. When you review a deployment, ask a plain question: can this container reach anything it has no reason to reach?
Secrets, accounts and patching
Three habits prevent a large share of damage.
Secrets stay out of images. An image is copied to registries, laptops and build caches. Inject secrets at runtime from the environment or a secret store, and rotate them when a person leaves.
Least privilege for every account. The application's database user should not be able to drop tables or create users. Workers, reporting and migrations can use different accounts with different rights. Staff get roles, not a shared admin login, and admin access gets a second factor.
Patch on a schedule. OWASP added Software Supply Chain Failures to its 2025 list for a reason: what you depend on is part of your attack surface. Rebuild images regularly, scan them for known vulnerabilities, and keep a short list of dependencies you actually use. You can check your server and container settings against the CIS Benchmarks for Docker and the common Linux distributions.
Finally, assume that something will eventually go wrong anyway. Keep backups that an attacker cannot reach with the same credentials, and practise a restore. I cover that in ransomware-ready backups for business systems; here I only add that a restore you have never tried is a hope, not a plan.
Shipping without downtime: build once, move the same artifact
A scary deploy makes people skip patches and avoid change, so the goal is a boring one. The first rule is to build once and ship the identical artifact. Build the image on one machine, push it, and have every node pull that exact image. Compare image IDs across nodes before you switch traffic. Never run a source update and a build on a serving node: two nodes built an hour apart can quietly differ, and a build that fails halfway can leave a live server broken.
Configuration travels separately from the image, so the same artifact moves from staging to production unchanged. That is what lets you say the thing you tested is the thing you released.
Migrations with the site up: expand and contract
Database changes are where "zero downtime" usually breaks. During a rolling deploy, old and new code run at the same time against one schema, so a migration must work for both.
The standard answer is the expand and contract pattern, also called parallel change:
Expand. Add the new column or table in a way old code ignores, for example a nullable column.
Migrate. Deploy code that writes to both the old and the new shape, and backfill existing rows in small batches in the background.
Switch. Move reads to the new shape and watch for errors.
Contract. Only when no running code uses the old shape, remove it in a later release.
A worked example, with illustrative numbers. You want to rename the phone column on a 2.4 million row customers table to contact_phone. Expand: add contact_phone, nullable. Migrate: new code writes both columns on every save, while a background job copies 5,000 rows per batch. That is 480 batches, and at about two seconds each it runs for roughly 16 minutes without locking the table for any user. Before switching reads, assert that the number of rows where phone is set but contact_phone is empty is zero. Only then contract.
The drop is the one irreversible step, so it waits. Run migrations while the site is serving, and assert the result: row counts before and after, and a check that nothing is left half-converted. Take a fresh database dump before the window, and keep a rollback reference for the code too.
Rolling backends, blue/green frontends, workers and scheduler
Different tiers need different strategies. For the stateless PHP backends I use a rolling deploy, one node at a time:
Every stage has a gate. If a gate fails, the pipeline stops and the previous version keeps serving.
Drain one node at the load balancer, so it finishes in-flight requests and receives no new ones.
Deploy the new image to that node and restart PHP workers so OPcache picks up the new code.
Run a smoke test against the node directly: log in, load a product, add to cart.
Rejoin the node to the load balancer and let it soak for a few minutes while you watch errors and latency.
Repeat with the next node.
The decision that matters happens while the node is still out of the load balancer: a failed smoke test costs nothing.
With two nodes the site always has one healthy server. For idempotent requests, a retry rule in the proxy can hide a transient failure from users, but do not enable it for requests that charge a card.
For server-rendered frontends I prefer blue/green. Start the new version beside the old one, test it privately, then switch with a single proxy file change and a graceful reload. Rollback is the same step in reverse, about one second. StoreConsole runs its own production this way.
Workers and the scheduler need care of their own. Drain queues before switching workers: stop taking new jobs, give in-flight jobs a generous grace period to finish, then start the new version. Start the scheduler last, on exactly one node, so a half-deployed system never runs a recurring task twice or against the wrong code.
Verification, soak and the go-live checklist
A deploy is finished when you have proven it, not when the command returns. Run an end-to-end check that covers a real purchase or a real submission, confirm the queues are moving, then soak: keep watching error rate, p95 latency and queue age for a defined time before you call it done. Use the dashboards from part 4 of this series, and verify a frightening alert against the real process before acting on it.
This is the checklist I use before a high-volume event:
Load test at the target rate with abort thresholds, and keep the results.
Pre-scale web and frontend capacity before the event, rather than waiting for reactive scaling.
Confirm the origin accepts traffic only from the edge, and that WAF, bot and rate-limit rules are active on login and checkout.
Confirm the database, cache and search node have no public address.
Check TLS settings and certificate expiry dates, and that upload limits match in every layer.
Take a fresh database dump, test that it restores, and record the rollback references.
Freeze changes except emergency fixes, and name who may approve one.
Confirm alerts reach a person who is awake, with a written escalation path.
Do a dry run of the rollback, including the one-reload frontend switch.
Check that the scheduler runs on exactly one node and queues are drained of stale jobs.
Back to 9:44. With this in place, the engineer says yes to the typo fix. The image was built and tested an hour earlier, the nodes roll one at a time, and the 4,000 login attempts a minute hit a challenge on the login path only, while buyers browse untouched. The database has no public address, so there is nothing for the script to find behind the edge. The fix is live and soaked by 9:55, five minutes before the sale opens.
The table below shows how each layer fails and how to check it quickly.
Layer
Typical failure
Quick check
Edge
Origin reachable directly
Request the origin address from outside the edge network
Network
Database open to the internet
Port scan from an external host
Host
Shared service alias
List which containers each name resolves to
Secrets
Credentials baked into an image
Search the image layers and repository history
Deploy
Nodes running different builds
Compare image IDs on every node
Database
Destructive migration under load
Review for expand and contract; rehearse on a copy
Key takeaways
Layer your defences: CDN and WAF, a locked origin, a private network, least-privilege accounts and secrets outside images.
Make only the edge and load balancer public, and review the allowed paths whenever the topology changes.
Isolate projects on a host with their own networks and unique service names, so a shared alias never sends traffic to the wrong code.
Build once, ship the identical artifact and verify image IDs; never build on a serving node.
Change the database in small steps with expand and contract, so old and new code both work throughout a deploy.
Roll backends one node at a time, switch frontends blue/green, drain workers, start the scheduler last, and finish with a soak and a rehearsed rollback.
Anichur Rahaman is a software architect and the creator of StoreConsole. He designs commerce and ERP systems for growing businesses, with a focus on event-driven architecture, data integrity and self-hosted operations.