Observability for High-Volume Systems: Logs, Metrics and Traces You Can Act On
A request ID, structured logs in OpenSearch, four golden signals, sampled traces and alerts tied to user impact: the small set that lets you find the slow layer in minutes, plus how to watch a load test and an event day.
Author
Anichur Rahaman
6 days ago12 min read1 views
"Customers say checkout hangs." The message reaches the platform lead of a ticketing site at 10:01 on release morning. Sales opened at 10:00, about a thousand people clicked in the same minute, and the plan said 440 requests per second. The dashboard shows CPU at 40% and nothing red.
Which layer is slow? One route, one node, one customer? With a console and a guess, the next ten minutes go to arguing. With a request ID and a few well-chosen numbers, they go to one search. (This scene is illustrative, not a real incident.)
You do not need a large platform for that. You need a request ID, structured logs, the four golden signals and alerts tied to user impact. In a platform I prepared for a scheduled spike of about a thousand users in the same minute, the servers were fine; what nearly hurt was that nobody could say quickly which layer was slow, or whether a scary number was real. This article builds the small set that answers those questions, then shows how to watch a load test and an event day with it.
This is part 4 of a five-part series, "Engineering for High Volume". Part 1 covers what actually scales, part 2 covers queues and workers, and part 3 covers the data tier. This part is about seeing all of it.
Observability in one paragraph
Monitoring tells you that something is wrong. Observability lets you ask why without shipping new code first. It rests on three kinds of data, often called signals:
Logs: one record per event, with detail. Best for "what exactly happened to this request?"
Metrics: numbers sampled over time. Best for "is it getting worse, and how fast?"
Traces: the path of one request through every service, with the time spent in each step. Best for "where did the time go?"
You do not need all three on day one, and you do not need an expensive platform. You need them to be connected, so you can jump from a red metric to the logs and the trace of one slow request. The glue is a request ID.
Start with a request ID that survives every hop
The cheapest improvement I know is one identifier that follows a request from the edge to the last queued job. In the setup I ran, nginx honours an incoming X-Request-Id header or generates one, passes it to PHP-FPM, and both write it in their logs. The application then adds it to every log line, and copies it into the payload of any job it dispatches, so a background task can be tied back to the click that caused it.
Return the same ID in the response header. When a customer or a support agent reports a problem, that one string finds the whole story.
One request ID, written by every layer, turns five separate log files into one timeline.
Where the ID must appear
Layer
What to do
Edge / CDN
Forward the incoming ID, or let the next layer create one
Host nginx and container nginx
Reuse the incoming ID, generate one if missing, add it to the access log
PHP-FPM and the app
Read it from the header, attach it to every log line as shared context
Queue jobs
Store it in the job payload and restore it when the job runs
Outgoing HTTP calls
Send it onward so partners and internal services can log it too
Response
Return it as a header so support can ask for it
Structured logs into an ELK or OpenSearch stack
Plain-text logs are written for humans to read one at a time. At high volume you need to search, count and filter them, which means one JSON object per line with named fields. A short, stable set of fields is enough to start.
Field
Why it earns its place
time
Ordering and correlation with metrics
request_id
The thread that joins every layer
route or path (without the query string)
Group slow requests by endpoint, not by raw URL
status
Error rate by endpoint
duration
The raw material for p95 and p99
upstream time
Separates "nginx was slow" from "PHP was slow"
user or tenant id (an internal ID, never an email)
Find out whether one customer is the problem
On naming: ELK is the old nickname for three tools, Elasticsearch, Logstash and Kibana, and the wider Elastic Stack added lightweight shippers called Beats. OpenSearch began in 2021 as an Apache-licensed fork of Elasticsearch and Kibana started by AWS, and in September 2024 it moved to the OpenSearch Software Foundation, under the Linux Foundation. The two are no longer identical, but for logs the workflow is the same: ship, index, search, chart. That is why people say "ELK-compatible".
You do not need a large cluster. In my setup a small OpenSearch node with 2 GB of memory and one vCPU held the logs of a platform handling a sharp spike. That is one setup, not a benchmark, but it shows the cost is small compared with the first outage you diagnose in minutes instead of hours.
What not to log
Logs are copied, indexed and kept for months, so they are a data-protection risk. Redact passwords, tokens, card data and full personal details before a line leaves the application. Log internal IDs, not names, emails or phone numbers. Strip query strings from paths, because they often carry tokens.
Retention and index lifecycle
Rolling indices by day, with an automatic lifecycle, keeps a small node healthy: recent days stay searchable, older indices are compressed or moved, and anything past your retention window is deleted. Decide the window with the people who own privacy and security, not by guessing. Thirty days of detailed logs is a common starting point; the right answer depends on your rules.
Metrics: the four golden signals
Google's Site Reliability Engineering book, in its chapter on monitoring distributed systems, names four signals that cover almost any user-facing service. If you can only measure four things, measure these.
Signal
Question it answers
For a web and worker platform
Latency
How long do requests take?
p50, p95 and p99 per route; track failed requests separately
Traffic
How much demand is there?
Requests per second at the edge, jobs dispatched per second
Two warnings. First, never alert or plan on averages: a healthy average hides a terrible tail. In a load test I ran on a server-rendered frontend, one process saturated at roughly 220 to 250 requests per second, and p95 jumped from about 260 ms to about 3 seconds while the host still had idle cores. The average would have looked acceptable for a while. The percentile told the truth.
Second, saturation is where trouble starts, and it is rarely CPU. For this kind of platform I watch these first:
PHP-FPM busy workers against the pool maximum. When it touches the limit, requests queue in nginx.
Queue depth and queue age per queue. Depth says how much waits; age says how late you already are.
Database connections and write throughput, not just CPU.
Redis memory and evictions, per role (cache, sessions, queues).
Memory per container, because an out-of-memory kill looks like a random error from outside.
Tracing with OpenTelemetry
Logs and metrics tell you a request was slow. A trace tells you it spent 40 ms in nginx, 90 ms in PHP, 1.2 seconds waiting for one query and 30 ms in Redis. For a system with a frontend tier, a backend tier and workers, that is the fastest way to find the guilty hop.
OpenTelemetry is the vendor-neutral standard for producing this data. According to its specification status page, the tracing and logs signals are stable and the metrics API, protocol and data model are stable, while the maturity of individual language SDKs varies. The PHP project documents traces, metrics and logs as available, with an optional extension for automatic instrumentation. Check the status of the specific libraries you use before you commit.
My advice is to adopt it in this order. Make sure the request ID exists first. Then trace only the critical path, such as checkout or submission, and sample it: keeping every trace at high volume costs more than it teaches. Keep all traces for errors and slow requests, and a small fraction of the rest. Put the trace ID into your log lines next to the request ID, so the two tools link to each other.
SLOs and alerts that mean something
An alert should answer one question: does a person need to act now? CPU at 85% rarely does. "Checkout is failing for 2% of users" always does. A service level objective makes that concrete: a target such as "99% of checkout requests succeed in under 2 seconds over 30 days", with an error budget for the rest.
Page a person for
Send to a ticket or chat for
Checkout error rate above the SLO burn threshold
Disk filling slowly
p95 on a critical route over target for several minutes
One node at high CPU, users unaffected
Queue age growing on the submissions queue
A single failed job that retried and passed
Failed jobs rising fast
Certificate expiring in three weeks
A worked example shows why. Say checkout takes 3,000,000 requests in 30 days and the objective is 99% success, so the error budget is 30,000 failed or too-slow requests. During a ten-minute spike at 440 requests per second you serve 264,000 requests. If 2% of them fail, that is 5,280 failures: about 18% of the whole month's budget gone in ten minutes. That is a page, whatever the CPU says. (Illustrative numbers.)
Every alert that wakes someone for nothing teaches the team to ignore the next one. Review alerts after each event and delete or downgrade the noisy ones.
The event-day dashboard
For a scheduled spike I build one screen, and only one. It is the golden signals for the critical routes on top, the saturation numbers beneath, queue depth and age beside them, and a line showing the target request rate. Nothing else. A dashboard that needs scrolling will not be read at the peak minute.
One screen for the peak minute: the four golden signals first, then the resources that run out.
Put a deploy marker on every chart. Half of all "mystery regressions" are a release that landed ten minutes earlier.
Observing a load test
A load test is the cheapest rehearsal you will ever get, and it is wasted if you cannot see inside it. Replay the real request mix of a previous peak rather than hitting one URL. In my tests I used k6 with abort thresholds: stop if p95 goes above 3 seconds, or on any 5xx or timeout. A small guard script stopped the run if CPU limits were hit, so a test could never become an outage.
Define abort thresholds before the first run, and write them down.
Run the dashboard you will use on the day, not a special test view.
Tag the test traffic with a header, so you can filter it in the logs.
Compare your own metrics with your cloud provider's. In my runs they agreed within a few percent; if yours do not, one of them is wrong.
Find the first saturated resource, fix it, and run again at the target rate.
Keep the results next to the code, so the next event starts from a baseline.
Verify the scary alert before you act
Under pressure it is tempting to act on whatever is red. Verify first. Some consoles mislabel containers, and a "100% CPU" alert may point at a different process than the one you are about to restart. Look at the real process list, cross-check with a second source, then decide.
Where a red alert goes next: most of the work is deciding whether it is real.
The same is true for your own background work. Health checks and schedulers consume resources too, and in one case a health check that spawned a process on every run, left orphaned under load, destabilised a database container. Metrics on the monitoring itself are not paranoia.
Back to 10:01. With this in place, the support message arrives with a request ID attached. One search shows nginx waiting 1.31 seconds on PHP, and the app logging a slow query on the same request. The dashboard agrees: p95 is climbing while 48 of 60 FPM workers are busy. The one red CPU alert belongs to a different container, so nobody restarts anything. The lead has a cause in about two minutes instead of ten, and the fix goes to the right place. (Still illustrative.)
Add a request ID at the edge and carry it through nginx, PHP-FPM, the app logs and queue jobs. It is the cheapest big win.
Write structured JSON logs with a short stable set of fields, redact personal data, and set a retention window on purpose.
Measure the four golden signals, use percentiles instead of averages, and watch saturation: workers, queue age, connections, memory.
Adopt OpenTelemetry tracing on the critical path only, sampled, and link trace IDs to your logs.
Page people only for user impact tied to an SLO, and prune noisy alerts after every event.
Rehearse with abort thresholds, compare with your provider's metrics, and verify a scary alert before acting on it.
Anichur Rahaman is a software architect and the creator of StoreConsole. He designs commerce and ERP systems for growing businesses, with a focus on event-driven architecture, data integrity and self-hosted operations.