Event-Driven Architecture Explained: Why Your Order Should Tell Your Stock and Books Itself
Nightly syncs and tangled calls leave stock and accounts out of step. See how events, the transactional outbox and idempotent listeners let one order update stock, ledger and loyalty reliably, with a worked example.
Author
Anichur Rahaman
2 months ago11 min read4 views
Over the sale weekend, a six-outlet retailer took 1,900 orders. On Monday morning its finance lead opens the reports, and the stock sheet says 14 jackets are left, but the outlet in the next town sold the last of them on Saturday. The ledger is worse: revenue is posted by a nightly batch, the batch timed out at order 1,412, and 488 orders are missing from the books. Nobody knows until the Tuesday reconciliation.
Nobody in this story made a mistake. The system asks every department to find out about an order by polling, exporting or being called at the wrong moment. This article argues for the opposite: when something happens, record it once as a fact, and let every other part of the business react to that fact on its own. That is event-driven architecture.
It is also the part of system design I have spent most of my career on, so expect opinions. The aim is that by the end you can tell a real event-driven system from a marketing claim, and know what to ask for.
Why nightly syncs and tangled calls fail
Most business software starts with a direct call. Checkout finishes, so the checkout code calls the stock code, then the accounting code, then the email code. It works in a demo and breaks in three predictable ways.
The slowest step owns the customer. If the email service takes eight seconds, the buyer waits eight seconds, or the order fails because email was down.
Every new need edits old code. Adding loyalty points means opening checkout, the most dangerous file in the company, and adding one more call.
Half-finished work has no owner. Stock was reduced, then accounting failed. Nothing records that the second step is still owed.
The common patch is a nightly sync: export orders at 2 a.m., import them elsewhere. That moves the failure to the hours when nobody is watching and makes every number up to 24 hours old. Teams then build reconciliation spreadsheets to chase the gaps, and the spreadsheet becomes the real system.
Commands and events: the vocabulary that matters
Two words do most of the work, and mixing them up is the usual first design error.
A command is a request: "place this order", "refund this payment". It is addressed to one owner, in the present tense, and it can be refused. An event is a fact: OrderPlaced, PaymentReceived, ParcelDelivered. It is named in the past tense, it already happened, and nobody can refuse it. The part of the system that owns the fact is the producer. Anything that cares is a listener (also called a consumer or subscriber).
The producer does not know who listens. That single property is the whole benefit. Adding loyalty points means writing one new listener for ParcelDelivered. Checkout is not touched, and it cannot break.
One order, three events, five listeners
An illustrative example makes this concrete. A customer buys two jackets, SKU JKT-M, at 40.00 each (unit cost 22.00), plus 5.00 shipping and 8.00 tax. The total is 93.00, paid by card. Over the next three days the order produces three events.
One fact, five independent reactions. The checkout never calls any of them.
Each event carries a small payload: enough for a listener to act without calling back to the producer.
stock_movements: type sale, qty -2, reservation released. On hand 46
ParcelDelivered
Ledger
Debit Customer deposits 93.00; credit Sales 80.00, Shipping income 5.00, Tax payable 8.00. Debit Cost of goods sold 44.00; credit Inventory 44.00
ParcelDelivered
Loyalty
loyalty_entries: customer 77, +80 points (1 per currency unit on goods), ref order 1042
ParcelDelivered
Notification
notifications: delivered message and review invitation, status queued
Check the arithmetic: the delivery journal has debits of 93.00 + 44.00 = 137.00 and credits of 80.00 + 5.00 + 8.00 + 44.00 = 137.00. It balances because the listener was written to post balanced entries, not because someone checked at month end. Revenue is recognised on delivery here; your accounting policy may differ, and that is a decision for the ledger listener alone.
The outbox: how an event is never lost
There is a trap in the producer. Placing an order means two writes: save the order in the database, and publish OrderPlaced to the queue. They are two different systems, so no single transaction covers both. If the database commits and the process dies before publishing, the order exists and the event does not. Stock never hears about it. If you publish first and the database rolls back, the world reacts to an order that was never saved. This is the dual-write problem.
The fix is the transactional outbox, described by Chris Richardson in his microservices pattern catalogue. The order and its event are written to the same database in the same transaction, the event into an outbox table. Either both rows exist or neither does. A separate relay process reads unsent outbox rows, publishes them to the queue and marks them sent.
One transaction, two rows. The relay carries the second row to the queue, as many times as it takes.
The relay can also crash between publishing and marking the row sent. On restart it publishes the event again. So the outbox buys you something precise: an event is never lost, and it may be delivered more than once. Which leads to the next point.
"Exactly once" is at least once, plus idempotent listeners
Vendors like to promise exactly-once delivery. Between separate machines connected by a network, you cannot have it in general: a sender that gets no acknowledgement cannot know whether the message arrived, so it must choose between sending again and risking a loss. Reliable systems choose to send again. Payment providers say it openly; Stripe's documentation, for example, tells you to expect the same webhook event more than once and to deduplicate by event ID.
So the working guarantee is at least once delivery plus idempotent handling. A listener is idempotent when processing the same event twice has the same effect as processing it once. The usual mechanism is small: a table processed_events with a unique key on (listener, event_id).
What a listener does with every event, including the one it has already seen.
Two details decide whether this works. First, the "apply the change" and "record the event ID" steps must commit in the same database transaction. If you apply and then record in two steps, a crash in the gap gives you a double posting on the retry. Second, the unique key must make the check atomic, so two copies of the same event arriving at once cannot both pass.
Run through the example: if ParcelDelivered arrives twice, the second Ledger run finds evt_5003 already recorded and skips. The journal is not posted twice, and the 80 loyalty points are not 160.
Ordering, retries and the history you get for free
Ordering
Events about different orders can be processed in any order. Events about the same order cannot: a RefundIssued processed before its PaymentReceived makes no sense. Two defences work together. Route events by order ID so one order's events travel through one lane in sequence, and give each event a per-order sequence number so a listener can recognise one that is too old and park or ignore it.
Retries and the dead-letter queue
When a listener fails, because the database was busy or a courier API was down, the event goes back for another try with increasing delays, for example 10 seconds, 1 minute, 5 minutes, 30 minutes, 2 hours. Backoff matters: instant retries turn a short outage into a self-inflicted overload. After a fixed number of tries, say five, the event moves to a dead-letter queue and someone is alerted. A dead letter is visible and recoverable. Compare that with a failed nightly export, which is simply absent.
The history
Because every business change is an event with a time, an ID and a payload, you get an audit trail without building one. "Why does this customer have 80 points?" has an answer: ParcelDelivered evt_5003, handled by Loyalty at a stated time. Keep events for a defined period, and when a listener has a bug, fix it and replay the affected events into it. Idempotency makes the replay safe.
When not to use it
Event-driven design adds moving parts: a queue, a relay, workers, monitoring and a habit of thinking in eventual consistency. For a brochure site, a single-user tool or a plain CRUD admin panel where one table is updated and one screen reads it, a direct function call is simpler and better. The same goes for a step that needs an answer now, such as "is this card valid?". That is a command with a reply, not an event.
The threshold I use: once an action has three or more independent consequences owned by different teams or modules (stock, money, messages, delivery), events start to pay for themselves.
Be honest about the cost too. Listeners run a moment after the order, so a screen that reads stock immediately after checkout can briefly show the old number. Good products hide this with reservations made inside the checkout transaction, and by showing "processing" rather than guessing.
Getting there, and how to tell it works
Steps
List the facts your business already talks about: order placed, payment received, parcel delivered, return approved, stock counted. Name each in the past tense.
Define each payload with enough fields that a listener never calls back, and give every event a unique ID and a timestamp.
Write events to an outbox table inside the same transaction as the change they describe.
Run a relay that publishes outbox rows to a queue and marks them sent.
Build each listener idempotent, with the processed-events record in the same transaction as its writes.
Add retries with backoff, a dead-letter queue and an alert on it before the first real order.
Move consumers over one at a time: stock first, then the ledger, then messages and loyalty. Retire the nightly job last, after a week of matching numbers.
Questions for a vendor
Is an event written in the same transaction as the business change, or published afterwards?
What happens if the same event is delivered twice? Ask them to show the deduplication key.
Where do failed events go, who is alerted, and can I replay them?
Can I see the event history for one order, with timestamps and which handler processed each?
One way to check these claims on a real system: a self-hosted platform such as StoreConsole handles its inventory, accounting and loyalty modules as separate listeners on the same order events.
What to measure
Outbox lag: age of the oldest unsent row. Seconds is healthy.
Queue depth and handling time per listener.
Dead-letter count: should be near zero and never silently growing.
Duplicate-skip rate: a nonzero number proves deduplication is working.
Daily drift check: units in the stock ledger against units implied by orders, and ledger revenue against order revenue. Both differences should be zero.
The same Monday, rebuilt
Return to the finance lead with 1,900 orders. Each order wrote its event in the same transaction as the sale, so there are 1,900 outbox rows and none are lost. At order 1,412 the ledger listener hit a timeout. Orders 1,412 to 1,900 waited in the queue, retried with backoff and posted within minutes. Two orders failed five times because of a bad tax code, landed in the dead-letter queue and raised an alert on Saturday afternoon, not on Tuesday.
The Saturday sale at the neighbouring outlet reserved stock the moment it happened, so the 14 jackets were really 14. The Monday drift check shows zero in both columns. The reconciliation spreadsheet is not opened.
Key takeaways
A command asks and can be refused; an event states a fact that already happened. Producers announce events and never know who listens.
Write the event in the same transaction as the business change (the outbox), so it is never lost.
Exactly-once delivery is not a realistic promise. Build at-least-once delivery with idempotent listeners.
Record the processed event ID in the same transaction as the listener's writes.
Retry with backoff, then dead-letter and alert. A visible failure beats a silent gap.
Use it when an action has three or more independent consequences. Skip it for simple CRUD.
Anichur Rahaman is a software architect and the creator of StoreConsole. He designs commerce and ERP systems for growing businesses, with a focus on event-driven architecture, data integrity and self-hosted operations.