One self-hosted console to run your entire business — commerce, ERP, HRM, CRM & manufacturing

Event-Driven Architecture Explained: Why Your Order Should Tell Your Stock and Books Itself

Nightly syncs and tangled calls leave stock and accounts out of step. See how events, the transactional outbox and idempotent listeners let one order update stock, ledger and loyalty reliably, with a worked example.

Author

Anichur Rahaman

2 months ago11 min read4 views
Event-Driven Architecture Explained: Why Your Order Should Tell Your Stock and Books Itself

Over the sale weekend, a six-outlet retailer took 1,900 orders. On Monday morning its finance lead opens the reports, and the stock sheet says 14 jackets are left, but the outlet in the next town sold the last of them on Saturday. The ledger is worse: revenue is posted by a nightly batch, the batch timed out at order 1,412, and 488 orders are missing from the books. Nobody knows until the Tuesday reconciliation.

Nobody in this story made a mistake. The system asks every department to find out about an order by polling, exporting or being called at the wrong moment. This article argues for the opposite: when something happens, record it once as a fact, and let every other part of the business react to that fact on its own. That is event-driven architecture.

It is also the part of system design I have spent most of my career on, so expect opinions. The aim is that by the end you can tell a real event-driven system from a marketing claim, and know what to ask for.

Why nightly syncs and tangled calls fail

Most business software starts with a direct call. Checkout finishes, so the checkout code calls the stock code, then the accounting code, then the email code. It works in a demo and breaks in three predictable ways.

  • The slowest step owns the customer. If the email service takes eight seconds, the buyer waits eight seconds, or the order fails because email was down.
  • Every new need edits old code. Adding loyalty points means opening checkout, the most dangerous file in the company, and adding one more call.
  • Half-finished work has no owner. Stock was reduced, then accounting failed. Nothing records that the second step is still owed.

The common patch is a nightly sync: export orders at 2 a.m., import them elsewhere. That moves the failure to the hours when nobody is watching and makes every number up to 24 hours old. Teams then build reconciliation spreadsheets to chase the gaps, and the spreadsheet becomes the real system.

Commands and events: the vocabulary that matters

Two words do most of the work, and mixing them up is the usual first design error.

A command is a request: "place this order", "refund this payment". It is addressed to one owner, in the present tense, and it can be refused. An event is a fact: OrderPlaced, PaymentReceived, ParcelDelivered. It is named in the past tense, it already happened, and nobody can refuse it. The part of the system that owns the fact is the producer. Anything that cares is a listener (also called a consumer or subscriber).

The producer does not know who listens. That single property is the whole benefit. Adding loyalty points means writing one new listener for ParcelDelivered. Checkout is not touched, and it cannot break.

One order, three events, five listeners

An illustrative example makes this concrete. A customer buys two jackets, SKU JKT-M, at 40.00 each (unit cost 22.00), plus 5.00 shipping and 8.00 tax. The total is 93.00, paid by card. Over the next three days the order produces three events.

Diagram of one OrderPlaced event fanning out to stock, ledger, loyalty, notifications and delivery listeners
One fact, five independent reactions. The checkout never calls any of them.

Each event carries a small payload: enough for a listener to act without calling back to the producer.

EventPayload fields
OrderPlacedevent_id evt_5001, order_id 1042, customer_id 77, lines [JKT-M, qty 2, price 40.00, cost 22.00], shipping 5.00, tax 8.00, total 93.00, USD, location WH-1, occurred_at
PaymentReceivedevent_id evt_5002, order_id 1042, payment_id 9001, amount 93.00, method card, occurred_at
ParcelDeliveredevent_id evt_5003, order_id 1042, delivery_id 311, delivered_at, lines [JKT-M, qty 2]

Now the part that matters to the owner: which rows exist afterwards, and who wrote them.

EventListenerRows written
OrderPlacedStockstock_movements: JKT-M, WH-1, type reserve, qty 2, ref order 1042. On hand stays 48, reserved +2, available 46
OrderPlacedNotificationnotifications: order 1042 confirmation to customer 77, status queued
PaymentReceivedLedgerDebit Card clearing 93.00; credit Customer deposits 93.00
ParcelDeliveredStockstock_movements: type sale, qty -2, reservation released. On hand 46
ParcelDeliveredLedgerDebit Customer deposits 93.00; credit Sales 80.00, Shipping income 5.00, Tax payable 8.00. Debit Cost of goods sold 44.00; credit Inventory 44.00
ParcelDeliveredLoyaltyloyalty_entries: customer 77, +80 points (1 per currency unit on goods), ref order 1042
ParcelDeliveredNotificationnotifications: delivered message and review invitation, status queued

Check the arithmetic: the delivery journal has debits of 93.00 + 44.00 = 137.00 and credits of 80.00 + 5.00 + 8.00 + 44.00 = 137.00. It balances because the listener was written to post balanced entries, not because someone checked at month end. Revenue is recognised on delivery here; your accounting policy may differ, and that is a decision for the ledger listener alone.

The outbox: how an event is never lost

There is a trap in the producer. Placing an order means two writes: save the order in the database, and publish OrderPlaced to the queue. They are two different systems, so no single transaction covers both. If the database commits and the process dies before publishing, the order exists and the event does not. Stock never hears about it. If you publish first and the database rolls back, the world reacts to an order that was never saved. This is the dual-write problem.

The fix is the transactional outbox, described by Chris Richardson in his microservices pattern catalogue. The order and its event are written to the same database in the same transaction, the event into an outbox table. Either both rows exist or neither does. A separate relay process reads unsent outbox rows, publishes them to the queue and marks them sent.

Diagram of the outbox pattern: one database transaction writes the order and an outbox row, a relay publishes to the queue, idempotent listeners consume
One transaction, two rows. The relay carries the second row to the queue, as many times as it takes.

The relay can also crash between publishing and marking the row sent. On restart it publishes the event again. So the outbox buys you something precise: an event is never lost, and it may be delivered more than once. Which leads to the next point.

"Exactly once" is at least once, plus idempotent listeners

Vendors like to promise exactly-once delivery. Between separate machines connected by a network, you cannot have it in general: a sender that gets no acknowledgement cannot know whether the message arrived, so it must choose between sending again and risking a loss. Reliable systems choose to send again. Payment providers say it openly; Stripe's documentation, for example, tells you to expect the same webhook event more than once and to deduplicate by event ID.

So the working guarantee is at least once delivery plus idempotent handling. A listener is idempotent when processing the same event twice has the same effect as processing it once. The usual mechanism is small: a table processed_events with a unique key on (listener, event_id).

Flowchart of an idempotent listener: event arrives, already processed check, skip or apply change, record event id, retry with backoff, dead-letter queue and alert
What a listener does with every event, including the one it has already seen.

Two details decide whether this works. First, the "apply the change" and "record the event ID" steps must commit in the same database transaction. If you apply and then record in two steps, a crash in the gap gives you a double posting on the retry. Second, the unique key must make the check atomic, so two copies of the same event arriving at once cannot both pass.

Run through the example: if ParcelDelivered arrives twice, the second Ledger run finds evt_5003 already recorded and skips. The journal is not posted twice, and the 80 loyalty points are not 160.

Ordering, retries and the history you get for free

Ordering

Events about different orders can be processed in any order. Events about the same order cannot: a RefundIssued processed before its PaymentReceived makes no sense. Two defences work together. Route events by order ID so one order's events travel through one lane in sequence, and give each event a per-order sequence number so a listener can recognise one that is too old and park or ignore it.

Retries and the dead-letter queue

When a listener fails, because the database was busy or a courier API was down, the event goes back for another try with increasing delays, for example 10 seconds, 1 minute, 5 minutes, 30 minutes, 2 hours. Backoff matters: instant retries turn a short outage into a self-inflicted overload. After a fixed number of tries, say five, the event moves to a dead-letter queue and someone is alerted. A dead letter is visible and recoverable. Compare that with a failed nightly export, which is simply absent.

The history

Because every business change is an event with a time, an ID and a payload, you get an audit trail without building one. "Why does this customer have 80 points?" has an answer: ParcelDelivered evt_5003, handled by Loyalty at a stated time. Keep events for a defined period, and when a listener has a bug, fix it and replay the affected events into it. Idempotency makes the replay safe.

When not to use it

Event-driven design adds moving parts: a queue, a relay, workers, monitoring and a habit of thinking in eventual consistency. For a brochure site, a single-user tool or a plain CRUD admin panel where one table is updated and one screen reads it, a direct function call is simpler and better. The same goes for a step that needs an answer now, such as "is this card valid?". That is a command with a reply, not an event.

The threshold I use: once an action has three or more independent consequences owned by different teams or modules (stock, money, messages, delivery), events start to pay for themselves.

Be honest about the cost too. Listeners run a moment after the order, so a screen that reads stock immediately after checkout can briefly show the old number. Good products hide this with reservations made inside the checkout transaction, and by showing "processing" rather than guessing.

Getting there, and how to tell it works

Steps

  1. List the facts your business already talks about: order placed, payment received, parcel delivered, return approved, stock counted. Name each in the past tense.
  2. Define each payload with enough fields that a listener never calls back, and give every event a unique ID and a timestamp.
  3. Write events to an outbox table inside the same transaction as the change they describe.
  4. Run a relay that publishes outbox rows to a queue and marks them sent.
  5. Build each listener idempotent, with the processed-events record in the same transaction as its writes.
  6. Add retries with backoff, a dead-letter queue and an alert on it before the first real order.
  7. Move consumers over one at a time: stock first, then the ledger, then messages and loyalty. Retire the nightly job last, after a week of matching numbers.

Questions for a vendor

  • Is an event written in the same transaction as the business change, or published afterwards?
  • What happens if the same event is delivered twice? Ask them to show the deduplication key.
  • Where do failed events go, who is alerted, and can I replay them?
  • Can I see the event history for one order, with timestamps and which handler processed each?

One way to check these claims on a real system: a self-hosted platform such as StoreConsole handles its inventory, accounting and loyalty modules as separate listeners on the same order events.

What to measure

  • Outbox lag: age of the oldest unsent row. Seconds is healthy.
  • Queue depth and handling time per listener.
  • Dead-letter count: should be near zero and never silently growing.
  • Duplicate-skip rate: a nonzero number proves deduplication is working.
  • Daily drift check: units in the stock ledger against units implied by orders, and ledger revenue against order revenue. Both differences should be zero.

The same Monday, rebuilt

Return to the finance lead with 1,900 orders. Each order wrote its event in the same transaction as the sale, so there are 1,900 outbox rows and none are lost. At order 1,412 the ledger listener hit a timeout. Orders 1,412 to 1,900 waited in the queue, retried with backoff and posted within minutes. Two orders failed five times because of a bad tax code, landed in the dead-letter queue and raised an alert on Saturday afternoon, not on Tuesday.

The Saturday sale at the neighbouring outlet reserved stock the moment it happened, so the 14 jackets were really 14. The Monday drift check shows zero in both columns. The reconciliation spreadsheet is not opened.

Key takeaways

  • A command asks and can be refused; an event states a fact that already happened. Producers announce events and never know who listens.
  • Write the event in the same transaction as the business change (the outbox), so it is never lost.
  • Exactly-once delivery is not a realistic promise. Build at-least-once delivery with idempotent listeners.
  • Record the processed event ID in the same transaction as the listener's writes.
  • Retry with backoff, then dead-letter and alert. A visible failure beats a silent gap.
  • Use it when an action has three or more independent consequences. Skip it for simple CRUD.

Anichur Rahaman is a software architect and the creator of StoreConsole. He designs commerce and ERP systems for growing businesses, with a focus on event-driven architecture, data integrity and self-hosted operations.

About the Author

Anichur Rahaman

Continue Reading