Benefits of event-driven architecture, and what they cost

Event-driven architecture is a way of connecting the parts of a system, or several systems, so that one part announces that something has happened and the others react to it, instead of each part calling the next and waiting for an answer. The benefits of event-driven architecture are real: the parts depend on each other less, a failure in one does not stop the rest, peaks in demand are absorbed and new features can be added without reopening old code. The price is a system that is harder to follow and to test. For most business software the right amount is some, in the places where it earns its keep.

What event-driven architecture means

An event is a record that something happened: “order 1042 was placed”, “invoice 77 was paid”, “a new customer signed up”.

In a request-driven design, the ordering screen calls the stock system, then the invoicing system, then the email service, one after another, and waits for each to answer. If the email service is slow, the customer waits. If the invoicing system is down, the order fails.

In an event-driven design, the ordering part saves the order and publishes one event, “order placed”. It does not know who is listening. Stock, invoicing and email each subscribe to that event and do their own work in their own time.

The vocabulary is short:

  • Producer: the part that publishes the event.
  • Consumer: a part that subscribes to the event and reacts to it.
  • Broker: the piece of infrastructure in the middle that receives events, holds them in a queue and delivers them to each consumer. Azure Service Bus, RabbitMQ and Apache Kafka are well-known examples.

The benefits of event-driven architecture

The parts are less tied together. The producer does not need to know which systems consume its events. When the business later wants a text message sent for every order, or a feed into a reporting database, a new consumer is added and the ordering code is left alone. Developers call this loose coupling, and it matters most when the systems belong to different teams or suppliers.

One failure does not stop everything. If the invoicing system is down for an hour, its events wait in the queue and are processed when it returns. The customer placing the order sees nothing wrong. In a request-driven design the same outage appears as an error on the screen.

Peaks are absorbed. A queue acts as a buffer. If a thousand orders arrive in ten minutes, the consumers work through them at the pace they can manage. Some work finishes later than usual, but nothing is lost and nothing falls over.

Users wait less. Slow jobs, such as producing a PDF, calling a courier’s system or sending email, move out of the user’s request. The screen responds as soon as the order is saved.

Systems stop polling each other. Without events, the usual way for one system to learn about changes in another is to ask at intervals, or to exchange a file overnight. With events, information moves when it changes and not at the next scheduled run.

There can be a record of what happened. If events are stored, they form a log of what happened and when, which helps with audit questions and with reporting. That is a decision someone has to make: many brokers remove a message once it has been processed.

What it costs

Each of those benefits has a cost attached, and a proposal that lists only the benefits is incomplete.

  • The flow is harder to see. No single piece of code shows everything that happens when an order is placed. Finding out why one order was never invoiced means following a message through several systems, so good logging and monitoring become essential.
  • Data is briefly out of step. For a few seconds, sometimes longer, the order exists and the stock figure has not yet changed. Developers call this eventual consistency. Screens and staff have to tolerate it, and some processes cannot.
  • Messages can arrive twice, or out of order. Most brokers promise to deliver each message at least once, which means occasionally more than once. Every consumer has to be written so that handling the same event twice does no harm. If it is not, a customer receives two invoices.
  • Failures become quiet. A message that cannot be processed is set aside in what is called a dead-letter queue. Nothing appears on anyone’s screen. Unless somebody is alerted and is responsible for looking, the failed work never happens.
  • There is more to run. The broker is another component to pay for, secure, monitor and keep up to date, and the team needs people who understand it.

Where it fits, and where it does not

SituationDoes event-driven fit?
Several systems need to react to the same business eventYes
Work that is slow or depends on a third party, such as emails, documents or courier bookingsYes
Demand arrives in burstsYes
Systems from different suppliers that exchange overnight files todayYes, one connection at a time
The user needs an answer now: a price, a stock check, a loginNo. Call and wait
Steps that must all succeed or all fail together, such as moving money between two accountsNo. Keep them in one database transaction
A small application with one database and one teamRarely worth it

Event-driven architecture is often proposed in the same breath as microservices, where an application is split into many small services. The two are separate decisions, and our article on microservices and web services explains the second. A single well-structured application can publish events at the handful of points where it helps and stay simple everywhere else.

Adding events to a system you already have

None of this requires a rebuild. The benefits come from individual connections, so they can be introduced into an existing system one at a time.

  • Background jobs. Move one slow task, such as sending email, behind a queue. This is the smallest useful step, and users notice it straight away.
  • Outgoing events. Have the existing system publish an event when something important changes, so that other systems can stop reading its database directly or waiting for an export. A common technique is the outbox: the event is saved in the same database transaction as the change itself and published a moment later, so the two cannot disagree.
  • Webhooks. Many packaged and cloud products can notify you when something happens by calling a web address you give them. Using that in place of a scheduled check is event-driven integration with very little new infrastructure.

For a system on the Microsoft stack hosted in Azure, Azure Service Bus is the usual broker for this kind of work. Apache Kafka is built for very high volumes of streaming data and is more than most business systems need.

If the underlying problem is that the same data is typed into several systems, start with our page on systems that do not talk to each other. Our system integration package connects two systems at a fixed price, with retries, alerts and a log of every transfer.

Questions to ask before you agree to it

If a developer or supplier proposes an event-driven design, these questions tell you how carefully it has been thought through:

  1. Which specific problem does it solve here, and what is the simpler alternative?
  2. What happens when a message cannot be processed, and who is told?
  3. What happens if the same event is delivered twice?
  4. How will we trace one order from start to finish when something goes wrong?
  5. Which parts will stay as direct calls, and why?
  6. What does the broker cost to run each month, and who keeps it patched?

Good answers are specific and mention monitoring. If the answers are vague, ask for the design to be tried on one connection first, and judge it by how easy that connection is to support a few months later.

Tell us about your system

Say what it does, what it is built on and what is worrying you. We will reply with what we would look at first and whether we are the right people to help.

Tell us about your system 0800 433 7990 Monday to Friday, 9am to 5pm. A first 20-minute call is free, and we reply to every enquiry within one working day. What happens after you get in touch