NULLBIT
NULLBIT
Blog
Author: ALFRED

Save Months of Migrations with Event Sourcing for Microservices

Microservices event sourcing: architecture and operations, with outbox, idempotent inbox, snapshots, and upcasters to avoid months of migrations.

Save Months of Migrations with Event Sourcing for Microservices

Save Months of Migrations with Event Sourcing for Microservices

Event sourcing title card with ledger sketches

Event sourcing stores every state change as an immutable, append-only event log, and current state becomes a deterministic projection of that log rather than a row you overwrite. The payoff is full auditability and the ability to replay history, but it comes with real operational overhead. It fits best in audit-heavy, regulated systems and workflows with complex business invariants, not in simple CRUD apps.


TL;DR:

  • Reserve event sourcing for aggregates that need audit trails, historical reconstruction, or complex invariants; use change data capture or ordinary event streaming for simple CRUD.
  • Projections can lag briefly after writes, so route read after write requests to the primary store or update projections near real time.
  • Consumers must record processed event IDs in an inbox because delivery can repeat; route unprocessable messages to a queue for later review.
  • Use upcasters to convert older event versions at read time, with compatible serializers and a schema registry to keep old logs readable.
  • Projection rebuilds can overload the event store and downstream services, so schedule and monitor replays, using snapshots and throttling to control load.

Nullbit
Build More Resilient Microservices
Nullbit develops custom software and AI solutions to help businesses address operational inefficiency and build scalable digital systems.
Explore Nullbit’s solutions

Table of Contents

What are the core concepts behind event sourcing?

An event is a past-tense, immutable fact about something that already happened: OrderPlaced, PaymentCaptured, InventoryReserved. Because the event already occurred, you never edit or delete it. If a mistake needs correcting, you append a compensating event instead, which keeps the log trustworthy as a historical record.

A correction appended to an unchanged event history

The event store is the system of record. It guarantees append-only writes, strict ordering per aggregate, and durability, which is why the event sourcing pattern as documented by Microsoft treats the store, not a mutable table, as the source of truth.

A few building blocks do the actual work:

  • Aggregates and command handlers validate incoming commands against current state and emit new events when a command is accepted.
  • Projections (read models) replay or subscribe to events and build the query-friendly views your application actually reads from.
  • Snapshots periodically freeze an aggregate’s state so replays do not have to start from event zero every time.

Projections can be rebuilt at any time because they are just derived data, which is what makes adding a new report or dashboard later a non-destructive operation.

Pro Tip: Keep your event schema and your read-model schema completely separate; coupling them is the fastest way to turn a simple projection change into a full migration.

How do you architect an event-sourced system?

The event store you pick determines your ordering and durability guarantees. Options range from specialized event stores to Kafka-based logs to DynamoDB-based patterns, and the AWS guidance on building a CQRS event store with DynamoDB treats per-aggregate ordering as a first-class requirement, not an afterthought.

A workable architecture generally includes:

  1. An event store that enforces append-only writes and strict ordering within each aggregate stream.
  2. An event bus or broker that fans events out to projections, analytics, and other services, separate from direct reads against the store.
  3. Sagas or process managers that coordinate multi-step workflows and cross-aggregate consistency without forcing distributed transactions.
  4. Sharding by aggregate ID, enforcing single-writer-per-aggregate semantics so concurrent commands cannot corrupt a stream.
  5. Observability: distributed tracing across event flows, replay monitoring, retention policies, and secured, tested backups.

Each of these is a correctness guarantee, not a convenience. Skip the ordering guarantee on your event store and projections drift silently out of sync with reality.

What benefits justify the added complexity?

The core payoff is a complete, forensic audit trail. Every change to every aggregate is recorded permanently, so you can answer “what happened and why” months after the fact without relying on application logs.

  • Time travel: you can reconstruct the exact state of any aggregate at any past point by replaying events up to that moment.
  • New projections without data loss: adding a new analytics view or report means writing a new projection and replaying history, never backfilling a mutable table.
  • Richer domain modeling: complex business logic, multi-step compensations, and workflow corrections map naturally onto a sequence of discrete events rather than a single overwritten record.
  • Debugging power: when a bug corrupts state, you can trace the exact sequence of events that produced it, something a CRUD table’s current-state snapshot cannot show you.

For regulated industries like finance, healthcare, and insurance, that audit trail is frequently the deciding factor over any performance consideration.

What challenges come with event sourcing in production?

Eventually consistent projections mean a user can write data and, for a brief window, read a stale view. Mitigations include near-real-time projection updates and routing read-after-write requests to the primary store rather than a lagging read model.

Event schemas also change over time, and old events do not disappear. The Akka documentation on schema evolution recommends upcasters that transform older event versions into the current shape at read time, paired with binary-compatible serializers like protobuf or Avro and a schema registry for long-lived logs.

  • Idempotency is mandatory: at-least-once delivery is the norm in distributed systems, so consumers need inbox-style deduplication to avoid processing the same event twice, a point the Microsoft event sourcing pattern treats as a baseline requirement, not an edge case.
  • Replay load spikes: rebuilding a projection from millions of events can hammer your infrastructure.
  • Testing and recovery complexity: replay tests and migration rehearsals need to be part of your release process, not an afterthought.

Replays can cause sharp, temporary load spikes on the event store and downstream projections; snapshotting and replay throttling are the standard mitigations. Treat a full rebuild as an operational event you schedule and monitor, not a background task you fire and forget.

Which implementation patterns prevent the classic failure modes?

Most event sourcing failures trace back to a handful of avoidable mistakes. These patterns close the gaps:

  1. Transactional outbox: write the domain change and the outgoing event in the same database transaction, then let a background publisher push from an outbox table to the broker. The AWS transactional outbox sample demonstrates exactly this shape, and it is the standard fix for the dual-write problem where a database commit succeeds but the broker publish fails.
  2. Optimistic concurrency: enforce a unique constraint on (aggregate_id, version). A write conflict triggers a retry, giving single-writer semantics per aggregate without locking.
  3. Snapshotting: take a snapshot every N events (hundreds, not thousands, for hot aggregates), and verify snapshot consistency against the underlying event stream periodically.
  4. Schema evolution: use upcasters plus an IDL format like protobuf or Avro and a schema registry so old events keep deserializing correctly years later.
  5. Idempotent consumers: back every consumer with an inbox table that records processed event IDs, and route poison messages to a dead-letter queue rather than retrying forever.

Pro Tip: Write a replay test that rebuilds a projection from scratch and diffs it against the live one before every schema change ships; it catches upcaster bugs long before production does.

When should you choose event sourcing, and when should you avoid it?

Event sourcing earns its complexity when the domain demands it, not by default.

  • Choose it when: you operate under audit or regulatory requirements, your domain has complex invariants that span multiple steps, or you genuinely need to replay history to reconstruct past state or feed new analytics.
  • Avoid it when: your domain is simple CRUD with no audit requirement; in that case, change data capture (CDC) or plain event streaming without full event sourcing solves integration needs with far less operational weight, a point AWS’s CQRS and event store guidance makes directly.
  • Consider a hybrid: apply event sourcing selectively to the one or two aggregates where it pays off, and use CDC or materialized views for everything else.

How does event sourcing fit with CQRS and microservices?

CQRS, splitting reads and writes into separate models, is orthogonal to event sourcing. You can run CQRS against a traditional write store with no event log at all, and you can run event sourcing without ever separating your read and write paths, though the two are frequently paired because projections are a natural fit for the “query” side.

In a microservices architecture, a few patterns recur:

  • Per-bounded-context event stores: each service owns its own event stream and projections; no service reaches into another’s store directly.
  • Cloud-managed logs versus self-hosted stores: a managed log reduces operational burden but hands you someone else’s durability and ordering guarantees, while self-hosting gives control at the cost of running the infrastructure yourself.
  • Outbox and inbox as the connective tissue: outbox on the publishing side, inbox with deduplication on the consuming side, with retry and backoff plus a dead-letter queue for messages that never process cleanly.
  • Sagas for orchestration: when a workflow spans services, a saga coordinates the steps instead of forcing a distributed transaction.

Teams designing for provenance and traceability outside of software, such as the approach described in AuksPort’s scouting database design, face the same core problem: deciding what counts as an immutable fact worth keeping forever versus a derived view you can safely regenerate.

What does practical experience with event-sourced systems look like?

Operational readiness matters more than the pattern itself. In reviewing systems where reliability and auditability were non-negotiable, like catching a misconfigured database before it becomes a breach, the deciding factor was always whether backups, retention, and replay paths had been tested, not just designed. The same held for consolidating a dozen disconnected data sources into one forecasting model, where projection design determined whether new reports could ship without touching history.

— Matija

How we help teams design and ship event-sourced systems

Getting the operational patterns right the first time, outbox, idempotent consumers, snapshotting, schema evolution, saves months of painful migrations later. We build custom software and review architectures with exactly these failure modes in mind, and we design cloud infrastructure that supports managed or self-hosted event stores depending on your durability and ordering requirements.

Nullbit

Our relevant services include:

  • System engineering for architecture reviews and integration design around event-sourced or hybrid systems.
  • Proof-of-concept development, from 5,000 EUR one-off, to validate an event sourcing approach before committing to a full build.
  • Custom software development for the aggregates, projections, and sagas your domain actually needs.
  • Cloud and infrastructure work to stand up or migrate a managed or self-hosted event store.

If you are weighing whether event sourcing fits your next system, start with a cooperation engagement and we will help you pressure-test the decision before any code gets written.

FAQ

What is the main difference between event sourcing and CQRS?

Event sourcing is a persistence pattern: it stores state as an append-only log of events. CQRS is an architectural separation of read and write models, and the two are often paired but neither requires the other.

Is event sourcing the same as event-driven architecture?

No. Event-driven architecture describes services communicating through events, while event sourcing specifically means using an event log as the system of record for an aggregate’s state. You can have one without the other.

How do you handle schema changes in an event-sourced system?

Use upcasters to transform older event versions into the current shape at read time, paired with a binary-compatible serializer like protobuf or Avro and, ideally, a schema registry, as the Akka schema evolution guide recommends for long-lived event logs.

Does event sourcing work for simple CRUD applications?

Generally no. The operational overhead of event stores, projections, and replay tooling outweighs the benefit when there is no audit or replay requirement, and practitioner guidance on event sourcing recommends CDC or plain event streaming instead for those cases.

What causes most production failures in event-sourced systems?

Non-idempotent consumers processing duplicate events and unhandled schema drift between old and new event versions are the most common causes. The Microsoft event sourcing pattern documentation treats idempotent handlers and versioning discipline as baseline requirements, not optional hardening.

Sources

Tags
event sourcing pattern
Stay ahead of the competition

Exclusive insights that drive change.

Get access to proven methodologies for digital growth, AI tool implementation, and AI product development.

  • Weekly digital strategy analyses
  • Advanced insights into AI trends and technology solutions

Your privacy is a priority. You can unsubscribe at any time.