MDM #16. Real-Time Master Data Architecture: CDC, Events, Consistency and Reconciliation
Real-time Master Data Management is often described too simply.
The usual architecture diagram looks like this:
→ Event Stream
→ Every System Updates Immediately
That diagram hides most of the difficult engineering decisions.
Should every master-data change really propagate immediately?
What happens when the same event is delivered twice?
What if version 103 reaches a consuming system before version 102?
What happens when one downstream application is unavailable for six hours?
How does the enterprise prove that all critical consumers eventually converged on the same mastered state?
And where does a broader concept such as Data Fabric actually fit?
Real-time MDM is not about making every master-data update instantaneous. It is about delivering the right mastered state to the right consumer within the business latency the process actually requires—and being able to recover when distributed systems disagree.
This article examines that architecture from a production perspective: Change Data Capture, business events, event contracts, ordering, idempotency, replay, eventual consistency and reconciliation.
Real-Time MDM and Data Fabric Are Not the Same Thing
The first distinction is conceptual.
A Data Fabric is broader than a real-time integration pipeline.
It is an architectural approach for connecting distributed enterprise data through capabilities such as integration, active metadata, governance, cataloging, security and multiple forms of data access.
A Data Fabric can use:
- batch integration,
- real-time integration,
- Change Data Capture,
- APIs,
- streaming,
- data virtualization,
- metadata, and
- governance services.
Real-time Master Data distribution can therefore be one capability inside a broader Data Fabric architecture.
It should not be used as a synonym for Data Fabric itself.
MDM and Data Fabric Solve Different Problems
| Capability | Primary Question | Typical Responsibility |
|---|---|---|
| MDM | What is the authoritative business entity? | Matching, mastering, hierarchy, stewardship and lifecycle control |
| Data Fabric | How can distributed enterprise data be discovered, governed, connected and consumed? | Metadata, integration, governance, security and access |
| Event Platform | How should state changes be communicated asynchronously? | Event transport, retention, partitioning and subscription |
| CDC | How can database-level changes be captured? | Change detection and low-latency propagation from source logs |
These capabilities can work together, but one does not replace the others.
1. Start with Business Freshness, Not “Real Time”
The first architecture question should not be:
It should be:
before a business process experiences unacceptable risk?”
Different master-data attributes can have radically different freshness requirements.
| Example Change | Business Consequence of Staleness | Possible Distribution Pattern |
|---|---|---|
| Supplier blocked for compliance reason | New transactions may continue with a supplier that should be blocked. | Event-driven or direct validation at transaction time |
| Customer delivery address | Orders or service interactions may use stale information. | Low-latency event or API-based retrieval |
| Material safety classification | Logistics or compliance processes may operate on incorrect classification. | Controlled event-driven propagation |
| Internal product description | Limited operational consequence. | Micro-batch or scheduled batch may be sufficient |
| Historical hierarchy snapshot | Primarily analytical rather than operational impact. | Batch |
“Real time” should therefore be an outcome of business requirements, not an architectural default.
A Better Decision Model: Freshness × Consequence × Change Frequency
A practical distribution decision can evaluate at least four dimensions.
× Business Consequence of Staleness
× Change Frequency
× Distribution Complexity
=
Appropriate Synchronization Pattern
This avoids arbitrary rules such as assigning a fixed percentage of enterprise master data to real-time, near-real-time or batch processing.
The appropriate mix is company-, domain- and workflow-specific.
2. Separate Commands, State and Events
Real-time MDM becomes easier to reason about when three concepts are separated.
| Concept | Meaning | Example |
|---|---|---|
| Command | A request to perform a business action. | Propose supplier address change |
| State | The current authoritative representation after processing. | Supplier master version 184 |
| Event | A statement that something material has already happened. | SupplierAddressChanged |
A consumer should not normally interpret an event as a request to decide whether the master-data change should have been approved.
That governance decision should occur before the authoritative event is published.
A Governed Change Flow
↓
Validation & Entity Resolution
↓
MDM Governance / Approval
↓
Authoritative Master State Updated
↓
Business Event Published
↓
Consumers Update Local State
↓
Reconciliation Confirms Convergence
This architecture avoids distributing unapproved source-system changes as if they were mastered truth.
3. CDC Is Change Capture — Not Master Data Governance
Change Data Capture detects changes in a source data store and makes those changes available for downstream processing.
Log-based CDC tools such as Debezium can capture create, update and delete operations from database transaction logs with relatively low latency.
Debezium change events can contain information such as:
- operation type,
- source information,
- timestamp,
- before state,
- after state, and
- optional transaction metadata.
Debezium — Change Event Structure
This makes CDC extremely useful.
But CDC answers:
MDM must still answer:
Is the change valid?
Which value should survive?
Should it become authoritative?”
CDC therefore does not replace matching, survivorship, validation or stewardship.
Database CDC and Business Events Should Not Be Confused
A CDC event may say:
supplier_status = ACTIVE → BLOCKED
A business event might say:
masterId: SUP-001234
effectiveAt: 2026-09-26T08:40:00Z
reasonCode: COMPLIANCE_RESTRICTION
entityVersion: 184
The first describes a physical source change.
The second communicates a governed business fact.
Large enterprises often need both, but they should serve different purposes.
4. Publish Mastered Business Events, Not Raw Internal Schema Changes
A common design mistake is exposing database-change structures directly to every downstream consumer.
This tightly couples consumers to:
- MDM vendor schemas,
- table design,
- internal technical attributes,
- survivorship implementation, and
- platform-specific change formats.
A stronger enterprise pattern introduces a stable event contract.
↓
Event Translation Layer
↓
Stable Business Event
↓
ERP · CRM · Data Products · AI · Analytics
A Master Data Event Contract
A useful event should provide enough information to determine what happened and whether a consumer has already processed it.
{
"eventId": "evt-82f61a",
"eventType": "SupplierMasterChanged",
"eventVersion": "2.0",
"occurredAt": "2026-09-26T08:40:00Z",
"entity": {
"masterId": "SUP-001234",
"entityType": "Supplier",
"entityVersion": 184
},
"change": {
"changedAttributes": [
"supplierStatus"
]
},
"context": {
"changeRequestId": "CR-89443",
"source": "ENTERPRISE_MDM"
}
}
The exact schema will differ by organization.
The important design elements are the stable entity identifier, unique event identifier, event type, schema version, entity version and relevant business context.
Use AsyncAPI to Document Event Interfaces
The AsyncAPI Specification provides a protocol-agnostic machine-readable method for describing message-driven APIs across technologies such as Kafka, AMQP, MQTT and WebSockets.
This makes event interfaces governable in much the same way that OpenAPI helps govern synchronous HTTP APIs.
5. Event Delivery Does Not Mean Every Consumer Is Immediately Consistent
Distributed systems create a period during which different applications can legitimately have different versions of the same master entity.
This is eventual consistency.
Microsoft's event-driven architecture guidance explicitly notes that during this window different parts of the system may have different views of current state.
Microsoft Azure Architecture Center — Event-Driven Architecture
This is not automatically a defect.
It is an architectural tradeoff.
Define Consistency Requirements by Business Process
For each consuming workflow, the architecture should determine whether it needs:
| Consistency Requirement | Example |
|---|---|
| Transaction-Time Validation | Check a critical supplier status directly before creating a purchase transaction. |
| Low-Latency Local Copy | Customer-service platform maintains a recently synchronized profile. |
| Eventually Consistent Copy | Analytics or downstream planning platform updates asynchronously. |
| Scheduled Snapshot | Historical reporting receives a periodic mastered snapshot. |
The most critical processes may therefore combine events with synchronous validation rather than relying solely on replicated state.
6. Duplicate Delivery Must Be Assumed
A resilient distributed architecture should assume that a consumer may receive an event more than once.
Retries, network failures and broker recovery can all produce duplicate processing scenarios.
Microsoft's event-sourcing guidance notes that event delivery is commonly at-least-once and recommends idempotent consumers so that duplicate processing does not change the final outcome.
Microsoft Azure Architecture Center — Event Sourcing Pattern
The design principle is:
=
Same Final Business State
Use Event Identity and Entity Version Together
A consumer can use:
- eventId to detect exact duplicate messages, and
- entityVersion to determine whether an event represents newer or older entity state.
For example:
Incoming event version = 184
→ already applied / duplicate
Incoming event version = 185
→ apply update
Incoming event version = 183
→ stale event; do not overwrite newer state
This can be more robust than assuming network arrival order always reflects business order.
7. Ordering Is Usually Per Entity, Not Global
Enterprises sometimes ask for absolute ordering across every master-data event.
That is often unnecessary and expensive.
What usually matters is preserving meaningful order for changes to the same business entity.
=
Stable Master Entity ID
Using the master entity identifier as the routing or partitioning key can help keep related entity events together where the messaging technology supports that model.
Apache Kafka, for example, provides ordering within a partition rather than promising one universal business order across all partitions.
Producer idempotence can also reduce duplicate writes caused by certain retry scenarios.
Ordering Cannot Fix Every Business Dependency
Entity-level ordering does not solve relationships across entities.
For example:
SupplierAssignedToParent
SupplierAddedToPurchasingOrganization
These may involve multiple aggregate boundaries and systems.
The architecture may need:
- explicit dependencies,
- workflow orchestration,
- temporary unresolved states, or
- reconciliation after processing.
Message ordering alone is not a substitute for business-process design.
8. Replay Is a Core Recovery Capability
An event architecture becomes substantially more resilient when consumers can rebuild state after downtime or defects.
Consider a CRM integration that fails at 10:00 and is repaired at 14:00.
A fragile architecture asks:
A replay-capable architecture can instead restore the consumer from a known position or rebuild relevant state from retained events and authoritative snapshots.
↓
Repair Consumer
↓
Resume / Replay Events
↓
Validate Latest Entity Versions
↓
Reconcile with MDM
Replay Is Not the Same as Blindly Reprocessing Everything
Replay strategy should specify:
- retention period,
- consumer checkpoints,
- event schema compatibility,
- idempotency behavior,
- historical business rules, and
- reconciliation after replay.
If an old event is reprocessed under today's business rules without careful design, the reconstructed state may differ from the state that originally occurred.
9. Reconciliation Is the Safety Net of Real-Time MDM
Real-time messaging reduces propagation delay.
It does not eliminate the possibility of divergence.
Consumers can:
- miss events,
- reject invalid messages,
- apply events incorrectly,
- experience schema incompatibility,
- remain offline beyond event-retention periods, or
- contain local changes that conflict with mastered state.
For this reason, production MDM still needs reconciliation.
A Real-Time Architecture Still Needs Periodic Truth Checks
MDM Change → Event → Consumer Update
+
CONTROL PATH
Periodic Snapshot / Version Comparison
→ Detect Divergence
→ Repair Consumer State
This is not a contradiction.
Batch and real-time mechanisms solve different problems.
Events provide speed.
Reconciliation provides assurance.
10. A Reconciliation Contract Should Be Explicit
For each critical consumer, define:
| Control | Question |
|---|---|
| Scope | Which entities and attributes must match MDM? |
| Frequency | How quickly must divergence be detected? |
| Comparison | Will the consumer compare entity version, checksum or selected business values? |
| Repair | Can divergence be automatically corrected or must it be reviewed? |
| Evidence | How is reconciliation status recorded and reported? |
The Reconciliation Contract is a Digital Future & Strategy practitioner framework.
11. Event Schemas Must Be Governed Like APIs
Once dozens of systems subscribe to master-data events, event schemas become enterprise contracts.
A seemingly simple change such as renaming:
to:
can break consumers if not handled through a compatible contract evolution process.
Event governance should therefore address:
- schema ownership,
- schema versioning,
- backward compatibility,
- required versus optional fields,
- deprecated fields,
- consumer inventory, and
- retirement of old event versions.
AsyncAPI provides a machine-readable structure for describing message-driven interfaces and their channels and messages, making it useful for event contract documentation and governance.
12. Separate Event Version from Entity Version
These two versions answer different questions.
| Version | Meaning |
|---|---|
| eventVersion | Which schema or contract format does the event use? |
| entityVersion | Which chronological state of the master entity does the event represent? |
Confusing the two makes both compatibility and ordering harder to manage.
13. Observability Should Measure Business Freshness, Not Only Broker Health
A streaming platform can be technically healthy while business data is stale.
CPU, throughput and broker availability are not sufficient MDM metrics.
The enterprise should observe the end-to-end data path.
↓
MDM RECEIVED
↓
MASTERED STATE COMMITTED
↓
EVENT PUBLISHED
↓
CONSUMER RECEIVED
↓
CONSUMER APPLIED
↓
RECONCILIATION VERIFIED
Useful operational metrics include:
- end-to-end propagation latency,
- consumer lag,
- failed event rate,
- duplicate event rate,
- stale-event rejection rate,
- dead-letter volume,
- replay backlog, and
- reconciliation mismatch rate.
Measure Freshness Against an SLO
Instead of promising that every master update is “real time,” define service objectives for specific business contexts.
For example:
99.9% of approved changes available to designated procurement controls within defined latency.
Customer Profile:
Approved address changes propagated to designated service channels within the agreed freshness objective.
The actual thresholds should be derived from business risk and measured production capability rather than copied from generic benchmarks.
14. Batch Is Not Legacy by Definition
The goal of modern MDM should not be to eliminate batch processing.
Batch remains appropriate when:
- business freshness requirements are low,
- large historical datasets are moved efficiently,
- reconciliation is required,
- snapshots are needed for reporting, or
- source systems cannot support reliable event integration.
A mature architecture chooses the integration mode according to the workload.
| Pattern | Best Suited For | Key Tradeoff |
|---|---|---|
| Synchronous API | Current-state validation at the moment of a transaction | Consumer becomes dependent on service availability and latency |
| Business Event | Low-latency distribution of approved state changes | Requires eventual-consistency and recovery design |
| CDC | Capturing database changes from systems that do not publish business events | Physical data changes may not represent governed business semantics |
| Micro-Batch | Moderate freshness requirements with simpler operations | Higher latency than streaming |
| Batch | Bulk distribution, snapshots, history and reconciliation | Not suitable for time-sensitive operational changes |
15. A Production Real-Time MDM Architecture
ERP · CRM · Supplier Portal · Product Systems
↓
CHANGE INTAKE
API · Business Event · CDC · Controlled Batch
↓
MDM PROCESSING
Identity Resolution · Validation · Match · Merge · Workflow
↓
AUTHORITATIVE MASTER STATE
Golden Record · Hierarchy · Entity Version
↓
DISTRIBUTION SERVICES
Synchronous API · Business Events · Batch Snapshots
↓
CONSUMERS
ERP · CRM · SaaS · Data Products · Analytics · AI Agents
↓
CONTROL PLANE
Schema Governance · Observability · Replay · Reconciliation
The architecture separates master-data governance from transport technology.
Kafka, Event Hubs, Kinesis, Pulsar or another broker can transport messages.
They do not determine which supplier is authoritative.
16. Real-Time MDM Failure Patterns
A production architecture should explicitly test the failure path, not only the successful path.
Failure 1 — Duplicate Event
Scenario: Consumer receives entity version 184 twice.
Required behavior: Second processing has no additional business effect.
Failure 2 — Out-of-Order Event
Scenario: Version 185 is applied, then version 184 arrives.
Required behavior: Consumer rejects or safely ignores the stale state.
Failure 3 — Missing Event
Scenario: Consumer moves from version 182 directly to version 184.
Required behavior: Detect version gap where relevant, retrieve current authoritative state or replay missing history.
Failure 4 — Consumer Outage
Scenario: Downstream ERP interface remains offline for several hours.
Required behavior: Resume from checkpoint or replay retained events and then reconcile state.
Failure 5 — Schema Change
Scenario: Event publisher introduces an incompatible field change.
Required behavior: Contract compatibility controls block or isolate the incompatible version.
Failure 6 — MDM Processing Delay
Scenario: Source changes arrive rapidly but MDM matching or stewardship creates backlog.
Required behavior: Freshness monitoring distinguishes source ingestion from mastered-state publication.
Failure 7 — Consumer Divergence
Scenario: Event processing reports success but local business state is incorrect.
Required behavior: Reconciliation identifies mismatch and triggers repair.
17. A Real-Time MDM Readiness Checklist
Before moving a master-data domain to event-driven distribution, answer the following questions.
1. What business decision requires fresher master data?
2. What is the maximum acceptable staleness?
3. Which system is authoritative before the event is published?
4. Is the message a raw CDC event or a governed business event?
5. Does every event have a unique event ID?
6. Does the event carry a stable master entity ID?
7. Can consumers distinguish entity versions?
8. Is duplicate processing idempotent?
9. How is meaningful event ordering preserved?
10. What happens when a consumer misses an event?
11. Can events be replayed safely?
12. Is the event schema versioned and governed?
13. Can critical consumers reconcile their state against authoritative MDM?
14. Are latency and data freshness observed end to end?
15. Is batch retained where it provides simpler or stronger control?
The Real-Time Master Data Position
The strategic goal of real-time MDM is not instantaneous replication.
It is controlled propagation of authoritative business state.
This leads to several architectural principles.
First, not all master data should be real time.
Freshness should reflect business consequence.
Second, CDC and business events are different.
CDC captures physical changes. MDM determines authoritative business state.
Third, asynchronous systems are not instantly consistent.
Consumers must tolerate and control periods of divergence.
Fourth, duplicate and out-of-order messages are normal design conditions.
Idempotency and entity versioning should be deliberate.
Fifth, replay is necessary but not sufficient.
Reconciliation confirms that consumers actually converged.
Sixth, batch remains part of a modern architecture.
It is often the appropriate mechanism for snapshots, bulk processing and control checks.
The target architecture is therefore not:
→ Replace Everything with Streaming
It is:
× Fit-for-Purpose Distribution
× Stable Event Contracts
× Idempotent Processing
× Version Control
× Replay
× Reconciliation
=
Reliable Real-Time MDM
The mature real-time MDM architecture is not the one that moves data fastest. It is the one that can explain which version is authoritative, deliver it within the required business latency and prove that every critical consumer eventually received the right state.
Sources & Further Reading
- IBM — What Is a Data Fabric?
- Debezium — Documentation
- Debezium — Change Event Structure and New Record State Extraction
- AsyncAPI Initiative — AsyncAPI Specification 3.1
- Microsoft Azure Architecture Center — Event-Driven Architecture
- Microsoft Azure Architecture Center — Event Sourcing Pattern
- Apache Kafka — Documentation
This article separates the broader concept of Data Fabric from the narrower technical problem of real-time master-data distribution. IBM material is used as a reference for the broader Data Fabric concept; Debezium documentation illustrates log-based Change Data Capture; AsyncAPI provides an open specification for message-driven API contracts; Microsoft and Apache Kafka documentation support event-driven architecture, eventual consistency, idempotency and messaging design concepts. The freshness decision model, Master Data Event Contract, Reconciliation Contract, failure-pattern taxonomy and readiness checklist are Digital Future & Strategy practitioner frameworks rather than official vendor standards. Specific latency targets, percentages of domains requiring real-time synchronization and implementation timelines should be derived from each organization's business requirements and measured production capability rather than generic benchmarks.
Reviewed: September 2026
MDM Strategy Series
Part 4 — Architecture & Integration
MDM #14. Master Data as a Data Product
MDM #15. Hybrid Federated MDM
MDM #16. Real-Time Master Data Architecture: CDC, Events, Consistency and Reconciliation
MDM #17. Designing MDM APIs for AI Agents: Contracts, Context, Security and MCP
Related Articles
MDM #14. Master Data as a Data Product
MDM #17. Designing MDM APIs for AI Agents
MDM #18. Regulatory-Ready Master Data
MDM #19. Zero Trust for Master Data
Previous: Hybrid Federated MDM
Comments
Post a Comment