MDM #5. Knowledge Graph-Based Entity Resolution — Beyond Fuzzy Matching for Enterprise MDM

Two records can look different and still describe the same real-world entity.

“International Business Machines Corporation” and “IBM” are an obvious example to a person.

Enterprise data is rarely that easy.

A supplier may appear under a legal name in one ERP, a trading name in a procurement platform, an abbreviated name in a regional system and a former name in historical transactions.

Addresses may differ. Phone numbers may have changed. Registration identifiers may be missing. Corporate ownership may have changed after an acquisition.

The opposite problem also occurs.

Two companies can have very similar names while being completely different legal entities.

This is why entity resolution is not simply a search for similar strings.

The real question is not “Do these records look alike?” It is “Does the available evidence support the conclusion that these records represent the same real-world entity?”

Traditional exact and fuzzy matching remain useful parts of that decision.

Knowledge graphs add another dimension: relationships and context.

That can make them particularly valuable for difficult entity-resolution problems in enterprise MDM.

Entity resolution comes before the golden record

MDM is often described in terms of creating a trusted or mastered record.

But before an enterprise can decide which values should survive, it first has to decide which records belong to the same entity.

Informatica defines matching as determining whether records should be merged automatically or treated as candidates for manual consolidation because specified attributes are identical or sufficiently similar.

Informatica — Resolving Duplicates Overview

This distinction is fundamental.

There are at least four different activities that are often discussed together:

Activity What it answers
Candidate generation Which records are plausible enough to compare?
Matching How strong is the evidence that two records represent the same entity?
Clustering / linking Which records belong to the same resolved entity?
Survivorship / consolidation Which values should represent that entity in the mastered record?

The first three establish identity.

The fourth establishes representation.

Confusing those decisions can create serious problems because a record can be matched correctly but consolidated incorrectly — or merged incorrectly before survivorship rules even become relevant.

Exact matching is powerful when the identifier is trustworthy

It would be a mistake to assume that modern entity resolution should replace deterministic rules.

If two records contain the same validated government registration number, globally unique identifier or other authoritative key, that evidence may be stronger than an elaborate AI model.

Exact matching is valuable when:

  • the identifier is stable,
  • the source is trusted,
  • the identifier has the same semantics across systems, and
  • reuse or data-entry errors are sufficiently controlled.

The difficulty is that enterprise data often lacks one universally available identifier.

A customer may have no registration number.

A supplier identifier may exist only in one country.

An external identifier may have been entered incorrectly.

A product may have several identifiers at different levels of its lifecycle.

At that point, the enterprise needs to combine evidence.

Fuzzy matching solves a different problem

Fuzzy matching helps when equivalent values are written differently.

Names can be abbreviated.

Addresses can be formatted differently.

Words can contain typographical errors.

Phone numbers can include different country-code conventions.

This is why mature MDM platforms support exact and fuzzy matching rather than only exact keys.

Informatica, for example, supports match configurations based on exact or fuzzy attributes and separates matches that can be consolidated automatically from those requiring manual review.

Informatica — Automatic and Manual Record Consolidation

But string similarity alone cannot answer every identity question.

Two suppliers may have similar names but different registration numbers and ownership structures.

Another pair may have quite different names but share a registration number, address history and corporate parent.

The challenge becomes contextual.

What the graph adds: relationships as evidence

A knowledge graph represents entities and the relationships among them.

For enterprise MDM, the relevant graph might connect:

Organization → Address → Registration ID → Parent Company → Contact → Bank Account → Location

Instead of asking only whether two organization names look similar, an entity-resolution process can examine the surrounding evidence.

Do both records connect to the same verified registration identifier?

Do they share an address history?

Do they belong to the same corporate parent?

Are the relationships consistent with what is already known about the entity?

Neo4j describes entity resolution as determining when different records represent the same real-world entity and notes that graph approaches can use both traditional identifiers and relationships between data points.

Neo4j — What Is Entity Resolution?

Google Cloud's Enterprise Knowledge Graph similarly provides an Entity Reconciliation capability that clusters records considered to represent the same entity and assigns a common machine identifier.

Google Cloud — Enterprise Knowledge Graph Overview

The value of the graph is therefore not that relationships magically prove identity. It is that relationships provide additional evidence that can strengthen, weaken or explain a proposed match.

A practical entity-resolution pipeline

I would not design enterprise entity resolution as one AI model that returns “same” or “different.”

A more controllable architecture separates the problem into stages.

Standardize → Generate Candidates → Compare Evidence → Score → Cluster → Review → Consolidate → Monitor

1. Standardize

Normalize values enough to make comparison meaningful.

Examples include:

  • name normalization,
  • address standardization,
  • phone formatting,
  • country and language normalization, and
  • reference-code mapping.

2. Generate plausible candidates

Comparing every record against every other record becomes impractical at enterprise scale.

Candidate-generation or blocking rules reduce the comparison space.

A supplier record might be compared only with entities sharing one or more useful signals, such as country, registration prefix, address region or normalized name pattern.

3. Compare multiple forms of evidence

The candidate pair can then be evaluated using several signals.

Evidence type Example
Deterministic identifier Verified registration, tax or global identifier
Attribute similarity Name, address, email or phone similarity
Relationship evidence Shared parent, location, contact or other connected entity
Historical evidence Prior names, former addresses or ownership changes
External reference Trusted registry or licensed third-party reference data

4. Score the evidence

Evidence should not necessarily carry equal weight.

A verified legal identifier may be more persuasive than a similar company name.

A shared address may be weak evidence when hundreds of organizations use the same office complex.

A relationship can therefore support a match without automatically proving one.

5. Build resolved entities or clusters

The next step converts pairwise evidence into an entity-level view.

This matters because identity is rarely only a sequence of independent pairs.

If A matches B and B matches C, the system must determine whether all three belong in one entity cluster or whether the transitive relationship has produced an incorrect grouping.

Neo4j's current discussion of production entity resolution similarly separates candidate blocking, matching, clustering and merging as distinct parts of the pipeline.

6. Route ambiguous cases to human review

The most important output is not always an automatic merge.

Sometimes the correct outcome is:

“There is enough evidence to review this pair, but not enough evidence to merge it automatically.”

This is where MDM stewardship remains essential.

False merges deserve more attention than duplicate-rate targets

A duplicate that remains unresolved is inconvenient.

An incorrect merge can be more damaging.

Imagine two companies with similar names being merged into one supplier.

Purchase history, risk information, payment information and compliance attributes may now appear to belong to the same organization.

Or imagine two individual customers incorrectly consolidated.

The privacy and service consequences can be significant.

This creates a basic trade-off:

Error type What happened Possible consequence
False negative The same entity remains as separate records. Duplicate processing, fragmented reporting or missed relationships.
False positive Different entities are incorrectly treated as one. Incorrect consolidation, privacy issues, financial error or corrupted history.

The acceptable balance depends on the domain.

A marketing use case may tolerate a different threshold from supplier payments, sanctions screening or regulated customer identity.

There is no universal “correct” match threshold. The threshold is a business-risk decision as much as a technical tuning parameter.

Do not turn a probability score into an automatic-merge policy

It is tempting to create a rule such as:

“If match confidence is above 95%, merge automatically.”

I would avoid using a universal number like that.

A production policy should also consider:

  • the quality of the underlying evidence,
  • the importance of the entity,
  • regulatory or privacy sensitivity,
  • downstream consequences,
  • whether the merge can be reversed safely, and
  • whether an authoritative identifier is present.

A more practical policy separates cases by risk.

Decision band Typical treatment Example condition
Deterministic match Governed automatic linking or merge Authoritative identifier plus consistent supporting evidence
Strong probabilistic match Automation only where domain policy permits Multiple independent signals with low ambiguity
Ambiguous match Steward review Conflicting identifiers or mixed evidence
Insufficient evidence Keep records separate Similarity without enough identity evidence

This decision model is an illustrative Digital Future & Strategy framework, not a universal industry threshold.

Knowledge graphs are especially useful when identity depends on relationships

Not every MDM domain needs a graph-based matching architecture.

If authoritative identifiers are strong and the entity model is simple, traditional matching may be sufficient.

I would look more seriously at graph-assisted resolution when:

  • corporate hierarchies matter,
  • entities have many aliases,
  • relationships provide meaningful evidence,
  • ownership or beneficial ownership matters,
  • records come from many heterogeneous sources,
  • historical relationships need to be preserved, or
  • the mastered entity will also support graph analytics or GraphRAG.

Neo4j's entity-resolved knowledge-graph material illustrates this point: entity resolution improves the usefulness of a graph by resolving records that represent the same underlying entity before downstream analytics or AI uses the graph.

Neo4j — Entity-Resolved Knowledge Graphs

This leads to an important architectural distinction.

A knowledge graph should not be introduced merely because the technology is fashionable. It should be introduced when relationships materially improve the identity decision or create downstream value.

A supplier example: why relationship context matters

Consider three supplier records.

Record A Record B Record C
Acme Korea Ltd.
Seoul
Registration ID: 12345
Parent: Acme Global
ACME KOR
Seoul
Registration ID: 12345
Parent: ACME Holdings
Acme Korea Services
Seoul
Registration ID: 98765
Parent: Acme Global

A name-based model may consider all three records highly similar.

Relationship and identifier evidence changes the picture.

A and B share the same verified registration identifier, strongly supporting the conclusion that they represent the same legal entity.

C has a similar name, location and parent relationship but a different verified registration identifier.

That suggests a related company rather than necessarily the same company.

The graph helps preserve both types of knowledge:

A = B     while     C → member of the same corporate group

This is an important advantage.

Enterprise MDM often needs to know not only whether two records are identical, but also how distinct entities are related.

Entity resolution and hierarchy management should not be confused

This distinction is easy to miss.

Entity resolution answers:

“Are these records the same entity?”

Hierarchy and relationship management answer:

“How are these separate entities related?”

A parent company and its subsidiary are not duplicates.

A headquarters and a branch may not be duplicates.

Two products belonging to the same product family are not duplicates.

If a matching model treats strong relationships as proof of identity, it can create false merges.

Graph models are useful precisely because they can represent both:

  • sameAs relationships for resolved identity, and
  • parentOf, subsidiaryOf, locatedAt, suppliedBy and other relationships between distinct entities.

Do not forget survivorship

Resolving identity does not automatically tell the MDM platform which value should become authoritative.

Suppose two records are correctly determined to describe the same supplier.

One has the newer address.

Another has the verified registration number.

A third source has the procurement classification approved by the business owner.

The master record may need values from several source records.

Informatica's MDM documentation describes survivorship as the process by which trusted source values contribute to the consolidated master record.

Informatica — Record Value Survivorship

This is why I would keep the architecture conceptually separate:

Entity Resolution → Entity Cluster → Survivorship → Master Record → Relationship Graph

The implementation may combine these functions, but the decisions are different.

A safer rollout starts with observation, not automatic merging

For a new entity-resolution model, I would resist the temptation to automate merges immediately.

A phased approach makes it easier to understand how the model behaves on real enterprise data.

Phase Operating mode What to learn
1. Observe Generate candidate matches only What duplicate patterns actually exist?
2. Recommend Provide match recommendations to stewards Where are precision and recall weak?
3. Govern Automate only clearly defined low-risk cases Are evidence, audit and rollback controls reliable?
4. Expand Broaden approved automation scope Does performance remain acceptable across new populations and sources?

The move between phases should be based on evidence rather than a fixed calendar.

Important exit criteria can include:

  • false-merge behavior,
  • missed-match behavior,
  • steward override patterns,
  • explainability of the evidence,
  • rollback capability,
  • performance at production scale, and
  • domain-owner acceptance.

Monitor the population after deployment

Entity resolution is not something I would tune once and then forget.

The data population changes.

New countries enter the system.

Acquisitions introduce different naming patterns.

New source systems arrive.

Corporate structures change.

Address formats evolve.

The distribution of match evidence can therefore drift.

I would monitor at least:

  • match and non-match volumes,
  • manual-review volumes,
  • steward acceptance and rejection patterns,
  • merge reversals,
  • performance by source and region, and
  • new recurring ambiguity patterns.

Neo4j's current treatment of entity resolution likewise emphasizes that the problem involves ambiguous and conflicting evidence, precision-versus-recall trade-offs and judgment cases rather than a one-time perfect matching operation.

Where knowledge graphs become more valuable in the AI era

Entity resolution has always mattered for MDM.

AI creates an additional reason to get it right.

An enterprise AI agent asked about a customer, supplier or product needs to retrieve information about the correct entity.

If several duplicate nodes represent the same organization, relevant information can remain fragmented.

If two distinct organizations have been merged incorrectly, the AI may combine information that should remain separate.

Entity-resolved knowledge graphs can therefore provide a useful foundation for downstream graph analytics and GraphRAG by making entity identity more coherent before retrieval and reasoning occur.

Neo4j specifically notes the importance of entity resolution for the quality and trust of downstream graph analytics and AI applications.

AI does not remove the old MDM identity problem. It increases the number of systems that depend on getting identity right.

My practical takeaway

Knowledge graphs do not make fuzzy matching obsolete.

They add context to it.

A strong enterprise entity-resolution strategy can combine:

  • authoritative identifiers where available,
  • exact matching for deterministic evidence,
  • fuzzy matching for variation in attributes,
  • relationship evidence from graphs,
  • trusted external reference data,
  • risk-based automation policy, and
  • human review for ambiguous cases.

The objective should not be a marketing claim such as “zero duplicates.”

Some identity cases are genuinely ambiguous.

Some source data is incomplete.

Some matches require business judgment.

And a false merge may be more damaging than leaving a duplicate unresolved.

The mature goal is not zero duplicates at any cost. It is explainable, risk-controlled entity resolution that gives the enterprise a more trustworthy view of who and what its data actually represents.

That is where knowledge graphs can make MDM more powerful: not by replacing established matching techniques, but by adding relationship context where identity cannot be determined from attributes alone.


Sources & Further Reading

Editorial Note
The entity-resolution pipeline, decision bands, rollout model and risk controls in this article are Digital Future & Strategy's practitioner framework. They are not universal industry thresholds. Matching and automation policies should be calibrated by domain, source-data quality, regulatory requirements, false-merge impact, reversibility and the strength of available evidence.

Reviewed: September 2026


Global MDM Strategy Series

Part 1 — AI & Agentic MDM

MDM #3. The Core of Agentic Data Management: The Role and Future of the Data Steward Agent
MDM #4. Self-Healing Master Data — What AI Can Fix Automatically and What Still Needs Human Review
MDM #5. Knowledge Graph-Based Entity Resolution — Beyond Fuzzy Matching for Enterprise MDM

Previous: Self-Healing Master Data — What AI Can Fix Automatically and What Still Needs Human Review

Next: Why Digital Transformation Initiatives Struggle — The Hidden Role of Master Data

Comments

Popular posts from this blog

AI Strategy #1. AI Agents: Chatbots, RPA and Agentic AI Explained

MDM #9. Why Enterprise MDM Governance Fails After Go-Live — and How to Make Ownership Real

AI Strategy #17. Hybrid Cloud and GenAI: Designing Enterprise AI Infrastructure