MDM #5. Knowledge Graph-Based Entity Resolution — Beyond Fuzzy Matching for Enterprise MDM
Two records can look different and still describe the same real-world entity.
“International Business Machines Corporation” and “IBM” are an obvious example to a person.
Enterprise data is rarely that easy.
A supplier may appear under a legal name in one ERP, a trading name in a procurement platform, an abbreviated name in a regional system and a former name in historical transactions.
Addresses may differ. Phone numbers may have changed. Registration identifiers may be missing. Corporate ownership may have changed after an acquisition.
The opposite problem also occurs.
Two companies can have very similar names while being completely different legal entities.
This is why entity resolution is not simply a search for similar strings.
The real question is not “Do these records look alike?” It is “Does the available evidence support the conclusion that these records represent the same real-world entity?”
Traditional exact and fuzzy matching remain useful parts of that decision.
Knowledge graphs add another dimension: relationships and context.
That can make them particularly valuable for difficult entity-resolution problems in enterprise MDM.
Entity resolution comes before the golden record
MDM is often described in terms of creating a trusted or mastered record.
But before an enterprise can decide which values should survive, it first has to decide which records belong to the same entity.
Informatica defines matching as determining whether records should be merged automatically or treated as candidates for manual consolidation because specified attributes are identical or sufficiently similar.
Informatica — Resolving Duplicates Overview
This distinction is fundamental.
There are at least four different activities that are often discussed together:
| Activity | What it answers |
|---|---|
| Candidate generation | Which records are plausible enough to compare? |
| Matching | How strong is the evidence that two records represent the same entity? |
| Clustering / linking | Which records belong to the same resolved entity? |
| Survivorship / consolidation | Which values should represent that entity in the mastered record? |
The first three establish identity.
The fourth establishes representation.
Confusing those decisions can create serious problems because a record can be matched correctly but consolidated incorrectly — or merged incorrectly before survivorship rules even become relevant.
Exact matching is powerful when the identifier is trustworthy
It would be a mistake to assume that modern entity resolution should replace deterministic rules.
If two records contain the same validated government registration number, globally unique identifier or other authoritative key, that evidence may be stronger than an elaborate AI model.
Exact matching is valuable when:
- the identifier is stable,
- the source is trusted,
- the identifier has the same semantics across systems, and
- reuse or data-entry errors are sufficiently controlled.
The difficulty is that enterprise data often lacks one universally available identifier.
A customer may have no registration number.
A supplier identifier may exist only in one country.
An external identifier may have been entered incorrectly.
A product may have several identifiers at different levels of its lifecycle.
At that point, the enterprise needs to combine evidence.
Fuzzy matching solves a different problem
Fuzzy matching helps when equivalent values are written differently.
Names can be abbreviated.
Addresses can be formatted differently.
Words can contain typographical errors.
Phone numbers can include different country-code conventions.
This is why mature MDM platforms support exact and fuzzy matching rather than only exact keys.
Informatica, for example, supports match configurations based on exact or fuzzy attributes and separates matches that can be consolidated automatically from those requiring manual review.
Informatica — Automatic and Manual Record Consolidation
But string similarity alone cannot answer every identity question.
Two suppliers may have similar names but different registration numbers and ownership structures.
Another pair may have quite different names but share a registration number, address history and corporate parent.
The challenge becomes contextual.
What the graph adds: relationships as evidence
A knowledge graph represents entities and the relationships among them.
For enterprise MDM, the relevant graph might connect:
Instead of asking only whether two organization names look similar, an entity-resolution process can examine the surrounding evidence.
Do both records connect to the same verified registration identifier?
Do they share an address history?
Do they belong to the same corporate parent?
Are the relationships consistent with what is already known about the entity?
Neo4j describes entity resolution as determining when different records represent the same real-world entity and notes that graph approaches can use both traditional identifiers and relationships between data points.
Neo4j — What Is Entity Resolution?
Google Cloud's Enterprise Knowledge Graph similarly provides an Entity Reconciliation capability that clusters records considered to represent the same entity and assigns a common machine identifier.
Google Cloud — Enterprise Knowledge Graph Overview
The value of the graph is therefore not that relationships magically prove identity. It is that relationships provide additional evidence that can strengthen, weaken or explain a proposed match.
A practical entity-resolution pipeline
I would not design enterprise entity resolution as one AI model that returns “same” or “different.”
A more controllable architecture separates the problem into stages.
1. Standardize
Normalize values enough to make comparison meaningful.
Examples include:
- name normalization,
- address standardization,
- phone formatting,
- country and language normalization, and
- reference-code mapping.
2. Generate plausible candidates
Comparing every record against every other record becomes impractical at enterprise scale.
Candidate-generation or blocking rules reduce the comparison space.
A supplier record might be compared only with entities sharing one or more useful signals, such as country, registration prefix, address region or normalized name pattern.
3. Compare multiple forms of evidence
The candidate pair can then be evaluated using several signals.
| Evidence type | Example |
|---|---|
| Deterministic identifier | Verified registration, tax or global identifier |
| Attribute similarity | Name, address, email or phone similarity |
| Relationship evidence | Shared parent, location, contact or other connected entity |
| Historical evidence | Prior names, former addresses or ownership changes |
| External reference | Trusted registry or licensed third-party reference data |
4. Score the evidence
Evidence should not necessarily carry equal weight.
A verified legal identifier may be more persuasive than a similar company name.
A shared address may be weak evidence when hundreds of organizations use the same office complex.
A relationship can therefore support a match without automatically proving one.
5. Build resolved entities or clusters
The next step converts pairwise evidence into an entity-level view.
This matters because identity is rarely only a sequence of independent pairs.
If A matches B and B matches C, the system must determine whether all three belong in one entity cluster or whether the transitive relationship has produced an incorrect grouping.
Neo4j's current discussion of production entity resolution similarly separates candidate blocking, matching, clustering and merging as distinct parts of the pipeline.
6. Route ambiguous cases to human review
The most important output is not always an automatic merge.
Sometimes the correct outcome is:
“There is enough evidence to review this pair, but not enough evidence to merge it automatically.”
This is where MDM stewardship remains essential.
False merges deserve more attention than duplicate-rate targets
A duplicate that remains unresolved is inconvenient.
An incorrect merge can be more damaging.
Imagine two companies with similar names being merged into one supplier.
Purchase history, risk information, payment information and compliance attributes may now appear to belong to the same organization.
Or imagine two individual customers incorrectly consolidated.
The privacy and service consequences can be significant.
This creates a basic trade-off:
| Error type | What happened | Possible consequence |
|---|---|---|
| False negative | The same entity remains as separate records. | Duplicate processing, fragmented reporting or missed relationships. |
| False positive | Different entities are incorrectly treated as one. | Incorrect consolidation, privacy issues, financial error or corrupted history. |
The acceptable balance depends on the domain.
A marketing use case may tolerate a different threshold from supplier payments, sanctions screening or regulated customer identity.
There is no universal “correct” match threshold. The threshold is a business-risk decision as much as a technical tuning parameter.
Do not turn a probability score into an automatic-merge policy
It is tempting to create a rule such as:
I would avoid using a universal number like that.
A production policy should also consider:
- the quality of the underlying evidence,
- the importance of the entity,
- regulatory or privacy sensitivity,
- downstream consequences,
- whether the merge can be reversed safely, and
- whether an authoritative identifier is present.
A more practical policy separates cases by risk.
| Decision band | Typical treatment | Example condition |
|---|---|---|
| Deterministic match | Governed automatic linking or merge | Authoritative identifier plus consistent supporting evidence |
| Strong probabilistic match | Automation only where domain policy permits | Multiple independent signals with low ambiguity |
| Ambiguous match | Steward review | Conflicting identifiers or mixed evidence |
| Insufficient evidence | Keep records separate | Similarity without enough identity evidence |
This decision model is an illustrative Digital Future & Strategy framework, not a universal industry threshold.
Knowledge graphs are especially useful when identity depends on relationships
Not every MDM domain needs a graph-based matching architecture.
If authoritative identifiers are strong and the entity model is simple, traditional matching may be sufficient.
I would look more seriously at graph-assisted resolution when:
- corporate hierarchies matter,
- entities have many aliases,
- relationships provide meaningful evidence,
- ownership or beneficial ownership matters,
- records come from many heterogeneous sources,
- historical relationships need to be preserved, or
- the mastered entity will also support graph analytics or GraphRAG.
Neo4j's entity-resolved knowledge-graph material illustrates this point: entity resolution improves the usefulness of a graph by resolving records that represent the same underlying entity before downstream analytics or AI uses the graph.
Neo4j — Entity-Resolved Knowledge Graphs
This leads to an important architectural distinction.
A knowledge graph should not be introduced merely because the technology is fashionable. It should be introduced when relationships materially improve the identity decision or create downstream value.
A supplier example: why relationship context matters
Consider three supplier records.
| Record A | Record B | Record C |
|---|---|---|
|
Acme Korea Ltd. Seoul Registration ID: 12345 Parent: Acme Global |
ACME KOR Seoul Registration ID: 12345 Parent: ACME Holdings |
Acme Korea Services Seoul Registration ID: 98765 Parent: Acme Global |
A name-based model may consider all three records highly similar.
Relationship and identifier evidence changes the picture.
A and B share the same verified registration identifier, strongly supporting the conclusion that they represent the same legal entity.
C has a similar name, location and parent relationship but a different verified registration identifier.
That suggests a related company rather than necessarily the same company.
The graph helps preserve both types of knowledge:
This is an important advantage.
Enterprise MDM often needs to know not only whether two records are identical, but also how distinct entities are related.
Entity resolution and hierarchy management should not be confused
This distinction is easy to miss.
Entity resolution answers:
“Are these records the same entity?”
Hierarchy and relationship management answer:
“How are these separate entities related?”
A parent company and its subsidiary are not duplicates.
A headquarters and a branch may not be duplicates.
Two products belonging to the same product family are not duplicates.
If a matching model treats strong relationships as proof of identity, it can create false merges.
Graph models are useful precisely because they can represent both:
- sameAs relationships for resolved identity, and
- parentOf, subsidiaryOf, locatedAt, suppliedBy and other relationships between distinct entities.
Do not forget survivorship
Resolving identity does not automatically tell the MDM platform which value should become authoritative.
Suppose two records are correctly determined to describe the same supplier.
One has the newer address.
Another has the verified registration number.
A third source has the procurement classification approved by the business owner.
The master record may need values from several source records.
Informatica's MDM documentation describes survivorship as the process by which trusted source values contribute to the consolidated master record.
Informatica — Record Value Survivorship
This is why I would keep the architecture conceptually separate:
The implementation may combine these functions, but the decisions are different.
A safer rollout starts with observation, not automatic merging
For a new entity-resolution model, I would resist the temptation to automate merges immediately.
A phased approach makes it easier to understand how the model behaves on real enterprise data.
| Phase | Operating mode | What to learn |
|---|---|---|
| 1. Observe | Generate candidate matches only | What duplicate patterns actually exist? |
| 2. Recommend | Provide match recommendations to stewards | Where are precision and recall weak? |
| 3. Govern | Automate only clearly defined low-risk cases | Are evidence, audit and rollback controls reliable? |
| 4. Expand | Broaden approved automation scope | Does performance remain acceptable across new populations and sources? |
The move between phases should be based on evidence rather than a fixed calendar.
Important exit criteria can include:
- false-merge behavior,
- missed-match behavior,
- steward override patterns,
- explainability of the evidence,
- rollback capability,
- performance at production scale, and
- domain-owner acceptance.
Monitor the population after deployment
Entity resolution is not something I would tune once and then forget.
The data population changes.
New countries enter the system.
Acquisitions introduce different naming patterns.
New source systems arrive.
Corporate structures change.
Address formats evolve.
The distribution of match evidence can therefore drift.
I would monitor at least:
- match and non-match volumes,
- manual-review volumes,
- steward acceptance and rejection patterns,
- merge reversals,
- performance by source and region, and
- new recurring ambiguity patterns.
Neo4j's current treatment of entity resolution likewise emphasizes that the problem involves ambiguous and conflicting evidence, precision-versus-recall trade-offs and judgment cases rather than a one-time perfect matching operation.
Where knowledge graphs become more valuable in the AI era
Entity resolution has always mattered for MDM.
AI creates an additional reason to get it right.
An enterprise AI agent asked about a customer, supplier or product needs to retrieve information about the correct entity.
If several duplicate nodes represent the same organization, relevant information can remain fragmented.
If two distinct organizations have been merged incorrectly, the AI may combine information that should remain separate.
Entity-resolved knowledge graphs can therefore provide a useful foundation for downstream graph analytics and GraphRAG by making entity identity more coherent before retrieval and reasoning occur.
Neo4j specifically notes the importance of entity resolution for the quality and trust of downstream graph analytics and AI applications.
AI does not remove the old MDM identity problem. It increases the number of systems that depend on getting identity right.
My practical takeaway
Knowledge graphs do not make fuzzy matching obsolete.
They add context to it.
A strong enterprise entity-resolution strategy can combine:
- authoritative identifiers where available,
- exact matching for deterministic evidence,
- fuzzy matching for variation in attributes,
- relationship evidence from graphs,
- trusted external reference data,
- risk-based automation policy, and
- human review for ambiguous cases.
The objective should not be a marketing claim such as “zero duplicates.”
Some identity cases are genuinely ambiguous.
Some source data is incomplete.
Some matches require business judgment.
And a false merge may be more damaging than leaving a duplicate unresolved.
The mature goal is not zero duplicates at any cost. It is explainable, risk-controlled entity resolution that gives the enterprise a more trustworthy view of who and what its data actually represents.
That is where knowledge graphs can make MDM more powerful: not by replacing established matching techniques, but by adding relationship context where identity cannot be determined from attributes alone.
Sources & Further Reading
- Informatica — Resolving Duplicates Overview
- Informatica — Automatic and Manual Record Consolidation
- Informatica — Record Value Survivorship
- Neo4j — What Is Entity Resolution?
- Neo4j — Entity-Resolved Knowledge Graphs
- Google Cloud — Enterprise Knowledge Graph Overview
The entity-resolution pipeline, decision bands, rollout model and risk controls in this article are Digital Future & Strategy's practitioner framework. They are not universal industry thresholds. Matching and automation policies should be calibrated by domain, source-data quality, regulatory requirements, false-merge impact, reversibility and the strength of available evidence.
Reviewed: September 2026
Global MDM Strategy Series
Part 1 — AI & Agentic MDM
MDM #3. The Core of Agentic Data Management: The Role and Future of the Data Steward Agent
MDM #4. Self-Healing Master Data — What AI Can Fix Automatically and What Still Needs Human Review
MDM #5. Knowledge Graph-Based Entity Resolution — Beyond Fuzzy Matching for Enterprise MDM
Previous: Self-Healing Master Data — What AI Can Fix Automatically and What Still Needs Human Review
Next: Why Digital Transformation Initiatives Struggle — The Hidden Role of Master Data
Comments
Post a Comment