MDM #3. The Core of Agentic Data Management: The Role and Future of the Data Steward Agent
The Data Steward has traditionally been the human control point inside enterprise MDM.
When duplicate records appear, a steward reviews them.
When a change request violates a rule, a steward investigates it.
When two business units disagree about an attribute, somebody eventually needs to understand the context and decide what happens next.
AI is beginning to change this operating model.
Not because every stewardship decision can suddenly be automated, but because much of the work surrounding the decision can now be detected, assembled, summarized, recommended and routed by software.
The important question is no longer “Can AI replace the Data Steward?” It is “Which parts of stewardship should an agent perform, and where must human accountability remain?”
I use the term Data Steward Agent for an AI-enabled operating role that assists with or executes selected stewardship activities inside defined policies, permissions and escalation rules.
It should not be interpreted as one universal industry role or as a fully autonomous MDM operator.
The more useful model is a partnership:
What has actually changed in 2026
The idea of applying AI to data stewardship is no longer purely conceptual.
SAP Master Data Governance now documents AI-assisted capabilities for change-request processing.
In SAP S/4HANA MDG 2025 FPS01, released in February 2026, SAP describes two relevant use cases:
- AI-assisted changes: a requester can describe a required master-data change in natural language instead of navigating the conventional user interface.
- AI-generated summaries: a reviewer can receive a natural-language summary of the changes contained in a master-data change request.
SAP Help Portal — Enabling Automated Tasks in Master Data Governance
These capabilities do not eliminate governance.
They reduce friction around how people create and review governed changes.
Informatica is moving in a similar direction.
Its Spring 2026 release introduced CLAIRE Copilot for Data Stewards, which can surface records, recommendations and match insights through natural-language interaction.
Informatica — IDMC Spring 2026 Data Quality and MDM Capabilities
Informatica has also announced Agentic Multidomain MDM and a Data Steward Agent for Q4 2026.
As of September 2026, that distinction is important: the Data Steward Agent is an announced capability rather than something I would treat as universally deployed production functionality today.
Informatica from Salesforce — Agentic Data Management Capabilities and Availability
The direction of travel is clear: AI is moving from helping stewards find information toward helping them analyze, recommend, orchestrate and eventually execute selected data-management actions.
The Data Steward's real workload is more than record correction
It is easy to describe stewardship as “cleaning data.”
That misses much of the role.
A steward may need to:
- triage quality exceptions,
- review suspected duplicates,
- investigate conflicting source values,
- understand business rules,
- evaluate change requests,
- identify the correct owner or approver,
- explain why a record was rejected,
- manage recurring exceptions,
- track unresolved issues, and
- coordinate with business and technology teams.
Some of this work is judgment.
Much of it is evidence gathering.
That distinction is where the opportunity for Agentic Data Management becomes clearer.
Five stewardship tasks where an agent can add practical value
1. Exception triage
A traditional quality queue may contain hundreds or thousands of issues from different systems and domains.
A steward has to determine:
- what type of issue occurred,
- how serious it is,
- which domain owns it,
- whether similar cases already exist, and
- what should happen next.
An agent can assist by classifying incoming cases, gathering related information and routing them according to policy.
The useful output is not simply:
It is closer to:
Source: regional onboarding system.
Business impact: record cannot proceed to final approval.
Related evidence: approved external registry contains a possible value.
Recommended route: Supplier Data Steward review.
2. Evidence gathering
Many stewardship decisions are slow because information is distributed across several systems.
For a suspected duplicate supplier, the steward may need to compare:
- legal names,
- addresses,
- registration identifiers,
- tax identifiers,
- corporate relationships,
- transaction history, and
- previous stewardship decisions.
An agent can assemble this evidence before the steward opens the case.
That changes the human task from searching for evidence to evaluating evidence.
3. Recommendation and explanation
An agent can also provide a recommendation.
For example:
Evidence supporting similarity: similar names, same city and shared corporate parent.
Evidence against merge: different verified legal registration identifiers.
Recommended action: maintain separate legal entities but establish a parent-subsidiary relationship.
This is more useful than a confidence score without explanation.
The steward can evaluate both the recommendation and the evidence behind it.
4. Workflow orchestration
Not every case should go to the same approver.
A simple descriptive correction may follow one path.
A legal-entity change may require another.
A cross-domain change may require several functions.
An agent can help determine which workflow is appropriate based on:
- the attribute being changed,
- business impact,
- domain ownership,
- geography,
- regulatory sensitivity, and
- the type of exception involved.
This is an area where agentic orchestration can reduce administrative work without transferring the final business decision to AI.
5. Decision documentation
Stewardship decisions generate valuable institutional knowledge.
Why was an exception approved?
Why were two records deliberately kept separate?
Why was a product moved to a different hierarchy?
In many organizations, that rationale disappears into email or meeting notes.
An agent can help capture:
- the issue,
- the evidence reviewed,
- the recommendation,
- the human decision,
- the reason for an override, and
- the final action.
This creates reusable precedent for future cases.
The important boundary: assistance, execution and authority are different
Three concepts are often mixed together in discussions of AI agents.
| Capability | What it means | Example |
|---|---|---|
| Assist | Provide information, analysis or a recommendation. | Summarize a change request and show relevant policy. |
| Execute | Perform an action permitted by existing policy. | Apply an approved reference-code mapping. |
| Decide | Determine the business policy or accept the consequences of ambiguity. | Approve an exception to the global supplier standard. |
AI can increasingly perform the first two.
The third requires much more caution.
The ability to execute a data change does not automatically create authority to decide what the business standard should be.
Do not govern an agent only with a confidence score
A common design is to let the AI act automatically above a certain confidence threshold and escalate everything below it.
That is too simple for enterprise master data.
Suppose two customer records receive a very high similarity score.
If an incorrect merge would affect privacy, financial history or contractual relationships, a high score alone may not justify automatic action.
Conversely, correcting a country code from an approved reference table may not require sophisticated probabilistic reasoning at all.
I would use a broader decision gate.
| Decision factor | Question |
|---|---|
| Evidence | Is the proposed action supported by an authoritative source or only probabilistic inference? |
| Business impact | What happens if the action is wrong? |
| Reversibility | Can the action be rolled back cleanly? |
| Policy | Has the business explicitly authorized automation for this case type? |
| Accountability | Who owns the outcome and reviews unexpected behavior? |
Confidence can be one input.
It should not be the entire governance model.
A practical human–agent division of work
I would define the split by decision characteristics rather than by attempting to maximize an automation percentage.
| Activity | Agent role | Human Steward role |
|---|---|---|
| Quality monitoring | Detect and classify recurring issues. | Prioritize business impact and root-cause action. |
| Evidence gathering | Collect relevant records, lineage and prior decisions. | Judge conflicting evidence and business meaning. |
| Standardization | Execute approved deterministic transformations. | Define and approve the standard. |
| Entity resolution | Generate candidates, evidence and recommendations. | Resolve ambiguous or high-impact identity decisions. |
| Change requests | Summarize, check policy and route. | Approve exceptions and material changes. |
| Governance rules | Identify recurring patterns and propose improvements. | Own policy and authorize rule changes. |
The Data Steward Agent needs context before autonomy
An AI agent cannot steward enterprise data effectively using only the current record.
It needs access to context.
That context can include:
- master-data definitions,
- business glossaries,
- quality rules,
- reference data,
- source-system authority,
- data lineage,
- entity relationships,
- approval policy,
- historical decisions, and
- permissions.
This is one reason Agentic MDM and AI-Ready Data are closely connected.
An agent can reason only over the context available to it.
If the organization's business definitions are contradictory or ownership is unclear, adding an AI agent does not resolve the underlying ambiguity.
An agent should not be expected to solve a governance decision that the organization itself has never made.
Permissions should be narrower than the agent's analytical reach
A Data Steward Agent may need broad read access to understand a case.
That does not mean it should have equally broad write authority.
I would separate at least four permission levels:
| Permission | Example |
|---|---|
| Read | Retrieve records, policies, lineage and relationship evidence. |
| Recommend | Generate a proposed correction or decision for human review. |
| Execute bounded action | Perform explicitly approved low-risk changes. |
| Escalate | Route ambiguous, high-impact or policy-conflicting cases to an accountable human. |
I would be very cautious about giving a general-purpose agent unrestricted update access across master-data domains.
The safer architecture is capability-based: the agent receives only the tools and actions required for its approved role.
Auditability has to be designed before autonomy
If an agent changes enterprise master data, the organization should be able to reconstruct what happened.
For material actions, I would retain:
- the original state,
- the proposed change,
- the evidence used,
- the governing rule or policy,
- the agent or service executing the action,
- the human approver where applicable,
- the final outcome, and
- any later override or reversal.
AI-generated reasoning text by itself is not enough.
The audit trail should link the action to identifiable evidence and policy.
“The agent decided” is not a governance explanation.
Human override is training data — but it is also governance evidence
Suppose an agent recommends merging two supplier records.
The steward rejects the recommendation because the records represent different subsidiaries.
That rejection contains valuable information.
It may indicate:
- a weakness in the matching logic,
- a missing relationship in the knowledge graph,
- an important domain rule, or
- a recurring ambiguity that should always require human review.
A mature operating model should therefore capture not only that the human overrode the agent, but why.
Repeated overrides are a signal that the automation policy or underlying model needs to change.
Do not start by trying to automate the whole steward role
I would introduce Agentic Data Management progressively.
| Stage | Agent responsibility | What should be proven |
|---|---|---|
| 1. Assist | Search, summarize and assemble evidence. | Is the information accurate, useful and traceable? |
| 2. Recommend | Generate proposed actions while humans decide. | Are recommendations consistently explainable and useful? |
| 3. Execute selected actions | Perform approved, reversible, low-risk actions. | Do controls, audit, rollback and monitoring work reliably? |
| 4. Orchestrate | Coordinate multiple tools, workflows and agents within policy. | Can the operating model scale without losing accountability? |
Progression should be based on evidence rather than a target automation percentage or a fixed calendar.
What should we measure?
The wrong KPI for Agentic Data Management is simply:
Automation can increase while quality becomes worse.
I would monitor a balanced set of measures.
| Measure family | Examples |
|---|---|
| Steward productivity | Queue age, review time, repeat manual work and backlog |
| Recommendation quality | Acceptance, rejection and override patterns |
| Automation safety | Reversals, incorrect actions and downstream incidents |
| Data quality | Recurring defects, unresolved exceptions and critical-rule failures |
| Governance | Actions with evidence, policy, owner and audit trail |
The most useful sign of progress may be that human stewards spend less time gathering routine evidence and more time resolving the business causes behind recurring data problems.
The steward role does not disappear — it moves up the decision stack
The traditional Data Steward can spend a large amount of time operating at record level.
Agentic Data Management can shift part of that workload upward.
The human role increasingly concentrates on:
- defining business meaning,
- resolving ambiguous identity and hierarchy cases,
- approving automation policy,
- reviewing systemic quality problems,
- managing exceptions,
- monitoring agent behavior,
- improving rules and context, and
- connecting data governance with business outcomes.
That does not make stewardship less important.
It makes the role more explicitly accountable for the operating policy under which automated stewardship occurs.
The future Data Steward is less likely to be the person who fixes every record and more likely to be the person who defines how records should be governed — including how AI is allowed to act.
My practical takeaway
The Data Steward Agent should not be designed as a digital employee with unrestricted authority over enterprise master data.
It is better understood as a governed execution layer around stewardship.
It can:
- monitor and triage cases,
- assemble evidence,
- summarize complex changes,
- recommend actions,
- route workflows,
- execute explicitly approved low-risk actions, and
- capture feedback for future improvement.
Human stewards should remain responsible for the decisions where business meaning, policy, ambiguity and risk matter.
The implementation principle I would use is therefore simple:
Give the agent enough context to understand the case.
Give it only the permissions required for its approved role.
Automate actions that are evidence-rich and governed.
Escalate ambiguity rather than hiding it behind a confidence score.
Keep human accountability attached to every material decision.
That is a more realistic foundation for Agentic Data Management than attempting to maximize autonomous execution from the beginning.
Sources & Further Reading
- SAP Help Portal — Enabling Automated Tasks in Master Data Governance
- Informatica — IDMC Spring 2026 Data Quality and MDM Capabilities
- Informatica from Salesforce — Agentic Data Management Capabilities and Availability
- Informatica — CLAIRE Innovations: Agentic CLAIRE GPT and CLAIRE Agents
“Data Steward Agent” is used in this article as both an emerging vendor capability and a broader practitioner operating concept. The role model, human-agent decision boundary, permission model, rollout stages and metrics are Digital Future & Strategy frameworks rather than a universal industry standard. Vendor capabilities and availability should be verified before implementation.
Reviewed: September 2026
Global MDM Strategy Series
Part 1 — AI & Agentic MDM
MDM #2. AI-Ready Data — Why Enterprise AI Depends on Trusted Master Data
MDM #3. The Core of Agentic Data Management: The Role and Future of the Data Steward Agent
MDM #4. Self-Healing Master Data — What AI Can Fix Automatically and What Still Needs Human Review
Previous: AI-Ready Data — Why Enterprise AI Depends on Trusted Master Data
Next: Self-Healing Master Data — What AI Can Fix Automatically and What Still Needs Human Review
Comments
Post a Comment