AI-Ready #4. Five Core Capabilities for Enterprise AI Readiness

What does an enterprise actually need to become AI-Ready?

The easy answer is a technology shopping list:

  • a lakehouse,
  • a vector database,
  • a feature store,
  • Master Data Management,
  • a data catalog,
  • synthetic data, and
  • an AI governance platform.

But owning these technologies does not make an enterprise AI-Ready.

A company can have a modern data platform and still fail because the AI retrieves the wrong policy.

It can have an MDM platform and still fail because the customer or supplier context is not accessible to the AI application.

It can have a powerful foundation model and still fail because nobody has defined how to evaluate its answers or what the agent is authorized to do.

AI readiness is not defined by the number of AI technologies an enterprise owns. It is defined by whether an AI use case can repeatedly access trustworthy data and context, be evaluated against meaningful evidence, and operate within governed boundaries.

For practical enterprise planning, I organize these requirements into five capabilities:

1. Data Architecture & Retrieval
2. Data Quality & Master Context
3. Synthetic & Scenario Data
4. Curation, Labeling & Evaluation
5. Governance, Security & Evidence

This is a Digital Future & Strategy practitioner framework.

It is not an official Gartner, McKinsey, NIST, SAP, Microsoft or industry-standard “five-pillar AI-Ready model.”

Its purpose is to help enterprises identify the capabilities that a real AI use case depends on and then determine which gaps should be addressed first.

The Five Capabilities at a Glance

# Capability Core Question Typical Evidence / Deliverables
1 Data Architecture & Retrieval Can AI access the right structured and unstructured information using the right retrieval mechanism? APIs, data products, keyword / vector / hybrid search, pipelines
2 Data Quality & Master Context Can AI trust the identity, attributes, relationships and freshness of critical business data? MDM, Data Quality rules, critical-data controls, metadata, master context
3 Synthetic & Scenario Data Do real datasets provide enough coverage for training, testing and rare scenarios? Synthetic data, simulation, test scenarios, rare-event augmentation
4 Curation, Labeling & Evaluation Can the organization define what good AI performance means and measure it repeatedly? Curated corpus, labels, ground truth, evaluation sets, regression tests
5 Governance, Security & Evidence Who can use what data, what can AI do with it, and can important decisions be reconstructed? Authorization, audit, lineage, HITL, monitoring, incident evidence

These five capabilities are connected, but they are not five mandatory technology projects.

A document-search assistant may need strong retrieval, corpus curation and access control but no feature store or synthetic data.

A supplier-master change agent may depend heavily on MDM, Data Quality, authorization, workflow and audit.

A computer-vision model for manufacturing defects may have a major synthetic-data requirement.

AI-Ready architecture should follow the use case rather than forcing every use case through the same technology stack.

Capability 1 — Data Architecture & Retrieval

Enterprise AI needs a way to access data, but “access” can mean several different things.

An AI application may need:

  • an exact material record from MDM,
  • a current order from ERP,
  • a policy document from a knowledge repository,
  • historical transactions from a lakehouse, or
  • semantically related technical documentation.

These are not the same retrieval problem.

Choose Retrieval by the Question

Business Question Possible Access Pattern
What is the current status of supplier SUP-104? Authoritative API / structured lookup
Find document POLICY-KR-104. Keyword / exact search
Find guidance related to delayed supplier deliveries. Vector or hybrid retrieval
What were this customer's orders during the last 90 days? Transactional API, query or governed data product
Which tools can this procurement agent execute? Controlled tool / API interface with authorization policy

Vector Search Is Useful — but Not Universal

Microsoft Azure AI Search currently supports vector search for conceptual similarity and hybrid search that combines vector and full-text retrieval.

Its hybrid-search implementation executes vector and full-text queries together and merges results using Reciprocal Rank Fusion.

Microsoft Learn — Vector Search in Azure AI Search

Microsoft Learn — Hybrid Search in Azure AI Search

The architectural lesson is not “use vectors everywhere.”

It is:

Exact Business Fact → Structured / Exact Access

Semantic Knowledge Discovery → Vector / Hybrid Retrieval

Freshness Is a Business Requirement

AI-Ready does not mean every dataset must become real time.

The right sequence is:

Business Decision
→ Maximum Acceptable Data Age
→ Batch / API / Event Pattern
→ SLA
→ Monitoring

Inventory availability may require much shorter latency than a product hierarchy used for periodic analysis.

Real-time architecture should be introduced because the business decision needs it — not because real time sounds more advanced.

Capability 1 Test

☐ The AI use case can reach its required structured data.

☐ Exact and semantic retrieval patterns are differentiated.

☐ Retrieval respects user and agent authorization.

☐ Source freshness requirements are explicit.

☐ Representative queries are used to evaluate retrieval quality.

☐ Tool access to operational systems is bounded and auditable.

Capability 2 — Data Quality & Master Context

An AI system can retrieve information successfully and still make the wrong decision if the underlying enterprise context is wrong.

Examples include:

  • a supplier duplicated across several systems,
  • a customer hierarchy that is incomplete,
  • a product with inconsistent classifications,
  • a material status that is no longer current, or
  • a business relationship linked to the wrong entity.

SAP MDG Data Quality Management supports validation rules, derivation rules, Data Quality KPIs, quality evaluation and monitoring for governed master data.

SAP Help Portal — MDG Data Quality Management

AI introduces another requirement:

quality has to be evaluated in relation to the AI decision.

Six Practical Quality Questions

Quality Dimension AI-Ready Question
Contextual Accuracy Is the value correct for the business context in which AI is using it?
Representation Do important populations, scenarios and operating conditions appear sufficiently in the data?
Ground Truth Quality Are labels, expected answers and evaluation judgments trustworthy?
Provenance Can important information be traced to an authoritative source?
Freshness Is the data current enough for the decision being made?
Identity & Relationship Can AI correctly identify the customer, product, supplier, material or other business entity and its relationships?

These six questions are a Digital Future & Strategy practitioner classification. They are not an official SAP, DAMA or NIST AI-Ready Data Quality standard.

Avoid Universal Data-Quality Thresholds

Rules such as:

Accuracy ≥ 98%
Completeness ≥ 95%
Duplicate Rate ≤ 1%

may be valid internal requirements for a particular process.

They are not universal AI-Ready standards.

A missing product marketing attribute and an incorrect supplier payment attribute do not create equivalent business harm.

The stronger pattern is:

AI Decision
→ Critical Data
→ Failure Scenario
→ Business Consequence
→ Required Quality Control

Capability 2 Test

☐ Critical business entities are identifiable across relevant systems.

☐ Decision-critical attributes have defined owners and quality rules.

☐ Hierarchies and relationships required by the AI workflow are governed.

☐ Freshness is defined according to use-case risk.

☐ Provenance is available for important business context.

☐ Data-related AI exceptions can be identified and investigated.

Capability 3 — Synthetic & Scenario Data

Synthetic data is useful when real data does not adequately cover the problem.

It is not a mandatory third step of AI readiness.

Potential uses include:

  • rare manufacturing defects,
  • robotics and autonomous-system simulation,
  • privacy-constrained development,
  • edge-case testing,
  • LLM evaluation scenarios, and
  • structured test data that should not expose production records.

Different Problems Require Different Generation Methods

Method Potential Use Main Validation Issue
Simulation Robotics, vision, physical environments Reality gap
Statistical / Generative Tabular Structured-data testing and controlled analysis Utility, privacy leakage and distribution distortion
LLM-Generated Conversation and evaluation scenarios Factuality, diversity and model-generated bias
Rule-Based Augmentation Controlled transformation of existing samples Business realism and label consistency

NVIDIA's synthetic-data tooling for simulation illustrates how physical-AI training datasets can be generated by varying object, camera and environmental conditions.

NVIDIA — Synthetic Data Generation with Replicator

Synthetic Does Not Automatically Mean Safe or Useful

Before use, the organization should test:

  • structural validity,
  • statistical fidelity,
  • scenario coverage,
  • privacy risk,
  • downstream task utility,
  • bias across relevant slices, and
  • generation provenance.

Synthetic data is valuable when it adds information or scenarios that the real dataset cannot efficiently provide — not merely when it increases record count.

Capability 3 Test

☐ A real-data limitation has been identified.

☐ Synthetic generation solves that specific limitation.

☐ Privacy assumptions have been independently reviewed.

☐ Important slices and rare scenarios are evaluated.

☐ Real-world holdout evidence is retained where feasible.

☐ Synthetic records and generation provenance remain identifiable.

Capability 4 — Curation, Labeling & Evaluation

Data can be technically accessible and still be unsuitable for AI.

A document repository may contain:

  • obsolete policies,
  • duplicate documents,
  • multiple versions without clear effective dates,
  • poor metadata, and
  • material that should never be included in the AI corpus.

A predictive model can suffer from:

  • inconsistent labels,
  • future information leakage,
  • unrepresentative training samples, or
  • evaluation data that does not resemble production conditions.

That is why curation and evaluation should be treated as an operating capability rather than a one-time data-cleaning exercise.

A Four-Step Curation Cycle

1. Select
Use-case-relevant sources and scenarios
↓

2. Clean & Structure
Remove obsolete content, duplicates and obvious quality defects
↓

3. Define Ground Truth
Create labels, expected answers or decision criteria where needed
↓

4. Evaluate & Improve
Measure actual AI outcomes and update the corpus, labels or rules

Labeling Should Match the Risk

Method Potential Advantage Main Control
Expert Labeling Suitable for complex domain judgments Document guidelines and disagreement resolution
Human-in-the-Loop Combines automation with human verification Define sampling, escalation and reviewer responsibility
Active Learning Can focus human review on uncertain cases Confirm uncertainty measures align with business importance
LLM-Assisted Labeling Can accelerate candidate labeling or scenario generation Use expert verification where errors are consequential

Evaluation Must Exist Before Scale

AI projects often define evaluation too late.

A more disciplined sequence is:

Use Case
→ Expected Outcome
→ Evaluation Set
→ Baseline
→ Pilot
→ Regression Test
→ Production Monitoring

For RAG, evaluation might include:

  • retrieval relevance,
  • evidence recall,
  • answer correctness,
  • groundedness,
  • authorization behavior, and
  • latency.

For an AI agent, additional measures may include:

  • tool-selection errors,
  • unauthorized action attempts,
  • human overrides,
  • workflow completion,
  • rollback events, and
  • business KPI impact.

Capability 4 Test

☐ The AI corpus or training dataset has an explicit selection policy.

☐ Outdated and duplicate sources are controlled.

☐ Ground truth or expected outcomes are documented where needed.

☐ Important ambiguous cases receive appropriate domain review.

☐ Evaluation covers representative and difficult scenarios.

☐ Regression testing is available before major model, retrieval or workflow changes.

Capability 5 — Governance, Security & Evidence

AI governance should not exist only as a policy document.

It has to appear in daily system behavior.

Examples include:

  • which documents a user can retrieve,
  • which customer attributes an agent can access,
  • which tools an AI agent can invoke,
  • which changes require human approval,
  • which events must be logged, and
  • how an important AI decision can be reconstructed later.

NIST's AI Risk Management Framework is designed to help organizations incorporate trustworthiness and risk-management considerations throughout the design, development, use and evaluation of AI systems.

NIST — AI Risk Management Framework

Four Practical Governance Pillars

Pillar Operational Meaning
Provenance & Lineage Know which sources, transformations and versions contributed to important AI outputs.
Access & Action Control Apply least privilege to data access, tool use and business actions.
Evaluation & Monitoring Monitor data, retrieval, model and agent behavior according to risk.
Evidence & Accountability Preserve approvals, overrides, incidents and material AI actions so responsibility remains clear.

These four pillars are a Digital Future & Strategy practitioner framework, not NIST's official AI RMF structure.

AI Agent Authority Should Be Explicit

Agent Authority Example Possible Control
Read Retrieve supplier status. Authentication, authorization, minimum fields
Recommend Recommend a possible duplicate. Evidence, evaluation and reviewer visibility
Submit Submit a master-data change request. Workflow policy and audit
Execute Perform an approved bounded change. Policy gate, rollback, monitoring and stronger approval where consequence is high

A model's decision to call a tool is not equivalent to enterprise authorization to execute that tool.

Capability 5 Test

☐ Data and tool permissions follow least-privilege principles.

☐ High-impact actions have explicit approval or policy gates.

☐ Important AI decisions can be reconstructed from evidence.

☐ Data and AI incidents have defined escalation paths.

☐ Human intervention requirements are based on consequence rather than habit.

☐ Governance is monitored in production rather than checked only at project approval.

The Five Capabilities Do Not Form a Fixed Sequence

A common mistake is to interpret the five capabilities as five project phases.

For example:

Quality → Architecture → Synthetic Data → Evaluation → Governance

That sequence will not fit every project.

Governance may have to begin before any data is connected.

Synthetic data may not be needed at all.

Retrieval may be the first problem for an enterprise-search assistant.

MDM may be the first problem for a supplier-change agent.

A more useful structure is a set of decision gates.

Four Decision Gates for AI Readiness

Gate Question Typical Decision
Gate 1 — Use Case Which business decision or workflow should AI improve? Define value, owner, risk and baseline.
Gate 2 — Data & Context Are the required identity, data and knowledge reliable and accessible? Improve retrieval, MDM, Data Quality or freshness.
Gate 3 — Evaluation Can the system be tested against representative scenarios and expected outcomes? Build evaluation set and pilot.
Gate 4 — Control & Scale Are permissions, monitoring, audit and operations strong enough for production scale? Scale, restrict authority, redesign or stop.

This creates a more realistic sequence:

Use Case
→ Trusted Context
→ Evidence
→ Controlled Scale

Two Use Cases Can Require Completely Different AI-Ready Investments

Use Case A — Internal Policy RAG

The primary gaps may be:

  • document version control,
  • corpus curation,
  • keyword / vector / hybrid retrieval,
  • access control, and
  • evaluation of grounded answers.

MDM and synthetic data may be secondary.

Use Case B — Supplier Master Change Agent

The primary gaps may be:

  • supplier identity,
  • relationship and status quality,
  • MDM API access,
  • agent identity,
  • workflow and approval,
  • audit, and
  • rollback.

Semantic vector search may be secondary.

Two AI projects inside the same company can have completely different AI-Ready bottlenecks. Enterprise strategy should standardize reusable capabilities without forcing every use case into the same architecture.

Do Not Use Fixed Investment Percentages

A framework becomes misleading when it assigns universal budget ratios such as:

Architecture: 38%
Data Quality: 22%
Synthetic Data: 15%
Curation: 15%
Governance: 10%

There is no defensible reason every enterprise should invest in those proportions.

A company with mature MDM but weak evaluation may need to spend differently from a company with fragmented identities and strong ML operations.

I would use an evidence-based decision matrix instead.

Capability Business Impact Current Gap Risk Reuse Potential
Architecture & Retrieval Evaluate Evaluate Evaluate Evaluate
Data Quality & Master Context Evaluate Evaluate Evaluate Evaluate
Synthetic & Scenario Data Evaluate Evaluate Evaluate Evaluate
Curation & Evaluation Evaluate Evaluate Evaluate Evaluate
Governance & Evidence Evaluate Evaluate Evaluate Evaluate

The objective is not to create a mathematically precise winner.

The objective is to make the investment logic visible.

What Executives Should See

An executive presentation should not begin with:

“We need a vector database, feature store and AI governance platform.”

A stronger statement is:

“Our customer-support agent cannot consistently retrieve the latest policy while linking it to the correct customer status. We will connect governed customer context with permission-aware retrieval, test it against a representative evaluation set, and measure service resolution and exception rates before scaling.”

The second statement makes five things visible:

  • the business problem,
  • the data problem,
  • the architectural response,
  • the control model, and
  • the evidence required to justify further investment.

A Practical Executive Review Template

1. Business Use Case
Which business decision or workflow are we improving?

2. Current AI-Ready Gap
Which of the five capabilities is constraining the use case?

3. Baseline Evidence
What errors, delays, retrieval failures, overrides or business costs exist today?

4. Near-Term Investment
Which capability should be improved during the next implementation cycle?

5. KPI
Which Data, AI and business outcomes will show whether the investment worked?

6. Scale Decision
What evidence is required before the capability becomes an enterprise-wide platform or service?

An Illustrative 90-Day Implementation

The following plan is an example, not a universal implementation schedule.

Period Primary Work Output
Days 0–30 Select one or two production-relevant AI use cases. Map required data, business entities, retrieval, risk, owners and current baseline. Use-Case Map, Gap Matrix, Baseline
Days 31–60 Improve the one or two most material capability gaps: retrieval, MDM / Data Quality, scenario coverage, evaluation or governance. Pilot-Ready Context and Evaluation Set
Days 61–90 Run realistic scenarios. Analyze data, retrieval and agent exceptions. Validate permissions and audit. Identify reusable components. Scale / Revise / Stop Decision and Reuse Backlog

Platform Selection Questions

Before introducing another AI data platform, ask:

  • Can the capability integrate with the existing SAP, Azure, AWS, GCP or enterprise data environment?
  • Can identity, metadata, authorization, lineage and audit be connected rather than recreated in separate silos?
  • Does the current AI portfolio genuinely require vector search, feature serving or synthetic-data tooling?
  • Have latency, availability, data residency and cost been measured on representative workloads?
  • Can business semantics remain stable if a model, database or AI vendor changes?
  • Does the platform reduce repeated engineering work across multiple use cases?

Reusable infrastructure should be created because reuse is demonstrated — not merely because a capability appears on an AI reference architecture.

What Should Be Standardized Across the Enterprise?

Not every AI component needs to be centralized.

But several capabilities benefit from shared enterprise standards.

Good Candidate for Reuse Why
Enterprise Identity / Master Context Multiple AI applications may need the same customer, supplier, product or material identity.
Authentication & Authorization Agent permissions should not be reinvented independently by every project.
Evaluation Harness Common regression and monitoring patterns reduce repeated engineering effort.
Audit / Observability Cross-system reconstruction becomes easier when evidence standards are shared.
Metadata / Provenance Standards Consistent source and context information improves both governance and AI evaluation.

My Practical Takeaway

Enterprise AI readiness should not be reduced to a list of products or a universal maturity score.

The five capabilities in this framework are better understood as questions that every important AI use case must answer.

Architecture & Retrieval
Can the AI obtain the right information in the right way?

Data Quality & Master Context
Can it trust the identity, attributes and relationships behind that information?

Synthetic & Scenario Data
Does the available real data adequately cover the situations the AI must handle?

Curation, Labeling & Evaluation
Can the organization define and repeatedly measure what good performance means?

Governance, Security & Evidence
Can the enterprise control who and what the AI can access, decide and execute — and prove what happened afterward?

The objective is not to build all five capabilities at maximum maturity before launching AI.

The objective is to identify the capability that currently constrains a real use case, improve it to a defensible level, evaluate the result and then standardize what proves reusable.

AI-Ready is not a technology inventory. It is an operating capability that connects trusted data, business context, evaluation and controlled action around real AI use cases.

Sources & Further Reading

Editorial Note
The five AI-Ready capabilities, six practical Data Quality questions, four governance pillars, four decision gates, investment matrix, executive review template and 90-day implementation plan in this article are Digital Future & Strategy practitioner frameworks. They are not official models, maturity standards or benchmarks published by Gartner, McKinsey, NIST, SAP, Microsoft, NVIDIA or another external organization. No universal AI-Ready score, budget ratio, ROI, data-quality threshold or mandatory technology stack is assumed. Actual priorities should be derived from the organization's AI use cases, data environment, business risk, existing platforms, regulatory obligations and operating model.

Reviewed: September 2026


AI-Ready Strategy Series

Part 1 — Readiness Assessment / Part 2 — Data Foundations

AI-Ready #3. Assessing Enterprise AI Readiness: 30 Questions Across Six Capabilities
AI-Ready #4. Five Core Capabilities for Enterprise AI Readiness
AI-Ready #5. Vector Search and Feature Stores: When Each Belongs in Enterprise AI Architecture

Previous: Assessing Enterprise AI Readiness: 30 Questions Across Six Capabilities

Next: Vector Search and Feature Stores: When Each Belongs in Enterprise AI Architecture

Comments

Popular posts from this blog

AI Strategy #1. AI Agents: Chatbots, RPA and Agentic AI Explained

MDM #9. Why Enterprise MDM Governance Fails After Go-Live — and How to Make Ownership Real

AI Strategy #17. Hybrid Cloud and GenAI: Designing Enterprise AI Infrastructure