AI-Ready #12. From Assessment to Operations: A Four-Stage Roadmap for Enterprise AI Readiness
“Where should we start?” is one of the most common questions in enterprise AI.
It is also one of the easiest questions to answer too simplistically.
A company with strong master data but weak AI evaluation does not need the same roadmap as a company whose customer, supplier and product identities are fragmented across multiple systems.
A read-only knowledge assistant does not require the same controls as an AI agent that can update ERP or procurement systems.
A regulated decision process does not have the same evidence requirements as an internal productivity tool.
There is no universal AI-Ready implementation sequence. The useful roadmap is the one that connects a specific business outcome to the minimum data, architecture, evaluation and governance required to achieve it safely.
For that reason, I use a four-stage implementation structure:
This is not an industry standard and it is not a substitute for NIST, regulatory guidance or vendor implementation methodologies.
It is a practitioner framework for structuring enterprise decisions.
Governance, security, data quality and evaluation should operate across all four stages rather than appearing only at the end.
How This Roadmap Differs from the NIST AI RMF
The NIST AI Risk Management Framework organizes AI risk-management activity around four functions:
- Govern,
- Map,
- Measure, and
- Manage.
NIST does not define these functions as a simple sequential project checklist. Its Playbook provides suggested actions across the AI lifecycle.
NIST — AI Risk Management Framework
The four-stage model in this article serves a different purpose.
| Framework | Primary Purpose | How to Use It |
|---|---|---|
| NIST AI RMF | AI risk management and trustworthiness | Apply Govern, Map, Measure and Manage across the AI lifecycle. |
| This Four-Stage Roadmap | Implementation and investment decision flow | Use Assess, Design, Pilot / Build and Operate / Scale to decide whether and how to move forward. |
Governance should not wait for Stage 4. It should influence every design and investment decision from the beginning.
The Four Stages at a Glance
| Stage | Core Question | Primary Evidence | Decision |
|---|---|---|---|
| 1. Assess | Is this a problem worth solving with AI, and what currently limits success? | Business baseline, data map, risk scenarios, current gaps | Proceed, redefine or stop |
| 2. Design | What is the minimum architecture and operating model required? | Target design, evaluation plan, permissions, ownership | Pilot ready or not ready |
| 3. Pilot / Build | Does the solution create value with acceptable risk in a realistic workflow? | Evaluation results, user evidence, exceptions, business KPI | Scale, revise or stop |
| 4. Operate / Scale | Can the capability remain reliable, economical and governed in production? | Monitoring, incidents, cost, adoption, reuse evidence | Expand, constrain, redesign or retire |
Stage 1 — Assess: Define the Business Decision Before the Architecture
The purpose of assessment is not to calculate one enterprise-wide “AI-Ready score.”
The first task is to define what the AI is expected to improve.
For example:
is not yet a useful business scope.
A more actionable problem might be:
The second version makes it easier to identify:
- the business owner,
- the decision being improved,
- the required data,
- the baseline,
- the possible risk, and
- the evidence needed to judge whether the initiative worked.
Five Areas to Assess
| Area | Question | Evidence |
|---|---|---|
| Business Problem | What cost, delay, error, risk or opportunity should improve? | Process KPI, incidents, workflow data |
| Data Dependency | Which master data, transactions, documents and reference data are required? | Source inventory, data map |
| Current Quality | Which defects materially affect the use case? | Data profiling, samples, exception history |
| AI Risk | What happens if the AI retrieves, recommends or executes incorrectly? | Risk scenarios and severity assessment |
| Value Baseline | What is the current process performance before AI? | Lead time, rework, cost, backlog or error baseline |
Not every AI use case requires MDM as a major prerequisite.
A policy-document assistant may depend more heavily on document governance, retrieval quality and access permissions.
A supplier-change agent or customer-action agent, however, may depend directly on entity identity, hierarchy and master-data quality.
The purpose of Stage 1 is to identify the actual bottleneck rather than assuming that every AI problem is a model problem, a data problem or an MDM problem.
Stage 1 Decision Gate
☐ A business owner is accountable for the outcome.
☐ The decision or workflow is clearly defined.
☐ A current baseline has been measured, at least through a representative sample.
☐ Critical data and sources have been identified.
☐ Major failure or harm scenarios are understood.
☐ Pilot success, revision and stop criteria are explicit.
Stage 2 — Design: Build the Minimum Architecture Needed to Test the Use Case
The design stage should not begin with:
It should begin with:
What must be true for this use case to work safely?
The architecture can then be designed around those requirements.
Six Design Decisions
| Design Area | Decision Question |
|---|---|
| Data Foundation | Which master data, data products, transactions or documents are authoritative enough for the use case? |
| Access / Retrieval | Does the AI need structured APIs, SQL, keyword search, vector search, hybrid retrieval or several methods? |
| Freshness | Is batch sufficient, or does the workflow require API, event or CDC-based updates? |
| Evaluation | How will retrieval quality, AI behavior, business outcome and failure cases be evaluated? |
| Authority | What may the AI read, recommend, update or execute? |
| Accountability | Who owns the business outcome, data, AI product, risk decision and production operation? |
Do Not Assume One Data Pattern Fits Every Use Case
A structured operational agent may be better served by governed APIs than by a vector database.
A knowledge assistant may depend heavily on hybrid document retrieval.
A predictive model may depend on historical features and stable entity identifiers.
An MDM-focused agent may require golden records, matching evidence, workflow state and policy information.
The design should follow the data interaction.
MDM as a Governed Data Foundation
SAP Master Data Governance supports master-data validation, derivation, data-quality KPIs, monitoring and remediation.
These capabilities illustrate why MDM can become an important foundation when AI use cases depend on governed customers, suppliers, products or materials.
SAP Help Portal — MDG Data Quality Management
But the design principle remains:
Use MDM where trusted enterprise identity and controlled master-data change are genuinely required; do not force every AI workload through MDM simply because MDM exists.
Stage 2 Decision Gate
☐ The minimum pilot architecture is defined.
☐ Critical data authority and access are known.
☐ Evaluation data and test cases are prepared.
☐ Read, Recommend, Write and Execute permissions are explicit.
☐ Human review and escalation conditions are defined.
☐ Logging, audit and rollback expectations are defined where relevant.
☐ Business, Data, Technology, AI and Risk ownership is clear enough to run the pilot.
Stage 3 — Pilot / Build: Test the Operating System, Not Just the Model
A pilot should not exist only to prove that the model can generate impressive outputs.
It should test whether the complete operating chain works.
Each link can fail independently.
For example:
- the model may be capable but the wrong customer record is retrieved,
- the correct document may exist but poor metadata prevents retrieval,
- the recommendation may be good but the workflow is too slow for users to adopt it,
- the agent may reason correctly but receive excessive system permissions, or
- technical accuracy may improve while the business process remains unchanged.
Evaluate at Five Levels
| Evaluation Layer | Examples | Question |
|---|---|---|
| Data | Critical-field defects, duplicate identities, freshness | Is the input context trustworthy enough? |
| Retrieval / Context | Relevant evidence retrieved, missed evidence, incorrect entity | Did the AI receive the right information? |
| AI Behavior | Task success, unsupported claims, recommendation errors | Did the AI perform the task reliably? |
| Action / Control | Wrong actions, overrides, escalations, rollback | Did execution remain inside policy? |
| Business | Cycle time, rework, cost, service level, opportunity | Did the workflow actually improve? |
Test Uncomfortable Data
A pilot containing only perfectly curated examples is not enough.
Include realistic failure conditions.
For example:
- duplicate customer identities,
- missing supplier attributes,
- conflicting authoritative values,
- stale product status,
- documents with similar terminology but different meaning,
- ambiguous tool requests, and
- cases where the AI should refuse to act or escalate.
The objective is not to embarrass the system.
It is to discover the operating boundary before production users discover it.
Stage 3 Decision Gate
☐ Business performance was compared with the pre-pilot baseline.
☐ Evaluation included realistic failure cases.
☐ Material data-related failure modes are understood.
☐ Human overrides and exceptions have been reviewed.
☐ Production operating cost is understood well enough for the next investment decision.
☐ Major security, permission and control failures have been addressed.
☐ There is evidence supporting Scale, Revise or Stop.
Stage 4 — Operate / Scale: Treat Production Evidence as a New Input
Deployment is not the end of AI readiness.
It is where another form of evidence begins.
Production exposes:
- real user behavior,
- new data distributions,
- unanticipated exceptions,
- integration failures,
- operating costs,
- security events,
- model or retrieval drift, and
- business-process effects.
This should create a continuous learning loop.
Production Monitoring Should Cover More Than Model Quality
| Monitoring Area | Examples |
|---|---|
| Business | Lead time, rework, cost, service quality, exception volume |
| Adoption | Active use, abandonment, override reasons, workflow completion |
| Data | Critical-data failures, unresolved duplicates, freshness, lineage gaps |
| AI / Retrieval | Task success, retrieval errors, unsupported answers, exception patterns |
| Control | Wrong actions, escalations, policy violations, reversals and incidents |
| Economics | Run cost, cumulative benefit, support effort and infrastructure consumption |
Governance Has to Enter Daily Operations
Microsoft Purview's current governance model includes data mapping, governance domains, business concepts, data products, data quality and role-based permissions.
This illustrates an important operational principle: governance works when it is attached to data, roles and day-to-day processes rather than existing only as policy documentation.
Microsoft Learn — Data Governance in Microsoft Purview
Microsoft Learn — Plan for Data Governance
Scale What Is Reusable — Not Everything from the Pilot
One successful pilot does not mean every technical component should become an enterprise platform.
Some parts are use-case specific.
Others may be genuinely reusable.
I would separate them.
| Potentially Reusable | Examples |
|---|---|
| Trusted Context | Customer, supplier, product and material identity services |
| Access | Governed APIs, authentication, authorization and tool interfaces |
| Retrieval | Common document ingestion, metadata and retrieval services where several use cases share requirements |
| Evaluation | Evaluation harness, logging standards, test-data management |
| Control | Permission patterns, HITL workflows, audit and rollback services |
Standardize recurring requirements after reuse is demonstrated. Do not convert every pilot component into enterprise infrastructure automatically.
Do Not Use Fixed Budget Percentages
Another common roadmap mistake is prescribing a universal allocation such as:
- 30% for data,
- 25% for platforms,
- 20% for AI development,
- 15% for governance, and
- 10% for adoption.
That can create the appearance of precision without reflecting the organization's actual bottleneck.
I would use an investment decision matrix instead.
| Investment Candidate | Business Impact | Current Gap | Risk Reduction | Reuse Potential | Cost / Effort |
|---|---|---|---|---|---|
| Master Data / Identity | Evaluate | Evaluate | Evaluate | Evaluate | Evaluate |
| Retrieval / Data Access | Evaluate | Evaluate | Evaluate | Evaluate | Evaluate |
| Evaluation / Observability | Evaluate | Evaluate | Evaluate | Evaluate | Evaluate |
| Agent / AI Product | Evaluate | Evaluate | Evaluate | Evaluate | Evaluate |
| Security / Governance | Evaluate | Evaluate | Evaluate | Evaluate | Evaluate |
The table is deliberately not converted into an automatic weighted score.
Executives should review the evidence behind the dimensions rather than allowing a formula to create false certainty.
Complexity Matters More Than Company Size
Roadmaps are sometimes divided into:
“small company,” “mid-size company” and “large enterprise.”
That is only partially useful.
A small company deploying an AI agent that can make sensitive financial changes may require more rigorous evaluation than a large enterprise deploying a read-only internal Q&A assistant.
I would look at complexity drivers instead.
| Complexity Driver | Lower Complexity | Higher Complexity |
|---|---|---|
| Domain / Source Count | One domain, few sources | Multiple domains, ERPs, countries and data models |
| Agent Authority | Read / draft | Transactions or sensitive changes |
| Data Sensitivity | Low-sensitivity internal content | Personal, financial or regulated data |
| Integration | Few APIs | Legacy, event, hybrid-cloud and cross-company systems |
| Organization | Single team | Global, multi-business-unit or multi-company environment |
| Evidence Requirement | General productivity support | High-impact, audited or regulated decisions |
Use-Case KPIs Are Better Than Universal AI-Ready Targets
I would avoid targets such as:
The numbers may be too strict for one use case and dangerously weak for another.
A better KPI structure is:
| KPI Layer | Examples | Target Basis | Owner |
|---|---|---|---|
| Data | Critical-field failures, duplicates, freshness | Business harm and risk tolerance | Data Owner |
| Retrieval / AI | Grounded response, retrieval recall, exception | Evaluation set and task requirement | AI Product Owner |
| Action Risk | Wrong action, override, rollback, escalation | Action severity and policy | Process / Risk Owner |
| Adoption | Use, completion, abandonment | Workflow design and user need | Business Owner |
| Business | Lead time, rework, cost, service or opportunity | Pre-pilot baseline | Business Owner |
| Economics | Run cost, cumulative benefit, ROI / NPV where defensible | Company investment criteria | Finance / PMO |
An Illustrative 90-Day Implementation Sequence
This does not mean every organization can become AI-Ready in 90 days.
The purpose is narrower: test the core logic of the four-stage roadmap with one bounded use case.
| Period | Roadmap Stage | Primary Work | Decision |
|---|---|---|---|
| Days 0–30 | Assess | Define use case, baseline, critical data, risk and owners. | Proceed / stop |
| Days 31–60 | Design + Pilot Preparation | Build minimum architecture, evaluation, permissions and governance. | Pilot ready? |
| Days 61–90 | Pilot / Build | Test realistic workflow, exceptions and business KPI. | Scale / revise / stop |
| After Day 90 | Operate / Scale | Monitor production, identify reusable capability and consider expansion. | Periodic reassessment |
Again, the dates are illustrative.
The decision gates matter more than the calendar.
What Management Should See After the First Pilot
A useful executive review should be able to answer five questions.
1. Business Outcome
What changed relative to the pre-pilot baseline?
2. Data / AI Evidence
Which data gaps, retrieval issues or AI failure modes constrained performance?
3. Risk / Control
What important exceptions, incidents, overrides or escalation patterns occurred?
4. Economics
What did the pilot cost to build and operate, and what measurable value has been observed?
5. Decision
Should the organization Scale, Revise, Stop or Gather Additional Evidence?
This is more useful than reporting only a model score or a list of completed technical activities.
My Practical Takeaway
AI-Ready implementation should not be treated as a single transformation program with one maturity target and one fixed architecture.
It is better understood as a repeated decision cycle.
For each important use case:
Assess the business decision, data dependencies, baseline and risk.
Design the minimum data, architecture, evaluation and governance required.
Pilot the complete workflow using realistic data and failure cases.
Operate and Scale only when production evidence supports broader reuse or authority.
The value of the four-stage roadmap is not that every organization follows the same sequence on the same schedule.
Its value is that it forces each expansion decision to be supported by evidence.
The goal of an AI-Ready roadmap is not to reach the final stage as quickly as possible. It is to create a repeatable way to decide what should be built, what should be governed, what should be scaled — and what should be stopped.
That is the transition from AI experimentation to an enterprise operating capability.
Sources & Further Reading
- NIST — AI Risk Management Framework
- NIST — AI RMF Playbook
- SAP Help Portal — MDG Data Quality Management
- Microsoft Learn — Data Governance in Microsoft Purview
- Microsoft Learn — Plan for Data Governance
The Assess → Design → Pilot / Build → Operate / Scale structure, decision gates, investment matrix, KPI model and 90-day sequence are Digital Future & Strategy practitioner frameworks. They are not an official NIST, SAP, Microsoft or consulting-industry methodology. NIST's Govern, Map, Measure and Manage functions serve a different AI risk-management purpose and should not be treated as equivalent stages. Actual implementation should be adapted to each organization's use case, risk profile, data environment, architecture and operating model.
Reviewed: September 2026
AI-Ready Strategy Series
Part 4 — Implementation & Evidence
AI-Ready #11. Connecting MDM to AI Agents: APIs, Permissions and Human-in-the-Loop Controls
AI-Ready #12. From Assessment to Operations: A Four-Stage Roadmap for Enterprise AI Readiness
AI-Ready #13. How to Read Enterprise AI Case Studies: Separating Public Evidence from Practitioner Interpretation
Previous: Connecting MDM to AI Agents: APIs, Permissions and Human-in-the-Loop Controls
Next: How to Read Enterprise AI Case Studies: Separating Public Evidence from Practitioner Interpretation
Comments
Post a Comment