AI Strategy #3. Inside an AI Agent: Reasoning, Enterprise Context, Tools, Memory and Control

An AI agent is often described as an LLM with access to tools.

That description is technically convenient, but strategically incomplete.

A production enterprise agent must do much more than generate language. It has to understand an objective, determine what information is missing, retrieve the right enterprise context, decide whether an external action is necessary, invoke the correct tool with valid parameters, maintain state across multiple steps and operate within explicit business and security boundaries.

The useful architecture is therefore not:

LLM + API

It is closer to:

Reasoning
+ Enterprise Context
+ Retrieval
+ Tools
+ Memory
+ Identity & Authority
+ Evaluation & Runtime Control

This distinction has become more important as Agentic AI moves beyond chat interfaces and begins interacting with enterprise systems of record.

An enterprise AI agent is not a model that happens to know how to call APIs. It is a governed decision-and-action system built around a model.

This article explains the major architectural components inside an AI agent, how they interact, where each component fails and what changes when the architecture moves from a prototype to enterprise production.

The Architecture of an Enterprise AI Agent

The easiest way to understand an agent is to separate the major responsibilities rather than treat the LLM as the entire system.

Component Primary Role Typical Failure
LLM / Reasoning Interprets goals, reasons over context, plans steps and decides what to do next. Incorrect reasoning, unsupported assumptions or poor planning
Enterprise Context Provides business entities, policies, documents, transaction state and organizational meaning. Conflicting, stale or semantically inconsistent information
Retrieval / RAG Finds relevant external information when the model needs it. Wrong document, missing evidence or unauthorized retrieval
Tools Allow the agent to query systems, calculate, create records or execute actions. Wrong tool, invalid parameters or unsafe action
Memory Preserves useful state or knowledge across steps and, where appropriate, across sessions. Stale, excessive, incorrect or privacy-sensitive memory
Identity & Authority Determines whose authority the agent uses and which actions are permitted. Excessive privilege or ambiguous delegation
Evaluation & Control Measures behavior, enforces policy and detects when intervention is required. Failures become visible only after business impact occurs

The architectural mistake is to optimize one component independently and assume the overall agent will improve proportionally.

A more capable LLM cannot reliably compensate for an incorrect supplier identity. Better retrieval cannot make an unsafe transaction permission acceptable. Longer memory cannot repair an API that exposes the wrong business operation.

The quality of the agent therefore depends on the interaction among these components.

1. The LLM Is the Reasoning Engine — Not the System of Record

The language model is the most visible component of an AI agent because it interprets instructions and determines what should happen next.

Depending on the task, the model may:

  • interpret user intent,
  • decompose a goal into smaller tasks,
  • decide which information is missing,
  • select an appropriate tool,
  • generate tool parameters,
  • interpret tool results,
  • revise a plan after failure, and
  • decide whether the task is complete.

This is fundamentally different from traditional software.

In conventional software, developers explicitly encode most process logic.

Input
→ Predetermined Logic
→ Predetermined Action

An agent introduces model-mediated decisions:

Goal
→ Observe Context
→ Reason
→ Select Action
→ Observe Result
→ Reason Again
→ Continue or Stop

This flexibility is the source of both the value and the risk of Agentic AI.

The model can respond to situations that developers did not explicitly enumerate, but that also means the execution path cannot always be predicted in advance.

Reasoning Does Not Mean the Model Knows the Business

A powerful model may understand procurement concepts, accounting terminology or customer-service principles.

It does not automatically know:

  • which supplier record is authoritative inside your company,
  • which version of a contract is currently effective,
  • which employee can approve a particular transaction,
  • which products belong to a specific hierarchy,
  • which internal policy applies in a particular jurisdiction, or
  • what happened in an enterprise system five minutes ago.

Those facts exist outside the model.

This is why enterprise context increasingly matters as much as model intelligence.

The LLM provides general reasoning capability. The enterprise must provide the business reality over which that reasoning operates.

2. Enterprise Context Is More Than RAG

RAG—Retrieval-Augmented Generation—is often described as the mechanism that gives an LLM enterprise knowledge.

That is only partly correct.

Retrieval is a mechanism for obtaining context. Enterprise context is the larger business layer that determines what information is meaningful, authoritative, current and permitted.

For a supplier-management agent, relevant context might include:

  • supplier master identity,
  • legal-entity hierarchy,
  • approved supplier status,
  • contracts,
  • quality history,
  • purchase transactions,
  • risk information,
  • internal policy, and
  • the requester's organizational authority.

Some of this information belongs in documents. Some belongs in transactional systems. Some belongs in master data. Some may need to be calculated in real time.

RAG alone does not unify those layers.

RAG Is a Retrieval Pattern, Not a Truth Engine

A typical RAG architecture converts enterprise content into searchable representations and retrieves relevant information before or during model reasoning.

Enterprise Content
→ Ingestion & Parsing
→ Metadata / Indexing / Embeddings
→ Retrieval
→ Selected Context
→ Model Reasoning

This architecture can dramatically improve relevance, but retrieval does not establish truth automatically.

A retrieval system can return:

  • the wrong document,
  • an obsolete document,
  • a document from the wrong business unit,
  • information about a similarly named entity,
  • multiple documents containing conflicting rules, or
  • content the user or agent should not have been permitted to retrieve.

The enterprise problem therefore extends beyond vector similarity.

Retrieval Question Enterprise Requirement
Is this relevant? Search, ranking and retrieval quality
Is this authoritative? Source ownership and system-of-record rules
Is this current? Effective dates, refresh policies and event awareness
Is this about the correct entity? Master data, entity resolution and hierarchy
May this agent retrieve it? Identity-aware access control
Can we prove where it came from? Metadata, provenance and citation evidence

This is one reason enterprise RAG and enterprise data governance increasingly converge.

Context Engineering Is Replacing the Idea of “Put Everything in the Prompt”

As agentic systems become more capable, the design problem is shifting from prompt engineering toward context engineering.

The objective is not to provide the model with the largest possible amount of information.

It is to provide the right information at the right stage of the task.

Anthropic describes a growing shift toward just-in-time context, where agents dynamically retrieve information when required rather than loading every potentially useful document, tool definition and previous interaction into the context window in advance.

Anthropic — Effective Context Engineering for AI Agents

The difference can be represented as:

Approach Advantage Risk
Preload Everything Simple architecture for small tasks Noise, token cost and outdated context
Static RAG Efficient retrieval from indexed knowledge May miss live transactional or task-specific context
Just-in-Time Context Agent retrieves information as the task evolves Requires reliable tools, search and decision logic

This produces a more useful architecture for long-running enterprise agents.

Instead of trying to make the prompt contain the enterprise, the agent should know how to navigate the enterprise.

3. Tools Turn Reasoning into Action

Retrieval allows an agent to obtain information.

Tools allow it to change the world outside the model.

Examples include:

  • querying ERP inventory,
  • creating a CRM record,
  • submitting a purchase request,
  • checking an external data source,
  • running a calculation,
  • executing code,
  • sending a business notification, or
  • initiating an approval workflow.

The presence of tools is one of the major architectural differences between an assistant and an operational agent.

Reasoning without Tools
→ Recommendation

Reasoning + Governed Tools
→ Potential Business Execution

A Tool Is Not Just an API

Traditional APIs are usually designed for deterministic software written by developers.

Agent tools have a different consumer: a probabilistic model that must understand what the capability does, decide when to use it and construct valid inputs.

Anthropic describes this as a new software-design problem: tools create an interface between deterministic systems and non-deterministic agents.

Anthropic — Writing Effective Tools for AI Agents

A production-quality agent tool therefore needs more than endpoint connectivity.

Tool Design Requirement Why It Matters
Clear Purpose The agent must know when this tool should and should not be used.
Unambiguous Parameters Poor parameter definitions increase invalid or unintended actions.
Bounded Capability A narrow business action is safer than unrestricted system access.
Useful Response The model needs enough structured context to decide what to do next.
Error Semantics The agent must distinguish retryable errors from business-rule failures.
Authorization Tool availability must reflect user, agent and transaction authority.
Auditability Material actions need traceable inputs, decisions and outcomes.

This is why exposing hundreds of existing APIs directly to an agent does not automatically create an effective tool architecture.

The enterprise needs an agent-facing capability layer.

Tool Selection Is Itself an AI Decision

A traditional application knows which function will execute next because developers encoded that sequence.

An agent may need to choose.

For a customer refund request, an agent might need to decide whether to:

  • retrieve the order,
  • check delivery status,
  • read the return policy,
  • verify the customer's identity,
  • calculate refund eligibility,
  • request human approval, or
  • execute a permitted refund.

The workflow can branch depending on the information returned at each step.

Customer Request
↓
Retrieve Order
↓
Check Policy & Eligibility
↓
Determine Authority
↓
Prepare or Execute Action
↓
Confirm Outcome

The important point is not that the agent can call these tools.

It is that each tool should preserve business rules and authority boundaries regardless of what the model decides.

4. Memory Is Not the Same as Context

Memory is one of the most frequently misunderstood parts of agent architecture.

A large context window is not the same as durable memory.

RAG is not the same as memory.

And storing every past interaction indefinitely is not a mature memory strategy.

A useful practitioner distinction is:

Memory Type What It Contains Example
Working State Information needed to complete the current task. Current supplier, task status and pending checks
Session Memory Useful information retained throughout one extended interaction. Decisions made earlier in a multi-step case
Persistent Memory Selected information intentionally retained beyond the session. Approved user preference or durable workflow note
Enterprise Knowledge Authoritative facts managed outside the agent. Supplier master, policy, contract or product definition

This four-part classification is a Digital Future & Strategy practitioner model. Terminology varies across vendors and agent frameworks.

Enterprise Knowledge Should Not Be Replaced by Agent Memory

This distinction is strategically important.

If the authoritative payment term for a supplier changes in ERP, the agent should not continue relying on an old remembered value.

If an employee changes roles, persistent agent memory should not override the current identity and authorization system.

Enterprise facts should generally remain governed by their authoritative systems.

Enterprise System of Record
≠
Agent Memory

Memory should help an agent preserve useful task continuity—not create a parallel uncontrolled master-data system.

Structured Memory Becomes Important for Long-Horizon Work

Long-running agents face a practical constraint: every intermediate observation, tool result and conversation cannot remain in active context forever.

One emerging pattern is structured note-taking.

The agent periodically writes concise task state or important findings to persistent storage and retrieves them when needed later.

Anthropic describes this pattern as agentic memory and uses structured notes as one mechanism for preserving continuity when tasks extend beyond a single context window.

Anthropic — Effective Context Engineering for AI Agents

The design question is therefore not simply:

“How much can the agent remember?”

It is:

“What should be retained,
for how long,
under whose ownership,
and with what mechanism for correction or deletion?”

Memory Is Also a Governance Problem

Persistent memory creates lifecycle questions that simple chat applications can often avoid.

Enterprises need to decide:

  • what information may be remembered,
  • which information must never be persisted,
  • how long memory should remain valid,
  • who can access stored memory,
  • how outdated memory is refreshed,
  • how incorrect memory is corrected,
  • how retention policies apply, and
  • how memory is removed when business or legal requirements demand it.

A memory architecture without ownership and lifecycle controls eventually becomes another uncontrolled data store.

5. Identity and Authority Determine What the Agent Is Allowed to Do

Once an agent moves from information retrieval to action, identity becomes part of the architecture.

At least three identities can matter:

  • the human requesting the work,
  • the agent performing the reasoning, and
  • the service or application identity used to execute a tool call.

The enterprise also needs to know which authority is being delegated.

Consider a procurement agent.

An employee may have permission to view a supplier record but not approve the supplier. A manager may approve only within a particular region. A central procurement role may have different authority again.

The agent should not automatically receive the union of all available permissions.

Its authority should be explicitly bounded.

User Identity
+ Agent Identity
+ Delegated Authority
+ Tool Permission
+ Transaction Constraint
=
Allowed Action

This is the control boundary that separates a useful tool-using agent from an unsafe automation layer.

Capability and Permission Must Remain Separate

The model may be technically capable of generating the correct parameters for a payment change.

That does not mean it should be permitted to execute the change.

The architecture should separate:

Question Owner
Can the model do it? Model capability and evaluation
Should the agent do it? Business policy and risk classification
May this agent do it now? Runtime identity, authorization and transaction control

This separation becomes increasingly important as model capabilities improve faster than organizational governance processes.

6. The Orchestration Loop Turns Components into an Agent

LLM, retrieval, tools and memory are individual capabilities.

The agent emerges when these capabilities operate inside a loop.

A simplified agent loop is:

1. Understand Goal
↓
2. Inspect Current Context
↓
3. Decide Next Step
↓
4. Retrieve Information or Use Tool
↓
5. Observe Result
↓
6. Update State / Memory
↓
7. Continue, Escalate or Stop

This loop is the architectural difference between a one-shot LLM response and a genuine agentic system.

Anthropic uses a similarly simple conceptual definition: agents are systems in which models dynamically direct their own processes and tool usage.

Anthropic — Building Effective Agents

More Agentic Loops Do Not Automatically Mean a Better System

Autonomy creates flexibility, but each additional reasoning step also creates another opportunity for error, latency and cost.

For predictable processes, a deterministic workflow may remain superior.

Task Characteristic Likely Starting Architecture
Stable deterministic rules Application logic or workflow engine
Simple document question answering RAG-enabled assistant
Predictable multi-step process LLM-enabled workflow with explicit orchestration
Dynamic task with uncertain path Agent with bounded tool autonomy
High-consequence irreversible decision AI-assisted human decision unless stronger evidence justifies delegation

The architecture should be only as agentic as the workflow requires.

7. Evaluation and Observability Are Part of the Architecture

An agent should not be considered production-ready simply because the model performs well in a benchmark or a demonstration succeeds.

Agents are difficult to evaluate precisely because they operate across multiple turns, retrieve changing context, invoke tools and modify state.

Anthropic's 2026 guidance on agent evaluation emphasizes that the capabilities that make agents useful—autonomy, flexibility and multi-step execution—also create multiple failure points that do not exist in a single model response.

Anthropic — Demystifying Evals for AI Agents

NIST's 2026 TEVV-Athlon draft similarly treats evaluation as application-specific and explicitly includes agentic systems within its scope.

NIST — TEVV-Athlon Framework for Evaluating AI Systems

A useful evaluation chain is:

Intent
→ Context
→ Reasoning
→ Tool Selection
→ Tool Parameters
→ Action
→ Recovery / Escalation
→ Business Outcome

Each layer creates a different evaluation question.

What Should Be Measured?

Dimension Example Measure
Task Completion Was the requested workflow completed correctly?
Context Quality Did the agent retrieve the correct entity, source and current information?
Tool Selection Did the agent choose the appropriate tool?
Tool Parameters Were the action parameters complete and valid?
Policy Compliance Did the agent remain within the permitted action boundary?
Recovery Did it recover or escalate appropriately when a tool failed?
Economics What did a successfully completed workflow cost?
Business Outcome Did the completed action improve the target process or result?

The evaluation system should also retain enough execution evidence to reconstruct what happened.

Agent Observability Must Go Beyond Uptime

Traditional application monitoring asks questions such as:

  • Is the service available?
  • What is the latency?
  • Did the API return an error?

Agent observability requires additional questions:

  • Which context did the agent use?
  • Which tools did it consider?
  • Which tool did it select?
  • What parameters were submitted?
  • Which policy checks were triggered?
  • Was human approval requested?
  • What changed in the system of record?
  • What did the workflow cost?
  • Did the business outcome match the intended result?

The relevant production trace is therefore not simply an application log.

User / Event
→ Context Retrieved
→ Model Decision
→ Tool Call
→ Policy Decision
→ System Change
→ Final Outcome

Without this trace, debugging an agent becomes guesswork.

A Practical Example: An Enterprise Refund Agent

Consider a retail or service organization using an AI agent to handle a customer refund request.

The customer says:

“My order arrived damaged. Can you refund it?”

A simplistic chatbot can draft a sympathetic answer.

A production agent has to do significantly more.

Step Agent Capability Enterprise Dependency
Understand Request LLM reasoning Correct interpretation of customer intent
Identify Customer & Order Tool + enterprise context Reliable customer and order identity
Check Delivery Transactional tool Current fulfillment status
Retrieve Policy RAG / retrieval Correct and effective refund policy
Determine Eligibility Reasoning + policy Correct application of rules and exceptions
Check Authority Runtime control Refund amount and approval limit
Execute or Escalate Tool use Payment API and approval workflow
Record Outcome State / memory / audit Case history and traceability

The LLM is essential, but it is only one part of the operating system required to complete the refund safely.

Where Can the Refund Agent Fail?

Almost everywhere.

Failure Root Cause
Wrong order selected Identity or retrieval failure
Old refund policy applied Freshness or source-governance failure
Wrong refund amount calculated Reasoning or deterministic-calculation design failure
Agent issues refund above allowed limit Authorization and control failure
Refund submitted twice after retry Tool and transaction-design failure
Incorrect decision persists into next case Memory-management failure

This is why the architecture of an agent must be evaluated end to end.

The Enterprise Agent Control Plane

The components described above can be summarized as a three-layer architecture.

BUSINESS WORKFLOW
Goal · Outcome · Rules · Human Decision Rights

↓

AGENT INTELLIGENCE
Reasoning · Context · Retrieval · Memory · Planning

↓

CONTROLLED EXECUTION
Identity · Tools · APIs · Authorization · Validation

↓

ENTERPRISE SYSTEMS
ERP · CRM · SCM · MDM · Data · Documents · External Services

↓

OBSERVABILITY & GOVERNANCE
Evaluation · Trace · Policy · Audit · Escalation · Recovery

This architecture makes an important point.

Governance is not an external committee sitting above the agent.

Controls must exist inside the path between reasoning and business execution.

Do Not Confuse the Components

Many weak agent designs begin by asking one technology to solve another technology's problem.

Problem Wrong Response Better Response
Supplier identities do not match. Increase prompt detail. Fix entity identity and master-data resolution.
Policy changes frequently. Fine-tune the model with policy text. Retrieve the current authoritative policy at runtime.
Agent must calculate tax precisely. Ask the LLM to calculate it. Use deterministic calculation or an authoritative service.
Agent performs unauthorized action. Add a stronger system prompt. Enforce authorization at the tool and transaction layer.
Long-running task loses important context. Load the entire history every time. Use compaction, structured state and selective memory.

The architecture becomes more reliable when each component performs the job it is designed to perform.

From Prompt Engineering to Agent Engineering

The evolution from chatbot to enterprise agent changes the engineering discipline.

Generation Primary Design Question Primary Engineering Focus
Prompt-Centric AI How do we ask the model effectively? Prompt design and response quality
RAG-Centric AI How do we provide external knowledge? Ingestion, retrieval and grounding
Agentic AI How do we let AI complete work safely? Context, tools, state, identity, evaluation and control

Prompt engineering remains useful.

It is simply no longer sufficient.

A Production-Readiness Checklist for Agent Architecture

Before moving an agent from demonstration into a real business workflow, enterprises should be able to answer at least the following questions.

1. What business outcome is the agent responsible for?

2. Which decisions are delegated to the model and which remain deterministic?

3. Which enterprise entities must the agent identify correctly?

4. Which information sources are authoritative?

5. Does retrieval enforce freshness, ownership and access policy?

6. Are agent tools designed specifically for safe model use?

7. Can the agent distinguish a retryable technical error from a business-rule rejection?

8. What information is retained in memory, and why?

9. How are stale or incorrect memories corrected?

10. What identity does the agent use when invoking each tool?

11. What authority is delegated from the human user?

12. Which actions require explicit approval?

13. Can every material action be reconstructed after execution?

14. Are production failures converted into future evaluation scenarios?

15. Can the enterprise reduce or revoke agent authority without rebuilding the application?

The Architecture Decision That Matters Most

The central Agentic AI architecture question is not which model, vector database or agent framework is best.

Those decisions matter, but they sit below a more fundamental design choice:

Which decisions should remain deterministic,
which decisions may be delegated to AI,
and which business actions may follow from those decisions?

That decision determines the required data quality, tool design, memory architecture, identity controls, evaluation depth and human oversight.

The architecture should therefore begin with the business consequence of the agent's decisions rather than the sophistication of the model.

The Enterprise Agent Architecture Position

The LLM remains the reasoning engine of an AI agent, but it is not the complete agent.

RAG gives the system access to relevant knowledge, but retrieval does not guarantee that the knowledge is current, authoritative or associated with the correct enterprise entity.

Tools allow the model to act, but those tools must translate probabilistic reasoning into tightly controlled deterministic capabilities.

Memory gives an agent continuity, but uncontrolled memory can become stale, sensitive or inconsistent with enterprise systems of record.

Identity and authorization determine whether an intelligent recommendation can become a permitted business action.

Evaluation and observability determine whether the organization can prove that the complete system continues to behave as intended.

The useful architecture is therefore:

Reasoning
× Trusted Context
× Governed Tools
× Managed Memory
× Explicit Authority
× Production Evaluation

If one critical dependency is weak, adding more model intelligence may not solve the production problem.

The architecture of an enterprise agent is not defined by how much the model can do. It is defined by how reliably the enterprise can provide context, authorize action, preserve state and control what the model is allowed to do.

Sources & Further Reading

Method Note
The enterprise-agent component model, four-part memory classification, three-layer control architecture and production-readiness questions in this article are Digital Future & Strategy practitioner frameworks. They are not official Anthropic, OpenAI or NIST taxonomies. Terminology such as agent, memory, context and orchestration varies among model providers and software frameworks. This article therefore focuses on durable architectural responsibilities rather than vendor-specific product definitions. The earlier version of this article included an illustrative enterprise incident whose exact product configuration and source could not be independently verified; that example has been removed and replaced with generic workflow analysis. The architecture should always be adapted to the specific business process, data environment, authority level and consequence of agent actions.

Reviewed: September 2026


AI Strategy Series

Part 1 — Understanding Agentic AI

AI Strategy #1. AI Agents: Chatbots, RPA and Agentic AI Explained
AI Strategy #2. Multi-Agent Systems: When AI Agents Work Together
AI Strategy #3. Inside an AI Agent: Reasoning, Enterprise Context, Tools, Memory and Control
AI Strategy #4. Why AI Agents Fail: Seven Enterprise Failure Patterns Beyond the Model
AI Strategy #5. Agentic AI in 2026: From Market Hype to Enterprise Reality

Part 2 — Enterprise AI Adoption & Value

AI Strategy #6. Enterprise AI Maturity: Assessing Readiness Before Scaling
AI Strategy #7. Building an AI Power-User Organization: From Individual Skill to Enterprise Capability

Comments

Popular posts from this blog

AI Strategy #1. AI Agents: Chatbots, RPA and Agentic AI Explained

MDM #9. Why Enterprise MDM Governance Fails After Go-Live — and How to Make Ownership Real

AI Strategy #17. Hybrid Cloud and GenAI: Designing Enterprise AI Infrastructure