AI Strategy #3. Inside an AI Agent: Reasoning, Enterprise Context, Tools, Memory and Control
An AI agent is often described as an LLM with access to tools.
That description is technically convenient, but strategically incomplete.
A production enterprise agent must do much more than generate language. It has to understand an objective, determine what information is missing, retrieve the right enterprise context, decide whether an external action is necessary, invoke the correct tool with valid parameters, maintain state across multiple steps and operate within explicit business and security boundaries.
The useful architecture is therefore not:
It is closer to:
+ Enterprise Context
+ Retrieval
+ Tools
+ Memory
+ Identity & Authority
+ Evaluation & Runtime Control
This distinction has become more important as Agentic AI moves beyond chat interfaces and begins interacting with enterprise systems of record.
An enterprise AI agent is not a model that happens to know how to call APIs. It is a governed decision-and-action system built around a model.
This article explains the major architectural components inside an AI agent, how they interact, where each component fails and what changes when the architecture moves from a prototype to enterprise production.
The Architecture of an Enterprise AI Agent
The easiest way to understand an agent is to separate the major responsibilities rather than treat the LLM as the entire system.
| Component | Primary Role | Typical Failure |
|---|---|---|
| LLM / Reasoning | Interprets goals, reasons over context, plans steps and decides what to do next. | Incorrect reasoning, unsupported assumptions or poor planning |
| Enterprise Context | Provides business entities, policies, documents, transaction state and organizational meaning. | Conflicting, stale or semantically inconsistent information |
| Retrieval / RAG | Finds relevant external information when the model needs it. | Wrong document, missing evidence or unauthorized retrieval |
| Tools | Allow the agent to query systems, calculate, create records or execute actions. | Wrong tool, invalid parameters or unsafe action |
| Memory | Preserves useful state or knowledge across steps and, where appropriate, across sessions. | Stale, excessive, incorrect or privacy-sensitive memory |
| Identity & Authority | Determines whose authority the agent uses and which actions are permitted. | Excessive privilege or ambiguous delegation |
| Evaluation & Control | Measures behavior, enforces policy and detects when intervention is required. | Failures become visible only after business impact occurs |
The architectural mistake is to optimize one component independently and assume the overall agent will improve proportionally.
A more capable LLM cannot reliably compensate for an incorrect supplier identity. Better retrieval cannot make an unsafe transaction permission acceptable. Longer memory cannot repair an API that exposes the wrong business operation.
The quality of the agent therefore depends on the interaction among these components.
1. The LLM Is the Reasoning Engine — Not the System of Record
The language model is the most visible component of an AI agent because it interprets instructions and determines what should happen next.
Depending on the task, the model may:
- interpret user intent,
- decompose a goal into smaller tasks,
- decide which information is missing,
- select an appropriate tool,
- generate tool parameters,
- interpret tool results,
- revise a plan after failure, and
- decide whether the task is complete.
This is fundamentally different from traditional software.
In conventional software, developers explicitly encode most process logic.
→ Predetermined Logic
→ Predetermined Action
An agent introduces model-mediated decisions:
→ Observe Context
→ Reason
→ Select Action
→ Observe Result
→ Reason Again
→ Continue or Stop
This flexibility is the source of both the value and the risk of Agentic AI.
The model can respond to situations that developers did not explicitly enumerate, but that also means the execution path cannot always be predicted in advance.
Reasoning Does Not Mean the Model Knows the Business
A powerful model may understand procurement concepts, accounting terminology or customer-service principles.
It does not automatically know:
- which supplier record is authoritative inside your company,
- which version of a contract is currently effective,
- which employee can approve a particular transaction,
- which products belong to a specific hierarchy,
- which internal policy applies in a particular jurisdiction, or
- what happened in an enterprise system five minutes ago.
Those facts exist outside the model.
This is why enterprise context increasingly matters as much as model intelligence.
The LLM provides general reasoning capability. The enterprise must provide the business reality over which that reasoning operates.
2. Enterprise Context Is More Than RAG
RAG—Retrieval-Augmented Generation—is often described as the mechanism that gives an LLM enterprise knowledge.
That is only partly correct.
Retrieval is a mechanism for obtaining context. Enterprise context is the larger business layer that determines what information is meaningful, authoritative, current and permitted.
For a supplier-management agent, relevant context might include:
- supplier master identity,
- legal-entity hierarchy,
- approved supplier status,
- contracts,
- quality history,
- purchase transactions,
- risk information,
- internal policy, and
- the requester's organizational authority.
Some of this information belongs in documents. Some belongs in transactional systems. Some belongs in master data. Some may need to be calculated in real time.
RAG alone does not unify those layers.
RAG Is a Retrieval Pattern, Not a Truth Engine
A typical RAG architecture converts enterprise content into searchable representations and retrieves relevant information before or during model reasoning.
→ Ingestion & Parsing
→ Metadata / Indexing / Embeddings
→ Retrieval
→ Selected Context
→ Model Reasoning
This architecture can dramatically improve relevance, but retrieval does not establish truth automatically.
A retrieval system can return:
- the wrong document,
- an obsolete document,
- a document from the wrong business unit,
- information about a similarly named entity,
- multiple documents containing conflicting rules, or
- content the user or agent should not have been permitted to retrieve.
The enterprise problem therefore extends beyond vector similarity.
| Retrieval Question | Enterprise Requirement |
|---|---|
| Is this relevant? | Search, ranking and retrieval quality |
| Is this authoritative? | Source ownership and system-of-record rules |
| Is this current? | Effective dates, refresh policies and event awareness |
| Is this about the correct entity? | Master data, entity resolution and hierarchy |
| May this agent retrieve it? | Identity-aware access control |
| Can we prove where it came from? | Metadata, provenance and citation evidence |
This is one reason enterprise RAG and enterprise data governance increasingly converge.
Context Engineering Is Replacing the Idea of “Put Everything in the Prompt”
As agentic systems become more capable, the design problem is shifting from prompt engineering toward context engineering.
The objective is not to provide the model with the largest possible amount of information.
It is to provide the right information at the right stage of the task.
Anthropic describes a growing shift toward just-in-time context, where agents dynamically retrieve information when required rather than loading every potentially useful document, tool definition and previous interaction into the context window in advance.
Anthropic — Effective Context Engineering for AI Agents
The difference can be represented as:
| Approach | Advantage | Risk |
|---|---|---|
| Preload Everything | Simple architecture for small tasks | Noise, token cost and outdated context |
| Static RAG | Efficient retrieval from indexed knowledge | May miss live transactional or task-specific context |
| Just-in-Time Context | Agent retrieves information as the task evolves | Requires reliable tools, search and decision logic |
This produces a more useful architecture for long-running enterprise agents.
Instead of trying to make the prompt contain the enterprise, the agent should know how to navigate the enterprise.
3. Tools Turn Reasoning into Action
Retrieval allows an agent to obtain information.
Tools allow it to change the world outside the model.
Examples include:
- querying ERP inventory,
- creating a CRM record,
- submitting a purchase request,
- checking an external data source,
- running a calculation,
- executing code,
- sending a business notification, or
- initiating an approval workflow.
The presence of tools is one of the major architectural differences between an assistant and an operational agent.
→ Recommendation
Reasoning + Governed Tools
→ Potential Business Execution
A Tool Is Not Just an API
Traditional APIs are usually designed for deterministic software written by developers.
Agent tools have a different consumer: a probabilistic model that must understand what the capability does, decide when to use it and construct valid inputs.
Anthropic describes this as a new software-design problem: tools create an interface between deterministic systems and non-deterministic agents.
Anthropic — Writing Effective Tools for AI Agents
A production-quality agent tool therefore needs more than endpoint connectivity.
| Tool Design Requirement | Why It Matters |
|---|---|
| Clear Purpose | The agent must know when this tool should and should not be used. |
| Unambiguous Parameters | Poor parameter definitions increase invalid or unintended actions. |
| Bounded Capability | A narrow business action is safer than unrestricted system access. |
| Useful Response | The model needs enough structured context to decide what to do next. |
| Error Semantics | The agent must distinguish retryable errors from business-rule failures. |
| Authorization | Tool availability must reflect user, agent and transaction authority. |
| Auditability | Material actions need traceable inputs, decisions and outcomes. |
This is why exposing hundreds of existing APIs directly to an agent does not automatically create an effective tool architecture.
The enterprise needs an agent-facing capability layer.
Tool Selection Is Itself an AI Decision
A traditional application knows which function will execute next because developers encoded that sequence.
An agent may need to choose.
For a customer refund request, an agent might need to decide whether to:
- retrieve the order,
- check delivery status,
- read the return policy,
- verify the customer's identity,
- calculate refund eligibility,
- request human approval, or
- execute a permitted refund.
The workflow can branch depending on the information returned at each step.
↓
Retrieve Order
↓
Check Policy & Eligibility
↓
Determine Authority
↓
Prepare or Execute Action
↓
Confirm Outcome
The important point is not that the agent can call these tools.
It is that each tool should preserve business rules and authority boundaries regardless of what the model decides.
4. Memory Is Not the Same as Context
Memory is one of the most frequently misunderstood parts of agent architecture.
A large context window is not the same as durable memory.
RAG is not the same as memory.
And storing every past interaction indefinitely is not a mature memory strategy.
A useful practitioner distinction is:
| Memory Type | What It Contains | Example |
|---|---|---|
| Working State | Information needed to complete the current task. | Current supplier, task status and pending checks |
| Session Memory | Useful information retained throughout one extended interaction. | Decisions made earlier in a multi-step case |
| Persistent Memory | Selected information intentionally retained beyond the session. | Approved user preference or durable workflow note |
| Enterprise Knowledge | Authoritative facts managed outside the agent. | Supplier master, policy, contract or product definition |
This four-part classification is a Digital Future & Strategy practitioner model. Terminology varies across vendors and agent frameworks.
Enterprise Knowledge Should Not Be Replaced by Agent Memory
This distinction is strategically important.
If the authoritative payment term for a supplier changes in ERP, the agent should not continue relying on an old remembered value.
If an employee changes roles, persistent agent memory should not override the current identity and authorization system.
Enterprise facts should generally remain governed by their authoritative systems.
≠
Agent Memory
Memory should help an agent preserve useful task continuity—not create a parallel uncontrolled master-data system.
Structured Memory Becomes Important for Long-Horizon Work
Long-running agents face a practical constraint: every intermediate observation, tool result and conversation cannot remain in active context forever.
One emerging pattern is structured note-taking.
The agent periodically writes concise task state or important findings to persistent storage and retrieves them when needed later.
Anthropic describes this pattern as agentic memory and uses structured notes as one mechanism for preserving continuity when tasks extend beyond a single context window.
Anthropic — Effective Context Engineering for AI Agents
The design question is therefore not simply:
It is:
for how long,
under whose ownership,
and with what mechanism for correction or deletion?”
Memory Is Also a Governance Problem
Persistent memory creates lifecycle questions that simple chat applications can often avoid.
Enterprises need to decide:
- what information may be remembered,
- which information must never be persisted,
- how long memory should remain valid,
- who can access stored memory,
- how outdated memory is refreshed,
- how incorrect memory is corrected,
- how retention policies apply, and
- how memory is removed when business or legal requirements demand it.
A memory architecture without ownership and lifecycle controls eventually becomes another uncontrolled data store.
5. Identity and Authority Determine What the Agent Is Allowed to Do
Once an agent moves from information retrieval to action, identity becomes part of the architecture.
At least three identities can matter:
- the human requesting the work,
- the agent performing the reasoning, and
- the service or application identity used to execute a tool call.
The enterprise also needs to know which authority is being delegated.
Consider a procurement agent.
An employee may have permission to view a supplier record but not approve the supplier. A manager may approve only within a particular region. A central procurement role may have different authority again.
The agent should not automatically receive the union of all available permissions.
Its authority should be explicitly bounded.
+ Agent Identity
+ Delegated Authority
+ Tool Permission
+ Transaction Constraint
=
Allowed Action
This is the control boundary that separates a useful tool-using agent from an unsafe automation layer.
Capability and Permission Must Remain Separate
The model may be technically capable of generating the correct parameters for a payment change.
That does not mean it should be permitted to execute the change.
The architecture should separate:
| Question | Owner |
|---|---|
| Can the model do it? | Model capability and evaluation |
| Should the agent do it? | Business policy and risk classification |
| May this agent do it now? | Runtime identity, authorization and transaction control |
This separation becomes increasingly important as model capabilities improve faster than organizational governance processes.
6. The Orchestration Loop Turns Components into an Agent
LLM, retrieval, tools and memory are individual capabilities.
The agent emerges when these capabilities operate inside a loop.
A simplified agent loop is:
↓
2. Inspect Current Context
↓
3. Decide Next Step
↓
4. Retrieve Information or Use Tool
↓
5. Observe Result
↓
6. Update State / Memory
↓
7. Continue, Escalate or Stop
This loop is the architectural difference between a one-shot LLM response and a genuine agentic system.
Anthropic uses a similarly simple conceptual definition: agents are systems in which models dynamically direct their own processes and tool usage.
Anthropic — Building Effective Agents
More Agentic Loops Do Not Automatically Mean a Better System
Autonomy creates flexibility, but each additional reasoning step also creates another opportunity for error, latency and cost.
For predictable processes, a deterministic workflow may remain superior.
| Task Characteristic | Likely Starting Architecture |
|---|---|
| Stable deterministic rules | Application logic or workflow engine |
| Simple document question answering | RAG-enabled assistant |
| Predictable multi-step process | LLM-enabled workflow with explicit orchestration |
| Dynamic task with uncertain path | Agent with bounded tool autonomy |
| High-consequence irreversible decision | AI-assisted human decision unless stronger evidence justifies delegation |
The architecture should be only as agentic as the workflow requires.
7. Evaluation and Observability Are Part of the Architecture
An agent should not be considered production-ready simply because the model performs well in a benchmark or a demonstration succeeds.
Agents are difficult to evaluate precisely because they operate across multiple turns, retrieve changing context, invoke tools and modify state.
Anthropic's 2026 guidance on agent evaluation emphasizes that the capabilities that make agents useful—autonomy, flexibility and multi-step execution—also create multiple failure points that do not exist in a single model response.
Anthropic — Demystifying Evals for AI Agents
NIST's 2026 TEVV-Athlon draft similarly treats evaluation as application-specific and explicitly includes agentic systems within its scope.
NIST — TEVV-Athlon Framework for Evaluating AI Systems
A useful evaluation chain is:
→ Context
→ Reasoning
→ Tool Selection
→ Tool Parameters
→ Action
→ Recovery / Escalation
→ Business Outcome
Each layer creates a different evaluation question.
What Should Be Measured?
| Dimension | Example Measure |
|---|---|
| Task Completion | Was the requested workflow completed correctly? |
| Context Quality | Did the agent retrieve the correct entity, source and current information? |
| Tool Selection | Did the agent choose the appropriate tool? |
| Tool Parameters | Were the action parameters complete and valid? |
| Policy Compliance | Did the agent remain within the permitted action boundary? |
| Recovery | Did it recover or escalate appropriately when a tool failed? |
| Economics | What did a successfully completed workflow cost? |
| Business Outcome | Did the completed action improve the target process or result? |
The evaluation system should also retain enough execution evidence to reconstruct what happened.
Agent Observability Must Go Beyond Uptime
Traditional application monitoring asks questions such as:
- Is the service available?
- What is the latency?
- Did the API return an error?
Agent observability requires additional questions:
- Which context did the agent use?
- Which tools did it consider?
- Which tool did it select?
- What parameters were submitted?
- Which policy checks were triggered?
- Was human approval requested?
- What changed in the system of record?
- What did the workflow cost?
- Did the business outcome match the intended result?
The relevant production trace is therefore not simply an application log.
→ Context Retrieved
→ Model Decision
→ Tool Call
→ Policy Decision
→ System Change
→ Final Outcome
Without this trace, debugging an agent becomes guesswork.
A Practical Example: An Enterprise Refund Agent
Consider a retail or service organization using an AI agent to handle a customer refund request.
The customer says:
“My order arrived damaged. Can you refund it?”
A simplistic chatbot can draft a sympathetic answer.
A production agent has to do significantly more.
| Step | Agent Capability | Enterprise Dependency |
|---|---|---|
| Understand Request | LLM reasoning | Correct interpretation of customer intent |
| Identify Customer & Order | Tool + enterprise context | Reliable customer and order identity |
| Check Delivery | Transactional tool | Current fulfillment status |
| Retrieve Policy | RAG / retrieval | Correct and effective refund policy |
| Determine Eligibility | Reasoning + policy | Correct application of rules and exceptions |
| Check Authority | Runtime control | Refund amount and approval limit |
| Execute or Escalate | Tool use | Payment API and approval workflow |
| Record Outcome | State / memory / audit | Case history and traceability |
The LLM is essential, but it is only one part of the operating system required to complete the refund safely.
Where Can the Refund Agent Fail?
Almost everywhere.
| Failure | Root Cause |
|---|---|
| Wrong order selected | Identity or retrieval failure |
| Old refund policy applied | Freshness or source-governance failure |
| Wrong refund amount calculated | Reasoning or deterministic-calculation design failure |
| Agent issues refund above allowed limit | Authorization and control failure |
| Refund submitted twice after retry | Tool and transaction-design failure |
| Incorrect decision persists into next case | Memory-management failure |
This is why the architecture of an agent must be evaluated end to end.
The Enterprise Agent Control Plane
The components described above can be summarized as a three-layer architecture.
Goal · Outcome · Rules · Human Decision Rights
↓
AGENT INTELLIGENCE
Reasoning · Context · Retrieval · Memory · Planning
↓
CONTROLLED EXECUTION
Identity · Tools · APIs · Authorization · Validation
↓
ENTERPRISE SYSTEMS
ERP · CRM · SCM · MDM · Data · Documents · External Services
↓
OBSERVABILITY & GOVERNANCE
Evaluation · Trace · Policy · Audit · Escalation · Recovery
This architecture makes an important point.
Governance is not an external committee sitting above the agent.
Controls must exist inside the path between reasoning and business execution.
Do Not Confuse the Components
Many weak agent designs begin by asking one technology to solve another technology's problem.
| Problem | Wrong Response | Better Response |
|---|---|---|
| Supplier identities do not match. | Increase prompt detail. | Fix entity identity and master-data resolution. |
| Policy changes frequently. | Fine-tune the model with policy text. | Retrieve the current authoritative policy at runtime. |
| Agent must calculate tax precisely. | Ask the LLM to calculate it. | Use deterministic calculation or an authoritative service. |
| Agent performs unauthorized action. | Add a stronger system prompt. | Enforce authorization at the tool and transaction layer. |
| Long-running task loses important context. | Load the entire history every time. | Use compaction, structured state and selective memory. |
The architecture becomes more reliable when each component performs the job it is designed to perform.
From Prompt Engineering to Agent Engineering
The evolution from chatbot to enterprise agent changes the engineering discipline.
| Generation | Primary Design Question | Primary Engineering Focus |
|---|---|---|
| Prompt-Centric AI | How do we ask the model effectively? | Prompt design and response quality |
| RAG-Centric AI | How do we provide external knowledge? | Ingestion, retrieval and grounding |
| Agentic AI | How do we let AI complete work safely? | Context, tools, state, identity, evaluation and control |
Prompt engineering remains useful.
It is simply no longer sufficient.
A Production-Readiness Checklist for Agent Architecture
Before moving an agent from demonstration into a real business workflow, enterprises should be able to answer at least the following questions.
1. What business outcome is the agent responsible for?
2. Which decisions are delegated to the model and which remain deterministic?
3. Which enterprise entities must the agent identify correctly?
4. Which information sources are authoritative?
5. Does retrieval enforce freshness, ownership and access policy?
6. Are agent tools designed specifically for safe model use?
7. Can the agent distinguish a retryable technical error from a business-rule rejection?
8. What information is retained in memory, and why?
9. How are stale or incorrect memories corrected?
10. What identity does the agent use when invoking each tool?
11. What authority is delegated from the human user?
12. Which actions require explicit approval?
13. Can every material action be reconstructed after execution?
14. Are production failures converted into future evaluation scenarios?
15. Can the enterprise reduce or revoke agent authority without rebuilding the application?
The Architecture Decision That Matters Most
The central Agentic AI architecture question is not which model, vector database or agent framework is best.
Those decisions matter, but they sit below a more fundamental design choice:
which decisions may be delegated to AI,
and which business actions may follow from those decisions?
That decision determines the required data quality, tool design, memory architecture, identity controls, evaluation depth and human oversight.
The architecture should therefore begin with the business consequence of the agent's decisions rather than the sophistication of the model.
The Enterprise Agent Architecture Position
The LLM remains the reasoning engine of an AI agent, but it is not the complete agent.
RAG gives the system access to relevant knowledge, but retrieval does not guarantee that the knowledge is current, authoritative or associated with the correct enterprise entity.
Tools allow the model to act, but those tools must translate probabilistic reasoning into tightly controlled deterministic capabilities.
Memory gives an agent continuity, but uncontrolled memory can become stale, sensitive or inconsistent with enterprise systems of record.
Identity and authorization determine whether an intelligent recommendation can become a permitted business action.
Evaluation and observability determine whether the organization can prove that the complete system continues to behave as intended.
The useful architecture is therefore:
× Trusted Context
× Governed Tools
× Managed Memory
× Explicit Authority
× Production Evaluation
If one critical dependency is weak, adding more model intelligence may not solve the production problem.
The architecture of an enterprise agent is not defined by how much the model can do. It is defined by how reliably the enterprise can provide context, authorize action, preserve state and control what the model is allowed to do.
Sources & Further Reading
- Anthropic — Building Effective Agents
- Anthropic — Effective Context Engineering for AI Agents
- Anthropic — Writing Effective Tools for AI Agents
- Anthropic — Demystifying Evals for AI Agents
- NIST — TEVV-Athlon Framework for Evaluating AI Systems
- OpenAI — Introducing the Agents API
The enterprise-agent component model, four-part memory classification, three-layer control architecture and production-readiness questions in this article are Digital Future & Strategy practitioner frameworks. They are not official Anthropic, OpenAI or NIST taxonomies. Terminology such as agent, memory, context and orchestration varies among model providers and software frameworks. This article therefore focuses on durable architectural responsibilities rather than vendor-specific product definitions. The earlier version of this article included an illustrative enterprise incident whose exact product configuration and source could not be independently verified; that example has been removed and replaced with generic workflow analysis. The architecture should always be adapted to the specific business process, data environment, authority level and consequence of agent actions.
Reviewed: September 2026
AI Strategy Series
Part 1 — Understanding Agentic AI
AI Strategy #1. AI Agents: Chatbots, RPA and Agentic AI Explained
AI Strategy #2. Multi-Agent Systems: When AI Agents Work Together
AI Strategy #3. Inside an AI Agent: Reasoning, Enterprise Context, Tools, Memory and Control
AI Strategy #4. Why AI Agents Fail: Seven Enterprise Failure Patterns Beyond the Model
AI Strategy #5. Agentic AI in 2026: From Market Hype to Enterprise Reality
Part 2 — Enterprise AI Adoption & Value
AI Strategy #6. Enterprise AI Maturity: Assessing Readiness Before Scaling
AI Strategy #7. Building an AI Power-User Organization: From Individual Skill to Enterprise Capability
Comments
Post a Comment