AI Strategy #12. AI Security in the Agentic Era: From Prompt Injection to Identity and Tool Abuse

Agentic AI changes the security problem because a model is no longer limited to producing text. An enterprise agent can retrieve private data, maintain memory, choose tools, call APIs, communicate with other agents and alter the state of business systems.

A successful prompt injection against a chatbot may produce a bad answer. The same injection against an agent with access to email, source code, ERP or customer data can become an authorization, data-loss or transaction-integrity incident.

The important security boundary therefore moves beyond the model.

Agentic AI security is not primarily about stopping the model from receiving a malicious instruction. It is about ensuring that an untrusted instruction cannot acquire trusted identity, privileged data or dangerous execution authority.

This is why conventional cybersecurity remains essential but no longer sufficient by itself. Identity, least privilege, network segmentation, software supply-chain controls and incident response still apply. What changes is the execution path: probabilistic reasoning now sits between an external input and a potentially privileged enterprise action.

The Security Model Changes When AI Can Act

NIST's 2026 analysis of security considerations for AI agents found broad agreement among respondents that agentic systems introduce security concerns beyond those of conventional AI applications, while also concluding that established cybersecurity practices remain relevant and need to be adapted for agent environments.

NIST — Security Considerations for AI Agents

The fundamental difference is agency.

Security Dimension Conventional Application Agentic AI
Input Structured fields, API parameters, files Natural language, retrieved documents, webpages, email, images and outputs from other agents
Execution Predetermined program logic Dynamic planning and tool selection influenced by model reasoning
Identity Human or service identity Human, delegated agent and machine identity can interact in one workflow
State Application database and explicit session state Prompt context, RAG, agent memory, tool results and persistent state
Side Effects Defined application functions Agent may dynamically chain several legitimate functions into an unintended outcome
Failure Mode Software defect or compromised credential Also includes manipulated context, goal hijacking, excessive agency and unsafe tool orchestration

Prompt Injection Is an Architectural Problem, Not Just an Input-Filtering Problem

Prompt injection remains one of the most important LLM security risks. OWASP defines it as manipulation of an LLM through crafted input that causes unintended behavior. The injection can be direct, but it can also arrive indirectly through content the AI has been asked to process.

OWASP — LLM01:2025 Prompt Injection

Indirect prompt injection is particularly important for enterprise agents because the model routinely consumes content that the user did not author:

  • webpages,
  • email,
  • documents,
  • retrieved knowledge,
  • tickets and comments,
  • external APIs,
  • images and multimodal content, and
  • messages from other agents.

That content may be valid business data and still contain malicious instructions.

Trusted User Intent
+
Untrusted External Content
↓
LLM Reasoning
↓
Privileged Tool
=
Potential Confused-Deputy Problem

A content filter can reduce attack probability, but it should not be the only control. OWASP explicitly notes that prompt-injection defenses are not complete simply because RAG, fine-tuning or input filtering is present.

The more durable defense is to limit what a compromised reasoning process can do.

The Blast Radius Is Determined by Agency

OWASP's concept of Excessive Agency is especially useful for enterprise architecture. The risk arises when an LLM-based application is given more functionality, permissions or autonomy than the task requires.

OWASP — LLM06:2025 Excessive Agency

This means the severity of a prompt-injection vulnerability depends heavily on what sits behind the model.

Compromised Capability Potential Consequence
Read Public Documents Incorrect output or manipulation of the user's analysis
Read Private Enterprise Data Sensitive-information disclosure
External Communication Data exfiltration or unauthorized communication
Write Enterprise Data Integrity loss, unauthorized changes or workflow corruption
Execute Administrative Tools Privilege abuse, destructive changes or lateral impact

The security objective should therefore be:

Assume Reasoning Can Be Manipulated
↓
Constrain Identity, Data, Tools and Actions
↓
Limit Blast Radius

Eight Threat Patterns Matter More Than a Long List of AI Attacks

Rather than treating every AI security term as a separate problem, I would organize the enterprise threat model around eight attack patterns.

This is a Digital Future & Strategy practitioner grouping informed by OWASP, NIST and current cloud-security guidance. It is not an official threat taxonomy.

1. Direct and Indirect Prompt Injection

An attacker influences the agent through a user prompt or through content the agent retrieves and interprets.

The critical security question is not only whether the instruction can be detected. It is whether the resulting model behavior can cross a privileged boundary.

2. Agent Identity and Privilege Abuse

Agents frequently need access to enterprise resources. Using a broad shared service account creates a dangerous ambiguity: whose authority is the agent exercising, and which resources should it be able to reach?

Google Cloud's current guidance recommends treating agents as first-class identities and applying least privilege to their data and system access.

Google Cloud — Securing Agentic AI with Agent Identity and Perimeter Controls

3. Tool Misuse and Excessive Agency

An agent may use a legitimate tool in an unintended way, select the wrong tool, pass unsafe parameters or combine individually safe operations into an unsafe workflow.

A read-only research agent should not inherit delete, payment or administrative functions merely because they are available through the same integration.

4. Data Exfiltration

The most dangerous agent configurations combine access to confidential data with exposure to untrusted content and a route for external communication.

A compromised agent may not need to “break out” of the application. If its legitimate tools allow it to retrieve data and communicate externally, manipulation of the workflow may be enough.

5. Memory Poisoning

Agent memory creates persistence. Malicious or incorrect context written into long-lived memory can influence future behavior after the original attack is gone.

OWASP's agent-security guidance specifically identifies memory poisoning as a risk in systems that maintain persistent agent state.

OWASP — AI Agent Security Cheat Sheet

6. Agentic Supply-Chain Risk

Agents increasingly depend on third-party models, libraries, tool servers, plugins, skills and external services. These components can extend the effective privilege boundary of the AI system.

Model Context Protocol and similar tool interfaces can improve interoperability, but interoperability does not automatically establish trust. Each tool provider, server, package and capability still requires security evaluation.

7. Multi-Agent Goal and Trust Propagation

In a multi-agent workflow, one agent may consume another agent's output as trusted context. A compromised or malfunctioning agent can therefore influence downstream decisions without attacking every agent independently.

The system should not assume that “internal agent” means “trusted input.”

8. Model, Data and Evaluation Manipulation

Agent security does not replace conventional AI security. Data poisoning, model theft, malicious model artifacts and manipulated evaluation data remain relevant where enterprises train, fine-tune, host or procure models.

The important change is that these model-level weaknesses can now propagate into enterprise actions rather than remaining limited to prediction quality.

Agent Identity Should Be Separate from User Identity

Identity is becoming one of the most important design problems in agentic systems.

An agent can act:

  • on behalf of a specific user,
  • on behalf of an application,
  • as an enterprise-owned autonomous service, or
  • as part of another agent's delegated workflow.

Those cases should not all use the same credential model.

Identity Pattern Appropriate Use Security Requirement
User-Delegated Agent Agent performs work within a user's authority. Preserve user context and do not silently expand beyond the user's permitted scope.
Agent Service Identity Background or autonomous task owned by the enterprise. Dedicated identity, tightly scoped permissions and explicit business owner
Temporary Elevated Authority Specific operation requires additional privilege. Short-lived authorization, policy check and strong audit

Microsoft similarly emphasizes agent identity, data permissions, decision rights and continuous monitoring as core elements of mature agent governance and security.

Microsoft Learn — Agentic AI Governance and Security

Least Privilege Needs to Apply to Tools, Not Just Data

Traditional IAM tends to focus on which resource a principal may access. Agentic systems add another dimension: what action should the agent be able to perform through that resource?

An email summarization agent may need:

Read Message
✓

Draft Response
Maybe

Send Message
Not Automatically Required

Delete Mailbox Content
No

OWASP explicitly recommends minimizing both the number of tools available to an agent and the functionality exposed through each tool.

A broad tool such as “execute arbitrary SQL,” “run shell command,” or “call any internal API” dramatically increases blast radius. Purpose-specific tools with constrained parameters are easier to authorize, validate and audit.

Put a Deterministic Policy Gate Between Reasoning and Execution

A language model should not be the final authority on whether an enterprise action is permitted.

The agent may decide that a supplier should be blocked. A deterministic control should still verify whether the agent has authority to perform the change, whether required evidence exists, whether the supplier falls within the agent's scope and whether human approval is required.

Untrusted / Mixed Input
↓
Agent Reasoning
↓
Proposed Tool Action
↓
Deterministic Policy Gate
Identity · Scope · Parameters · Limits · Approval
↓
Enterprise Tool / API
↓
Audit + Outcome

This architecture assumes that the reasoning layer can be wrong or manipulated without allowing every reasoning error to become an enterprise transaction.

RAG Content Should Be Treated as Data, Not Instructions

RAG improves the factual grounding of AI responses, but it does not create a trusted security boundary.

A document retrieved from an approved enterprise repository can still contain malicious instructions, compromised content or text that was never intended to control an agent.

The security model should separate:

  • system policy — trusted instructions defining what the agent is allowed to do,
  • user intent — the authorized business request, and
  • retrieved content — evidence to analyze, not executable authority.

This distinction is conceptually simple and technically difficult because an LLM processes all three as contextual information. That is another reason action authorization cannot depend solely on prompt hierarchy.

Memory Requires Its Own Trust Model

Persistent agent memory creates value by allowing the system to retain preferences, history and working context. It also creates a new attack surface.

A durable-memory architecture should answer:

  • Who may write to memory?
  • Which agent or user can read it?
  • How long should it persist?
  • Where did the stored information originate?
  • Can untrusted content become durable memory automatically?
  • Can users inspect or correct stored state?
  • What event should invalidate previous memory?

I would separate at least three forms of state:

State Control Principle
Session Context Short-lived; isolated to the active interaction wherever practical.
User / Agent Memory Purpose-limited, access-controlled, traceable and correctable.
Authoritative Business State Should remain in governed enterprise systems rather than being inferred from agent memory.

The last distinction is particularly important. An agent remembering that a supplier “appears blocked” is not equivalent to reading the authoritative supplier status from the relevant enterprise system.

Multi-Agent Systems Need Explicit Trust Boundaries

Multi-agent architecture can distribute responsibilities across specialized agents, but it also creates transitive trust.

An orchestrator may assume that a research agent's answer is safe because the research agent is “internal.” A transaction agent may then act on that answer. If the research agent was influenced by malicious external content, the compromise can propagate through the system.

The safer model is:

Agent Output
≠
Authorization

Agent-to-agent messages should retain provenance, and downstream execution should still pass through the same authorization and policy boundaries that would apply to a direct request.

Third-Party Tools and MCP Servers Are Part of the Supply Chain

Agent interoperability is expanding quickly. Tool protocols and reusable agent skills can reduce integration cost, but they also make security review more important.

A third-party tool may gain access to:

  • enterprise credentials,
  • prompts and model context,
  • internal APIs,
  • files or databases, and
  • agent actions.

Organizations should therefore treat agent tools and tool servers similarly to other privileged software dependencies.

Security review should include:

  • publisher and package provenance,
  • requested permissions,
  • authentication model,
  • data flows,
  • network destinations,
  • update mechanism,
  • logging,
  • known vulnerabilities, and
  • ability to revoke access quickly.

Human Approval Helps — but Approval Fatigue Is a Security Risk

Human-in-the-loop review is often proposed as the answer to unsafe agent behavior. It is useful when the reviewer has enough information and the number of escalations remains manageable.

It becomes weaker when every routine action asks for confirmation.

Users may eventually approve actions without examining them carefully, especially when the AI usually behaves correctly.

The better design is risk-based:

Action Possible Control
Read Low-Sensitivity Data Automatic within least-privilege scope
Draft / Recommend Human consumes the output before execution
Low-Risk Reversible Action Bounded automation with deterministic policy and monitoring
Sensitive External Communication Explicit approval or other strong policy gate
Destructive / High-Impact Action Stronger authorization, confirmation, segregation of duties and recovery controls

Google's agent-security guidance similarly distinguishes human-approved execution from agent-only execution and notes that human approval itself can fail when users over-trust agent recommendations.

Google Cloud — AI Security and Safety for Agentic Tool Use

Security Testing Must Follow the Action Path

Traditional LLM testing often asks whether a model can be induced to produce prohibited or incorrect content.

An enterprise agent requires a wider test.

The red team should ask:

  • Can untrusted content alter the agent's objective?
  • Can the agent access information outside the user's legitimate scope?
  • Can an injected instruction trigger a privileged tool?
  • Can several individually allowed tools be chained into a prohibited outcome?
  • Can external data be written into durable memory?
  • Can one compromised agent influence another?
  • Can the agent export sensitive information through a legitimate communication channel?
  • Can security controls be bypassed after model, tool or prompt changes?

OWASP's 2026 red-teaming work explicitly treats agent identity, privilege escalation, data poisoning, prompt injection and emergent workflow behavior as lifecycle security concerns rather than only model-output problems.

OWASP — AI and Agentic Red Teaming

Observability Needs to Capture Authority, Not Just Prompts

Logging every prompt is not the same as creating an agent audit trail. It can also create a new sensitive-data repository if logging is implemented carelessly.

A production audit model should capture the information needed to reconstruct material actions without indiscriminately retaining confidential content.

Audit Element Question It Answers
User / Delegator Who initiated or authorized the workflow?
Agent Identity Which agent performed the action?
Model / Agent Version Which runtime configuration was active?
Context Provenance Which relevant enterprise or external sources influenced the decision?
Tool Invocation Which capability was requested and executed?
Policy Decision Why was the action allowed, blocked or escalated?
Result / Side Effect What actually changed in the enterprise system?
Human Intervention Was the action approved, overridden or reversed?

Every Production Agent Needs a Revocation Path

Security architecture should assume that an agent may eventually need to be contained.

The enterprise should be able to:

  • revoke the agent identity,
  • remove or narrow tool permissions,
  • disable external communication,
  • suspend memory writes,
  • switch the workflow to recommendation-only mode,
  • isolate the agent from critical data, and
  • stop execution without rebuilding the entire application.

The ability to reduce autonomy is as important as the ability to increase it.

Agent Authority Is Revocable Privilege,
Not a Permanent Property of the Model

A Seven-Layer Enterprise Agent Security Architecture

The following architecture is a Digital Future & Strategy practitioner framework. It is intended to connect conventional security controls with agent-specific risks.

Security Layer Primary Controls Threats Reduced
1. Content Trust Input screening, provenance, untrusted-content boundaries, output handling Direct / indirect prompt injection, malicious content
2. Identity Dedicated agent identity, user delegation, short-lived credentials, least privilege Privilege abuse, confused deputy, lateral access
3. Data & Memory Purpose-based access, data minimization, memory isolation, provenance and retention Exfiltration, memory poisoning, inappropriate cross-user context
4. Tool Boundary Purpose-specific tools, parameter validation, capability allowlisting Tool misuse, excessive functionality, command abuse
5. Runtime Policy Deterministic authorization, transaction limits, approval, segregation of duties Excessive autonomy, unsafe action execution
6. Observability & Response Agent telemetry, audit trail, anomaly detection, revocation, isolation and rollback Undetected abuse, cascading failures, prolonged compromise
7. Lifecycle & Supply Chain Threat modeling, red teaming, component inventory, vendor review, regression testing Model, tool and dependency compromise; security regression

A Supplier Agent Example

Consider an agent that evaluates supplier risk and can initiate a supplier-status review.

A weak architecture gives the agent:

  • access to the entire supplier database,
  • a general-purpose web browser,
  • write access to supplier master data, and
  • a privileged integration account.

The prompt may say, “Only recommend changes and never update supplier records.”

That is not a strong security boundary.

A stronger design would:

1. Give the agent a dedicated identity.

2. Expose only the supplier context required for the use case.

3. Separate external research from transaction execution.

4. Provide a specific "Create Supplier Review" tool rather than unrestricted master-data write access.

5. Validate supplier ID, scope and requester authority outside the LLM.

6. Require additional approval before high-impact master-data changes.

7. Log the source evidence, proposed action, policy decision and downstream result.

This design does not assume prompt injection has been eliminated. It assumes that the enterprise system should remain resilient even if the agent's reasoning is manipulated.

What Security Teams Should Measure

Counting blocked prompts is not enough. A secure agent architecture should measure both attempted abuse and the effectiveness of privilege boundaries.

Evidence Area Possible Measures
Prompt / Content Defense Detected adversarial inputs, bypass findings from red-team tests, regression failures
Identity Overprivileged agents, shared credentials, stale identities, failed authorization attempts
Tools Denied tool calls, out-of-scope parameters, unused privileged capabilities
Data Sensitive-data exposure, cross-user access, unauthorized external transfer attempts
Agent Behavior Policy violations, abnormal tool chains, escalation and rollback events
Response Time to identify, contain, revoke and recover from a material agent incident

No universal target applies across all agent workloads. A low-risk internal research agent and an autonomous financial or administrative agent should not share the same security thresholds.

Seven Questions for an Architecture Review

1. Which inputs can the agent consume that we do not fully trust?

2. Does the agent have its own identity, and can we explain whose authority it is exercising?

3. Which private data can the agent read, and does the task genuinely require all of it?

4. Which tools can the agent call, and are any of those tools more powerful than the intended task requires?

5. Which controls exist outside the model to prevent an unsafe action from executing?

6. Can persistent memory, third-party tools or another agent introduce untrusted state into future decisions?

7. Can we revoke the agent's authority and reconstruct every material action during an incident?

The Security Position

Agentic AI does not invalidate decades of cybersecurity practice. It makes several established principles more important: least privilege, complete mediation, separation of duties, secure defaults, defense in depth, strong identity and auditable execution.

What changes is where those controls must be applied.

Security can no longer stop at the application boundary because the agent dynamically interprets untrusted context and decides which capability to invoke. Enterprises therefore need controls around the entire chain from input to side effect.

Untrusted Content
→ Agent Reasoning
→ Identity
→ Data
→ Tool
→ Policy Gate
→ Business Action
→ Audit & Recovery

Microsoft's 2026 guidance reaches the same broader conclusion: secure agent adoption depends on identity, data access, policy, lifecycle governance and continuous monitoring—not only on model safeguards.

Microsoft Security — What Is Agentic AI Security?

The safest enterprise agent is not the agent that can never be manipulated. It is the agent whose manipulation cannot silently become unauthorized enterprise action.

Sources & Further Reading

Method Note
The eight threat patterns, identity model, seven-layer enterprise agent-security architecture and architecture-review questions in this article are Digital Future & Strategy practitioner frameworks informed by NIST, OWASP, Microsoft and Google Cloud security guidance. They are not official threat taxonomies or mandatory security standards. The original version of this article contained unsupported incident counts, growth percentages and loss estimates; these have been removed. Security controls should be selected according to the agent's data access, tools, authority, business consequence, reversibility, deployment environment and threat model.

Reviewed: September 2026


AI Strategy Series

Part 3 — AI Governance, Security & Regulation

AI Strategy #11. Enterprise AI Governance: From Policy to Runtime Control
AI Strategy #12. AI Security in the Agentic Era: From Prompt Injection to Identity and Tool Abuse
AI Strategy #13. Global AI Regulation in 2026: Building One Enterprise Control Architecture Across Jurisdictions

Previous: Enterprise AI Governance: From Policy to Runtime Control

Next: Global AI Regulation in 2026: Building One Enterprise Control Architecture Across Jurisdictions

Comments

Popular posts from this blog

AI Strategy #1. AI Agents: Chatbots, RPA and Agentic AI Explained

MDM #9. Why Enterprise MDM Governance Fails After Go-Live — and How to Make Ownership Real

AI Strategy #17. Hybrid Cloud and GenAI: Designing Enterprise AI Infrastructure