AI has changed how work gets done. Agents can plan tasks, query enterprise repositories, invoke tools, write code, call application programming interfaces (APIs), and take actions across connected systems. As they work, they create, copy, fragment, transform, and share data across workflows.
A perimeter and a one-time authorization decision do not provide enough control for this operating model. An agent can chain permissions, tools, integrations, and data sources while completing a task. When it inherits broad user or service-account access, it can make sensitive data easier to discover, retrieve, summarize, transform, and transmit.
Zero trust for AI agents means continuously verifying an agent’s identity, task, authorization, behavior, tool use, and data access throughout execution. Each meaningful action should receive a decision based on the current context.
How do AI agents access enterprise data? They commonly use delegated user permissions, service identities, API tokens, OAuth grants, Model Context Protocol (MCP) servers, software-as-a-service (SaaS) integrations, databases, cloud storage, developer tools, and retrieval-augmented generation (RAG) systems.
The Anthropic framework for AI agents identifies controls enterprises need for this environment. The implementation challenge is enforcing those controls across agents, AI applications, data stores, developer environments, and third-party connections. Cyberhaven participates in Anthropic’s Cyber Verification Program to advance defensive AI research related to sensitive-data exposure, insider risk, and data exfiltration.
Learn more about Cyberhaven’s involvement in Anthropic’s Cyber Verification Program.
The Six Pillars of Anthropic’s Zero Trust Framework
The Anthropic framework for AI agents provides a practical structure for securing autonomous systems. Teams should operationalize the framework as a continuous lifecycle, with runtime enforcement and evidence collection throughout each agent execution.
1. Agent Identity and Authentication
Every agent needs a distinct, verifiable non-human identity. That identity should link to the workload, owner, deployment context, software version, approved capabilities, and current session.
Shared service accounts and long-lived API keys make it difficult to distinguish one agent from another. They also make it difficult to determine whether access should continue after a task ends. Cryptographically rooted identities and short-lived credentials provide a stronger foundation for AI agent security.
Before an agent acts, security teams should confirm:
- Who deployed the agent
- Which version is running
- Which tools the agent may invoke
- Whether the identity is valid for the environment and session
- Whether the agent has authority for the requested task
A customer-support agent and a software-development agent may use the same model. Each still needs a separate identity, tool permissions, and policy scope.
2. Access Control and Privilege Management
AI agent access control should limit permissions to the current task, required tools, and minimum data needed. Access should expire when the task ends.
Broad delegated permissions increase exposure because agents can search, reason over, and act on enterprise data quickly. Task-scoped authorization constrains the effects of a compromised prompt, tool, session, or integration.
Consider an agent assigned to summarize a support case. It may read a defined ticket and an approved knowledge base. It should not export attachments, search unrelated customer records, modify customer relationship management (CRM) entries, or send files to an external destination without further authorization.
Apply these controls:
- Just-in-time authorization: Grant access for a specific task and time period.
- Tool-level permissions: Limit the tools an agent can call and the operations each tool permits.
- Data-aware policy: Evaluate data sensitivity before allowing retrieval, transformation, or transfer.
- Approval gates: Require human review for high-consequence actions.
- Automatic expiration: Remove access when the task, session, or approved window ends.
Static credentials and startup-only checks do not provide continuous zero trust.
3. Observability and Auditing
Generic AI usage logs show that a user or system invoked a model. Full agent observability shows what an agent did, under whose authority, with which data, and where the results went.
Security teams should be able to reconstruct:
- Agent identity and initiating user or system
- Model, prompt, and input
- Retrieved files, records, and other content
- Tool calls and MCP server calls
- Permissions used and policy decisions
- Actions taken and outputs generated
- Data destinations and outbound transfers
This evidence connects agent behavior to data access and transformation. It supports incident response, policy review, and governance.
Enterprise controls for AI code assistants need the same level of detail. Log repository access, code searches, secret retrieval attempts, shell commands, continuous integration and continuous delivery (CI/CD) tool calls, generated pull requests, and outbound actions.
4. Behavioral Monitoring and Response
An agent’s behavior can change during execution. It may receive an indirect prompt injection through retrieved content, call a new chain of tools, attempt access beyond its assigned scope, or move data to an unexpected destination.
Monitor for signals such as:
- Unusual data volume
- New or unexpected tool chains
- Privilege escalation attempts
- Repeated policy denials
- Access outside a task’s expected scope
- Requests to new destinations
- Instructions that redirect the agent’s original objective
An agent can complete multi-step actions quickly, so alerts need predefined response actions. Revoke task access, terminate sessions, block a tool or destination, require step-up approval, or disable a compromised integration when policy conditions require containment.
5. Input Validation and Output Controls
Agent inputs extend beyond the original user prompt. An agent may process instructions from documents, web pages, retrieved content, tool descriptions, API responses, memory, and other agents.
Controls should inspect inputs, tool instructions, tool outputs, retrieval content, generated responses, and downstream actions. This addresses prompt injection, indirect prompt injection, malicious tool descriptions, tool poisoning, jailbreak attempts, sensitive-data disclosure, unsafe code generation, and unauthorized data transfer.
Keyword filters cannot determine whether an action fits the assigned task. Context-aware controls should evaluate data sensitivity, the recipient, the requested action, and the agent’s active authorization.
For example, a document may contain hidden text instructing an agent to retrieve confidential files and send them to an external endpoint. A policy should identify that instruction as outside the assigned task and outside the agent’s approved destination scope.
6. Integrity and Recovery
Agents rely on instructions, memory, configurations, tool definitions, model connections, and policy settings. Each can be changed or poisoned.
Maintain the following two main controls:
- Approved tool registries: Restrict available tools, plugins, MCP servers, and external connections.
- Change monitoring: Record changes to an agent, tool definition, permission, or integration.
Dependencies, plugins, MCP servers, external tools, and third-party AI services expand the agent’s trust boundary. Integrity controls must account for every connection.
The Framework Is Sound. Enforcement Is What Matters.
Most enterprises do not operate one agent with one identity and one data source. They use sanctioned and unsanctioned AI applications, SaaS-embedded agents, custom copilots, AI code assistants, MCP servers, cloud-hosted models, collaboration platforms, databases, endpoints, and third-party integrations.
Point tools focused only on prompts, models, or inventories can leave data and execution paths outside the control decision. To enforce zero trust for agents, teams need runtime decisions informed by identity, task intent, tool chain, data classification, Data Lineage, destination, and current risk.
Runtime authorities and secure gateways are useful design patterns for continuous authorization. They place a policy decision point between an agent and a target system, allowing teams to inspect, authorize, and record requests as the agent operates.
What Good Enforcement Looks Like
- Discover agents and connections: Identify AI applications, agents, integrations, MCP servers, and their owners.
- Attribute every action: Link activity to a human, service, and agent identity.
- Authorize continuously: Evaluate actions throughout a session, rather than only at startup.
- Limit task privileges: Grant the smallest practical scope for individual tasks and tools.
- Inspect data movement: Evaluate inputs, retrieval, outputs, and destinations in context.
- Detect behavioral drift: Identify anomalous actions, bypass attempts, and unexpected tool chains.
- Contain elevated risk: Block, isolate, terminate, or revoke access when policy requires it.
- Retain execution evidence: Preserve an end-to-end record for investigation and audit.
This is the practical answer to how to secure autonomous AI agents in enterprise environments.
How Cyberhaven Enforces Zero Trust for AI Agents
Cyberhaven provides Data Security for the Agentic Enterprise. Cyberhaven traces the full lifecycle of data and adapts protection as context changes. The lifecycle includes where data comes from, who or what touches it, how it changes, and where it goes next.
Effective AI agent security requires that data context. An enforcement decision needs more than an agent name and prompt. It needs the data’s sensitivity, the authority behind the request, the tool in use, the workflow state, and the destination. Cyberhaven Agentic Data Security applies Data Lineage, risk scoring, visibility, and runtime guardrails across AI workflows.
Discover: Build a Living Inventory of AI Agents
Teams cannot govern unknown agents, locally installed tools, browser-based AI use, unapproved integrations, or MCP servers. Shadow agents extend the shadow IT problem because teams can deploy embedded and local agents before central governance records them.
Cyberhaven discovers AI activity across endpoints, SaaS, developer environments, and web. A useful agent inventory includes:
- Ownership and business purpose.
- Model or provider.
- Tools, data sources, and MCP connections.
- Identities and permission scopes.
- Risk level and sanctioned status..
Assess: Prioritize Agentic AI Security Risks With Data Context
Agentic AI security risks include exposed sensitive data, over-privileged tools, risky identity delegation, insecure external integrations, unapproved data destinations, prompt injection exposure, and insufficient auditability.
“The agent can access SharePoint” does not provide enough context for a risk decision. Teams need to know whether the agent can reach regulated records, source code, credentials, financial information, or sensitive employee data. They also need to know whether it can retrieve, summarize, modify, or transfer that data.
Risk scoring should consider data sensitivity, access scope, action capability, exposure path, user and agent identity, and destination. Data lineage turns a raw alert into an investigation path: What data did the agent access? Where did it come from? Under whose authority did the agent act? Where did the data go?
Enforce: Apply Runtime Guardrails to Data Access
Runtime guardrails should evaluate meaningful execution steps, including retrieval, tool use, file access, code actions, API calls, and outbound data movement.
Risk-based controls can allow, block, redact, require approval, limit data fields, redirect work to a safe workflow, or terminate an agent session. Decisions should use data and execution context instead of static keywords alone.
Govern: Extend Controls Across AI Applications
Governance includes third-party AI applications used by employees and business teams. Policies need to extend across SaaS AI, endpoints, browsers, web, and developer workflows.
During AI vendor evaluation, review:
- Data retention and training-use policies.
- Administrator controls and identity integration.
- Compliance APIs and audit records.
- Integration scope and data access paths.
- Controls for sensitive data before it reaches an AI application.
Cyberhaven announced compliance API integrations with ChatGPT Enterprise and Claude Enterprise. These integrations extend discovery, classification, monitoring, and incident response into approved AI application workflows.
Monitor: Reconstruct the Full Execution Lifecycle
Isolated alerts require analysts to infer what happened. A lifecycle record connects the original task, prompts, retrieved content, tool calls, policy decisions, agent actions, and data movement.
This context helps analysts determine whether activity was expected, excessive, or malicious. Connecting activity to Data Lineage and user context can reduce false positives and speed investigations.
Test: Validate Agents Against Attacks and Drift
Pre-deployment assessment is necessary, but agent risk changes as models, prompts, tools, data sources, integrations, and business workflows change.
Test recurring scenarios that include:
- Risky outbound actions and data exfiltration.
- Prompt injection and malicious retrieval content.
- Tool poisoning and unauthorized data extraction.
- Unsafe code generation and developer-tool misuse.
- Privilege escalation attempts.
Zero Trust for AI Agents Requires Data Context
AI agent security and data security must work together because the consequence of an agent failure depends on the data an agent can access and move.
An agent can authenticate successfully and still expose sensitive data through excessive underlying permissions. The model and prompt are parts of the attack surface. Identities, tools, permissions, integrations, endpoints, and data flows determine what the agent can do.
Effective decisions require:
- Data classification and sensitivity: Identify the type and risk of the data involved.
- Data Lineage and movement history: Trace where data originated, changed, and traveled.
- Source and destination: Evaluate the data store, recipient, application, and transfer path.
- Identity and authorization chain: Link the agent to the initiating user, service, workload, and active permissions.
- Business purpose and task intent: Confirm that access and actions match the assigned objective.
- Tools and action history: Evaluate calls, integrations, prior steps, and outputs.
- Environmental risk: Consider the user, device, application, and destination.
Learn more in Governing the Autonomous Enterprise: Agentic AI Security.


.avif)
.avif)
