HomeBlog

How to Build a Data Security Program for Generative AI

No items found.

October 1, 2026

•

1 min

|

Updated:

October 1, 2026

How to Build a Data Security Program for Generative AI
In This Article

Ask five security leaders what "data security for generative AI" covers and you'll get five different answers. Some may point to an existing DLP rule for ChatGPT, while others would point to an AI acceptable use policy. Some operate under the assumption that their DSPM vendor already handles genAI security.

The confusion isn't about effort, as most teams are actively trying to get ahead of GenAI risk. It's that no single capability answers the question on its own, and a vendor selling a point solution limited in scope has every incentive to call it comprehensive. Building the program means knowing which capability covers which part of the problem, and where the gaps between them sit.

What Is Data Security for Generative AI?

Data security for generative AI is the practice of classifying sensitive data, controlling how it flows into AI tools, and monitoring usage patterns that signal risk, all tied together by knowing where the data originated. This framework combines three capabilities that most programs build separately:

  1. Data security posture management (DSPM)
  2. Data loss prevention (DLP)
  3. Insider risk management (IRM)

Each one answers a different question, and none of them alone accounts for how employees use generative AI today.

Step 1: Classify Sensitive Data Before It Reaches AI Tools

A data security program for generative AI has to start before the AI tool ever enters the picture. GenAI adoption is no longer edge-case behavior.

Cyberhaven Labs research found that 37.7% of employees now use GenAI SaaS applications, and 39.7% of all AI interactions in the enterprise involve sensitive data. That combination makes classification the step that determines whether everything after it works. If security teams don't know which documents, records, or code repositories contain sensitive data, no downstream control has anything reliable to act on.

This is where DSPM comes in. DSPM discovers and classifies sensitive data across file stores, cloud repositories, and SaaS applications, then assigns a sensitivity level to each source. That classification becomes the foundation every later step relies on. When an employee later copies a paragraph from a document into an AI prompt, the system already knows whether that document was confidential, regulated, or public. It does not need to guess from the pasted text alone.

Most GenAI security programs skip this step and jump straight to blocking or monitoring AI tools. Programs that start with DSPM get a meaningful head start, since policy in the next step can be based on data origin rather than pattern matching against conversational text.

Step 2: Control What Data Can Flow Into Generative AI Applications

Once sensitive data is classified, the next step is controlling what data reaches an AI tool. This is the job of DLP, extended to cover the channels generative AI introduces, including chat prompts, browser-based uploads, and API calls that legacy DLP was never built to inspect.

That workflow has grown fast across enterprises. Cyberhaven Labs found an 80% year-over-year increase in data movement events into and out of GenAI SaaS, which is the volume any DLP-for-GenAI control now has to handle. Legacy DLP inspects file transfers and email attachments. It has no mechanism for a copy-paste action inside a browser tab, which is how most sensitive data reaches an AI tool today.

Effective DLP for GenAI needs endpoint-level visibility into that clipboard activity, plus the lineage classification from step one, so policy can apply to a pasted paragraph even when the text itself contains nothing a rule engine would flag.

DLP stops data from flowing where it shouldn't, but a security team also needs to see when an employee's overall pattern of AI usage signals risk, which is where insider risk management comes in.

Step 3: Monitor Usage Patterns for Insider Risk

Classification and enforcement cover most of the data that touches a GenAI tool, but the complementary functions don't capture user and AI behavior over time. Cyberhaven Labs research found that employees input sensitive data into AI tools on average once every three days, meaning an employee who pastes small excerpts from a strategic document into a personal AI account on that kind of cadence may never trip a single DLP threshold, even though the cumulative exposure is significant.

This is where IRM fits into the program. IRM looks at patterns across sessions and users rather than individual events, distinguishing routine AI-assisted work (i.e. drafting, summarizing, debugging) from behavior that looks like deliberate exfiltration. The distinction matters because most GenAI exposure is not malicious, and instead originates with employees using AI tools to work faster without intending to create a security incident.

Building this step requires the same data lineage foundation as steps one and two. Knowing that a piece of data originated in a confidential document, and tracking every place it has moved since, gives IRM the context to tell a one-time lapse from a sustained pattern.

Step 4: Extend Coverage as Agentic AI Usage Grows

Every step so far assumes a human is typing a prompt, and while critical, that assumption is being challenged by user behavior changes. Cyberhaven Labs research recorded a 509% year-over-year increase in endpoint-based AI agents, or tools that query data stores, call APIs, and pass outputs between systems on a schedule or trigger, without a person initiating each action.

A program built only for human-initiated prompts will miss most of what an agent does with sensitive data. Classification, enforcement, and monitoring all need to extend to tool calls and agent-generated outputs, connecting each automated action back to the data it touched. Otherwise, an alert on an agent's behavior arrives with no context: no way to confirm what the agent read, transformed, or sent onward.

This step is where most programs currently stop, largely because agentic AI adoption has outpaced the tooling built to observe it. Extending lineage tracking to agent workflows now costs less than retrofitting a program later, once dozens of agents are already embedded in daily operations.

How Cyberhaven Unifies Data Security for Generative AI

Cyberhaven built its platform around a single Data Lineage graph rather than creating separate point products for each step above. Data Lineage tracks a piece of data from the moment it is created, through every copy, transformation, or share, including any AI tool it eventually reaches.

  • DSPM inside Cyberhaven’s platform classifies sensitive data across file stores and SaaS applications, feeding that classification into policy the moment data moves toward an AI interface.
  • DLP enforces that policy at the endpoint, whether the action is a browser paste, a file upload, or an API call.
  • IRM applies behavioral context on top, so a security team sees a pattern of usage rather than a list of disconnected alerts.
  • AI Security extends visibility to the AI tools and agents themselves, including which ones employees are using, through which accounts, and whether each tool trains on submitted data.

This is Cyberhaven's approach to Data Security for the Agentic Enterprise: one platform and one lineage graph, instead of four point tools that each see a different slice of the same data.

Most data security gaps in generative AI show up as program gaps: classification without enforcement, enforcement without behavioral context, or human-focused controls that miss agentic activity entirely. Building the program in sequence, DSPM first, then DLP, then IRM, then agentic coverage, closes each gap before the next one opens.

Full understand how to evolve your data security with our ebook, “Securing The Agentic Enterprise”

Frequently Asked Questions

What is data security for generative AI?

Data security for generative AI is the combined practice of classifying sensitive data, controlling how it reaches AI tools, and monitoring usage patterns for risk. It draws on DSPM to locate sensitive data, DLP to control its flow into AI tools, and IRM to catch risky usage patterns, all connected through data lineage.

Do I need DSPM if I already have DLP for GenAI?

Yes. DLP controls what data flows into an AI tool, but it needs to know what is sensitive first. DSPM classifies data at rest, before an employee opens an AI interface, so DLP policy can act on origin rather than guessing from pasted text.

Where should a security team start if building this from scratch?

Start with data classification. Establishing which documents, records, and repositories contain sensitive data gives every later control something reliable to act on. From there, add DLP for enforcement, then insider risk management for behavioral context, and extend coverage to agentic workflows as AI agents become part of daily operations.

How is data security for generative AI different from AI security generally?

Data security for generative AI focuses specifically on classifying, controlling, and monitoring sensitive data as it moves into AI tools. AI security is broader and includes tool discovery, third-party model risk, and agentic workflow governance alongside data protection.

Can insider risk management replace DLP for catching GenAI data exposure?

DLP and IRM address different failure modes. DLP enforces policy at the point data tries to leave, blocking or flagging specific actions. IRM identifies patterns across sessions, such as small repeated transfers over time, that a single DLP rule would not catch. Effective programs run both together.