Data lineage is fast becoming one of the most important ways to understand data exfiltration, for both human and agentic workflows. More and more vendors now market data lineage, each claiming their approach is superior. That makes us proud parents: what was once an esoteric technology is finally going mainstream.
Cyberhaven pioneered data lineage in data security and has spent years perfecting it. In this post, I want to share a few lessons on why data lineage matters, and why it is so hard to build at scale.
Why Data Lineage Matters for both Human and Agentic Workflows.
Data protection depends on business context
Data protection is all about context: which actions are allowed, given the business context in which they happen? Malware attacks tend to look unambiguously malicious across every customer. Data protection, by contrast, has to be closely tuned to each customer's specific business workflows.
For example, the CEO may have perfectly valid reasons to share financial data with board members over Signal, while an engineer at the same company should not be sharing architecture designs over Signal.
Protecting data without false positives and false positives requires deep understanding of business context around the transaction: who is the user, what is the type and history of the data involved, who is the recipient.
Data Lineage allows one to track fragments and replication through the workflow
Lineage is what connects individual actions to the broader workflow they belong to. Understanding the context of the data, the user, and the transaction makes it possible to decide, at a fine-grained level, what is allowed and what is not.
For the Agentic world it matters even more, because of scale and the number of actions. Just like humans, an agent participates in multiple workflows, and just like humans, they are unpredictable.
Most of the data that flows through these workflows is fragmented and replicated, and this fragmentation is on a whole another level when agents are involved. One cannot rely only on data classification, as these tags do not follow along with the workflow. That is why tracking lineage across endpoints, SaaS, PaaS and Agentic tools become even more important as the enterprise workflows transcend all of these entities.
Why Is Lineage Hard to Build and Scale?
Scale: A purpose-built graph engine for a trillion nodes
Data security needs lineage at both the micro and the macro level. At the micro level, you need accurate forensics: tracing how data moved from a CRM tool to a personal chat app, across multiple hops and potentially multiple endpoints. At the macro level, you need to zoom out and reason about aggregate information: how does software development actually work for this customer? Answering such questions requires a graph engine that can connect many hops, at the scale of trillion nodes.
That’s several orders of magnitude beyond what off-the-shelf graph DBs can manage. We did not invent a general-purpose graph database that is a thousand times better than traditional ones. Instead, we designed a purpose-built engine tailored to our specific needs. It supports path traversal and aggregation over graphs spanning a trillion nodes, at hundreds of thousands of operations per second, and it answers both analytical queries (high-level aggregations) and low-latency blocking queries on the same graph. Our technology is proven at scale: many of our customers run deployments of more than 100,000 users.

AI that understands intent and workflows
Every company is unique, and data protection is where this matters the most. What data is important, where is it allowed to go, and who is allowed to access it in your company? It depends on your company's security architecture, internal policies and regulations, how you have set up your SOC 2 controls, and more. This makes traditional DLP a game of thousands of exceptions, cursed by never-ending FPs.
AI can solve that by deeply understanding your company's workflows to tailor the policies and handle the exceptions intelligently. But using AI in data security is not as simple as stapling an LLM onto an old-school DLP. You need the right data, and a large volume of it, properly structured. You need expertise to train cost-efficient models that can understand and reason about data, its lineage, users, and processes at enterprise scale.
Our product combines three ingredients:
- Precise, scalable AI classification labels. We label data, lineage, users, agents, locations, and apps. Together, these labels form an ontology: a shared understanding between humans and AI of what data means, what is relevant, and what is acceptable. They operate at a much higher level than traditional DLP rules, capturing intent rather than enumerating all the messiness of the underlying data.
- Access to the full data flow map. Our graph gives the AI access to the entire flow of data across the customer's environment, at every level of detail. The AI can zoom in or out and pull in data as needed when it evaluates different scenarios. This allows the AI to map the local data operation to the global cross-endpoint/SaaS context of the workflow. The model is small enough to run on endpoints without noticeable overhead.
- Semantic understanding with Linea. We built Linea, which combines our own in-house model of lineage at scale with a frontier LLM. This gives our AI a semantic understanding of data, so it can accurately triage incidents and create and tune policies for each enterprise.

Endpoint is hard but essential
Endpoints are a gnarly world, but endpoint is where the human and many agents interact with most software systems, so the endpoint is where you can get maximum visibility. When building endpoint security software, you encounter an enormous diversity of use cases and environments, and you have to support many versions of your own code running in production at once.
Blocking without slowing anyone down
Blocking is hard for several reasons. First, blocking requires low-latency decisions. That takes a fast database (see the graph engine above), but often that alone is not enough, so you need caches too. How do you keep caches consistent across millions of endpoints that may span dozens of different code versions?
Second, blocking requires intercepting operations inline, which risks performance and reliability problems for the very operations you intercept. There is a reason DLP has a bad reputation in the industry.
Our architecture switches dynamically between inline and out-of-band inspection, only as needed (read more on our engineering blog). This is not just theory: we have proven how lightweight it is by deploying at some of the most innovative companies in the world, where any interference would not be tolerated. The graphs below show real-world performance at such a customer, aggregated over their entire fleet of over 100K seats. Median CPU usage is 0.2%, while the 99 percentile is 0.5%.

Coverage: keep up with the speed of change
The fundamental problem with coverage is that the world changes faster than monitoring tools can keep up. In the agentic world, data is accessed in multiple forms: via UI, via agents, via agent tools, via scripts, via custom vibe-coded apps. We combine highly specific sensors (for example, understanding exactly how users interact with ChatGPT in the browser) with generic sensors (for example, seeing every upload and download to any website, and any data opened or saved by any app) and use AI to glue information where possible.
Coverage also means tracking the data wherever it is. We cover Windows, Mac OS, Linux, ChromeOS and iOS, we have specific sensors for Chrome, Chromium, Edge, Firefox, Safari, but also generic sensors for any Chromium-based variant, we have endpoint sensors, browser sensors, AI agent sensors, mobile device sensors, network proxy sensors, SaaS connectors and PaaS connectors to more than 30 service providers. We cover data in motion, but also discover it at rest. And we bring it all together in a single, uniform, graph database to build the ultimate data knowledge graph.
The Bottom Line
We don't claim our lineage is perfect. But it is the most mature, scaled, and proven lineage technology available, and we keep advancing it. That is why some of the world's leading innovators and some of the largest enterprises trust us.
Learn more about our mission at cyberhaven.com/product




.avif)
.avif)
