For decades, enterprise security assumed that a strong enough perimeter would keep everything inside it safe. Firewalls, intrusion prevention systems, and VPNs were the bedrock of that model. But employees now work from anywhere, contractors connect directly to shared systems, and applications live in SaaS environments the organization doesn't fully control. The perimeter, as a security boundary, no longer holds. Data itself has to become the anchor.
What Does It Mean to Build Security Around Data Instead of the Network?
Data-centric security means anchoring protection to the data itself, wherever it travels, instead of to the network boundary around it. Instead of asking whether a device is inside the network, the question becomes what data is being accessed, by whom, and under what conditions. This shifts defenses from a fixed location to a moving asset: data flows across endpoints, on-premises systems, and clouds, so protection has to follow it rather than wait for it to cross a gate.
This distinction matters because it changes what a security program measures. A perimeter model tracks who is inside or outside the network. A data-centric model tracks what happens to sensitive information itself, which is the thing attackers and careless insiders both ultimately act on.
Why Is the Network Perimeter No Longer a Reliable Security Boundary?
The perimeter eroded gradually, then all at once with the rise of cloud and remote work. SaaS applications such as Salesforce, Slack, and Google Workspace run outside the corporate data center. Employees reach them from unmanaged devices. Sensitive files move continuously between cloud storage, email, and collaboration tools, none of which a firewall was built to inspect.
Attackers no longer need to breach a perimeter to reach valuable data. Phishing, credential theft, and misconfigurations bypass network defenses entirely, and insiders with legitimate access can move data out without triggering any of them. Data breached across multiple environments, such as public cloud, private cloud, and on-premises combined, cost an average of USD 5.05 million, compared with USD 4.01 million for breaches confined to on-premises systems alone (IBM Cost of a Data Breach Report, 2025). The cost gap tracks almost exactly with how much visibility a perimeter-only model has into each environment.
How Does Cyberhaven Stop Data Loss Across Cloud, SaaS, and Endpoints?
Cyberhaven stops data loss by tracing the origin and movement of every file and using that lineage to distinguish legitimate use from exfiltration in real time, across endpoints, cloud storage, SaaS applications, and email. Rather than scanning content in isolation at a single exit point, it follows data continuously from creation through every copy, edit, and transfer, so a policy decision reflects where data actually came from and where it is actually going.
That continuous view lets security teams act before data leaves rather than only after the fact. If continuous monitoring shows an employee suddenly downloading a large volume of intellectual property, or sensitive files moving to an external domain through a SaaS app, the activity is visible and actionable at the moment it happens, not weeks later during an investigation.
Structured vs. Unstructured Data Coverage
- Structured data: Credit card numbers, Social Security numbers, and other pattern-matched fields, the traditional territory of DLP.
- Unstructured data: Source code, product designs, contracts, and other files with no fixed pattern, tracked through data lineage rather than keyword or regex matching.
- Derivative data: Content copied, reformatted, or pasted into a new file, which inherits the sensitivity of its origin even after the original file is gone.
How Does Data Lineage Distinguish Legitimate Use From Exfiltration?
Data lineage answers the question rule-based DLP cannot: is this specific action, by this specific person, on this specific piece of data, consistent with how the data has been used before?A rule that flags every large file transfer treats a designer exporting an approved asset the same as an employee exfiltrating a customer list. Lineage instead tracks a file's full history, including where it originated, who has touched it, and how it has been transformed, so a policy can evaluate the action in that context rather than in isolation.
For a security architect evaluating tools, this is the practical difference between a system that generates a rule for every new data type and one that reasons about data behavior directly. It also means policies hold up as data moves and changes form, since the lineage travels with a file's derivatives, not just the original.
How Does a Data-Centric Model Support Zero Trust and Reduce Insider Risk?
Zero trust assumes no user, device, or application is trusted by default, and that every access request is validated on context. A data-centric approach gives zero trust something concrete to validate against: not just who is requesting access, but what data they are requesting and how they intend to use it. A contractor might be allowed to log into a SaaS app but blocked from downloading sensitive files unless specific conditions are met.
This combination is also where insider risk and generative AI data loss converge. Malicious-insider incidents are the most expensive initial breach vector IBM tracks, reflecting how much access an insider already holds before any policy engages (IBM Cost of a Data Breach Report, 2025). The same lineage that flags an insider moving data to a personal drive also flags an employee pasting a customer list into an AI chat tool, since both actions represent sensitive data leaving its approved path regardless of intent.
What Should Security Teams Do to Shift From Perimeter to Data-Centric Defense?
Moving from a perimeter model to a data-centric one is not a single tool swap. It is a shift in what the security program measures and where it intervenes.
- Inventory where sensitive data actually lives across SaaS, cloud storage, and endpoints, not just what the network topology assumes.
- Replace pattern-only detection with lineage-based classification for unstructured and derivative data that regex-based DLP cannot see.
- Tie access policy to data context, not just identity, so zero trust decisions account for what is being requested, not only who is requesting it.
- Monitor continuously rather than at fixed checkpoints, since data moves between email, cloud storage, and AI tools throughout the day, not only at defined exit points.
- Treat insider activity and AI tool use as the same category of risk, since both bypass a perimeter that was never built to see them.
Each step reduces the same blind spot: a security posture that only knows where a device is, not what is happening to the data on it.
Frequently Asked Questions
How does Cyberhaven stop data loss?
Cyberhaven traces the origin and full movement history of every file, called data lineage, and uses that context to distinguish legitimate use from exfiltration in real time across endpoints, cloud storage, SaaS applications, and email, rather than relying only on pattern matching at a single exit point.
Does data-centric security replace the need for a firewall or VPN?
No. Firewalls and VPNs still control network access. Data-centric security addresses what those tools cannot see: data movement across SaaS, cloud, and endpoints once a user is already authenticated and inside.
How is data lineage different from standard content inspection in DLP?
Content inspection scans a file's contents for patterns at a point in time. Data lineage tracks a file's full history, including origin, edits, copies, and transformations, so policy decisions reflect context rather than a single snapshot.
Can this approach stop insider threats as well as external attackers?
Yes. Because lineage evaluates the data and the action rather than the identity of the user alone, it flags risky movement of sensitive data regardless of whether it originates from a compromised external account or a legitimate insider.
Does moving to data-centric security disrupt existing workflows?
Not when policy is based on data context rather than blanket rules. Distinguishing routine, approved data use from risky movement means normal business activity proceeds without added friction, while only genuinely anomalous actions trigger intervention.


.avif)
.avif)
