-
The Third Annual State of Data Compliance and Security Report: 2026 Edition
- A Letter from the Authors
- High Confidence, Persistent Risk: Why Data Protection Isn’t Adding Up
- Why Data Protection Breaks Down in Practice
- How Organizations Are Responding to Growing Data Risk
- Key Takeaways: What It Takes to Protect Data at Scale and AI Speed
- Respondents Snapshot: Segments, Industries, & Job Titles
- Key Terms to Know
Report > The Third Annual State of Data Compliance and Security Report: 2026 Edition
Why Data Protection Breaks Down in Practice
Compliance Efforts Are Constrained by the Need for Quality and Scale
Data Quality Is the #1 Barrier to Protecting Sensitive Data
The #1 reason that stops organizations from protecting ALL sensitive data in non-production is one we have heard time and again: They are concerned that the quality of data becomes too degraded. This was a barrier for 24% of leaders, according to our survey.
What is preventing organizations from protecting all sensitive data in non-production environments?
| Category | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| The quality of data is degraded when we try to protect it | 24 | |||||||||
| It's too big of an effort to implement | 23 | |||||||||
| We protect data in some but not all | 20 | |||||||||
| It's cost prohibitive | 18 | |||||||||
| It slows down innovation | 14 | |||||||||
| It's difficult to locate all instances of sensitive data fields throughout non-production datasets | 10 | |||||||||
| Our data consumers want a full production copy of the data | 10 | |||||||||
| We have a compliance exception | 6 | |||||||||
| It's not a priority | 1 | |||||||||
| Nothing | 16 |
In our conversations with enterprise leaders, many have told us they are worried that in the process of transforming data to protect sensitive information, the data becomes less realistic. They fear that testing will break down and become unreliable, analytics use cases will be disrupted, and so on.
The truth? You can have high-fidelity, production-like data even when you mask it to remove PII. (We’ll talk more about this in Key Takeaways: What It Takes to Protect Data at Scale and AI Speed.)
This aligns with findings from our 2026 Test Data Management Report for AI-Ready Enterprises, where data quality is also the top challenge and a leading priority for test data automation, cited by 43% of organizations.
The High Effort Involved is the #2 Barrier
The second most-cited reason for not protecting all sensitive data in non-production is “It’s too big of an effort to implement” (23%).
For many organizations, protecting sensitive data in non-production requires coordinating multiple manual steps. It involves identifying sensitive data across complex environments, masking it using sub-optimal approaches, and hoping that the resulting data remains usable for development and analytics.
It’s difficult to solve these challenges in a scalable way without the right tools, especially as data volumes and the number of environments continue to grow.
Together, the perceived effort and fears about reducing the quality of data help explain why 84% of organizations still allow compliance exceptions in non-production — even when formal policies or masking mandates are in place.
Back to top
Sensitive Data Growth is Increasing Risk
Like we said, as data volumes grow, organizations find it harder to protect sensitive information across non-production environments. In our research, 57% of respondents reported an increase in the volume of sensitive data in non-production over the past year — further expanding the potential exposure surface.
How has the overall volume (number of records) of sensitive data that your organization stores in non-production environments changed over the past 12 months?
57% say sensitive data volume increased in non-production over past year
| Category | Significantly increased | Somewhat increased | Stayed the same | Somewhat decreased | Significantly decreased |
|---|---|---|---|---|---|
| 10 | 48 | 29 | 13 | 0 |
This growth is being driven by how organizations are using these environments. The top reason is the increased use of non-production environments (31%), followed closely by the growing use of data to support decision-making (30%).
As data is copied, shared, and used more widely, it becomes more difficult to consistently govern it and apply protection controls across every dataset and environment.
What do you believe is the biggest reason that the overall volume of sensitive data that your organization stores in non-production environments has increased?
| Category | ||||||
|---|---|---|---|---|---|---|
| Greater use of non-production environments | 31 | |||||
| Increased use of data to drive decision-making | 30 | |||||
| Increasing digital transformation | 21 | |||||
| AI and ML advances | 12 | |||||
| Regulatory requirements for data that must collected/stored | 4 | |||||
| Cloud adoption | 1 |
Our Take: Data Growth, AI Adoption, and the Challenges of Scale Are Driving Risk
Here’s our take on the findings we just covered:
Data Growth Is Outpacing Protection
Sensitive data in non-production environments is growing, largely due to faster release cycles and increased use of data for analytics and decision-making. We also see AI and machine learning emerging as major new drivers of data usage.
This growth expands the attack surface. More data exists across more environments and is accessed by more teams. As a result, it becomes harder to apply consistent protection at scale.
Across test data workflows, this growth is also contributing to longer wait times for usable data. Most organizations wait more than a day — and often weeks — for fresh datasets, according to our 2026 Test Data Management Report for AI-Ready Enterprises.
Confidence Often Aligns to Policy — Not Execution
At the same time, there is strong pressure to move quickly. Organizations need to deliver data for development, testing, analytics, and increasingly AI. And they often need to do it on tight timelines.
It is possible that many organizations equate having policies or partial controls in place with confidence in their ability to protect sensitive data — even when exceptions are common. This interpretation of the conflicting data we found could help explain why confidence levels remain high, even as concerns and risks persist.
Risk Is Showing Up in Multiple Ways
In this data, we’re seeing a concerningly high percentage of organizations that are experiencing two closely related categories of risk in non-production environments:
- Regulatory and compliance challenges (including audit failures and regulatory exposure)
- Security threats (including data breaches, theft, and ransomware)
These risks point to gaps in how data protection controls are applied and enforced across environments.
In AI, the Challenge is Even Greater
With the rapid adoption of AI, the attack surface is expanding to include training datasets, model inputs and outputs, and unstructured data, as well as increased demand for data at faster speeds for agentic software development and testing.
AI also introduces a new type of data vulnerability due to its inherent behavior. It is not possible to control what an AI model does with PII (personally identifiable information) once it is put in the model. This creates a new risk: AI data leaks.
Many organizations are still working to understand and remediate these risks.
The Consequences Go Beyond Compliance
Non-production environments are widely recognized as the ‘weak link’ of an organization. The impact of storing sensitive data in them reaches far beyond regulatory fines. A breach or failed audit can disrupt operations, damage customer trust, and affect the long-term stability of the business.
Even though organizations face many competing priorities and challenges to protecting data, the risks and their concerns are driving action. In the next section, we’ll examine how organizations are responding and what they prioritize in a data protection solution.