Report
The 2026 State of AI and Data Privacy Report
AI,
Data Management
Executive Summary: AI/ML Investment, Data Privacy Concerns, Confidence, & Solutions
Amid innovative agentic development initiatives, the 2026 State of AI and Data Privacy Report shines a light on the concerning trends related to the perceived give-and-take relationship between compliance and innovation. This research evaluates enterprises’ outlook on balancing the two, including AI best practices, overall outlook, and solutions for protecting your data.
Key Findings: Highlights on Data Privacy for AI/ML Environments
- 80% of surveyed organizations will invest in AI/ML data privacy in 2026–2027.
- Organizations are showing a high amount of confidence related to sensitive data identification and protection, even in AI/ML.
- 80% are highly confident in their ability to identify all sensitive data.
- 98% are confident in their ability to protect sensitive data in AI/ML.
- However, 84% have data compliance exceptions and 51% are concerned about privacy compliance and audits.
- The #1 challenge in protecting AI/ML and analytics: unstructured data.
- 51% also cite data quality challenges.
- 86% have a data privacy mandate for AI/ML and analytics.
- Synthetic was ranked most purpose-fit for AI/ML workflows.
Table of Contents
- A Letter From the Authors
- False Confidence & Data Privacy in AI/ML: How Enterprises Can Put Themselves at Risk
- The Most Concerning Data Risks in AI/ML Environments
- Future-Proofing Data for AI/ML: Solutions for Risk Mitigation & Agentic Efficiency
- Key Takeaways: Confidence in AI/ML & Data Privacy Must Be Earned
- Who Took the AI and Data Privacy Report Survey?
A Letter from the Authors
As the adoption of AI rises, so do the concerns regarding data privacy. Enterprises want to find the sweet spot between compliance and innovation, but one is being weighed more heavily than the other. We need to balance the scale.
It’s not just about protecting training data that goes into AI; it’s about ensuring your data is secure, compliant, and can move at AI speeds for use cases like application development. In this Perforce Delphix 2026 State of AI and Data Privacy Report, 98% say they are confident in their ability to protect sensitive data in AI/ML workflows. However, 84% still allow data compliance exceptions. We want these enterprises to earn their confidence, especially since 51% are concerned about privacy compliance and audits.
Based on this and other core findings from 518 enterprise organizations, we’ll break down current trends in data privacy related to AI/ML and where you’re most at risk. With our help and research, you can learn how to use the tools at your disposal (masking, synthetic data, virtualization, etc.) to scale agentic development while maintaining compliance.
Discover how you can solve the top challenges of data protection in AI/ML environments, from improving quality to maintaining AI/ML speed. Continue reading the report in full or use the navigation on the left to jump to a specific section.
Sincerely,
Matthew Yeh, Mayank Ahluwalia, Rod Cope
About the Authors
Contributing Editor
Back to top
False Confidence & Data Privacy in AI/ML: How Enterprises Can Put Themselves at Risk
The future of AI/ML is bright, but there are data privacy concerns that will dim it if enterprises are not careful. With 80% of surveyed enterprises investing in AI/ML data privacy in 2026–2027, they must set parameters to ensure full compliance amid growing initiatives, including critical agentic development programs.
Most enterprises know they should be prioritizing data privacy for AI/ML, agentic development, and analytics workflows. 86% have a data privacy mandate for just that reason.
86%
of enterprises have a data privacy mandate for AI/ML and analytics workflows.
Do you have a corporate data privacy compliance mandate for AI/ML and analytics workflows?
| Label | Value |
|---|---|
| Yes | 86 |
| No | 13 |
| I'm not sure | 1 |
While 86% have a mandate for AI/ML and analytics environments, that's still less than the 99% who have a mandate for non-production environments in general. This could be because of three reasons:
- AI programs are new, and security/compliance teams haven't caught up.
- AI innovation is seen as so important/strategic that no one wants to stand in the way.
- Regulations haven't caught up yet, or teams don't yet understand implications of existing laws on AI programs.
The top use case (cited by 64%) for mandating the protection of sensitive data is analysis and reporting to inform business operations. Only 37% cited AI/ML model training and fine-tuning as a use case for protecting sensitive data. This is likely because while everyone has well-established analysis and reporting workflows in place, not everyone is doing significant model training and tuning at scale.
For which of the following use cases does your organization mandate the protection of sensitive data? (Select all that apply.)
| Category | Analysis and reporting to inform business operations | Data exploration with production-fidelity data | Development and testing of business reports and dashboards | AI/ML model training and fine tuning | Development and testing of ETL workflows | Offshore development and testing of analytics and AI workflows | RAG application development | My organization does not mandate data protection based on use case |
|---|---|---|---|---|---|---|---|---|
| 64 | 47 | 38 | 37 | 24 | 15 | 8 | 4 |
Confidence in AI/ML & Data Privacy is Good — But Only if You Can Back It Up
Our research proves that enterprises are making efforts to support data privacy, and as a result, enterprises are growing increasingly confident that they’re compliant. 97% say that they’re moderately to extremely confident in their ability to identify every instance of sensitive data residing in non-production environments. The remaining 3% describe themselves slightly confident, meaning all enterprises have at least somewhat confident in how they maintain data privacy.
Another surprising amount of confidence: 98% are confident in their enterprise’s ability to protect sensitive data in AI/ML.
How confident are you that your organization is able to identify (e.g., discover, profile) every instance of sensitive data residing in non-production environments?
| Category | Extremely confident | Very confident | Moderately confident | Slightly confident | Not at all confident |
|---|---|---|---|---|---|
| 9 | 71 | 17 | 3 | 0 |
How confident are you that your organization is prepared with the necessary tools and approaches to protect sensitive data in AI/ML workflows?
| Category | Extremely confident | Very confident | Moderately confident | Slightly confident | Not at all confident |
|---|---|---|---|---|---|
| 6 | 65 | 27 | 2 | 0 |
But confidence doesn’t mean your organization is safe from risk, especially if your data privacy practices aren’t airtight. Confidence could mean enterprises aren’t maintaining a critical eye when it comes to assessing their data privacy risk. While our research shows that these organizations do have compliance and masking policies, we also found that many organizations don’t stick to those mandates.
“When I read that 98% confidence figure, I had to chuckle. The pace of change and urgency around AI/ML is putting real pressure on foundational areas like data security; no one wants to be seen as slowing down the AI/ML freight train. But that level of confidence is only justified if data protection is deliberately built into the strategy from the start.”
All enterprises we surveyed (100%) have data subject to regulation in their non-production environments, and 57% report an increase in the volume of sensitive data in them. That’s not too surprising to us, considering 84% have data compliance exceptions.
But what does that mean for these confident enterprise leaders? Many of them sacrifice data privacy in the name of speed and innovation. In their eyes, masking and data protection processes slow down processes. However, it’s only the manual processes that will. If you leverage these — rather than automated processes — you’ll likely experience the kind of slowdown innovative organizations fear. This only compounds for AI/ML.
84%
of enterprises have data compliance exceptions.
Here at Perforce Delphix, we regularly hear from our customers about balancing the need to protect sensitive data across their environments — without sacrificing speed or quality. As agentic development accelerates, data privacy needs to keep up. Data compliance processes can become a bottleneck within agentic workflows where AI-generated change is happening faster than ever before. Spoiler alert: Delphix can help. More on this later on.
Get Your Guide to Stopping AI Trade-Offs
See how our best practices can help you can address key challenges organizations to protect sensitive data, including the unique risks that AI introduces like internalizing sensitive data, and the misguided or overly-simplistic compliance efforts that put your organization at greater risk.
Back to top
The Most Concerning Data Risks in AI/ML Environments
When we talk about concerns related to agentic SDLC, AI/ML, and analytics, it often comes down to a few core things: data quality, speed, and security. We've been hearing this for customers for years, and it's only increasing in importance.
Organizations, despite being confident in their abilities to protect data, are still concerned about risk. For example, 84% of enterprises are at least moderately concerned about expanded data exposure in non-production. So how can enterprises address this fear and others like it?
Going the Extra Mile to Prevent Data Breaches, Audit Failures, & Re-Identification
Enterprises worry most about data leaks (68% at least moderately concerned), theft or breach of model training data (62%), unauthorized data access (53%), privacy compliance and audits (51%), and personal data re-identification (51%) in AI/ML development and training environments.
What is your level of concern with the following involving AI/ML development and training environments at your organization?
| Category | Extremely concerned | Very concerned | Moderately concerned | Slightly concerned | Not at all concerned |
|---|---|---|---|---|---|
| Data Leakage | 3 | 23 | 42 | 28 | 4 |
| Theft or breach of model training data | 2 | 19 | 41 | 31 | 7 |
| Unauthorized data access | 2 | 16 | 35 | 41 | 6 |
| Privacy compliance and audits | 4 | 14 | 33 | 44 | 5 |
| Personal data re-identification | 2 | 10 | 38 | 42 | 7 |
When it comes to addressing these fears, preparation is key. Let’s discuss the core solutions to minimizing potential damage:
1.
Data Masking & Synthetic Data
Static data masking and synthetic data are the ideal pairing for a well-rounded data compliance strategy. Masking sensitive data is a foundational data protection approach, and synthetic data is perfect for early development and scenarios where production data does not exist yet. Synthetic can even fill in the gaps where you may just not have enough production data.
Using this two-fold masking/synthetic approach will allow you to maintain delivery speeds while meeting compliance standards. Since static data masking is irreversible and synthetic isn’t real data, your data leakage, re-identification, and theft risk are all significantly reduced. (For more info, check out the Future-Proofing Data for AI/ML section.)
2.
Cyberattack Recovery Strategy
Having a cyberattack recovery strategy can reduce your worries about breaches and leaks, because you’ll know your data is safe and you can recover quickly. What’s your plan when one occurs?
For example, you can use data virtualization to immutably protect and recover data to any point in time. In a downtime event — such as a ransomware attack — your organization needs to rebuild applications and recover data quickly to a known good state. With an automated approach to recovery using point-in-time virtual data copies, you gain rapid recovery with an RTO/RPO advantage over traditional backups.
3.
Data Governance
If you worry about regulatory compliance audits like 51% of surveyed enterprises, having proper data governance will give you peace of mind that you can identify and protect data at scale.
You can reduce risk in development and testing with centralized visibility and control of sensitive data and compliance traceability. In Delphix’s case, this includes automated policy-driven masking and comprehensive records of compliance activity, so you’re ready with documentation when an audit comes around.
4.
Access Control
Establishing access control can ease concerns about unauthorized users getting their data. The 53% of enterprises that worry about it could benefit from placing restrictions on who can access data and when.
Consider tag-enabled, role-based access control. Along with centralized data discovery, enterprises can find their data then secure it accordingly.
Data Delivery, Masking, Governance, & Control — All in One
Want to get all these capabilities at your fingertips? Delphix unites them so you can get AI-ready data when and where you need it. See how our platform works to meet and elevate your data compliance and speed standards.
Identifying the “Why” Behind Non-Compliance
While there’s never a “perfect” excuse for not protecting sensitive data, enterprises do have their reasons. When asked why they don’t, our respondents cited concerns such as data quality degradation (24%), implementation burden (23%), cost (18%), and slowing down innovation (14%).
What is preventing your organization from protecting all your sensitive data in non-production? (Select all that apply.)
| Category | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| The quality of data is degraded when we try to protect it | 24 | |||||||||
| It's too big of an effort to implement | 23 | |||||||||
| We protect data in some but not all | 20 | |||||||||
| It's cost prohibitive | 18 | |||||||||
| It slows down innovation | 14 | |||||||||
| It's difficult to locate all instances of sensitive data fields throughout non-production datasets | 10 | |||||||||
| Our data consumers want a full production copy of the data | 10 | |||||||||
| We have a compliance exception | 6 | |||||||||
| It's not a priority | 1 | |||||||||
| Nothing | 16 |
Especially as organizations continue to ramp up agentic development, they need high-quality data and effective management that also includes protection. They need to know that they have compliant, production-like data that will move at the speed of AI.
Many organizations often fall victim to two pitfalls:
1.
They don't invest in automating these data workflows, and they have manual processes that prevent them from achieving the speed/compliance/quality trifecta.
2.
They might cobble together point solutions that involve tradeoffs and don't scale.
In short, organizations take shortcuts vs. adopting a comprehensive approach. Data ends up being a bottleneck that only gets worse as AI workflows demand more data, faster.
But why do enterprises make exceptions and end up non-compliant in AI/ML and analytics workflows, specifically? Our research found that 32% say it’s hard to protect unstructured data sources, 26% say protecting data prohibits getting realistic, production-like data, and 25% say protecting data does not preserve relationships — among other reasons.
What challenges do you experience or foresee in trying to achieve data privacy in AI/ML and analytics workflows? (Select all that apply.)
| Category | ||||||||
|---|---|---|---|---|---|---|---|---|
| Hard to protect unstructured data sources | 32 | |||||||
| Protecting data prohibits getting production-quality data that is realistic enough | 26 | |||||||
| Protecting data does not preserve relationships across data entities | 25 | |||||||
| Pushback from the business consumers of the data | 20 | |||||||
| Hard to scale protection to larger analytical sources without disrupting business timelines | 18 | |||||||
| Slows down analytical and AI/ML processes | 16 | |||||||
| Hard to find all individual occurrences of PII/PHI in my data | 6 | |||||||
| Do not experience or foresee any challenges | 11 |
Why is Quality Such a Concern?
Data quality is important in many facets. 26% of enterprises say protecting data in AI/ML and analytics workflows interferes with quality. While this is a concern for AI and analytics, it's also one that's been dealt with in test environments for decades. In our 2026 Test Data Management Report for AI-Ready Enterprises, survey respondents named it as their #1 test data automation priority and their #1 challenge.
As a result of bad data, organizations have experienced:
- Increased production defects and application instability.
- Incorrectly trained AI models.
- Slower release cycles due to rework and retesting.
- Business analysts received bad data, impacting decision-making.
- Higher operational costs.
- Greater compliance and data exposure risk.
- Reduced customer satisfaction.
If AI agents take in bad data, they’ll continue to cause issues like these. Despite this, enterprises need to find a sweet spot between protecting data and streamlining its delivery. That’s where data automation comes in. By automating consistent data masking across databases and supplementing that data with synthetic data, you can ensure high-quality data and improve AI agents’ output.
Why is Scalability Essential to Organizations?
18% of enterprises shared that it’s hard to scale protection to larger analytical sources without disrupting business timelines — and that’s a valid concern.
Enterprises increasingly depend on large sources (warehouses, lakes, etc.), such as Snowflake and Databricks. Consequently, some key sources may not be protected (i.e., there are gaps and/or holes in protection). They may also struggle to protect data consistently (i.e., can't adopt a unified approach) or have a solution for large sources that performs poorly (e.g., takes too long to work) and impacts delivery speed.
Enterprises must vet solutions to find the best possible one fit for their data sources. If a solution is automated and can scale with ease, you’ll be able to get your data at speeds that fit your timelines.
Why Does Innovation Take Priority?
As it stands, 18% of enterprises worry that their workflows will slow down if they protect data in AI/ML and analytics environments.
AI has ushered in an era where things can be created very quickly and automatically: code is created more easily, and analyses can be executed at high speed. Even so, in our 2026 Test Data Management Report for AI-Ready Enterprises, we found that 99% of respondents are waiting longer than 1 business day for a fresh full production copy of test data, with 42% waiting weeks or months.
Enterprises want to take advantage of this AI-enabled level of efficiency, but this can only happen when trusted data is readily available. This accessible data needs to be safe — i.e., protected — for AI, agents, and people to use. Pairing automated data masking with virtualization and self-service access to data can help quash both AI bottlenecks. By using both, you can provision compliant data at will and have it ready for your AI workflows.
Back to top
Future-Proofing Data for AI/ML: Solutions for Risk Mitigation & Agentic Efficiency
When it comes to preparing your organization for the future, there’s the perceived solution and the actual one. If you’re hoping to implement more AI workflows and achieve agentic speeds, some AI data protection and management options will fare better than others.
For example, surveyed enterprises named static data masking the top data protection approach across multiple use cases: 45% chose it for software development and testing, 32% for data analytics workflows, and 30% for integration and unit testing.
The only one it was not #1 for was AI/ML workflows, where synthetic data was chosen by 56% of respondents. Irreversible static data masking, however, can be a great asset to ensure that enterprises get production-like data for AI workflows at scale. Even better, synthetic and masking together can provide a well-rounded approach to ensuring you have high-quality data for all your use cases.
Which solution do you consider the most purpose-fit approach to protect data for the following use cases?
| Category | AI/ML workflows | Data analytics workflows | Integration and unit testing | Software development and testing | Software development and testing when no production data exists |
|---|---|---|---|---|---|
| Static Data Masking | 27 | 32 | 30 | 45 | 38 |
| Dynamic data masking | 8 | 27 | 28 | 21 | 24 |
| Synthetic data | 56 | 14 | 12 | 13 | 22 |
| Data subsetting | 6 | 15 | 14 | 10 | 8 |
| Tokenization | 1 | 12 | 15 | 10 | 8 |
| Don't know | 1 | 0 | 0.5 | 0 | 0 |
Static masking also tops the charts for AI/ML and analytics across criteria and above counterparts such as synthetic data and dynamic masking. Respondents cited it as top in terms of referential integrity (41%), realism (38%), cost (38%), scalability (32%), ease of use (31%), and speed (29%).
Which of the following approaches do you believe supports your needs best on each criterion for delivering privacy-compliant data to AI/ML and analytics workflows and data sources?
| Category | Cost | Ease of use | Realism | Referential integrity | Scalability | Speed |
|---|---|---|---|---|---|---|
| Static Data Masking | 38 | 31 | 38 | 41 | 32 | 29 |
| Synthetic data | 15 | 21 | 20 | 15 | 27 | 23 |
“Enterprise leaders may be divided, but static data masking emerges as the leading approach for delivering compliant data. Our position aligns with this result: statically masked production data in non‑production environments provides the best balance of privacy compliance and real‑world fidelity — preserving data relationships, realism, and scale without exposing sensitive information.”
Analytics teams are specifically looking for certain qualities in their data protection solution, including security (46%), ease of us (28%), and scale/speed (22%), among others.
What are the most important criteria you use to evaluate a data protection solution for your organization?
| Category | |||||||
|---|---|---|---|---|---|---|---|
| Security | 46 | ||||||
| Ease of use/UI | 28 | ||||||
| Scale/Speed | 22 | ||||||
| Cost efficiency | 16 | ||||||
| Data quality (e.g., production-like realism, completeness | 13 | ||||||
| Sensitive data discovery | 13 | ||||||
| Integration with data delivery | 13 |
Many organizations use a portfolio approach for data protection — selecting between and combining tokenization, synthetic data, dynamic/static masking, etc. — to meet their specific needs. That way, they may be able to meet their security and scalability requirements.
Analytics professionals are specifically looking for the right solution to protect and mask their platforms like Microsoft SQL Server (47%), Microsoft Azure (44%), Databricks (41%), and Oracle (30%).
What data sources does your organization need to mask when used in non-production environments? (Select all that apply.)
| Category | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Microsoft SQL server | 47 | ||||||||||||
| Microsoft Azure data sources | 44 | ||||||||||||
| Databricks | 41 | ||||||||||||
| Oracle | 30 | ||||||||||||
| Snowflake | 29 | ||||||||||||
| Microsoft Dynamics | 17 | ||||||||||||
| Salesforce | 12 | ||||||||||||
| SAP | 9 | ||||||||||||
| Unstructured text data (e.g., documents, email, PDFS) | 6 | ||||||||||||
| Microsoft Fabric data sources | 3 | ||||||||||||
| Workday | 3 | ||||||||||||
| Trizetto | 2 | ||||||||||||
| Guidewire | 1 |
This just goes to show how many different sources could hold data relevant to AI/ML workflows — and how important it is that they’re secured. The tools and techniques used to do so should work in a platform-agnostic manner, so you can achieve your analytics team’s top priorities (ease of use, scale/speed, sensitive data discovery across sources, etc.) across your organization.
Perception vs. Reality: What Tools Enterprises Actually Use, Prefer, & Need
Teams have a wide assortment of approaches to choose from, and many enterprises leverage a plethora of approaches to meet their needs. But which is their best or preferred approach for their use cases?
Respondents ranked synthetic data as most purpose-fit for AI/ML workflows, according to our research, but as we’ve seen, actual use of it in AI/ML lags perception. Only 28% use synthetic data extensively in AI/ML workflows. Despite this, most organizations (77%) across all focuses have used synthetic in some capacity.
To what extent has your organization used synthetic data in AI/ML workflows?
| Category | |||||
|---|---|---|---|---|---|
| We use synthetic data extensively as a core part of our AI/ML workflows | 28 | ||||
| We use synthetic data on a limited basis for specific use cases | 20 | ||||
| We have experimented with synthetic data in AI/ML workflows but are not currently using it | 15 | ||||
| We evaluated synthetic data but found it did not meet our needs for AI/ML workflows | 13 | ||||
| We have not used synthetic data in AI/ML workflows | 23 |
How to Curate a Test Data Management Approach Fit for Modern SDLCs
Want to know how to keep up with the fast-paced tech landscape, including agentic development? It’s all about having the right test data management method that can scale over time.
This Perforce Delphix webinar examines common test data challenges, innovative approaches that solve them, and why your current data management strategy might not cut it.
Evaluating Data Protection Outlook & Needs
Static data masking leads the charts across all facets. Synthetic unexpectedly ranks on the lower end in some cases. Nevertheless, synthetic and data masking together can cover bases for data protection. For providing irreversible protection, 65% of enterprises cite static data masking and 42% cite synthetic data as proficient. They also rank highly for data realism: 60% say static masking and 36% say synthetic provides it.
Which solution(s) would be sufficient to meet your organization’s needs if you required a data protection method that ...
| Category | … provides irreversible protection | … meets SLAs to data consumers when protecting large-scale data sources | … is cost efficient | … prevents sensitive data breach or theft | … provides data realism | … provides referential integrity within and across databases |
|---|---|---|---|---|---|---|
| Static Data Masking | 65 | 57 | 64 | 62 | 60 | 63 |
| Dynamic data masking | 30 | 41 | 36 | 38 | 38 | 36 |
| Synthetic data | 42 | 34 | 27 | 30 | 36 | 34 |
| Data subsetting | 15 | 19 | 22 | 16 | 19 | 18 |
| Tokenization | 13 | 17 | 19 | 21 | 16 | 17 |
Some of these stats aren’t too surprising. Subsetting is driven by legacy tools and not purpose-built for security or compliance, so it erodes data quality and delivery speeds. Tokenization is also more niche and specifically used for things like sharing data with partners.
However, the numbers surrounding synthetic data and data realism give us pause. Let’s talk about them.
The Synthetic Data Disconnect: What’s Your Solution Missing?
Despite the optimistic outlook on synthetic data for AI workflows and high adoption rates, enterprises aren’t as confident in its overall data protection qualities:
42%
say it provides irreversible protection.
30%
say it prevents sensitive data breach or theft.
34%
believe it meets SLAs to data consumers when protecting large-scale data sources.
36%
agree it provides data realism.
27%
consider it cost efficient.
34%
believe it provides referential integrity within and across databases.
With a somewhat negative outlook on these, we worry that enterprises are simply not getting effective solutions that will meet their needs. The right synthetic data solution should give enterprises confidence that the data is high quality, a great tool for preventing data breach/theft, and able to maintain referential integrity across databases.
Especially if it’s considered the top option for AI/ML workflows, synthetic data must maintain high standards for quality. What happens if bad data goes into AI? You get bad outcomes like software defects. AI workflows therefore become useless without good data. If you don’t have the right solution, that’s what will happen.
Referential Integrity Isn’t Just a Positive. It’s Essential.
Let's go back to a stat shared earlier in the report: 25% of enterprises say that preserving relationships is one of the top challenges they face when trying to achieve data privacy in AI/ML and analytics environments.
This data reflects a similar sentiment and outlook. When asked which solution(s) would be sufficient to meet their needs if they required a data protection method that provides referential integrity within and across databases, here’s what enterprises said:
63%
static data masking
18%
data subsetting
36%
dynamic data masking
17%
tokenization
34%
synthetic data
Again, the right scalable solution should be able to maintain referential integrity across large-scale databases. And if it doesn’t, it can greatly impact AI/ML and analytics:
- Since analytics and AI pipelines rely on consistent identifiers to join data across systems. Pipelines can still run with broken referential integrity, but they’ll show incomplete populations and metrics drift.
- AI and machine-learning workflows/models can intake flawed or incomplete data, resulting in bias increases, performance degradation, and harder-to-diagnose failures.
- Broken referential integrity also erodes trust for analytics and AI. Teams will turn back to older, more manual processes, rather than improving models that haven’t met their quality standards.
If you’re not trusting your solution to maintain referential integrity, you should be rethinking how you use it for AI/ML workflows.
Want to see the full impact of referential integrity? Read our guide
Back to top
Key Takeaways: Confidence in AI/ML & Data Privacy Must Be Earned
The 2026 State of AI and Data Privacy Report reveals a striking paradox. Enterprises are moving fast on AI, yet confidence has outpaced control. The level of confidence from surveyed enterprises is only justified when their data protection is deliberately built into the strategy from the start. Right now, too many organizations are trusting the freight train without checking the brakes.
The blockers are consistent, and they are solvable. Focus on these core steps next:
- Automate sensitive data discovery so nothing slips through manual, error-prone declarations.
- Pair static data masking with synthetic data to secure real data while filling gaps for edge cases and new schemas.
- Take advantage of data virtualization to deliver compliant, production-like data in minutes, not weeks.
Share the key findings with your organization >> 2026 State of AI and Data Privacy Report Data Sheet
Get Your Roadmap to Effective, Compliant AI Workflows
Here’s what we’ll leave you with: You do not have to choose between innovation and compliance. Using the right approach, Perforce Delphix will protect data, preserve quality, and move at agentic speed, all at once. The question is no longer whether you can keep up with AI. It is whether your data privacy efforts are on par with the confidence you already have.
Delphix merges data masking, virtualization, and synthetic data to efficiently deliver anonymized data that is compliant and efficient for AI workflows. Masked, virtual data behaves like actual copies but with only a fraction of the storage requirements, ensuring seamless delivery for AI processes without compromising speed or compliance.
When organizations desire security, scale, and cost efficiency, Delphix can deliver. We’ve helped customers achieve:
- 58% faster time to develop an application.*
- 77.2% more data and data environments masked and protected.*
- $8.4 million in additional revenue from improved software development productivity.*
- 408% 3-year ROI.*
Explore how Delphix supports your data privacy needs in the AI landscape with rapid, automated processes. Request a zero-pressure demo today and understand why top enterprises choose Delphix to manage data risks and accelerate AI innovation.
Please consult Perforce's Citation and Usage Guidelines for any inquiries regarding citation, usage rights, or requests for extended use of this and any other Perforce Software reports.
*IDC Business Value White Paper, sponsored by Delphix, by Perforce, The Business Value of Delphix, #US52560824, December 2024
Who Took the AI and Data Privacy Report Survey?
Perforce Delphix partnered with a third-party research firm to survey 518 enterprise leaders worldwide. This report was created from the data collected — see full demographics below.
Age
| Label | Value |
|---|---|
| 25 to 34 | 2 |
| 35 to 44 | 80 |
| 45 to 54 | 18 |
| 55+ | 1 |
Years of Experience
| Label | Value |
|---|---|
| 1 year or less | 1 |
| 2 to 5 years | 4 |
| 6 to 10 years | 40 |
| 11 to 15 years | 41 |
| 16 - 20 years | 9 |
| More than 20 years | 5 |
Industry
| Label | Value |
|---|---|
| Accommodation and Food Services | 2 |
| Arts, Entertainment, and Recreation | 2 |
| Banking | 12 |
| Consumer Products (CPG) | 2 |
| Education | 4 |
| Financial Services | 10 |
| Healthcare/Medical | 13 |
| Information Technology | 6 |
| Insurance | 10 |
| Manufacturing | 14 |
| Retail | 15 |
| Telecommunications | 7 |
| Utilities | 3 |
Country
| Label | Value |
|---|---|
| Australia | 15 |
| Brazil | 14 |
| Canada | 10 |
| France | 6 |
| Germany | 6 |
| Mexico | 6 |
| New Zealand | 4 |
| United Kingdom | 8 |
| United States | 31 |
Frequency of Working with Sensitive Consumer Data
| Label | Value |
|---|---|
| Sometimes | 14 |
| Often | 52 |
| Very Often | 34 |
Decision-Making Status
| Label | Value |
|---|---|
| I am the primary decision maker | 2 |
| I own the budget, but do not make the final selection | 18 |
| I share the decision-making authority | 65 |
| I participate by giving input/feedback/technical evaluation but have no decision-making authority | 15 |
Job Role
| Label | Value |
|---|---|
| Technology Architect | 4 |
| Manager/Sr. Manager | 57 |
| Director | 32 |
| Vice President/Sr. Vice President | 7 |
| C-Suite Executive | 1 |
Job Function
| Label | Value |
|---|---|
| Cybersecurity | 11 |
| Data Compliance | 10 |
| Data Engineering | 12 |
| Data Science/Analytics | 8 |
| DevOps | 1 |
| IT Operations | 20 |
| Information Security | 10 |
| Privacy | 3 |
| Risk Management | 4 |
| Software Engineering | 18 |
| Software Testing | 2 |
Organization Size
| Label | Value |
|---|---|
| 1,000 to 4,999 employees | 11 |
| 5,000 to 24,999 employees | 35 |
| 25,000 to 49,999 employees | 30 |
| 50,000 employees or more | 24 |
Organization Annual Revenue
| Label | Value |
|---|---|
| $0 to $49,999,999 | 1 |
| $50,000,000 to $249,999,999 | 4 |
| $250,000,000 to $499,999,999 | 4 |
| $500,000,000 to $999,999,999 | 7 |
| $1,000,000,000 to $4,999,999,999 | 25 |
| $5,000,000,000 or more | 58 |
| I don't know/Prefer not to respond | 1 |
Budget Allocated to New Technology Initiatives
| Label | Value |
|---|---|
| Less than 10% | 47 |
| 10% to 24% | 50 |
| 25% to 49% | 2 |
| 50% or more | 1 |
| I'm not sure | 1 |
Budget Allocated to Data Compliance and Test Data
| Label | Value |
|---|---|
| Less than 10% | 52 |
| 10% to 24% | 42 |
| 25% to 49% | 5 |
| 50% or more | 0 |
| I'm not sure | 1 |
How Overall Budget Changed from 2025 to 2026
| Label | Value |
|---|---|
| Decreased | 10 |
| Stayed the same | 36 |
| Increased | 53 |
| I'm not sure | 1 |



