Report
The Synthetic Data Market Gap: Findings from our 2026 Survey
Data Management
Executive Summary
Synthetic data is generating strong interest among enterprise leaders, but our survey of 518 enterprise technology leaders found that current adoption is lower than the market hype would suggest. This disconnect helped us uncover a market gap in early 2026: Organizations increasingly need to generate realistic, customizable, privacy-safe data, but many did not believe synthetic data solutions available at that time could consistently deliver the quality and referential integrity required for enterprise environments.
Key Findings: A Contradiction Between Perception and Reality Reveals a Market Gap
- Synthetic data was ranked the most purpose-fit approach for protecting data in AI/ML workflows by 56% of leaders.
- Yet, synthetic data’s adoption in AI and ML lags behind perception.
- 51% of organizations are not using synthetic data in AI/ML workflows.
- 66% of Software leaders say their organizations do not use synthetic data in development and testing.
- Many leaders do not feel synthetic data met enterprise requirements.
- Only 38% said synthetic data provides data realism.
- Only 34% said it provides referential integrity.
- These findings suggest that at the time of this survey, there was a market gap and a need for an enterprise-ready synthetic data solution.
Table of Contents
- About the Authors
- A Quick Note on Who We Surveyed
- A Contradiction Emerges: Synthetic Data is Preferred for AI/ML, But Is Not Mainstream Yet
- Most Software Leaders Don’t Use Synthetic Data in Development or Testing
- The Need for Quality Data Limits Legacy Synthetic’s Usability
- Our Interpretation: This Data Reflects a Market Gap and Need for Next-Gen Solutions
- Key Takeaways: Realism and Quality Will Define Synthetic Data's Future
- Who Took the Synthetic Data Survey?
A Quick Note on Who We Surveyed
These findings come from our survey of 518 global enterprise leaders in early 2026. We surveyed leaders in functions ranging from cybersecurity, to compliance, to software engineering, and more.
Some of our findings below are from the full 500+ respondent pool.
In some cases, we wanted to zero in on the perspectives of two specific groups. We’ll refer to them as:
Analytics Leaders
Job functions = Data engineering or data science/analytics. There were 104 respondents in this group, about 20% of the total response pool.
Software Leaders
Job functions = Software engineering or software testing. There were 103 respondents in this group, also about 20% of the total response pool.
You can find more information about who we surveyed here: Who Took the Synthetic Data Survey?
A Contradiction Emerges: Synthetic Data is Preferred for AI/ML, But Is Not Mainstream Yet
Synthetic data has captured significant mindshare as organizations invest in AI and machine learning. But when we compared leaders' perceptions with actual usage patterns, we found a notable disconnect between how organizations think synthetic data should be used and how they are using it today.
What Do We Mean by "AI/ML Workflows"?
For the purposes of this report, AI/ML workflows include a broad range of activities that use data to build, test, or operate AI-enabled systems. This includes traditional AI and machine learning use cases, such as model training and analytics, as well as emerging use cases like agentic software development, AI-assisted testing, and other workflows where AI systems interact with enterprise data.
Most Leaders Say Synthetic Data is Well-Suited for AI/ML Workflows
When reviewing responses from all 518 survey participants, we found that synthetic data was ranked the most purpose-fit approach for protecting data in AI/ML workflows (56%). Static data masking was ranked the next most purpose-fit approach.
Which approach is the most purpose-fit to protect data in AI/ML workflows?
| AI/ML Workflows | ||||||
|---|---|---|---|---|---|---|
| Synthetic data | 56 | |||||
| Static data masking | 27 | |||||
| Dynamic data masking | 8 | |||||
| Data subsetting | 6 | |||||
| Tokenization | 1 | |||||
| Don't know | 1 |
Analytics Leaders Especially Prefer Synthetic Data for AI/ML
When we looked at the responses of leaders in data engineering, analytics, and data science functions (which from here on out we’ll refer to simply as “Analytics leaders”), we found that this number was even higher: 59% chose synthetic data as the most purpose-fit approach for protecting data in AI/ML.
Analytics Leaders
| Category | ||||||
|---|---|---|---|---|---|---|
| Synthetic data | 59 | |||||
| Static data masking | 23 | |||||
| Data subsetting | 9 | |||||
| Dynamic data masking | 7 | |||||
| Tokenization | 2 | |||||
| Don't know | 1 |
Interestingly, however, this preference does not translate into widespread adoption.
Most Organizations Are Not Using Synthetic Data in AI/ML
Across the organizations we surveyed, the actual use of synthetic data in AI/ML workflows lags far behind perceptions.
More than half (51%) of organizations are not currently using synthetic data in AI/ML workflows. This group includes:
- 23%: Have not used synthetic data in AI/ML.
- 15%: Experimented with it in AI/ML but are not currently using it.
- 13%: Evaluated it for AI/ML but found it did not meet their needs.
Only 28% report using synthetic data extensively as a core part of their AI/ML workflows.
51%
are not using synthetic data in AI/ML workflows.
To what extent has your organization used synthetic data in AI/ML workflows?
| Category | |||||
|---|---|---|---|---|---|
| We use synthetic data extensively as a core part of our AI/ML Workflows | 28 | ||||
| We use synthetic data on a limited basis for specific use cases | 20 | ||||
| We have experimented with synthetic data in AI/ML workflows but are not currently using it | 15 | ||||
| We evaluated synthetic data but found it did not meet our needs for AI/ML workflows | 13 | ||||
| We have not used synthetic data in AI/ML workflows | 23 |
Synthetic Data’s Use in AI/ML, By Role
We were surprised to find that the gap was even larger among the teams most closely connected to AI and data initiatives.
Before we dive into those numbers, we want to provide some context: We see two distinct and disparate needs for synthetic data.
The first need: generating synthetic data for development and testing, including agentic or AI-assisted workflows. Referential integrity and business realism are critical in this use case, and statistical preservation takes a backseat.
The second need: synthetic data for AI/ML model training. Statistical preservation is critical for this use case, as is referential integrity.
Synthetic for AI/ML: Software Leaders vs. Analytics Leaders
A third (33%) of Analytics leaders — the group that would use synthetic data for AI/ML model training — said their organizations have not used synthetic data in AI/ML workflows. Among Software leaders — the group that would use synthetic for AI-assisted and agentic development and testing — that number rises to 38%.
Analytics Leaders
| Category | |||||
|---|---|---|---|---|---|
| We use synthetic data extensively as a core part of our AI/ML workflows | 28 | ||||
| We use synthetic data on a limited basis for specific use cases (e.g. data augmentation, rare-event-modeling) | 20 | ||||
| We have experimented with synthetic data AI/ML workflows but are not currently using it | 12 | ||||
| We evaluated synthetic data but found it did not meet our needs for AI/ML workflows (e.g. quality, speed, realism, scale etc.) | 8 | ||||
| We have not used synthetic data in AI/ML workflows | 33 |
Software Leaders
| Category | |||||
|---|---|---|---|---|---|
| We use synthetic data extensively as a core part of our AI/ML workflows | 18 | ||||
| We use synthetic data on a limited basis for specific use cases (e.g. data augmentation, rare-event-modeling) | 13 | ||||
| We have experimented with synthetic data AI/ML workflows but are not currently using it | 17 | ||||
| We evaluated synthetic data but found it did not meet our needs for AI/ML workflows | 14 | ||||
| We have not used synthetic data in AI/ML workflows | 38 |
Want to know why we think actual use of synthetic data in AI/ML workflows is lagging?
Jump ahead to Why Does the Use of Synthetic in AI/ML Workflows Lag Behind Market Perception?
Back to top
Most Software Leaders Don’t Use Synthetic Data in Development or Testing
This is one of the findings that surprised us the most: 66% of Software leaders say they don’t use synthetic data tools or approaches in development and testing.
Software leaders: What synthetic data tools or approaches does your organization use for software development and testing?
For this question, we focused on Software leaders because they are closest to development and testing workflows. They are the people most likely to know whether synthetic data is actually being used day-to-day.
That is why this finding stands out.
When we look across all 500+ leaders surveyed, 46% say their organizations do not use synthetic data for development or testing.
At the same time, our broader survey findings show that synthetic data is already part of many organizations' data protection strategies. Organizations take a portfolio approach: More than half (51%) use synthetic data as one method for protecting sensitive data alongside approaches such as static and dynamic masking.
So, synthetic data is already widely recognized and used as part of organizations' data protection strategies, yet most Software leaders are not using it in development and testing.
WEBINAR
How Should Enterprises Be Using Synthetic Data?
Learn about the key benefits and use cases for synthetic data, so that you can use it effectively. Perforce Delphix expert Mayank Ahluwalia walks through best practices and provides a Delphix Synthetic Data demo in this insightful webinar.
How Does Synthetic Data Rank Against Other Methods for Development and Testing?
To add more color to this data, let’s look at how all the enterprise leaders we surveyed ranked synthetic data against other data protection methods, like static masking:
Which solution do you consider the most purpose-fit to protect data for the following use cases?
| Category | Integration and unit testing | Software development and testing | Software development and testing when no production data exists |
|---|---|---|---|
| Static data masking | 30 | 45 | 38 |
| Dynamic data masking | 28 | 21 | 24 |
| Synthetic data | 12 | 13 | 22 |
| Data subsetting | 14 | 10 | 8 |
| Tokenization | 15 | 10 | 8 |
Only 12-13% chose synthetic data as the most purpose-fit for development, testing and integration and unit testing. Even when data does not exist, it was only chosen as the most purpose-fit by 22% of leaders. That would explain why so few people were actually using the solutions available in the market in early 2026.
To understand this gap, we need to look beyond synthetic data itself and examine the challenges and priorities enterprise teams face today.
Why This Finding Matters: The Time to Get Ahead is Now
While most Software leaders are not yet using synthetic data in development and testing, we increasingly hear a different story from our customers who are leading innovation in industries like healthcare, retail, and financial services. For these teams, synthetic data is becoming a foundational capability for delivering high-quality software faster.
As AI-driven and agentic development accelerate software delivery, organizations that adopt scalable synthetic data now will be better positioned to innovate quickly while maintaining quality and compliance.
Back to top
The Need for Quality Data Limits Legacy Synthetic’s Usability
So, what’s the reason for these contradictions? Why is synthetic data seen as the most viable option for AI and ML workflows, when most organizations are not actually using it for that use case? And why do most software teams not use synthetic data for development and testing at all?
We think the data indicates a market gap that existed at the time this survey was taken. Before we explain, here’s some background information:
For Enterprises, High-Quality Data is the Top Challenge AND Priority
In Delphix’s third annual Data Compliance & Security Report, we found that data quality is the #1 barrier to protecting sensitive data in non-production.
What is preventing organizations from protecting all sensitive data in non-production environments?
| Category | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| The quality of data is degraded when we try to protect it | 24 | |||||||||
| It's too big of an effort to implement | 23 | |||||||||
| We protect data in some but not all | 20 | |||||||||
| It's cost prohibitive | 18 | |||||||||
| It slows down innovation | 14 | |||||||||
| It's difficult to locate all instances of sensitive data fields throughout non-production datasets | 10 | |||||||||
| Our data consumers want a full production copy of the data | 10 | |||||||||
| We have a compliance exception | 6 | |||||||||
| It's not a priority | 1 | |||||||||
| Nothing | 16 |
And as our colleagues reported in Delphix’s 2026 Test Data Management Report for AI-Ready Enterprises, “consistent, high-quality test data” is Software leaders’ #1 priority in test data automation.
Software leaders: Which of the following do you prioritize in your test data automation?
Data quality is also the #1 challenge leaders across industries and roles have with test data.
What are your organization's biggest challenges with test data?
At the Time of This Survey, Synthetic Data Solutions Fell Short
When asked which data protection solutions meet their needs, organizations ranked synthetic data relatively low across several criteria.
- Only 36% said synthetic data provides data realism.
- Only 34% said synthetic data provides referential integrity.
Which solution(s) would be sufficient to meet your organization’s needs if you required a data protection method that ...
Realism and referential integrity are two critical elements of data quality.
It’s especially noteworthy that synthetic data was ranked relatively low for referential integrity because testing across complex systems is the second biggest challenge organizations face with test data.
Back to top
Our Interpretation: This Data Reflects a Market Gap and Need for Next-Gen Solutions
Why Isn't Synthetic Data Used More for Development and Testing?
Taken together, our findings help explain why synthetic data adoption may not be higher among enterprises. Organizations say they need high-quality, realistic test data, yet relatively few believe synthetic data provides the realism and referential integrity required for enterprise development and testing.
We think this is a reflection of the limitations of many synthetic data options available at the time of the survey in early 2026.
The Issue with Legacy Synthetic Solutions
For years, synthetic data tools have often been delivered either as standalone point solutions or part of a limited test data management platform. We have found in our conversations with enterprise leaders that legacy synthetic tools struggle with end-to-end testing across multiple systems and creating enough variety to accurately model complex business scenarios.
Many synthetic data approaches also require significant manual effort to implement and maintain. Teams may need to define rules field by field, configure relationships manually, and rely on specialized expertise to make generated data usable. While these approaches may work for individual applications or isolated projects, they can become difficult to scale across complex enterprise environments, particularly as organizations adopt AI-driven and agentic development workflows.
The survey findings appear to reflect these realities. More than half of respondents viewed synthetic data as the most purpose-fit approach for AI/ML workflows, yet fewer than a third reported using it extensively. At the same time, survey respondents gave synthetic data relatively low marks for two capabilities that are essential in enterprise environments: data realism and referential integrity.
These findings suggest that many organizations saw value in synthetic data but did not believe the solutions available at the time consistently met enterprise requirements.
Why Does the Use of Synthetic in AI/ML Workflows Lag Behind Market Perception?
As we saw in our survey findings, leaders view synthetic data as the most purpose-fit approach for protecting data in AI/ML initiatives, yet actual usage remains relatively low.
One reason for this contradiction may be that there is a lot of hype and controversy in the market around the use of synthetic data for AI/ML use cases, especially for analytics use cases and model training. We have found that the market is very divided; get two experts in the room, and they will have two opinions.
The survey findings may indicate that organizations recognize the promise of synthetic data but have struggled to find solutions capable of producing the quality, realism, scale, and governance required for enterprise AI initiatives.
In other words, synthetic data may be ahead in mindshare compared to adoption. Enterprises appear to understand where the technology can create value. What they are still looking for are solutions purpose-built to meet enterprise requirements.
A New Approach to Synthetic Data that Enables Agentic Development
To keep pace with modern software delivery, enterprise teams need realistic data that can be quickly and easily generated by developers and agents. This data has to reflect production environments, preserve relationships within and across systems, and be customized for each testing scenario.
Yet, our survey findings suggest many organizations did not believe existing synthetic data solutions could consistently deliver these capabilities.
Leaders saw the potential of synthetic data, but many felt that it didn’t perform as needed in enterprise environments.
Closing the Gap Between Synthetic Data's Promise and Reality
This gap is one of the reasons Delphix decided to enter the synthetic data market.
Through our work with some of the world's largest enterprises, we repeatedly heard the same challenge: teams needed to quickly generate synthetic data for agile development and testing. But they could not sacrifice data quality, realism, or their ability to test complex, interconnected systems. They also wanted a tool that could complement the data masking solution they were already using.
The survey data points to a clear enterprise requirement: teams need realistic, scenario-specific data that preserves relationships within and across systems, supports complex testing, and can be governed at scale. As AI agents become more involved in development and testing, that data must also be available at agentic speed.
Delphix Synthetic Data was designed to fulfill these requirements. Delphix combines AI-first discovery and configuration with governed synthetic data generation, making it easier for teams to create realistic, scenario-specific data without the manual setup and specialized expertise often associated with traditional synthetic data tools.
It lets teams generate realistic data for specific testing scenarios, maintain referential integrity within and across data sources, and support enterprise-scale development and testing workflows.
Delphix lets teams move beyond standalone synthetic data generation and provide safe, realistic data for enterprise development and testing workflows. By bringing together AI-powered synthetic data generation, automated data discovery, masking, data delivery, and governance in a unified platform, Delphix helps teams deliver trusted data faster while maintaining quality, consistency, and control across the enterprise.
See How Delphix Synthetic Data Works
Key Takeaways: Realism and Quality Will Define Synthetic Data's Future
Our findings suggest that enterprise teams are not simply looking for a synthetic data generator. They are looking for realistic, consistent, customizable synthetic data that can hold up in the complexity of enterprise environments but still be easy to use — qualities that legacy synthetic tools have not been able to provide.
- Synthetic data’s adoption remains uneven. Many organizations have evaluated or experimented with synthetic data without making it a core part of their workflows.
- Data quality drives decisions. IT leaders ranked data quality among their biggest challenges and highest priorities.
- Testing realistic and complex business scenarios is critical. Teams need data that reflects how applications and processes behave in the enterprise but can also be customized to new business scenarios.
- Enterprise requirements and agentic development are reshaping the market. The next generation of synthetic data solutions will need to support realism, referential integrity, and governance while operating at agentic speed and integrating into broader test data management systems.
Generate Realistic, Referentially Intact Data On-Demand with Delphix Synthetic Data
Our survey found that enterprise teams need realistic, customizable data that supports complex testing, preserves relationships across systems, and can be delivered at the speed modern software development demands. Delphix Synthetic Data was built for exactly that. It uses AI-first discovery and configuration to help teams quickly generate realistic, scenario-specific data for developers and AI agents, while maintaining governance and referential integrity at enterprise scale.
Delphix combines synthetic data generation with masking, data delivery, and governance in a unified platform. The result is trusted data that helps teams move faster without compromising quality, security, or compliance.
According to IDC research, organizations that use Delphix have achieved:
- *58% faster application development
- *77.2% more data and data environments protected
- *408% three-year ROI
- *Payback in less than six months
See how Delphix Synthetic Data helps teams generate realistic, enterprise-ready data at agentic speed. Request a personalized demo and explore how your organization can close the synthetic data gap.
Generate Synthetic Data at Agentic Speed
Please consult Perforce's Citation and Usage Guidelines for any inquiries regarding citation, usage rights, or requests for extended use of this and any other Perforce Software reports.
*IDC Business Value White Paper, sponsored by Delphix, by Perforce, The Business Value of Delphix, #US52560824, December 2024


