Non-production data typically accounts for around 80% of an enterprise’s total data footprint. However, teams often overlook its protection compared to the 20% of production data above the waterline. This creates a massive governance risk that many organisations are failing to address. As AI adoption accelerates, including the adoption of agentic AI that operates autonomously, the stakes are becoming even higher.
Organisations rigorously scrutinise production data environments, such as customer-facing applications, and enforce strong security and privacy measures. They rarely apply the same standards to non-production environments. Teams often treat software development, testing, QA, data analytics, and AI or machine learning as operational necessities. They must become strategic governance concerns.
Furthermore, teams typically copy production data, such as customer information, eight to twelve times into non-production environments. This creates an untracked, vast data sprawl. Security controls either vanish, are applied inconsistently, or are bypassed for sensitive data.
Every new copy increases the attack surface, and every compliance waiver another blind spot. An engineer could have a copy of a customer database on their laptop that no one else in the organisation knows about. Even that engineer may have forgotten about it.
The Perforce 2025 State of Data Compliance and Security highlights this problem. 84% of respondents said that they allow some form of data compliance exception within their non-production data environments. Yet 60% reported data breaches or theft in these same environments. Additionally, 32% had data compliance audit issues or failures, and 22% incurred regulatory fines.
AI and Cloud Compound the Problem
Cloud adoption amplifies the challenge. A huge, unseen part of the data iceberg is now distributed across multiple locations, geographies, providers, and teams. That makes it increasingly difficult to locate, govern, and secure.
Historically, enterprises could at least assume that most development and testing environments remained within a controlled on-premises estate. Today, data is often distributed across hybrid environments, amid growing data legislation and sovereignty concerns.
AI adds another layer of complexity. Data moves through workflows at unprecedented speed. Agentic AI agents engage with datasets without direct human involvement. A developer might come into work to find that overnight, an AI agent has run thousands of tests and deployed new product features, with users already accessing those.
Suddenly, the governance risk reaches a whole new level, yet organisations are not yet prepared. The 2026 Perforce State of DevOps Report found that of its 820 respondents, only 39% have fully automated audit trails.
The Demand for Realistic Data
The underlying challenge is fourfold.
- Development, testing and AI teams need realistic data, and they need it quickly
- Developers and QA teams want datasets that accurately replicate real-world conditions
- AI initiatives require large volumes of data to train, validate, and refine models
- Cloud transformation programmes depend on continuous development and testing as applications are modernised and migrated.
The temptation to allow the use of production data is understandable. Provisioning compliant datasets takes days or even weeks. From the requesting team’s perspective, this lag can feel completely incompatible with agile delivery cycles, sprint timelines, or urgent debugging needs.
There is also the perception that data realism and complexity are hard to replicate, whether synthetic or anonymised. Teams are drawn to production data because they believe sanitised datasets are not good enough. As a result, using production data is the path of least resistance, and the risk feels abstract until an incident occurs.
That approach is becoming increasingly unsustainable. Also, in many countries, organisations will increasingly have to guarantee that data remains within specific jurisdictions or regional boundaries. As AI regulation evolves, it will bring in further compliance obligations.
Make Secure Data Easier to Use
The answer lies in making secure alternatives faster and easier to consume than production data itself. For instance, automated self-service delivers datasets to users within minutes rather than days. That removes one of the biggest drivers behind production data misuse.
At the same time, technologies for masking data and generating realistic synthetic data have improved significantly. There are a variety of techniques and tools now available. Each has its own relative merits and suitability for different use cases.
For example, dynamic masking intercepts queries and obscures sensitive information at the point of access. However, the sensitive data remains within the repository. That leaves it at risk from a hacker. Dynamic masking is generally better suited to selected use cases, such as analytics.
Static data masking
For many development and testing scenarios, especially in highly regulated environments, static data masking is a safer approach. Transform sensitive data within the dataset itself. Even if attackers compromise the environment, they cannot access the underlying information because the system has already desensitised it.
Modern masking technologies also ensure referential integrity across the board. Testing scenarios change the integrity of non-production data on every system they run on.
Synthetic data is also playing an increasingly important role, particularly for greenfield application development and certain AI initiatives. Rather than copying production records, synthetic datasets replicate the statistical properties and behavioural patterns of real data without containing actual customer information. This allows organisations to support innovation and experimentation while reducing exposure to sensitive data.
Regardless of the technology, teams must remove friction for users by embedding masking and synthetic data generation into CI/CD pipelines and data delivery workflows.
Organisations must also have a good grasp of what happens to data. What data was saved and shared elsewhere, for instance, in a saved backup or on a database? Or, if a new sensitive data column has been added by the CISO or regulatory body, is it clear where all that data currently lives?
Data sovereignty is a hot topic in the EMEA region. It is fast becoming vital to find ways to have end-to-end visibility of the origins and history of datasets, and the ability to discern between real and artificial data.
Culture Matters
In IT, it is often the case that success depends as much on culture and processes as tooling. Organisations must adopt a governance-first mindset that treats non-production data as a first-class engineering concern. It requires the same ownership, monitoring, and operational discipline as production data.
Governance should not just be seen as defensive but also offensive. Well-governed data helps organisations to move faster, reduce operational friction and improve development efficiency.
There are also direct cost benefits. Companies using virtualisation and smaller but richer masked datasets can reduce the amount of redundant copies of non-production data. This also reduces storage overheads and compute consumption.
Non-production data may sit below the waterline of enterprise IT, but it has become too important to ignore. The security, privacy, compliance and operational risks are now too significant. AI is accelerating software delivery and data movement across increasingly distributed environments. It is time for organisations to give non-production data the levels of protection it merits.
Perforce delivers a DevOps Tech Stack for teams building and running high-stakes software systems and revenue-critical applications, where failure is not an option. As a trusted partner helping organizations govern software delivery for AI, Perforce solutions enforce guardrails across code quality, infrastructure, and data — enabling innovation without introducing risk. With customers in over 80 countries — including more than 75% of the Fortune 100 and 50% of the Global 500 — Perforce is trusted by the world’s most innovative teams to build, test, secure, and deliver critical software at scale.

















