Perforce has released Delphix Synthetic Data, an ML‑driven service that automatically discovers data structures, relationships, and business context across multiple sources and produces synthetic data sets for application testing. The change replaces manual, legacy data‑access processes with a self‑service platform that can feed both human developers and AI agents via GUI, API, or the Model Context Protocol (MCP).
Why Synthetic Data Generation Matters to Practitioners
AI engineers, cloud/platform engineers, DevOps/SRE teams, and security engineers all rely on realistic test data to validate code, train models, and verify compliance. By generating synthetic data that mirrors production schemas without exposing actual records, teams can eliminate a common bottleneck—granting controlled access to production databases. The service also integrates with large language models (LLMs), enabling automated test creation while keeping data‑related costs in check.
How the ML Engine Operates
The tool applies machine‑learning algorithms to infer data structures and relationships from existing sources. It then synthesizes data that aligns with specific use cases, addressing the low referential integrity (34 %) and realism (36 %) scores reported in a recent Perforce survey. Practitioners can choose a mix of synthetic and masked data to better reflect the environments where code will run.
Operational Implications for Testing Pipelines
- Integration points: GUI, REST‑style APIs, and MCP provide flexible access for CI/CD jobs, test harnesses, or custom scripts.
- Workflow shift: Teams move from manual data provisioning to on‑demand synthetic data generation, reducing lead time for test environment setup.
- Cost control: Loading synthetic data into an LLM of choice allows teams to generate test cases without repeatedly pulling large production snapshots.
- Scalability: Because the service discovers schemas across multiple sources automatically, it can support heterogeneous data landscapes without additional configuration.
Security and Governance Considerations
Replacing production data with synthetic equivalents reduces the attack surface associated with data leakage, but it also introduces new governance questions. Practitioners should evaluate:
- Whether the synthetic data maintains sufficient referential integrity for downstream validation.
- How access to the Delphix Synthetic Data service is controlled—especially when APIs are used in automated pipelines.
- Potential compliance implications of mixing synthetic and masked data, ensuring that any residual sensitive attributes are properly sanitized.
Related CloudNinjas coverage: DevOps.
What This Means For Practitioners
Adopting ML‑based synthetic data generation can streamline test data provisioning, lower the risk of exposing production information, and enable AI‑assisted test creation. Teams should pilot the service in a non‑critical pipeline, measure realism and referential integrity against existing test suites, and establish access controls around the API and MCP endpoints. Ongoing evaluation of how well the synthetic data supports both human‑written and LLM‑generated tests will determine the true operational benefit.

