Live
AI Agent Data: Production Realities That Break Demo SuccessPerforce Delphix Synthetic Data: ML‑Based Synthetic Data Generation Cuts Production Data Exposure for DevOps TestingAlways‑On AI Agents Show Rising Boundary Errors in Longer Task ChainsDeploy‑anywhere Spanner: GA brings on‑prem and multi‑cloud capabilities to AI workloadsTurning Kubernetes Policy Gates into Guardrails for Faster, Safer DeploymentsAzure Container Apps Express GA: Sub‑second startup on microVM sandbox, but key integrations missingBasin GA unlocks serverless pipelines, Iceberg catalog, and SQL for Cloudflare analyticsAI Search GA: What Engineers Need to Adjust for Hybrid, Multimodal, and OCR ChangesAI Agent Data: Production Realities That Break Demo SuccessPerforce Delphix Synthetic Data: ML‑Based Synthetic Data Generation Cuts Production Data Exposure for DevOps TestingAlways‑On AI Agents Show Rising Boundary Errors in Longer Task ChainsDeploy‑anywhere Spanner: GA brings on‑prem and multi‑cloud capabilities to AI workloadsTurning Kubernetes Policy Gates into Guardrails for Faster, Safer DeploymentsAzure Container Apps Express GA: Sub‑second startup on microVM sandbox, but key integrations missingBasin GA unlocks serverless pipelines, Iceberg catalog, and SQL for Cloudflare analyticsAI Search GA: What Engineers Need to Adjust for Hybrid, Multimodal, and OCR Changes

Perforce Delphix Synthetic Data: ML‑Based Synthetic Data Generation Cuts Production Data Exposure for DevOps Testing

AI SummaryPowered by AI

Perforce introduced Delphix Synthetic Data, an ML‑powered tool that automatically discovers data structures and creates synthetic data for testing. It lets engineers replace production data with realistic synthetic sets, reducing data‑access bottlenecks, supporting AI‑driven test generation, and lowering security risk.

Perforce has released Delphix Synthetic Data, an ML‑driven service that automatically discovers data structures, relationships, and business context across multiple sources and produces synthetic data sets for application testing. The change replaces manual, legacy data‑access processes with a self‑service platform that can feed both human developers and AI agents via GUI, API, or the Model Context Protocol (MCP).

Why Synthetic Data Generation Matters to Practitioners

AI engineers, cloud/platform engineers, DevOps/SRE teams, and security engineers all rely on realistic test data to validate code, train models, and verify compliance. By generating synthetic data that mirrors production schemas without exposing actual records, teams can eliminate a common bottleneck—granting controlled access to production databases. The service also integrates with large language models (LLMs), enabling automated test creation while keeping data‑related costs in check.

How the ML Engine Operates

The tool applies machine‑learning algorithms to infer data structures and relationships from existing sources. It then synthesizes data that aligns with specific use cases, addressing the low referential integrity (34 %) and realism (36 %) scores reported in a recent Perforce survey. Practitioners can choose a mix of synthetic and masked data to better reflect the environments where code will run.

Operational Implications for Testing Pipelines

  • Integration points: GUI, REST‑style APIs, and MCP provide flexible access for CI/CD jobs, test harnesses, or custom scripts.
  • Workflow shift: Teams move from manual data provisioning to on‑demand synthetic data generation, reducing lead time for test environment setup.
  • Cost control: Loading synthetic data into an LLM of choice allows teams to generate test cases without repeatedly pulling large production snapshots.
  • Scalability: Because the service discovers schemas across multiple sources automatically, it can support heterogeneous data landscapes without additional configuration.

Security and Governance Considerations

Replacing production data with synthetic equivalents reduces the attack surface associated with data leakage, but it also introduces new governance questions. Practitioners should evaluate:

  • Whether the synthetic data maintains sufficient referential integrity for downstream validation.
  • How access to the Delphix Synthetic Data service is controlled—especially when APIs are used in automated pipelines.
  • Potential compliance implications of mixing synthetic and masked data, ensuring that any residual sensitive attributes are properly sanitized.

Related CloudNinjas coverage: DevOps.

What This Means For Practitioners

Adopting ML‑based synthetic data generation can streamline test data provisioning, lower the risk of exposing production information, and enable AI‑assisted test creation. Teams should pilot the service in a non‑critical pipeline, measure realism and referential integrity against existing test suites, and establish access controls around the API and MCP endpoints. Ongoing evaluation of how well the synthetic data supports both human‑written and LLM‑generated tests will determine the true operational benefit.

Originally published atDevOps.com