Operational resilience in financial services has moved beyond static checklists into a dynamic engineering challenge driven by regulatory mandates like the EU's Digital Operational Resiliency Act (DORA). To meet these stringent requirements, Deutsche Bank is deploying an agentic AI platform that transforms manual preparation for critical business service disruptions. The core change involves migrating from inputs based on assumptions to simulations grounded in actual operational data flows, logs, and incident history.
From Static Inputs to Context-Aware Scenarios
The primary engineering shift is the replacement of static scenario definitions with context-aware generation. Previously, resilience exercises relied on predefined paths that often failed to reflect true system dependencies or failure patterns in complex environments spanning interdependent applications and third-party services.
The new architecture ingests enterprise signals—architecture diagrams, data flows, logs, alerting signals, and operational telemetry—to construct scenarios with clear timelines. This ensures every simulation reflects real business context rather than theoretical assumptions. For platform teams, this implies that the definition of a "scenario" is no longer just a document but an artifact generated by agents reasoning over live system behavior.Dual Orchestration for Control and Adaptability
To satisfy both regulatory audit needs and operational flexibility, Deutsche Bank employs a dual-orchestration architecture. This separation prevents the conflation of governed execution with adaptive investigation:
- Regulator-Aligned Execution (LangGraph): For scenarios requiring strict traceability for supervisory review, workflows use LangGraph to ensure deterministic processes. Every step maintains clear lineage from input context to output artifacts.
- Adaptive Investigation (Google ADK): When analyzing real operational conditions or investigating incidents without predefined paths, the platform utilizes Google Agent Development Kit (ADK). This allows agents to dynamically analyze data and generate responses based on current system states rather than rigid scripts.
This architectural distinction is critical. It ensures that while an agent can explore unknown failure modes adaptively using ADK, it cannot bypass audit requirements when generating evidence for regulators via LangGraph. The same intelligence layer supports both workflows without merging their distinct execution models.
Infrastructure and Evidence Persistence
The implementation relies on specific Google Cloud services to support the lifecycle of these resilience artifacts:
- Gemini Enterprise Agent Platform: Transforms raw operational context into structured scenarios with decision points.
- Cloud Run: Provides elastic execution for scenario generation and evidence creation services, scaling as needed during exercise windows or incident response events.
- Google ADK: Enables the adaptive coordination layer described above.
- Cloud SQL: Offers durable persistence specifically required to store session artifacts, review records, and scenario definitions for compliance audits.
This stack supports a consistent resilience model where evidence is generated continuously rather than assembled manually after an event. For security engineers, this means the audit trail must be engineered into the agent's reasoning process from the start of execution to ensure it meets "regulator-ready" standards automatically.
What This Means For Practitioners
The transition to agentic resilience testing requires a fundamental rethinking of how operational data is structured for consumption. Engineers must design systems where telemetry and architecture metadata are accessible in formats that agents can reason over without compromising the integrity of audit trails.
Furthermore, teams need to architect their orchestration layers carefully. If you separate deterministic workflows (for compliance) from adaptive ones (for investigation), your infrastructure must support distinct execution paths for each. This prevents a single failure mode or configuration error in an agent's reasoning process from invalidating regulatory evidence generated by that same system.Finally, the reliance on Cloud SQL for persistent session artifacts highlights that resilience is not just about compute elasticity but also data durability and retrieval speed during audits. Practitioners should evaluate whether their current logging strategies can support this level of structured persistence without impacting production performance or violating privacy constraints inherent in financial services.
As supervisory expectations evolve, the ability to demonstrate controlled response at scale will depend less on manual documentation and more on engineering systems that produce evidence as a byproduct of normal operations. This aligns closely with broader trends discussed under AI engineering, where agents are increasingly expected to handle complex reasoning tasks autonomously.




