Live
Dynatrace integrates Arize’s AI observability into its monitoring platformEnabling Node Swap in Kubernetes 1.34: Practical Impact on AI‑Heavy WorkloadsModel Context Protocol trust gaps enable cascading prompt attacksCutting MCP Token Overhead with Codemode: Practical Implications for AI EngineersGitHub secret scanning now detects Lovable Labs, Pydantic, and Supabase credentialsAutonomous code security gains 23‑point boost on CyberGym‑E2E benchmarkGLM 5.3 on Amazon Bedrock: coding‑optimized MoE model with cross‑region inference and prompt cachingAdd SageMaker inference optimization to any coding agent with the aws‑ai‑ml skillDynatrace integrates Arize’s AI observability into its monitoring platformEnabling Node Swap in Kubernetes 1.34: Practical Impact on AI‑Heavy WorkloadsModel Context Protocol trust gaps enable cascading prompt attacksCutting MCP Token Overhead with Codemode: Practical Implications for AI EngineersGitHub secret scanning now detects Lovable Labs, Pydantic, and Supabase credentialsAutonomous code security gains 23‑point boost on CyberGym‑E2E benchmarkGLM 5.3 on Amazon Bedrock: coding‑optimized MoE model with cross‑region inference and prompt cachingAdd SageMaker inference optimization to any coding agent with the aws‑ai‑ml skill
AWS

Agentic AI Workflows for Fleet-Wide IFEC Diagnostics on AWS

AI SummaryPowered by AI

Panasonic Avionics has deployed a multi-agent system using Amazon Bedrock to automate root cause analysis across thousands of unique in-flight entertainment configurations. This shift from manual log correlation to automated agent workflows allows engineering teams to reduce Mean Time to Detect and resolve issues while scaling institutional knowledge without requiring deep familiarity with every fleet variant.

From Manual Correlation to Agentic Workflows

Panasonic Avionics Corporation manages IFEC systems for a global fleet where manual diagnosis of system failures is no longer viable at scale. Engineers previously spent hours correlating logs, metrics, and ticketing data across diverse deployment configurations using deep institutional knowledge that was difficult to replicate or share.

The solution replaces this bottleneck with an agentic AI architecture built on AWS services. This approach processes raw operational data through a structured pipeline: ingestion via Apache Iceberg in Amazon S3, transformation by AWS Glue and Amazon EMR, and analysis driven by distinct agent roles powered by Amazon Bedrock. The result is a system that identifies anomalies faster than human review while maintaining diagnostic rigor.

The Multi-Agent Architecture Pattern

The implementation relies on three specialized layers working in harmony to handle the complexity of fleet-wide diagnostics:
  • Trend Analyzer: This agent monitors key performance indicators and service degradation metrics across the data lakehouse.
  • Parallel Diagnostic Agents: These agents execute specific functions such as correlation analysis, system checks, and log pattern matching against unique fleet variants.
  • Summarizer Agent: Powered by a large language model (LLM), this component integrates outputs from the parallel workers into coherent diagnostic reports containing root cause analysis and recommended actions.
The use of Apache Iceberg for data storage ensures that standardized service metrics are maintained across different fleet variants, allowing agents to normalize terminology through domain ontology. This consistency is critical when processing large volumes of daily operational data gathered from the entire fleet.

Operational Implications and Efficiency Gains

The primary benefit observed in this architecture is a significant reduction in Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). By automating repetitive investigative tasks, engineers are freed from routine diagnostic overhead. This bandwidth reallocation allows teams to focus on solution design, optimization, feature development, and long-term reliability improvements rather than manual log reviews.

Furthermore, this architecture addresses the challenge of knowledge scaling. Previously, effective diagnosis required deep system familiarity with specific configurations. The agentic workflow scales institutional knowledge broadly across engineering teams without requiring every engineer to possess identical expertise in every fleet variant. The solution also supports proactive health monitoring and pattern recognition at scale. By processing data through distinct layers—ingestion/normalization, parallel analysis, and summarization—the architecture handles the complexity of unique log patterns generated by individual deployments while preserving analytical depth.

Related CloudNinjas coverage: AWS.

What This Means For Practitioners

The deployment described here establishes a reference pattern for agentic AI in operational technology. It demonstrates that complex diagnostic workflows can be decomposed into specialized agents: one for anomaly detection, multiple parallel workers for specific analysis tasks (like log matching), and an LLM-driven summarizer to synthesize findings. For practitioners building similar systems on AWS, the key takeaway is the separation of concerns within the agent layer. The Trend Analyzer handles metric aggregation, while Parallel Diagnostic Agents The data pipeline relies heavily on ETL capabilities provided by AWS Glue, which is essential for normalizing raw operational data into a standardized format before agents consume it. This normalization step, often overlooked in generic AI implementations, ensures that the LLM receives consistent inputs regardless of fleet variant differences. Security and governance considerations are inherent to this design; however, practitioners must ensure that agent outputs from unstructured text generation (the Summarizer) do not inadvertently expose sensitive operational data or hallucinate root causes. The architecture assumes high accuracy is maintained by grounding agents in the normalized metrics stored in Amazon S3 rather than relying solely on generative capabilities for factual claims.

Originally published atAWS Machine Learning Blog