Live
EU Cyber Resilience Act expands software supply‑chain responsibilities for digital product manufacturersTyped Probability Model Jev Shifts AI Output from Text to Structured DecisionsBasin Pipelines per‑stream ingest capacity jumps to 1 GB/s – what engineers need to knowAI‑driven vulnerability management: moving from CVE counts to contextual riskDynamic Tier in Google Cloud Managed Lustre: Cost‑Effective, Low‑Latency Storage for AI and HPCArgo CD 4.0 Visioning and Scaling Lessons from ArgoCon NA 2026Always‑On OpenAI Dots: Free Baseline, Metered Delegation, and What It Means for Cost and GovernanceConfidential Advisory Comments Enable Secure In‑Repo Vulnerability CollaborationEU Cyber Resilience Act expands software supply‑chain responsibilities for digital product manufacturersTyped Probability Model Jev Shifts AI Output from Text to Structured DecisionsBasin Pipelines per‑stream ingest capacity jumps to 1 GB/s – what engineers need to knowAI‑driven vulnerability management: moving from CVE counts to contextual riskDynamic Tier in Google Cloud Managed Lustre: Cost‑Effective, Low‑Latency Storage for AI and HPCArgo CD 4.0 Visioning and Scaling Lessons from ArgoCon NA 2026Always‑On OpenAI Dots: Free Baseline, Metered Delegation, and What It Means for Cost and GovernanceConfidential Advisory Comments Enable Secure In‑Repo Vulnerability Collaboration
Kubernetes

KubeCon North America AI Inference Track

AI SummaryPowered by AI

The Cloud Native Computing Foundation has released the schedule for KubeCon + CloudNativeCon, introducing a dedicated track focused on production-grade AI inference and agentic workflows. This event highlights critical tools like vLLM and KServe that are essential for DevOps professionals preparing to operationalize large language models within Kubernetes clusters.

The industry is shifting rapidly from experimental AI pilots to robust, scalable deployments in cloud-native environments. The Cloud Native Computing Foundation (CNCF) has officially released the schedule for KubeCon + CloudNativeCon North America 2026, an event taking place November 9–12 in Salt Lake City, Utah. A significant addition this year is a new track dedicated to **AI inference** and agentic workflows. For engineers preparing for advanced certifications or managing production systems, understanding how cloud-native infrastructure supports these workloads is no longer optional; it has become central to modern platform engineering.

Operationalizing AI Inference with Kubernetes

  • vLLM: High-throughput inference engine optimized for GPU clusters.
    KServe: Standardized model serving on K8s

The new AI inference track addresses the immediate challenges of deploying large language models (LLMs) and other AI agents in production. The sessions feature deep dives into projects like vLLM, which is designed to maximize throughput for high-volume requests using GPU scheduling optimizations. Engineers will also explore KServe's role as a standardization layer that simplifies model serving across heterogeneous environments.

A critical architectural consideration discussed at the event involves managing memory and compute resources efficiently when running multiple models simultaneously on shared clusters. This directly impacts how you design your Kubernetes resource quotas for AI workloads, ensuring cost efficiency without sacrificing performance metrics like tokens per second (TPS). Observability tools such as OpenTelemetry are highlighted to provide necessary visibility into model latency and error rates during inference.

Platform Engineering at Scale

Beyond the new **AI** focus, established sessions on platform engineering will explore how internal developer platforms can accelerate software delivery. The goal is moving teams away from repetitive manual tasks toward self-service workflows that allow developers to provision infrastructure with minimal friction.

In a production setting, this translates to building robust Internal Developer Platforms (IDPs) using tools like Backstage or custom GitOps pipelines defined in ArgoCD manifests. These platforms enforce policy-as-code principles before code even reaches the cluster, ensuring compliance and security standards are met automatically.

Security for Cloud Native AI Systems

The event also addresses supply chain security specifically regarding containerized models and dependencies used by agentic workflows.

The rise of complex AI applications introduces new attack vectors. Sessions will cover runtime protection strategies that monitor model behavior to detect prompt injection attacks or unauthorized data exfiltration attempts from agents interacting with external APIs.

Vulnerability management for AI pipelines requires scanning not just the container image but also validating input prompts and ensuring third-party models are hosted on trusted registries rather than unverified public endpoints. Policy enforcement mechanisms must be updated to handle dynamic model updates without disrupting live inference services.

What This Means For You

This event signals a maturation of the ecosystem where AI is no longer an add-on but core infrastructure workloads for many organizations. If you are pursuing certifications like CKA, CKS, or specialized tracks in cloud security and DevOps engineering (such as AWS ML Specialty), understanding these operational patterns will be vital.

The integration of AI inference into standard Kubernetes operations means that future exams may test your ability to configure GPU scheduling policies alongside traditional CPU-based workloads. Familiarize yourself with the tools mentioned, such as Ray for distributed computing and Cilium for network visibility in AI clusters.

Originally published atCNCF