Flipkart has officially won the End User Case Study Contest at KubeCon + CloudNativeCon India 2026, recognizing their specialized work in chaos engineering within a Kubernetes-native architecture. The central reliability engineering (CRE) team constructed a centralized platform that orchestrates experiments using LitmusChaos to harden microservices estates against failures.
Architecting for Production Reliability
The core challenge addressed by Flipkart involved managing hundreds of tightly coupled services across both Kubernetes and virtual machine workloads. By integrating Litmus Chaos Engine (LitmusChaos), the team achieved a critical operational milestone: executing approximately 90% of chaos experiments within staging environments prior to major festive sales events.
This approach ensures that production systems are stress-tested without risking live customer transactions or data integrity. The architecture leverages Kubernetes-native orchestration, allowing for granular control over failure injection scenarios such as pod crashes, network latency simulation, and disk I/O errors. This capability is essential for DevOps professionals preparing for advanced reliability certifications like the Kubernetes Security Specialist (CKS), where understanding fault tolerance mechanisms is a primary competency.
LitmusChaos Integration Strategy
The technical implementation relies on Flipkart's custom, multi-tenant chaos engineering platform. This solution abstracts the complexity of running experiments across heterogeneous infrastructure layers. By utilizing Kubernetes Chaos Engineering Scale (KCE) principles embedded in LitmusChaos, engineers can define experiment templates that automatically propagate to specific namespaces or labels.
The team contributed five core fixes and enhancements back upstream to the CNCF incubating project. This open-source contribution model accelerates feature development for the entire community while ensuring Flipkart's proprietary requirements are met through custom configurations rather than waiting on generic releases. The integration allows engineers to schedule experiments dynamically, triggering tests based on CI/CD pipelines or specific traffic thresholds.
Operational Impact and Upstream Contributions
The shift from ad-hoc testing to a systematic chaos engineering platform has significantly reduced the mean time to recovery (MTTR) for critical services. By identifying weak links in microservice dependencies during staging, Flipkart prevents catastrophic outages that could occur under high load conditions typical of Indian festive seasons.
- Staging Validation: 90% of experiments run safely before production deployment
- Microservices Hardening: Systematic failure injection across diverse workload types (VMs and K8s)
- Community Contribution: Five upstream fixes enhancing the LitmusChaos project stability
This operational model serves as a blueprint for organizations adopting cloud-native practices. It highlights that reliability engineering is not merely about redundancy but actively breaking systems to prove their robustness.
Certification Relevance and Best Practices
For engineers studying for industry-recognized credentials, this case study underscores the practical application of concepts found in Kubernetes certifications (CKA/CKAD). Specifically, it demonstrates how to configure admission controllers like LitmusChaos within a cluster's lifecycle management strategy.
What This Means For You
The Flipkart case study validates that investing time into chaos engineering tooling yields immediate returns in system stability. By adopting similar patterns with Kubernetes Chaos Engineering Scale (KCE), your organization can proactively identify and resolve latent defects before they impact end-users.


