We swapped a pure‑LLM incident‑diagnosis approach for a retrieval‑augmented generation (RAG) pipeline that harvests existing tickets, runbooks, postmortems, and live alerts, then feeds the most relevant fragments to a language model for reasoning. The shift matters because it turns the dominant knowledge‑retrieval bottleneck into a fast, repeatable service, cutting average diagnosis time by roughly 60 % and pushing root‑cause accuracy above 87 % for on‑call engineers.
Architecture Overview
The system consists of five layers that map directly onto assets already present in most organisations:
- Data sources: incident tickets, runbooks, postmortems, and alert streams are pulled from
ServiceNow,PagerDuty,Confluence,GitHub, andPrometheus/Dynatrace. No new storage or migration is required. - Ingestion and indexing: documents are split into chunks using three strategies – fixed‑window for long‑form docs, boundary‑aware chunking for tickets (preserving summary, timeline, resolution sections), and sentence‑level chunking for short alert notes. Each chunk is stored with provenance metadata (source, section, timestamp, service labels).
- Retrieval: chunks are embedded and placed in a
FAISSindex for approximate nearest‑neighbor search. A cross‑encoder re‑ranks the top candidates before they reach the language model, a step that contributed a measurable lift in accuracy. - LLM reasoning: the selected context, together with incident metadata, is assembled into a structured prompt that asks the model to enumerate ranked root causes, confidence scores, remediation steps, escalation flags, and links back to the source fragments. The design works with
GPT‑4‑Turbo,Claude 3 Opus, andLlama 3.1 70B, demonstrating model‑agnostic behaviour. - Feedback loop: once an incident is resolved, the confirmed root cause is annotated and fed back into the knowledge base, keeping the retrieval set current as infrastructure evolves.
Key Engineering Decisions
Two implementation choices proved decisive during evaluation:
- Cross‑encoder re‑ranking – Skipping this step and feeding raw vector results directly to the LLM reduced accuracy by 7.7 percentage points in an ablation study.
- Semantic chunking per document type – Using a one‑size‑fits‑all window degraded performance; the boundary‑aware approach added 10.9 percentage points of accuracy, making provenance metadata useful for downstream reasoning.
Both choices are relatively lightweight to add to an existing pipeline but have outsized impact on diagnostic quality.
Operational Impact
Testing on 2,400 annotated incidents – a mix of real financial‑services tickets, the DeathStarBench microservice benchmark with injected faults, and synthetic edge cases – showed consistent gains:
- Root‑cause identification accuracy: 87.3 % versus 71.8 % for the next‑best BM25 + LLM baseline.
- Mean diagnosis time reduced by 59 % overall.
- For P1 incidents, average manual search time fell from 48.2 minutes to 19.8 minutes.
These figures indicate that the RAG pipeline not only speeds up response but also improves the reliability of the suggested remediation, which is critical for high‑severity outages.
Related CloudNinjas coverage: DevOps.
What This Means For Practitioners
Adopting a RAG‑based incident diagnosis stack is practical for teams that already maintain ticketing, documentation, and observability tools. The primary effort lies in building robust ingestion pipelines and selecting appropriate chunking logic for each source type. Adding a cross‑encoder re‑ranker and preserving provenance metadata are low‑cost steps that deliver measurable accuracy gains. Once in place, the feedback loop ensures the knowledge base stays current, reducing the risk of model drift as services evolve. Teams should evaluate their existing data pipelines for compatibility, prototype the retrieval layer with FAISS, and run a small‑scale ablation to confirm the value of re‑ranking in their environment before scaling to production.
