Live
AI Agent Inbox: Deploy Pizza Bot for Background Task ExecutionOpenAPPA delivers zero‑success prompt‑injection protection in benchmark tests – what AI engineers need to knowEU Cyber Resilience Act expands software supply‑chain responsibilities for digital product manufacturersTyped Probability Model Jev Shifts AI Output from Text to Structured DecisionsBasin Pipelines per‑stream ingest capacity jumps to 1 GB/s – what engineers need to knowAI‑driven vulnerability management: moving from CVE counts to contextual riskDynamic Tier in Google Cloud Managed Lustre: Cost‑Effective, Low‑Latency Storage for AI and HPCArgo CD 4.0 Visioning and Scaling Lessons from ArgoCon NA 2026AI Agent Inbox: Deploy Pizza Bot for Background Task ExecutionOpenAPPA delivers zero‑success prompt‑injection protection in benchmark tests – what AI engineers need to knowEU Cyber Resilience Act expands software supply‑chain responsibilities for digital product manufacturersTyped Probability Model Jev Shifts AI Output from Text to Structured DecisionsBasin Pipelines per‑stream ingest capacity jumps to 1 GB/s – what engineers need to knowAI‑driven vulnerability management: moving from CVE counts to contextual riskDynamic Tier in Google Cloud Managed Lustre: Cost‑Effective, Low‑Latency Storage for AI and HPCArgo CD 4.0 Visioning and Scaling Lessons from ArgoCon NA 2026
AWS

Closing the Loop on ITSM Data with Bedrock, S3 Vectors, and ECS

AI SummaryPowered by AI

A new architecture pattern demonstrates how to mine resolved incident tickets for knowledge generation while simultaneously curating existing articles via deduplication and quality scoring. This approach matters because it transforms static ticket history into a dynamic RAG grounding layer that reduces hallucinations in support workflows.

Enterprise IT teams often face a paradox: thousands of resolved incidents contain valuable troubleshooting data, yet this knowledge remains trapped within unstructured ticket histories while the official Knowledge Base (KB) suffers from stale content and duplicates. The solution presented introduces an automated pipeline that treats these two systems as parts of a single closed-loop lifecycle.

Architecture Pattern: Generation with RAG Grounding

The core engineering challenge is generating high-quality articles without inventing procedures or using outdated terminology. To solve this, the architecture implements Retrieval Augmented Generation (RAG) before writing occurs. When a cluster of related tickets arrives in Amazon S3, an upstream process groups them by theme and drops the result into a bucket scoped to one customer.

Before generating new content using Anthropic Claude Sonnet 4.5 via Amazon Bedrock, the system queries its existing knowledge base stored as vectors in Amazon S3 Vectors. It retrieves the five most similar articles from that index and passes them into the model prompt as reference context.

This step is critical for practitioners building LLM applications. By grounding generation against a vector store of approved content, engineers ensure terminology consistency across different product versions before an article ever reaches human review. If no existing vectors match the theme, the system generates from ticket data alone but flags procedures specifically for manual verification.

Compute Strategy: ECS Fargate vs Lambda

The architecture distinguishes between generation and curation workloads based on compute requirements. Generation runs on Amazon Elastic Container Service (Amazon ECS) with AWS Fargate rather than serverless functions like Lambda.

This distinction exists because generating two full documents for a single theme can take several minutes, creating long-running tasks that exceed typical function timeout limits or require significant memory allocation during token streaming. The containerized approach allows the system to poll an Amazon SQS queue and process up to five themes simultaneously without cold starts impacting throughput.

Operational Implications: Curation as a Feedback Loop

The curation subsystem operates on every article, whether newly generated or existing. It executes four sequential steps via AWS Step Functions orchestrated by Lambda functions:

  • Classification: Sorting articles into specific types.
  • Deduplication: Removing near-identical drafts to prevent KB bloat.
  • Quality Scoring: Evaluating content accuracy and currency.
  • Improvement: Rewriting weak or stale sections using generative models.

The loop closes because curation embeds every approved article back into the Amazon S3 Vectors index. This ensures that generation always has fresh vectors to ground against on subsequent runs, preventing knowledge drift over time.

What This Means For Practitioners

This pattern offers a blueprint for building large-scale document-processing pipelines where data quality is paramount before it reaches the model layer. Platform teams should consider whether their own RAG implementations rely on static datasets or if they can automate vector updates to maintain grounding accuracy.

For security and compliance, note that this architecture relies heavily on human-in-the-loop review via ServiceNow for final approval. This manual gate ensures a person still owns what goes live into production environments before the decision is written back to Amazon DynamoDB. Engineers should evaluate whether their own pipelines can similarly enforce strict ownership boundaries between automated generation and public-facing knowledge assets.

Finally, practitioners using AWS services must ensure they have access enabled for specific models like Anthropic Claude Sonnet 4.5 and Titan Text Embeddings V2 before deploying this pattern to production environments.

Originally published atAWS Machine Learning Blog