Enterprise IT teams often face a paradox: thousands of resolved incidents contain valuable troubleshooting data, yet this knowledge remains trapped within unstructured ticket histories while the official Knowledge Base (KB) suffers from stale content and duplicates. The solution presented introduces an automated pipeline that treats these two systems as parts of a single closed-loop lifecycle.
Architecture Pattern: Generation with RAG Grounding
The core engineering challenge is generating high-quality articles without inventing procedures or using outdated terminology. To solve this, the architecture implements Retrieval Augmented Generation (RAG) before writing occurs. When a cluster of related tickets arrives in Amazon S3, an upstream process groups them by theme and drops the result into a bucket scoped to one customer.
Before generating new content using Anthropic Claude Sonnet 4.5 via Amazon Bedrock, the system queries its existing knowledge base stored as vectors in Amazon S3 Vectors. It retrieves the five most similar articles from that index and passes them into the model prompt as reference context.
This step is critical for practitioners building LLM applications. By grounding generation against a vector store of approved content, engineers ensure terminology consistency across different product versions before an article ever reaches human review. If no existing vectors match the theme, the system generates from ticket data alone but flags procedures specifically for manual verification.
Compute Strategy: ECS Fargate vs Lambda
The architecture distinguishes between generation and curation workloads based on compute requirements. Generation runs on Amazon Elastic Container Service (Amazon ECS) with AWS Fargate rather than serverless functions like Lambda.
This distinction exists because generating two full documents for a single theme can take several minutes, creating long-running tasks that exceed typical function timeout limits or require significant memory allocation during token streaming. The containerized approach allows the system to poll an Amazon SQS queue and process up to five themes simultaneously without cold starts impacting throughput.
Operational Implications: Curation as a Feedback Loop
The curation subsystem operates on every article, whether newly generated or existing. It executes four sequential steps via AWS Step Functions orchestrated by Lambda functions:
- Classification: Sorting articles into specific types.
- Deduplication: Removing near-identical drafts to prevent KB bloat.
- Quality Scoring: Evaluating content accuracy and currency.
- Improvement: Rewriting weak or stale sections using generative models.
The loop closes because curation embeds every approved article back into the Amazon S3 Vectors index. This ensures that generation always has fresh vectors to ground against on subsequent runs, preventing knowledge drift over time.
What This Means For Practitioners
This pattern offers a blueprint for building large-scale document-processing pipelines where data quality is paramount before it reaches the model layer. Platform teams should consider whether their own RAG implementations rely on static datasets or if they can automate vector updates to maintain grounding accuracy.
For security and compliance, note that this architecture relies heavily on human-in-the-loop review via ServiceNow for final approval. This manual gate ensures a person still owns what goes live into production environments before the decision is written back to Amazon DynamoDB. Engineers should evaluate whether their own pipelines can similarly enforce strict ownership boundaries between automated generation and public-facing knowledge assets.
Finally, practitioners using AWS services must ensure they have access enabled for specific models like Anthropic Claude Sonnet 4.5 and Titan Text Embeddings V2 before deploying this pattern to production environments.



