Live
Improved timeline accessibility: GitHub now presents issue and PR histories as navigable listsBatch‑Creating Cloudflare Workflow Instances Reduces Calls and Improves Type SafetyScaling Irish Workloads with Gemini Enterprise: Architecture and Ops ImplicationsDocsy Introduces AI‑Ready Documentation Features After Joining Linux FoundationProactive AI Incident Automation: Architectural Shifts and Operational GuardrailsWhen an AI Agent Inherits Your Azure Credential: Risks and Architecture ImplicationsGround Truth CLI Brings Headless Observability to AI‑Assisted TroubleshootingImplementing Multi‑Tenant GPU Sharing on SageMaker HyperPod with EKSImproved timeline accessibility: GitHub now presents issue and PR histories as navigable listsBatch‑Creating Cloudflare Workflow Instances Reduces Calls and Improves Type SafetyScaling Irish Workloads with Gemini Enterprise: Architecture and Ops ImplicationsDocsy Introduces AI‑Ready Documentation Features After Joining Linux FoundationProactive AI Incident Automation: Architectural Shifts and Operational GuardrailsWhen an AI Agent Inherits Your Azure Credential: Risks and Architecture ImplicationsGround Truth CLI Brings Headless Observability to AI‑Assisted TroubleshootingImplementing Multi‑Tenant GPU Sharing on SageMaker HyperPod with EKS
AWS

Architecting Grounded AI at Scale with Amazon Bedrock

AI SummaryPowered by AI

Qlik re‑architected Qlik Answers with a layered design that routes requests to specialist agents and accesses Amazon Bedrock via a dedicated gateway, adding OpenSearch retrieval and fallback to SageMaker. This change lets engineers build regulated, region‑aware, high‑throughput generative AI while keeping capacity planning and security controls explicit.

Qlik has re‑engineered its Qlik Answers product to run on Amazon Bedrock using a multi‑layered architecture that separates conversational entry, routing, answer generation, specialist agents, analytics handling, document retrieval, and model access. The redesign delivers a grounded, source‑backed AI experience that scales across regions, respects data‑sovereignty, and provides predictable capacity planning, which directly impacts AI engineers, platform teams, SREs, and security practitioners.

Why the Change Matters

Traditional monolithic assistants become slower and less accurate as capabilities grow. By breaking the system into focused layers, Qlik can add new specialist agents without touching the user‑facing interface, keep latency low for simple queries, and enforce Bedrock Guardrails consistently. For operators, this means clearer service boundaries, easier monitoring, and a path to meet regulatory requirements without duplicating entire deployments per region.

Layered Architecture Overview

  1. Entry layer – a stable conversational endpoint inside Qlik Cloud that shields downstream changes from the client.
  2. Routing layer – a lightweight component that inspects the incoming message and context, then decides which downstream path to invoke. Its sole responsibility is fast, accurate routing.
  3. Answer layer – orchestrates response creation. It chooses a fast path for trivial look‑ups or a deliberative path that decomposes the request into sub‑questions and aggregates results.
  4. Specialist agent layer – a shared swarm runtime that hosts domain‑specific agents, tools, and optional human‑in‑the‑loop steps. New agents plug into this runtime without redefining orchestration logic.
  5. Conversational analytics layer – routes structured data questions to an app‑aware reasoning path built for analytics rather than free‑form text generation.
  6. Retrieval layer – indexes unstructured documents in Amazon OpenSearch Service, providing searchable content that grounds answers to knowledge‑base and document queries.
  7. Model access layer – a Qlik‑owned LLM gateway that forwards chat, streaming, embedding, and reranking calls to Amazon Bedrock. It applies Bedrock Guardrails for prompt‑injection protection, PII filtering, secret redaction, and denied‑topic enforcement, plus a grounding‑validation step on generated text. When a required model is unavailable in a region, the request is temporarily handled by Amazon SageMaker AI and later migrated back to Bedrock once regional support arrives.

Operational Implications

Capacity forecasting now focuses on token consumption and model availability 3–6 months ahead of major releases. Qlik validates these forecasts against actual usage after launch, allowing the team to adjust provisioning before a capacity shortfall occurs. The separation of retrieval (OpenSearch) and generation (Bedrock) enables independent scaling: indexing pipelines can be tuned without affecting LLM throughput.

Region‑specific deployments are achieved by reusing the same layered codebase while configuring the model access layer to point at the appropriate Bedrock endpoint. This avoids the operational overhead of maintaining eleven divergent stacks, yet satisfies data‑sovereignty constraints.

Security Considerations

Guardrails are applied at the model access layer for every request and response, providing content‑level filtering without acting as an authorization boundary. The retrieval layer stores documents in OpenSearch, so standard OpenSearch security controls (IAM policies, encryption at rest, VPC isolation) remain applicable. The fallback to SageMaker introduces an additional surface that must be secured with the same IAM and network controls used for Bedrock.

Related CloudNinjas coverage: AWS.

What This Means For Practitioners

Engineers should evaluate whether a layered approach can replace monolithic LLM integrations in their own products, especially when they need to meet regional data‑residency rules or enforce consistent content filtering. Monitoring token usage early helps avoid capacity surprises, and separating retrieval from generation simplifies scaling. Finally, applying Guardrails at the gateway level provides a uniform safety net, but teams must still enforce network and IAM boundaries around each service component.

Originally published atAWS Machine Learning Blog