Input tokens sent to a foundation model (FM) on every inference call represent a meaningful portion of total cost when running Retrieval Augmented Generation at scale. A new architectural pattern addresses this by introducing query-aware compression, which reduces the volume of input data reaching the primary model while maintaining answer fidelity.
The Compression Pattern
Traditional RAG workflows typically retrieve a broad set of chunks—often 5 to 20—to ensure high recall. While confident that source material is available at inference time, this design causes technical-documentation and legal workloads to consume several thousand input tokens per query.
The proposed solution inserts an intermediate processing step between retrieval and the final answer generation. After retrieving chunks but before invoking the primary model for synthesis, a smaller foundation model reads both the retrieved context and the user's original query. This secondary filter outputs only verbatim spans relevant to the specific question. The primary model then receives this filtered subset of context rather than the full raw dump.Operational Implications
This pattern relies on a cascade architecture where both the compression call and the final answer generation occur within a single AWS Lambda function execution environment. While Anthropic Claude Haiku is used as an example for this smaller model, practitioners can substitute other small/primary model pairs available in their region.
The primary benefit extends beyond cost reduction; removing irrelevant context also reduces the surface area where hallucinations might originate from noise or misaligned data sources. This filtering step effectively refines what reaches the inference engine without altering upstream retrieval logic.Integration Considerations
This architecture is compatible with existing Amazon Bedrock capabilities, including Knowledge Bases and Intelligent Prompt Routing. It can also layer on top of prompt caching mechanisms to compound savings further.
Practitioners should evaluate the latency tradeoff inherent in adding this secondary filtering step against their specific service level objectives (SLOs). The cost model shifts from raw token volume per query toward a more optimized balance between retrieval breadth and generation efficiency. This approach supports custom post-retrieval processing steps that refine data quality before it hits expensive inference endpoints.Related CloudNinjas coverage: AWS.
What This Means For Practitioners
To implement this solution, teams must ensure they have access to the specific models required for both stages of the cascade in their AWS Region. Additionally, an IAM role with appropriate permissions is necessary for any Lambda function orchestrating these calls.
Engineers should assess whether current RAG workloads are hitting token limits or cost ceilings that justify this architectural shift. By filtering context at inference time rather than relying solely on retrieval parameters like top-k settings, teams can achieve significant input-token reduction while preserving the thoroughness required for complex answers.
