Amazon SageMaker AI now supports streamed speech inference by pairing the vLLM‑Omni deep‑learning container with the Qwen3‑TTS model, allowing audio chunks to be returned over a single persistent bidirectional connection. This change lets engineers replace the traditional request‑response pattern with real‑time voice output, which is critical for low‑latency voice agents and accessibility tools.
Architecture Overview
The solution builds on the vLLM‑Omni project, an extension of vLLM that adds multimodal routing and streaming capabilities. The container image bundles the Qwen3‑TTS model and middleware that translates SageMaker’s bidirectional streaming API into OpenAI‑compatible calls. A SageMaker endpoint exposes a full‑duplex WebSocket, enabling a client to send text and receive incremental audio payloads without waiting for the full synthesis to finish.
Implementation Steps
All required artifacts reside in the 03-features/bidirectional-streaming-vLLM-Omni directory of the aws-samples/sagemaker-genai-hosting-examples repository. Practitioners clone the repo, run the provided deployment script to create a SageMaker endpoint using the vLLM‑Omni DLC, and then launch the bundled Gradio client to test the end‑to‑end flow. The client opens a WebSocket, streams text prompts to the endpoint, and plays back the returned audio chunks as they arrive.
Operational and Security Implications
Because the endpoint maintains a persistent connection, monitoring must include WebSocket health metrics and latency of audio chunk delivery. Scaling policies should consider the continuous nature of the stream rather than isolated request spikes. The specialized DLC abstracts framework dependencies, reducing the surface area for version‑drift bugs, but it also means that any security updates to the container must be applied through the standard SageMaker image‑update process.
Access to the streaming endpoint is governed by SageMaker IAM permissions; practitioners should grant the minimum set of actions required for endpoint invocation. Network controls need to allow outbound WebSocket traffic from client environments and inbound traffic to the SageMaker endpoint, which may affect VPC security group configurations. Since the container handles both text and audio payloads, data‑in‑transit encryption (TLS) remains essential, and any logging of audio data should respect privacy requirements.
Related CloudNinjas coverage: AWS.
What This Means For Practitioners
Engineers can now prototype real‑time voice experiences without building custom streaming infrastructure, leveraging a managed SageMaker endpoint and a pre‑packaged multimodal container. The primary considerations are ensuring WebSocket connectivity, applying least‑privilege IAM policies, and monitoring streaming health. Future evaluations should compare latency and cost against alternative streaming solutions and verify that the container’s update cadence aligns with organizational security patching cycles.

