Live
Mitigating the New NetScaler ADC Zero‑Day Exploits in Production EnvironmentsNew Mesh and Workers VPC logging fields improve Cloudflare traffic observabilityAutomating Resource Ownership Tracking to Eliminate Orphaned Cloud AssetsFrom RAG to Structured Extraction: Building an AI Contract Intelligence Pipeline on AWSFabric‑Copilot Integration Shifts Data Foundations for AI‑Driven AppsEnv Zero’s EZ Control adds a policy‑driven control plane for agentic DevOps workflowsDecoupled Multimodal Video Search Using Bedrock Embeddings and OpenSearchGKE Agent Sandbox cuts RL sandbox startup to seconds, easing GPU idle and control‑plane loadMitigating the New NetScaler ADC Zero‑Day Exploits in Production EnvironmentsNew Mesh and Workers VPC logging fields improve Cloudflare traffic observabilityAutomating Resource Ownership Tracking to Eliminate Orphaned Cloud AssetsFrom RAG to Structured Extraction: Building an AI Contract Intelligence Pipeline on AWSFabric‑Copilot Integration Shifts Data Foundations for AI‑Driven AppsEnv Zero’s EZ Control adds a policy‑driven control plane for agentic DevOps workflowsDecoupled Multimodal Video Search Using Bedrock Embeddings and OpenSearchGKE Agent Sandbox cuts RL sandbox startup to seconds, easing GPU idle and control‑plane load
AWS

Deploying Streamed Speech Inference on SageMaker with vLLM‑Omni

AI SummaryPowered by AI

SageMaker AI now offers streamed speech inference using the vLLM‑Omni container and Qwen3‑TTS, delivering audio chunks over a persistent bidirectional connection. This enables low‑latency voice applications while shifting operational focus to streaming health, IAM scoping, and container update management.

Amazon SageMaker AI now supports streamed speech inference by pairing the vLLM‑Omni deep‑learning container with the Qwen3‑TTS model, allowing audio chunks to be returned over a single persistent bidirectional connection. This change lets engineers replace the traditional request‑response pattern with real‑time voice output, which is critical for low‑latency voice agents and accessibility tools.

Architecture Overview

The solution builds on the vLLM‑Omni project, an extension of vLLM that adds multimodal routing and streaming capabilities. The container image bundles the Qwen3‑TTS model and middleware that translates SageMaker’s bidirectional streaming API into OpenAI‑compatible calls. A SageMaker endpoint exposes a full‑duplex WebSocket, enabling a client to send text and receive incremental audio payloads without waiting for the full synthesis to finish.

Implementation Steps

All required artifacts reside in the 03-features/bidirectional-streaming-vLLM-Omni directory of the aws-samples/sagemaker-genai-hosting-examples repository. Practitioners clone the repo, run the provided deployment script to create a SageMaker endpoint using the vLLM‑Omni DLC, and then launch the bundled Gradio client to test the end‑to‑end flow. The client opens a WebSocket, streams text prompts to the endpoint, and plays back the returned audio chunks as they arrive.

Operational and Security Implications

Because the endpoint maintains a persistent connection, monitoring must include WebSocket health metrics and latency of audio chunk delivery. Scaling policies should consider the continuous nature of the stream rather than isolated request spikes. The specialized DLC abstracts framework dependencies, reducing the surface area for version‑drift bugs, but it also means that any security updates to the container must be applied through the standard SageMaker image‑update process.

Access to the streaming endpoint is governed by SageMaker IAM permissions; practitioners should grant the minimum set of actions required for endpoint invocation. Network controls need to allow outbound WebSocket traffic from client environments and inbound traffic to the SageMaker endpoint, which may affect VPC security group configurations. Since the container handles both text and audio payloads, data‑in‑transit encryption (TLS) remains essential, and any logging of audio data should respect privacy requirements.

Related CloudNinjas coverage: AWS.

What This Means For Practitioners

Engineers can now prototype real‑time voice experiences without building custom streaming infrastructure, leveraging a managed SageMaker endpoint and a pre‑packaged multimodal container. The primary considerations are ensuring WebSocket connectivity, applying least‑privilege IAM policies, and monitoring streaming health. Future evaluations should compare latency and cost against alternative streaming solutions and verify that the container’s update cadence aligns with organizational security patching cycles.

Originally published atAWS Machine Learning Blog