The latest example shows how to host two distinct generative models—FLUX.2‑klein‑4B for image creation and Wan2.1‑VACE‑1.3B for video synthesis—using a single AWS vLLM‑Omni deep‑learning container on SageMaker AI. One endpoint runs in real‑time, returning a base64‑encoded PNG directly, while the other runs as an asynchronous inference job that stores the resulting MP4 in Amazon S3. This split lets engineers match latency expectations and instance sizing to each workload without maintaining separate container images.
Solution overview
The architecture reuses the same pinned vLLM‑Omni DLC image for both endpoints. Deployment swaps the SM_VLLM_MODEL environment variable to load either the FLUX.2‑klein model or the Wan VACE model. The real‑time endpoint exposes /v1/images/generations and returns the generated image inline. The asynchronous endpoint exposes /v1/videos/sync, reads a multipart request from S3 that contains the image reference and a motion prompt, and writes the final MP4 back to a caller‑provided S3 location.
Implementation steps
- Clone the sample repository and run the provided
deploy.pyscript twice, adjustingSM_VLLM_MODELfor each model. - Select instance types that suit the model’s compute profile; the image model can use a lower‑cost real‑time instance, while the video model may require a larger instance for the longer inference job.
- For the image call, the application sends a text prompt, receives a base64 PNG, resizes it, converts it to a compact JPEG data URL, and embeds that URL in the multipart payload for the video request.
- The video request is uploaded to S3; SageMaker Async Inference reads the payload, runs Wan VACE, and writes the MP4 to the output S3 path. The application polls or receives a notification to retrieve the file.
Operational and security implications
Running two endpoints from the same container reduces the surface area of container management but introduces separate IAM permissions for real‑time and async endpoints, as each interacts with S3 in different ways. Practitioners should ensure the S3 bucket policies grant write access only to the async endpoint’s execution role and read access only to the consuming application. Monitoring should track both endpoint latency (real‑time) and job completion time (async) to detect performance regressions. Because the video model writes large binary artifacts, storage lifecycle policies may be needed to avoid uncontrolled growth.
Related CloudNinjas coverage: AWS.
What This Means For Practitioners
Adopting a shared vLLM‑Omni container simplifies version control and reduces operational overhead, while separate endpoint types let you optimize cost and performance per model. Evaluate your workload to decide which models merit real‑time versus asynchronous handling, and configure S3 permissions and monitoring accordingly. The pattern also provides a clear migration path for adding additional multimodal models without duplicating container images.


