AWS has demonstrated an image-to-video workflow on SageMaker AI using two models served from the same vLLM-Omni container image. A real-time endpoint runs FLUX.2-klein-4B to turn a text prompt into a still image, then an asynchronous endpoint uses Wan2.1-VACE-1.3B to animate that image from a motion prompt.

The image endpoint returns a base64-encoded PNG directly. The sample resizes and converts it before passing it to the video model, whose longer-running job writes an MP4 to Amazon S3. AWS includes command-line code and an optional Streamlit interface. Keeping the models on separate endpoints lets operators choose different instance types and response patterns while retaining a common serving stack.

The vLLM-Omni project adapts vLLM’s serving approach to models that process or produce audio, images, and video through OpenAI-compatible APIs. Reusing a container can reduce software variation, but it does not combine the two models into one endpoint or remove the compute demands of video generation. Teams adopting the sample still need to manage endpoint capacity, S3 access, model licenses, and asynchronous job handling. The tutorial provides a reproducible deployment pattern rather than a managed consumer video product.