AI Audio · 28 Sep 2026 · 18:15 CEST
Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1

Publisher preview · OZZZER analysis pending editorial review.
PUBLISHER ARTICLE PREVIEW
From the original article
Voice agents, interactive learning applications, accessibility tools, and customer service assistants need to respond without long silent pauses. In this tutorial, you deploy a text-to-speech (TTS) model on Amazon SageMaker AI that can start playing speech before it finishes generating the full response. You use the AWS vLLM-Omni Deep Learning Container (DLC) to deploy Qwen3-TTS, stream text in and audio out over one persistent bidirectional connection, and try the workflow through a Gradio application.
AWS Deep Learning Containers provide Docker images with deep learning frameworks and dependencies for training and inference on AWS. AWS provides deployment guidance for broadly adopted serving frameworks such as vLLM and SGLang. This post is Part 1 of a series about specialized DLCs, including vLLM-Omni, WhisperX, and llama.cpp. It focuses on streamed speech for real-time voice applications.
Part 2 applies the vLLM-Omni DLC to image and video generation. The series pairs focused use cases with deployment examples and reproducible benchmarks where they add useful evidence.
The vLLM-Omni project extends vLLM beyond text generation to serve models that process or generate text, audio, images, and video. The AWS vLLM-Omni DLC packages tracked vLLM-Omni releases in AWS images and adds routing middleware for SageMaker AI. You use SageMaker bidirectional streaming to send text and receive audio chunks over the persistent connection.
The earlier post, Build real-time voice applications with Amazon SageMaker AI and vLLM, demonstrates the input side of a voice pipeline. It streams microphone audio to the Voxtral-Mini-4B Realtime speech-to-text (STT) model and returns transcription events. This post adds the output side by sending text to Qwen3-TTS and streaming generated speech back through the vLLM-Omni DLC.
You will clone the code sample, deploy a Qwen3-TTS model, and try streaming speech through a Gradio application.
Specialized inference runtimes support model architectures and media pipelines that differ from general text generation. AWS DLCs package these runtimes with
Source
AWS AI · 28 Sep 2026 · 18:15 CEST
Open the original at AWS AI ↗