AI Audio · 25 Sep 2026 · 18:09 CEST
Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AI

Publisher preview · OZZZER analysis pending editorial review.
PUBLISHER ARTICLE PREVIEW
From the original article
With voice cloning, you can generate new speech in a target speaker’s voice from a short reference recording, without retraining a model. You can now deploy the publicly available Qwen3-TTS-12Hz-1.7B-Base text-to-speech model from Amazon SageMaker JumpStart to a fully managed, real-time inference endpoint.
Voice cloning reproduces the vocal identity of a specific speaker. Start with a short recording of the speaker and its transcript. Then supply the new text to synthesize. The model speaks that text in the reference speaker’s voice, without retraining. Media teams, educators, and application developers can use this capability to create personalized voice experiences and localize multilingual content.
They can also support accessible communication and preserve a speaker’s identity across languages.
With a self-hosted, publicly available voice cloning model, you control cost and keep audio data within your AWS environment. You can also adapt the model to your domain. With Amazon SageMaker AI, you can run the model on a fully managed real-time endpoint and handle infrastructure provisioning, health monitoring, and automatic scaling. You don’t manage the underlying GPU servers.
This post shows how to deploy Qwen3-TTS-12Hz-1.7B-Base from Amazon SageMaker JumpStart using the Amazon SageMaker Python SDK, and how to invoke the resulting endpoint to clone a voice from a reference clip. It also covers the configuration settings that make this deployment work in practice, along with the Amazon CloudWatch metrics you can use to monitor and right-size the endpoint.
Qwen3-TTS is a publicly available text-to-speech model family developed by the Qwen team at Alibaba Cloud. It covers 10 languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian. The models use the Qwen3-TTS-Tokenizer-12Hz speech tokenizer and support streaming generation for low-latency, interactive scenarios.
This post uses the Base variant, Qwen3-TTS-12Hz-1.7B-Base. It performs voice cloning from only a few seconds of user audio. It can also serve as a base for fine-tuning. For voice cloning, the model takes a reference audio clip and its transcript. It captures the speaker’s vocal characteristics, such as timbre, pitch, and cadence, and applies them to new text.
This differs from the CustomVoice variant, which generates speech from a fixed set of predefined speakers rather than a user-supplied reference.
The model also supports cross-lingual cloning: You can capture a voice from a reference in one language and generate speech in another while preserving the speaker’s vocal identity.
Qwen3-TTS-12Hz-1.7B-Base is available in Amazon SageMaker JumpStart alongside Qwen3-TTS-12Hz-1.7B-CustomVoice and Qwen3-ASR-1.7B. With these JumpStart options, you can use the streamlined deployment shown in this post.
With voice cloning, applications can reproduce the vocal identity of a chosen speaker from a reference recording,
Source
AWS AI · 25 Sep 2026 · 18:09 CEST
Open the original at AWS AI ↗