Unclassified · 30 Sep 2026 · 23:19 CEST
Deploying an HSTU Generative Recommender with NVIDIA Dynamo-Triton

Publisher preview · OZZZER analysis pending editorial review.
Generative recommender (GR) systems are emerging as a powerful new approach for large-scale personalization. Instead of treating recommendation as a set of isolated retrieval, ranking, and prediction stages, GRs reformulate recommendation as sequence modeling over user behavior. A user’s interactions, context, candidate items, and actions become tokens in a high-cardinality event stream, and the model learns to generate or score the next relevant items from that sequence.
This approach is especially attractive for modern recommendation workloads, where user histories can be long, item catalogs are constantly changing, and personalization quality depends on modeling rich sequential behavior. But it also introduces a serving challenge: GR models need low-latency inference despite long histories, large embedding tables, and sequence-heavy model architectures. NVIDIA Dynamo-Triton (formerly NVIDIA Triton Inference Server) now supports an end-to-end Hierarchical Sequential Transduction Unit (HSTU) GR inference workflow through the NVIDIA recsys-examples repository.
The workflow combines HTSUs, PyTorch Ahead-of-Time Inductor compilation, FlexKV-backed KV caching, native C++ validation, NV embedding cache, and Dynamo-Triton deployment. The result is a practical path for serving HSTU ranking models with strong latency performance. At dynamic batch…
Excerpt supplied by the publisher.
Source
NVIDIA · 30 Sep 2026 · 23:19 CEST
Open the original at NVIDIA ↗