OZZZER · AI NEWS1 of 3 free stories opened
← Back to AI News

Governance · 25 Sep 2026 · 18:18 CEST

Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

AWS AI · 25 Sep 2026 · 18:18 CESTRead original at AWS AI ↗
Share
LinkedInX
Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

Publisher preview · OZZZER analysis pending editorial review.

PUBLISHER ARTICLE PREVIEW

From the original article

Reinforcement learning (RL) post-training is becoming a standard step in building capable language model agents. Models learn to reason and act across sequences of steps by generating trajectories, receiving rewards, and updating their policy based on outcomes. Running this at scale, across multiple nodes with hundreds of GPU-hours of rollouts per training run, requires persistent cluster infrastructure.

That infrastructure needs to sustain long jobs, recover from hardware failures without losing progress, and provide visibility into training dynamics as they unfold.

Amazon SageMaker HyperPod provides this infrastructure for large-scale machine learning (ML) workloads on Amazon Elastic Kubernetes Service (Amazon EKS). Through its cluster resiliency features, it continuously monitors node health and automatically replaces faulty nodes, so a hardware failure does not take the cluster down with it. Paired with checkpointing, a training job can pick up from its last saved step instead of restarting from scratch.

This matters for long multi-node RL runs, where a single hardware failure would otherwise cost hours of rollout progress. Combined with the Ray capabilities on HyperPod, you can create Ray clusters from SageMaker Studio, submit jobs remotely using secure connections, and monitor training through pre-built Amazon Managed Grafana dashboards that the HyperPod Observability EKS add-on provisions for you.

In this post, we show how to use these capabilities to run SkyRL, an open-source RL framework, to train a Qwen3-VL-8B vision-language model to navigate visual mazes using Group Relative Policy Optimization (GRPO) on SageMaker HyperPod. Starting from the VisGym SFT checkpoint, a supervised fine-tuning (SFT) starting point, GRPO post-training on HyperPod improves the maze solve rate from 43.75% to more than 95% on a fixed 64-maze evaluation set.

This section reviews the reinforcement learning concepts behind the training and the cluster topology the walkthrough uses.

Standard single-turn RL assigns a reward to a single model output. Multi-turn RL instead trains an agent over a whole sequence of steps, where it observes a state, acts, gets feedback, and moves on to the next state. The policy learns from the reward accumulated over the entire episode rather than from any one step.

Consider the example problem of navigating a 2D maze. One episode is a single run at a maze, and each turn is one move: the model looks at the current picture of the maze, chooses a direction or decides to stop, and the environment sends back the updated view. Rewards are sparse, so the model earns 1.0 only when it actually reaches the goal within the move limit and nothing otherwise.

There is no move-by-move answer key to train against, since whether a move was good depends on the moves around it.

This is where SkyRL’s Group Relative Policy Optimization (GRPO) comes in. For each starting position, the agent runs the maze several times under the current policy, and GRPO grades those runs against one another, reinforcing the ones that beat the group’s average and pushing down the ones that trail it. That within-group comparison is the whole training signal, which lets GRPO work without a separate critic or value model.

The solution discussed here runs SkyRL on a HyperPod Ray cluster with three GPU worker nodes and a CPU head node. SkyRL colocates inference and training on the same GPUs: vLLM engines generate rollouts (complete maze episodes) while a policy model sharded with Fully Sharded Data Parallel (FSDP) handles gradient updates. After each optimizer step, updated LoRA adapter weights sync from the training ranks to the inference engines through Amazon FSx for Lustre shared storage.

These are the instance types we used. Other GPU instances and cluster sizes work as well, provided the workers have enough GPU memory for the model.

HyperPod provides the cluster infrastructure: the Ray cluster is created from SageMaker Studio, job submission uses the sagemaker_ray:// protocol, and training metrics flow automatically into pre-built Amazon Managed Grafana dashboards through the HyperPod Observability add-on.

Figure 1: RayCluster topology on Amazon SageMaker HyperPod, with one CPU head node and three GPU worker nodes that colocate FSDP policy shards and vLLM rollout engines over a shared Amazon FSx for Lustre filesystem

The following steps walk through preparing the training environment, launching the cluster, running the job, monitoring progress, and hosting the trained model.

To get started quickly, use the following Dockerfile to build a container image with SkyRL, VisGym, and their dependencies pre-installed. This is the image you will specify when launching your Ray cluster on HyperPod in the next step. It builds on the official NovaSky-AI SkyRL base and pins both SkyRL and VisGym to specific commit SHAs so the build is reproducible:

Build the image and push it to an Amazon Elastic Container Registry (Amazon ECR) repository in your account. Note the full image URI, as you will use it when creating the Ray cluster in the next step:

Navigate to SageMaker Studio, choose HyperPod, select your cluster, then go to the Tasks tab. From the task type list, choose RayCluster, then choose Create Ray Cluster.

In the

Source

AWS AI · 25 Sep 2026 · 18:18 CEST

Open the original at AWS AI ↗