AWS Ships Multi-Turn RL Infrastructure for Amazon Nova on SageMaker HyperPod
AWS has released infrastructure for deploying multi-turn reinforcement learning to train Amazon Nova models on SageMaker HyperPod. The system requires a minimum of 10 ml.p5.48xlarge instances and costs approximately $786-$1,180 per hour when running.
AWS Ships Multi-Turn RL Infrastructure for Amazon Nova on SageMaker HyperPod
AWS has released infrastructure for deploying multi-turn reinforcement learning (RL) to train Amazon Nova models on Amazon SageMaker HyperPod. The system addresses the training challenge of enterprise agents that execute multi-step workflows requiring sequential decision-making across multiple turns.
Technical Architecture
The infrastructure operates as an event-driven pipeline with three compute layers:
- SageMaker HyperPod (EKS): Training pods and vLLM generation replicas run on P5 instances, generating responses and performing GRPO (Group Relative Policy Optimization) weight updates
- ECS on AWS Fargate: Reward workers execute custom environments, receive model responses via SQS, and return reward signals
- Amazon Nova Forge SDK: Routes messages between the model and reward environment while tracking conversation state
The system uses a two-phase deployment model. A one-time AWS CDK deployment provisions the foundation (VPC, EKS/HyperPod, ECS, S3, IAM), while each training run spins up ephemeral resources. AWS Step Functions orchestrates runs, triggered by Amazon EventBridge when data lands in S3.
Infrastructure Requirements
Minimum specifications:
- 10 ml.p5.48xlarge instances (12-14 recommended for production)
- Amazon Nova Forge subscription
- Python 3.12+
- AWS CDK v2
- Docker for Lambda container image builds
Operating costs: $786-$1,180 per hour when running, with deployment taking 30-40 minutes.
Configuration Parameters
Key settings in cdk.json:
instance_type: ml.p5.48xlarge (default) or ml.p5en.48xlargeinstance_count: 10 (default)nova_model: NOVA_MICRO, NOVA_LITE, NOVA_LITE_2, or NOVA_PROgeneration_replicas: 4 (default)global_batch_size: 64 samples per training stepmax_steps: 10 (default)
The system supports two training methods: RFT_MULTITURN_FULL and RFT_MULTITURN_LORA. Users can override parameters at deploy time via CDK command-line flags.
Training Approach
Unlike standard RLHF, which optimizes single responses in isolation, multi-turn RL optimizes over entire interaction sequences. According to AWS, this approach teaches agents tool orchestration, error recovery, and multi-step reasoning through trial and error—capabilities that supervised fine-tuning, RAG, and continued pre-training typically do not provide on their own.
The infrastructure includes a built-in Wordle environment for validation, with support for custom Bring Your Own Orchestrator (BYOO) environments.
Availability
The infrastructure is available now via GitHub at aws-samples/nova-multi-turn-rl-infra. AWS also offers multi-turn RL as a fully managed, serverless capability in SageMaker AI for users who don't require full control over the training stack.
What This Means
This release gives enterprises infrastructure to train agents on complex, multi-step workflows without managing the orchestration layer themselves. The $786-$1,180 hourly cost puts this firmly in the enterprise category, but the event-driven architecture and ephemeral resource model prevent idle GPU waste. The two-phase deployment separates foundation from compute, reducing iteration cycles for teams that need custom reward environments or specific instance configurations beyond what AWS's managed service provides.
Related Articles
AWS releases open-source MCP server to automate cross-account promotion of Amazon Quick agents
AWS has published the Quick Resource Migrator, a sample MCP server on Amazon Bedrock AgentCore that promotes Amazon Quick resources between AWS accounts in a single tool call. It is idempotent, never deletes from the target, and writes versioned S3 backups before every update.
AWS adds managed Web Search to Claude Desktop via Bedrock AgentCore Gateway in three Regions
AWS published a walkthrough for connecting Claude Desktop on Amazon Bedrock to a managed, MCP-compatible Web Search capability through Amazon Bedrock AgentCore Gateway. According to AWS, the search is backed by an Amazon web index spanning tens of billions of documents, and query traffic stays within AWS infrastructure. Web Search is available in three AWS Regions; pricing is not disclosed in the post.
Amazon Bedrock adds Z.ai's 753B-parameter GLM 5.3 for eligible enterprise customers
Amazon Bedrock now offers GLM 5.3, Z.ai's 753B-parameter mixture-of-experts model, through managed APIs with cross-Region inference, prompt caching and service tiers. Access is limited to eligible enterprise customers. Pricing and context window were not disclosed in AWS's announcement.
AWS ships aws-ai-ml skill so Kiro, Claude Code and Codex can benchmark SageMaker inference endpoints
Amazon SageMaker AI has released the aws-ai-ml skill, distributed through the Agent Toolkit for AWS. It lets MCP-compatible coding agents such as Kiro, Claude Code and Codex benchmark endpoints, recommend deployment configurations and generate executable SageMaker Python SDK v3 code.
Comments
Loading...