reinforcement learning

6 articles tagged with reinforcement learning

August 21, 2026

Google DeepMind Extends Game AI Research to EVE Online, Building on SIMA 2 Agent

Google DeepMind published a retrospective on 15 years of game-based AI research, tracing a line from 2015's Atari-playing DQN through AlphaGo, AlphaZero, MuZero, and AlphaStar to its current generalist agent, SIMA 2. The post also details a new research partnership with Fenris Creations, the studio behind EVE Online, to study continual learning, memory, and long-horizon planning in a persistent multiplayer universe.

August 20, 2026
analysis

Z.ai CEO Jie Tang: Parameter Count Alone No Longer Predicts Model Capability

Z.ai CEO Jie Tang says raw parameter counts no longer predict model quality, pointing to GLM 5.3's benchmark gains that came entirely from reinforcement learning on synthetic long-horizon environments rather than scaling weights. The claim lands alongside a broader reshuffling of agent and legal benchmark leaderboards.

August 14, 2026
product updateAmazon Web Services

AWS Details Custom Reward Function Design for Multi-Turn RL on Amazon Nova Forge

AWS published a technical guide on designing custom composite reward functions for multi-turn reinforcement fine-tuning (RFT) of Amazon Nova models via Nova Forge's Bring Your Own Orchestration (BYOO) capability. The post covers GRPO-based reward scoring, combining outcome rewards, behavioral rewards, and penalties, plus a serverless multi-turn RL option now generally available.

July 18, 2026
model releaseOpenAI+1

OpenAI's GPT-5.6 Sol Adds Five Reasoning Effort Settings, Follows DeepSeep-R1 RLVR Training Method

OpenAI released GPT-5.6 Sol, a new reasoning model family that comes in three sizes with roughly five to six reasoning-effort settings each. The release follows the DeepSeek-R1 methodology of using reinforcement learning with verifiable rewards (RLVR), nearly two years after OpenAI's original o1 model popularized LLM-based reasoning.

July 6, 2026
product updateAmazon Web Services

AWS Ships Multi-Turn RL Infrastructure for Amazon Nova on SageMaker HyperPod

AWS has released infrastructure for deploying multi-turn reinforcement learning to train Amazon Nova models on SageMaker HyperPod. The system requires a minimum of 10 ml.p5.48xlarge instances and costs approximately $786-$1,180 per hour when running.

May 28, 2026
product updateMistral AI

Mistral AI launches Forge, enterprise platform for training custom models on proprietary data

Mistral AI has launched Forge, a platform for enterprises to train custom AI models on proprietary data including codebases, compliance policies, and operational records. Early partners include ASML, DSO National Laboratories Singapore, Ericsson, European Space Agency, and HTX Singapore.