reinforcement learning
10 articles tagged with reinforcement learning
Microsoft releases FrogNano-4B, an Apache 2.0 coding agent trained with RL on 1,500 synthetic tasks
Microsoft has released FrogNano-4B-2609, a repository-level coding agent derived from Qwen3.5-4B and published under Apache 2.0 with open weights. Microsoft says it was post-trained only with reinforcement learning on about 1,500 synthetic software-engineering tasks, with no stronger-model trajectories. It is evaluated at roughly 131K tokens of context.
Xiaomi's MiMo-V2.6-Pro Tops Open Model Rankings at $0.13 Per Task, But Anthropic Says It Used Claude to Get There
Xiaomi's new MiMo-V2.6-Pro, a 1.02 trillion parameter mixture-of-experts model, now leads open-model rankings with a 46 on Artificial Analysis's Intelligence Index while costing roughly $0.13 per task. Anthropic simultaneously accuses Xiaomi of funneling user conversations through Claude to train the model, part of a broader pattern the company calls illegal distillation.
Xiaomi's MiMo-V2.6-Pro Becomes Top Open-Weights Model, Trained for $3M According to Xiaomi
Xiaomi released MiMo-V2.6-Pro, a 1.02T-parameter mixture-of-experts model with 42B active parameters, which debuted as the top-scoring open-weights model on Artificial Analysis' Intelligence Index (46). The company claims the model's RL training run cost roughly $2.6M and completed in 130 hours.
Xiaomi Releases MiMo-V2.6-Pro-RL, a 1.02T-Parameter Omnimodal Model with 1M-Token Context
Xiaomi's MiMo team has released MiMo-V2.6-Pro-RL, a 1.02-trillion-parameter sparse mixture-of-experts model with 42B active parameters, 1M-token context, and native text/image/video/audio processing. The model was trained via a single mixed reinforcement learning run spanning coding, agentic, visual, and cybersecurity tasks, with benchmark scores that Xiaomi claims approach or match Claude Opus 5 and GPT-5.6 on several agentic and coding tests.
Google DeepMind Extends Game AI Research to EVE Online, Building on SIMA 2 Agent
Google DeepMind published a retrospective on 15 years of game-based AI research, tracing a line from 2015's Atari-playing DQN through AlphaGo, AlphaZero, MuZero, and AlphaStar to its current generalist agent, SIMA 2. The post also details a new research partnership with Fenris Creations, the studio behind EVE Online, to study continual learning, memory, and long-horizon planning in a persistent multiplayer universe.
Z.ai CEO Jie Tang: Parameter Count Alone No Longer Predicts Model Capability
Z.ai CEO Jie Tang says raw parameter counts no longer predict model quality, pointing to GLM 5.3's benchmark gains that came entirely from reinforcement learning on synthetic long-horizon environments rather than scaling weights. The claim lands alongside a broader reshuffling of agent and legal benchmark leaderboards.
AWS Details Custom Reward Function Design for Multi-Turn RL on Amazon Nova Forge
AWS published a technical guide on designing custom composite reward functions for multi-turn reinforcement fine-tuning (RFT) of Amazon Nova models via Nova Forge's Bring Your Own Orchestration (BYOO) capability. The post covers GRPO-based reward scoring, combining outcome rewards, behavioral rewards, and penalties, plus a serverless multi-turn RL option now generally available.
OpenAI's GPT-5.6 Sol Adds Five Reasoning Effort Settings, Follows DeepSeep-R1 RLVR Training Method
OpenAI released GPT-5.6 Sol, a new reasoning model family that comes in three sizes with roughly five to six reasoning-effort settings each. The release follows the DeepSeek-R1 methodology of using reinforcement learning with verifiable rewards (RLVR), nearly two years after OpenAI's original o1 model popularized LLM-based reasoning.
AWS Ships Multi-Turn RL Infrastructure for Amazon Nova on SageMaker HyperPod
AWS has released infrastructure for deploying multi-turn reinforcement learning to train Amazon Nova models on SageMaker HyperPod. The system requires a minimum of 10 ml.p5.48xlarge instances and costs approximately $786-$1,180 per hour when running.
Mistral AI launches Forge, enterprise platform for training custom models on proprietary data
Mistral AI has launched Forge, a platform for enterprises to train custom AI models on proprietary data including codebases, compliance policies, and operational records. Early partners include ASML, DSO National Laboratories Singapore, Ericsson, European Space Agency, and HTX Singapore.