model releaseMicrosoft

Microsoft releases FrogNano-4B, an Apache 2.0 coding agent trained with RL on 1,500 synthetic tasks

TL;DR

Microsoft has released FrogNano-4B-2609, a repository-level coding agent derived from Qwen3.5-4B and published under Apache 2.0 with open weights. Microsoft says it was post-trained only with reinforcement learning on about 1,500 synthetic software-engineering tasks, with no stronger-model trajectories. It is evaluated at roughly 131K tokens of context.

3 min read
0

Microsoft has released FrogNano-4B-2609, an open-weight coding agent derived from Qwen/Qwen3.5-4B, under the Apache License 2.0. The model card lists a release date of September 22, 2026. Microsoft says the model was post-trained exclusively with reinforcement learning (RL) on approximately 1,500 synthetic software-engineering tasks.

Key specifications

  • Base model: Qwen/Qwen3.5-4B
  • Architecture: Dense, 32-layer hybrid Gated DeltaNet / gated-attention causal language model
  • Parameters: The model card lists a range of 500M-5B. The "4B" in the name matches the base model.
  • Context: Approximately 131K tokens in the evaluated configuration, including generated output
  • Output limit: Up to 8,192 generated tokens per assistant turn in the validated configuration
  • Inputs: Text only (instructions, source code, tool outputs)
  • Training dates: June 2026 to August 2026
  • License: Apache 2.0, with upstream Qwen copyright and attribution notices retained
  • Pricing: No hosted API pricing disclosed. Weights are downloadable from Hugging Face.
  • Knowledge cutoff: Not disclosed

How it was trained

FrogNano's RL environments were generated, validated and calibrated against the evolving policy using a tool called TaskPilot. Rewards come from executable tests over complete multi-turn coding trajectories. They cover functional correctness, reliable tool use and preservation of existing behavior.

The model runs through Leaf, a lightweight five-tool harness. According to Microsoft, the agent-specific post-training uses no solution trajectories, actions, reasoning traces or patch targets from stronger models. That sets it apart from distillation-based approaches. This is Microsoft's description of its own training process and has not been independently verified.

What it does

The model generates structured Leaf tool calls to search and inspect files, edit code, run shell commands and run tests. The goal is to produce multi-file candidate patches from natural-language issue descriptions. Leaf executes the calls in an isolated repository environment. FrogNano does not deploy changes itself.

Limitations Microsoft discloses

  • Performance is sensitive to the Leaf harness and to test quality.
  • Training data are Python-heavy and primarily English. Non-English use and very different repositories are unevaluated.
  • Vision components inherited from Qwen3.5-4B were not post-trained or evaluated, and image and video input is not supported.
  • No dedicated safety-preference, refusal or adversarial datasets were used. Microsoft says the model should not be considered independently safety-aligned for unrestricted autonomous deployment.
  • Generated patches may be incorrect or insecure even when they pass tests, and require human review.

Benchmarks

The model card material reviewed does not include benchmark scores. It says main evaluations use a combined interaction context of about 131K tokens and points to a technical report. Numbers will be added once the report's results are confirmed.

Distribution

Weights, configuration files, tokenizer assets and the model card are on Hugging Face (microsoft/FrogNano-4B-2609). The technical report and any released harness or evaluation code are on GitHub (microsoft/FrogNano).

What this means

FrogNano is a test of whether a 4B-class model can handle long-horizon repository work through RL on a small set of well-calibrated synthetic environments, rather than by imitating larger models. The 1,500-task scale is small, so the technical report's benchmark results will determine whether the approach holds up.

The Apache 2.0 license and small size make it practical for local and research use. The narrow scope limits it: it is a sandbox-bound, Python-focused agent, not a general assistant. Its dependence on the Leaf harness also means results from other scaffolds may not match Microsoft's.

Related Articles

model release

Musubi Releases Open-Weights PolicyLM-1.7B, a Sub-50ms Content Moderation Model

Musubi released PolicyLM-1.7B on Tuesday, a 1.7-billion-parameter open-weights decision model built for real-time content moderation. The company claims it applies plain-English policies to messages in under 50 milliseconds and needs no retraining when policies change.

model release

Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0

Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says the quantized model needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Google also launched a Mac app, AI Edge Foresight, to demonstrate it.

product update

Microsoft's Copilot gets access to local Windows files and OS-level actions under 'Hybrid Intelligence'

Microsoft announced an upgrade to Copilot at its Windows and Surface event that gives the assistant access to local files and the ability to take actions across Windows. The company calls the underlying approach "Hybrid Intelligence," which combines local and cloud AI models. Pricing, model details, and availability were not disclosed in the available reporting.

product update

Microsoft will turn Windows Search into a Copilot-connected command interface this fall

Microsoft announced a redesigned Windows Search menu that accepts short typed commands to change system settings and can converse with the new Copilot app without launching it. The company says the update arrives this fall. Model, pricing and availability details were not disclosed.

Comments

Loading...