model releaseNex Agi

Nex AGI releases Nex-N2-Mini: open-source agentic MoE model with 262K context window

TL;DR

Nex AGI has released Nex-N2-Mini, an open-source agentic mixture-of-experts model with a 262K-token context window. The model accepts text and image inputs and is priced at $0.025 per 1M input tokens and $0.10 per 1M output tokens.

2 min read
0

Nex-N2-Mini — Quick Specs

Context window262K tokens
Input$0.025/1M tokens
Output$0.1/1M tokens

Nex AGI releases Nex-N2-Mini: open-source agentic MoE model with 262K context window

Nex AGI has released Nex-N2-Mini, an open-source agentic mixture-of-experts model with a 262K-token context window, now available through OpenRouter.

Specifications

The model is priced at $0.025 per 1M input tokens and $0.10 per 1M output tokens. It accepts both text and image inputs and produces text output, making it a multimodal model.

Nex-N2-Mini is described as the "smaller sibling in the Nex-N2 series," indicating a larger model exists or is planned in the same family.

Capabilities

According to Nex AGI, the model is built for:

  • Coding tasks
  • Tool use
  • Deep research
  • Long-horizon agentic workflows
  • Native reasoning support

The 262K-token context window places it among models with extended context capabilities, though below the largest available windows from providers like Anthropic (Claude with 200K standard) and Google (Gemini 1.5 Pro with 2M tokens).

Architecture

The model uses a mixture-of-experts (MoE) architecture, a design pattern that activates only subsets of parameters for each inference, potentially improving efficiency compared to dense models of similar capability.

Nex AGI has made the model weights available as open-source, allowing researchers and developers to download and deploy the model independently.

Availability

The model is currently hosted exclusively through OpenRouter, which forwards requests directly to the provider without routing decisions. OpenRouter reports that prompt caching can reduce effective costs by 60-80% below list prices for workloads with repeated context.

The official release date is listed as June 24, 2026, though this appears to be an error in the source data given the current date.

What this means

Nex-N2-Mini enters a competitive space for coding and agentic models, with its 262K context window and MoE architecture potentially offering cost-performance advantages for long-context tasks. The open-source release allows developers to self-host, though the lack of disclosed benchmark scores makes direct capability comparisons difficult. The pricing sits in the mid-range: cheaper than frontier models like GPT-4 but more expensive than some open models. Whether the model delivers on its claimed agentic capabilities will depend on independent testing and real-world usage data.

Related Articles

model release

Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter

Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.

model release

Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta

Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.

model release

inclusionAI releases Ling 3.1 Flash: 560B MoE, 25B active, 262K context, free on OpenRouter

inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters and a 262K-token context window. It is listed as free on OpenRouter through NovitaAI. No benchmark scores have been published on the listing.

model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

Comments

Loading...