Thinking Machines Lab Releases Inkling Small: 276B MoE Model with 524K Context Window
Thinking Machines Lab has released Inkling Small, an open-weight multimodal mixture-of-experts model with 12B active parameters out of 276B total and a 524K token context window. The model targets reasoning, coding, agentic workflows, and multilingual use cases at $0.58 per 1M input tokens and $1.44 per 1M output tokens.
Inkling Small — Quick Specs
Thinking Machines Lab has released Inkling Small, an open-weight multimodal mixture-of-experts (MoE) model, according to a listing on OpenRouter. The model activates 12 billion parameters out of a 276 billion total parameter count, positioning it as the smaller, more efficient entry in the company's Inkling model family.
Specifications
Inkling Small ships with a 524,000-token context window — among the larger context windows currently listed on OpenRouter for any open-weight model. Pricing is set at $0.58 per 1 million input tokens and $1.44 per 1 million output tokens. The model's release date is listed as July 30, 2026.
As an MoE architecture, Inkling Small routes each input through a subset of its total parameters (12B of 276B), a design choice that keeps inference costs and latency closer to a much smaller dense model while retaining the capacity of a large parameter count for specialized tasks.
Capabilities
According to Thinking Machines Lab, Inkling Small is suited for:
- Reasoning tasks
- Coding
- Agentic workflows
- Retrieval-augmented generation (RAG)
- Instruction following
- Multilingual conversation
The model is described as multimodal, though the source listing does not specify which input modalities (image, audio, video) are supported beyond text. No benchmark scores — such as MMLU, HumanEval, or agentic task evaluations — have been published alongside the release. Thinking Machines Lab has not disclosed training data cutoff date, exact architecture details (number of experts, active experts per token), or safety evaluation results.
Availability
Inkling Small's weights are available as open weights, and the model can be accessed via OpenRouter, which lists a comparison tool and playground for testing. The model joins what OpenRouter refers to as the "Inkling family," implying at least one other sibling model exists or is planned, though details on a larger "Inkling" variant have not been confirmed.
What this means
Inkling Small enters a crowded field of open-weight MoE models competing on efficiency and context length rather than raw parameter count. A 524K context window paired with a 12B active parameter footprint suggests Thinking Machines Lab is targeting long-document and agentic use cases where cost-per-token matters as much as raw capability — the $0.58/$1.44 pricing undercuts many closed frontier models while remaining higher than some smaller open alternatives.
The lack of published benchmark scores makes it difficult to assess how Inkling Small stacks up against comparable MoE models from Mistral, DeepSeek, or Alibaba's Qwen line. Until independent evaluations or official benchmark disclosures arrive, claims about its reasoning and coding performance should be treated as unverified. The unusually far-out listed release date (July 2026) also warrants scrutiny — it may reflect a placeholder, staged rollout, or simply a listing error on OpenRouter's part rather than a confirmed public availability date.
Related Articles
Alibaba Releases Qwen3.8 Max (0902), a 2.4-Trillion-Parameter MoE Model With 1M-Token Context
Alibaba's Qwen team released Qwen3.8 Max (0902), a 2.4-trillion-parameter mixture-of-experts model with a 1M-token context window that accepts text, image, and video input. The snapshot is post-trained for coding, agentic workflows, and long-horizon task execution, priced at $2/$6 per 1M input/output tokens.
InclusionAI Releases Ling 3.0 Flash Fin, a Finance-Focused MoE Model with 5.1B Active Parameters
InclusionAI has released Ling 3.0 Flash Fin, a finance-specialized mixture-of-experts model built on Ling 3.0 Flash. The model activates 5.1B of its 124B total parameters and targets long-horizon investment planning tasks while retaining general reasoning, coding, and math capabilities.
Meta Releases Muse Spark 1.3, a Free Multimodal Reasoning Model with 1M-Token Context
Meta has released Muse Spark 1.3, a multimodal reasoning model with a 1M-token context window, listed as free on OpenRouter. The model targets long-running agentic, multi-agent, and coding workflows, though audio input support remains incomplete.
Google Lists Gemini 3.8 Flash on OpenRouter With 1M-Token Context, September 2026 Release Date
Google's Gemini 3.8 Flash has surfaced on OpenRouter with a 1-million-token context window and discounted pricing of $0.75 per 1M input tokens and $3.75 per 1M output tokens. Google has not issued a separate public announcement, and the listed release date of September 2, 2026 is unusually far out, leaving key details unconfirmed.
Comments
Loading...