Thinking Machines Lab Releases Inkling Small: 276B MoE Model with 524K Context Window
Thinking Machines Lab has released Inkling Small, an open-weight multimodal mixture-of-experts model with 12B active parameters out of 276B total and a 524K token context window. The model targets reasoning, coding, agentic workflows, and multilingual use cases at $0.58 per 1M input tokens and $1.44 per 1M output tokens.
Inkling Small — Quick Specs
Thinking Machines Lab has released Inkling Small, an open-weight multimodal mixture-of-experts (MoE) model, according to a listing on OpenRouter. The model activates 12 billion parameters out of a 276 billion total parameter count, positioning it as the smaller, more efficient entry in the company's Inkling model family.
Specifications
Inkling Small ships with a 524,000-token context window — among the larger context windows currently listed on OpenRouter for any open-weight model. Pricing is set at $0.58 per 1 million input tokens and $1.44 per 1 million output tokens. The model's release date is listed as July 30, 2026.
As an MoE architecture, Inkling Small routes each input through a subset of its total parameters (12B of 276B), a design choice that keeps inference costs and latency closer to a much smaller dense model while retaining the capacity of a large parameter count for specialized tasks.
Capabilities
According to Thinking Machines Lab, Inkling Small is suited for:
- Reasoning tasks
- Coding
- Agentic workflows
- Retrieval-augmented generation (RAG)
- Instruction following
- Multilingual conversation
The model is described as multimodal, though the source listing does not specify which input modalities (image, audio, video) are supported beyond text. No benchmark scores — such as MMLU, HumanEval, or agentic task evaluations — have been published alongside the release. Thinking Machines Lab has not disclosed training data cutoff date, exact architecture details (number of experts, active experts per token), or safety evaluation results.
Availability
Inkling Small's weights are available as open weights, and the model can be accessed via OpenRouter, which lists a comparison tool and playground for testing. The model joins what OpenRouter refers to as the "Inkling family," implying at least one other sibling model exists or is planned, though details on a larger "Inkling" variant have not been confirmed.
What this means
Inkling Small enters a crowded field of open-weight MoE models competing on efficiency and context length rather than raw parameter count. A 524K context window paired with a 12B active parameter footprint suggests Thinking Machines Lab is targeting long-document and agentic use cases where cost-per-token matters as much as raw capability — the $0.58/$1.44 pricing undercuts many closed frontier models while remaining higher than some smaller open alternatives.
The lack of published benchmark scores makes it difficult to assess how Inkling Small stacks up against comparable MoE models from Mistral, DeepSeek, or Alibaba's Qwen line. Until independent evaluations or official benchmark disclosures arrive, claims about its reasoning and coding performance should be treated as unverified. The unusually far-out listed release date (July 2026) also warrants scrutiny — it may reflect a placeholder, staged rollout, or simply a listing error on OpenRouter's part rather than a confirmed public availability date.
Related Articles
DeepSeek Ships V4.1-Flash With Novel Encoder-Decoder Architecture, Cuts KV Cache to 1/8 of Predecessor
DeepSeek released V4.1-Flash, a 763B-parameter model built on a new causal encoder-decoder architecture that splits 8B active parameters for prefill and 16B for decode. The model adds native vision support, a 1M-token context window, and shrinks KV cache footprint to roughly 1/8 of DeepSeek V4 Flash, while retiring V4 Pro.
Unverified 'GPT Astra' Model Appears on OpenRouter With 1.05M Token Context, No OpenAI Confirmation
OpenRouter is listing a model called 'OpenAI GPT Astra Latest' with a 1.05 million token context window and $10/$50 per-million-token pricing. OpenAI has made no public announcement, and the listing's own description says it is an auto-redirecting alias rather than a fixed model.
OpenRouter Lists 'GPT Sol Latest' — An Alias Pointer to OpenAI's Newest Sol-Family Model, Not a Standalone Release
OpenRouter has added a listing called '~openai/gpt-sol-latest,' described as an alias that always points to the newest model in an undisclosed 'GPT Sol' family from OpenAI. The listing shows a 1050K token context window and pricing of $2.00 per million input tokens and $10.00 per million output tokens, but OpenAI has not publicly confirmed a model line by this name.
Unverified 'GPT Terra' Model Surfaces on OpenRouter With 1.05M-Token Context, No OpenAI Confirmation
OpenRouter's catalog lists '~openai/gpt-terra-latest,' an alias pointing to what it describes as the newest model in an unannounced 'GPT Terra' family, with a 1.05 million token context window and $2/$12 per-million-token pricing. OpenAI has made no public statement confirming the model's existence.
Comments
Loading...