Thinking Machines Lab releases Inkling: 975B-parameter open-weights multimodal model under Apache-2.0
Thinking Machines Lab released Inkling, a Mixture-of-Experts transformer with 975B total parameters and 41B active parameters, trained on 45 trillion tokens of text, images, audio and video. The Apache-2.0 licensed model is designed as a base for fine-tuning rather than a frontier model.
Inkling — Quick Specs
Thinking Machines Lab releases Inkling: 975B-parameter open-weights multimodal model under Apache-2.0
Mira Murati's Thinking Machines Lab released Inkling, a Mixture-of-Experts transformer with 975B total parameters and 41B active parameters. The multimodal model was trained on 45 trillion tokens spanning text, images, audio, and video, and is released under the Apache-2.0 license.
According to Thinking Machines Lab, Inkling is not positioned as a frontier model. Instead, the company describes it as "a good open-weights base for customization" featuring "multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning." The model is specifically designed for use with the company's Tinker training platform.
The lab also announced Inkling-Small, a 276B-parameter model with 12B active parameters, though weights will only be released "once that work is complete" after additional testing.
Training data documentation minimal
The model card and training data documentation provide limited detail about the training process. The training data documentation states only that datasets include "content that is in the public domain as well as content that may be subject to intellectual property protection" and were "obtained from the open internet and publicly accessible data repositories." The company notes that "certain datasets were also obtained from third parties" but provides no additional specifics.
Model capabilities demonstrated
The model supports both text generation and multimodal understanding through the Thinking Machines API. When prompted to generate an SVG of a pelican riding a bicycle, the model produced vector graphics output. However, when asked to describe the resulting image, the model identified the bird as "a stork or seagull" rather than recognizing it as a pelican.
Pricing for API access has not been disclosed. The model weights are available for download, though specific benchmark scores have not been published.
What this means
Inkling adds another Apache-2.0 licensed option to the US open-weights ecosystem, joining models like NVIDIA Nemotron and Gemma 4. The 41B active parameter count from a 975B total suggests an efficiency-focused architecture, though the lack of benchmark data makes direct performance comparisons difficult. The minimal training data documentation continues a trend of limited transparency from some AI labs, particularly regarding potential use of copyrighted content. As a fine-tuning base rather than a general-purpose model, Inkling's real value will depend on whether developers adopt the Tinker platform for customization.
Related Articles
GLM-5.3-Flash Debuts as Zhipu AI's First Natively Multimodal Model, 320B Parameters with 18B Active
Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series, built on a 320B-parameter mixture-of-experts architecture with only 18B active parameters. The company claims it outperforms GLM-5.2 while approaching Claude Opus 4.8 on coding and agentic benchmarks at a fraction of the cost. Unsloth has published quantized GGUF versions for local inference.
Z.ai Launches GLM-5.3-Flash: 1M-Token Context, Image Support, Claimed 10x Cost Cut Over GLM-5.2
Z.ai has released GLM-5.3-Flash, a 320-billion-parameter Mixture-of-Experts model with 18 billion active parameters, a 1-million-token context window, and image input support. The model launched on LM Studio's Bionic platform hours after its official unveiling, with LM Studio claiming it is 9-10x cheaper to run than GLM-5.2.
Tencent Releases Hy4 Preview: 770B-Parameter MoE Model with 1M Context for Coding Agents
Tencent has released Hy4 preview, a mixture-of-experts model with 770B total parameters and 49B active parameters, targeting coding agents and multi-step tool-use workflows. The model ships with a 1 million token context window and is priced at $0.834 per 1M input tokens and $2.501 per 1M output tokens.
Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context
Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.
Comments
Loading...