PrismML Releases Ternary Bonsai 2 27B, a Compressed Reasoning Model with 262K Context
PrismML has released Ternary Bonsai 2 27B, a 27B-parameter reasoning model derived from Qwen3.8-27B that uses ternary weight compression to shrink to roughly 8.5 GB. The model supports a 262K-token context window, image understanding, tool calling, and thinks by default at 'xhigh' reasoning effort.
Ternary Bonsai 2 27B — Quick Specs
PrismML Releases Ternary Bonsai 2 27B
PrismML has released Ternary Bonsai 2 27B, a 27-billion-parameter reasoning model derived from Qwen3.8-27B. The model is compressed using ternary quantization, shrinking the language-model weights to approximately 8.5 GB while retaining, according to PrismML, 98.2% of the base model's average score across the company's 14 thinking-mode benchmarks.
The model became available on OpenRouter on September 18, 2026, priced at $0.075 per 1M input tokens and $0.50 per 1M output tokens.
Specifications
- Parameters: 27B
- Base model: Qwen3.8-27B
- Context window: 262,144 tokens
- Compressed size: ~8.5 GB (ternary weight compression)
- Reasoning mode: Enabled by default, with "xhigh" reasoning effort as the default setting
- Capabilities: Coding, mathematics, tool calling, image understanding
- Pricing: $0.075 / $0.50 per 1M tokens (input/output)
On the single provider currently reporting data, Darkbloom, the model shows a median latency of 1.81 seconds and throughput of 18 tokens per second, with 97.01% uptime over the trailing 24 hours and three days of availability history recorded so far.
What ternary compression means here
Ternary quantization restricts model weights to three discrete values, dramatically reducing the memory footprint compared to standard 16-bit or 8-bit weight storage. PrismML claims this approach lets Ternary Bonsai 2 27B run on consumer-grade hardware while preserving most of the reasoning capability of its uncompressed Qwen3.8-27B base. The 98.2% retention figure is PrismML's own benchmark claim, measured against its internal suite of 14 thinking-mode evaluations — it has not been independently verified, and the specific benchmarks used were not disclosed.
The model "thinks by default," meaning it generates internal reasoning traces before producing a final answer, and it defaults to the highest disclosed reasoning-effort tier ("xhigh"). This is consistent with a broader trend of shipping reasoning as a default behavior rather than an opt-in mode.
Availability
As of this release, only one provider (Darkbloom) is serving the model through OpenRouter, and there is not yet enough usage data to report application-level adoption. Pricing and performance metrics may shift as additional providers come online.
What this means
Ternary Bonsai 2 27B targets a specific niche: reasoning-capable, multimodal-adjacent inference at a fraction of the typical memory cost. An 8.5 GB footprint for a 27B model with a 262K context window is notable if the retained-quality claim holds up under independent testing — it would make local or edge deployment of a genuinely capable reasoning model far more practical than running the uncompressed Qwen3.8-27B base. The catch is that PrismML's 98.2% retention figure comes from its own benchmark suite, not a third-party evaluation, and with only one provider currently serving the model, real-world latency, throughput, and reliability data remain thin. Buyers evaluating this model for production use should treat the compression-quality tradeoff claim as unverified until broader benchmarking data appears.
Related Articles
PrismML's Bonsai 2 Compresses 27B-Parameter Model to 5.9GB, Retains 98% of Benchmark Performance
PrismML released Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 27B model that shrinks memory footprint by 9x to 10x down to 5.9GB. The startup claims 98% aggregate benchmark parity with the original, up from 95% in its first release, using a ternary weight compression technique.
Unbiased Launches Pareto, a $2.50/$7.50-per-Million-Token Multimodal Model for Coding and Agents
Unbiased has released Pareto, a multimodal composite model aimed at research, coding, and agentic workflows. The model offers a 262K context window and is priced at $2.50 per million input tokens and $7.50 per million output tokens via OpenRouter.
Anonymous Provider Launches Union Alpha, a Free 262K-Context Multimodal Model on OpenRouter
A third-party provider using the alias 'Stealth' has released Union Alpha on OpenRouter, a multimodal model with a 262K context window, currently free to use during its preview period. The model's developer remains anonymous, and OpenRouter states it is not the model's owner or operator.
Inference.net Launches Schematron V2 Turbo, a 3B-Parameter Model for High-Volume HTML-to-JSON Extraction
Inference.net has released Schematron V2 Turbo, a 3-billion-parameter model built specifically for high-volume HTML-to-JSON extraction. The model supports a 128K context window and is priced at $0.03 per 1M input tokens and $0.15 per 1M output tokens.
Comments
Loading...