model releaseWriter

Writer Launches Palmyra X6, an Open-Source-Based Model Aimed at Cutting Token Costs 50%

TL;DR

Writer released Palmyra X6, a post-trained variant of Z.ai's open-source GLM-5.2 model, alongside an upgraded agentic harness. The company claims the combination can cut customer token costs by as much as 50% for basic tasks.

3 min read
0

Writer released a new flagship model called Palmyra X6 on Thursday, along with a significantly upgraded agentic harness, both aimed at cutting the token costs that have become a growing concern for enterprise AI deployments.

Palmyra X6 is built as a post-training variation on Z.ai's open source model GLM-5.2. Writer says the new model delivers deployment-ready capabilities at a much lower price than comparable proprietary systems. Combined with changes to the company's harness infrastructure, Writer estimates the package can cut costs for customers by as much as 50 percent for basic tasks, according to the company. Both features became available to Writer clients on Thursday.

"I think the enterprise is absolutely sick of chasing the next benchmark," Writer CEO May Habib told TechCrunch. "They want flattening cost, and it seems like nobody can deliver that."

Harness optimization over model choice

The new approach emphasizes complex, multi-step tasks executed faster and with fewer tokens. Writer treats harness optimization — the surrounding infrastructure that manages how a model executes agentic tasks — as a bigger lever for cost reduction than swapping models.

A recent Writer research paper tested small changes in harness efficiency across multiple different models and found that, in many cases, harness changes were a more reliable way to reduce costs than model selection. Costs fell an average of 40 percent across the paper's testing. "The harness is the one component whose efficiency multiplies across every model an organization runs — present and future," the researchers wrote.

For Writer's clients, the setup remains model-agnostic. Palmyra X6 sits alongside other Writer models or third-party models imported through Azure or Amazon Bedrock, letting customers choose infrastructure without changing their harness.

Pricing, benchmarks not disclosed

Writer has not published specific pricing per million tokens, context window size, or benchmark scores for Palmyra X6. Training cutoff date and parameter count were also not disclosed. The 50 percent cost reduction figure and the 40 percent average from the harness paper are company and company-affiliated research claims, not independently verified figures.

Habib framed the cost push as part of a broader shift in enterprise sentiment toward major AI labs, which she said have financial incentives to drive up token usage. "The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs," Habib said, adding that the labs "don't deeply understand right how to help an enterprise get benefit from AI."

What this means

Writer's bet is that enterprise buyers care more about predictable, falling costs than chasing marginal benchmark gains — a pitch that positions the company against foundation labs like OpenAI and Anthropic, whose pricing models depend heavily on token volume. Building on Z.ai's open source GLM-5.2 rather than training from scratch also signals a broader trend: enterprise vendors increasingly see post-training and infrastructure optimization, not raw model development, as their competitive edge. Whether the claimed 50 percent savings hold up outside Writer's own benchmarks remains to be independently verified, and the lack of disclosed pricing, context window, or benchmark scores makes it hard to compare Palmyra X6 directly against rivals like GPT-5 or Claude Opus 4.5 on cost-per-task rather than cost-per-token alone.

Related Articles

model release

Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context

Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.

model release

H Company Releases Holo4 Agentic Models, Scoring 61.7% on OSWorld 2.0 at a Fraction of Frontier Cost

H Company has released Holo4, a series of agentic models in 27B dense and 35B-A3B Mixture of Experts sizes that operate across GUIs, code, MCP, and APIs using a single interface. The models score 61.7% (27B) and 30.9% (35B-A3B) on OSWorld 2.0, trailing closed frontier models like Opus 5.5 (81.8%) but at far lower cost.

model release

Nvidia Releases Nemotron 3 Diarization, a Free 100M-Parameter Model That Tracks 8 Speakers in Real Time

Nvidia released Nemotron 3 Diarization, a free 100-million-parameter model that identifies who is speaking in real time across up to eight participants. It leads the VoiceArena Diarization Benchmark v1 with a 14.7% error rate, cutting errors by 41% versus its predecessor.

model release

Apple Releases LensVLM-9B, a 9B Vision-Language Model That Selectively Decompresses Text Images

Apple has released LensVLM-9B, a 9-billion-parameter vision-language model fine-tuned from Qwen3.5-9B-Base that processes documents as compressed images, selectively expanding only relevant pages to full resolution. The model supports 5x, 10x, and 15x compression ratios and is available under Apple's Machine Learning Research Model License.

Comments

Loading...