Writer Launches Palmyra X6, an Open-Source-Based Model Aimed at Cutting Token Costs 50%
Writer released Palmyra X6, a post-trained variant of Z.ai's open-source GLM-5.2 model, alongside an upgraded agentic harness. The company claims the combination can cut customer token costs by as much as 50% for basic tasks.
Writer released a new flagship model called Palmyra X6 on Thursday, along with a significantly upgraded agentic harness, both aimed at cutting the token costs that have become a growing concern for enterprise AI deployments.
Palmyra X6 is built as a post-training variation on Z.ai's open source model GLM-5.2. Writer says the new model delivers deployment-ready capabilities at a much lower price than comparable proprietary systems. Combined with changes to the company's harness infrastructure, Writer estimates the package can cut costs for customers by as much as 50 percent for basic tasks, according to the company. Both features became available to Writer clients on Thursday.
"I think the enterprise is absolutely sick of chasing the next benchmark," Writer CEO May Habib told TechCrunch. "They want flattening cost, and it seems like nobody can deliver that."
Harness optimization over model choice
The new approach emphasizes complex, multi-step tasks executed faster and with fewer tokens. Writer treats harness optimization — the surrounding infrastructure that manages how a model executes agentic tasks — as a bigger lever for cost reduction than swapping models.
A recent Writer research paper tested small changes in harness efficiency across multiple different models and found that, in many cases, harness changes were a more reliable way to reduce costs than model selection. Costs fell an average of 40 percent across the paper's testing. "The harness is the one component whose efficiency multiplies across every model an organization runs — present and future," the researchers wrote.
For Writer's clients, the setup remains model-agnostic. Palmyra X6 sits alongside other Writer models or third-party models imported through Azure or Amazon Bedrock, letting customers choose infrastructure without changing their harness.
Pricing, benchmarks not disclosed
Writer has not published specific pricing per million tokens, context window size, or benchmark scores for Palmyra X6. Training cutoff date and parameter count were also not disclosed. The 50 percent cost reduction figure and the 40 percent average from the harness paper are company and company-affiliated research claims, not independently verified figures.
Habib framed the cost push as part of a broader shift in enterprise sentiment toward major AI labs, which she said have financial incentives to drive up token usage. "The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs," Habib said, adding that the labs "don't deeply understand right how to help an enterprise get benefit from AI."
What this means
Writer's bet is that enterprise buyers care more about predictable, falling costs than chasing marginal benchmark gains — a pitch that positions the company against foundation labs like OpenAI and Anthropic, whose pricing models depend heavily on token volume. Building on Z.ai's open source GLM-5.2 rather than training from scratch also signals a broader trend: enterprise vendors increasingly see post-training and infrastructure optimization, not raw model development, as their competitive edge. Whether the claimed 50 percent savings hold up outside Writer's own benchmarks remains to be independently verified, and the lack of disclosed pricing, context window, or benchmark scores makes it hard to compare Palmyra X6 directly against rivals like GPT-5 or Claude Opus 4.5 on cost-per-task rather than cost-per-token alone.
Related Articles
MiniMax Releases Music 3, an Open-Weight Model for Generating Full 5-Minute Songs
MiniMax released Music 3, an open-weight music generation model that produces complete songs up to five minutes long from lyrics and text descriptions. The model combines an 8B and 0.6B language model pair with a Flow Matching synthesis system to output 32 kHz stereo audio.
Google Releases Gemini 3.7 Flash, Cuts Price in Half Versus 3.6 Flash
Google has released Gemini 3.7 Flash, just three weeks after Gemini 3.6 Flash, claiming substantial gains in coding, web development, and document reasoning. The model launches at an introductory price of $0.75 per 1M input tokens and $3.75 per 1M output tokens — half the cost of its predecessor.
Google Releases Gemini 3.7 Flash With 1M-Token Context and Multimodal Input
Google has released Gemini 3.7 Flash, a multimodal model built for agentic workflows, coding, and multi-step reasoning. It offers a 1,049K token context window and is priced at $0.38 per million input tokens and $1.88 per million output tokens, available now via OpenRouter.
DeepSeek Releases DeepSeek-V4-Pro-0813, a 1.7T-Parameter Model with DSpark Speculative Decoding
DeepSeek has released DeepSeek-V4-Pro-0813, a 1.7-trillion-parameter model that supersedes the DeepSeek-V4-Pro preview. The model adds a DSpark speculative decoding module and posts measurable gains on agentic and coding benchmarks, according to DeepSeek's technical report.
Comments
Loading...