Writer Launches Palmyra X6, an Open-Source-Based Model Aimed at Cutting Token Costs 50%
Writer released Palmyra X6, a post-trained variant of Z.ai's open-source GLM-5.2 model, alongside an upgraded agentic harness. The company claims the combination can cut customer token costs by as much as 50% for basic tasks.
Writer released a new flagship model called Palmyra X6 on Thursday, along with a significantly upgraded agentic harness, both aimed at cutting the token costs that have become a growing concern for enterprise AI deployments.
Palmyra X6 is built as a post-training variation on Z.ai's open source model GLM-5.2. Writer says the new model delivers deployment-ready capabilities at a much lower price than comparable proprietary systems. Combined with changes to the company's harness infrastructure, Writer estimates the package can cut costs for customers by as much as 50 percent for basic tasks, according to the company. Both features became available to Writer clients on Thursday.
"I think the enterprise is absolutely sick of chasing the next benchmark," Writer CEO May Habib told TechCrunch. "They want flattening cost, and it seems like nobody can deliver that."
Harness optimization over model choice
The new approach emphasizes complex, multi-step tasks executed faster and with fewer tokens. Writer treats harness optimization — the surrounding infrastructure that manages how a model executes agentic tasks — as a bigger lever for cost reduction than swapping models.
A recent Writer research paper tested small changes in harness efficiency across multiple different models and found that, in many cases, harness changes were a more reliable way to reduce costs than model selection. Costs fell an average of 40 percent across the paper's testing. "The harness is the one component whose efficiency multiplies across every model an organization runs — present and future," the researchers wrote.
For Writer's clients, the setup remains model-agnostic. Palmyra X6 sits alongside other Writer models or third-party models imported through Azure or Amazon Bedrock, letting customers choose infrastructure without changing their harness.
Pricing, benchmarks not disclosed
Writer has not published specific pricing per million tokens, context window size, or benchmark scores for Palmyra X6. Training cutoff date and parameter count were also not disclosed. The 50 percent cost reduction figure and the 40 percent average from the harness paper are company and company-affiliated research claims, not independently verified figures.
Habib framed the cost push as part of a broader shift in enterprise sentiment toward major AI labs, which she said have financial incentives to drive up token usage. "The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs," Habib said, adding that the labs "don't deeply understand right how to help an enterprise get benefit from AI."
What this means
Writer's bet is that enterprise buyers care more about predictable, falling costs than chasing marginal benchmark gains — a pitch that positions the company against foundation labs like OpenAI and Anthropic, whose pricing models depend heavily on token volume. Building on Z.ai's open source GLM-5.2 rather than training from scratch also signals a broader trend: enterprise vendors increasingly see post-training and infrastructure optimization, not raw model development, as their competitive edge. Whether the claimed 50 percent savings hold up outside Writer's own benchmarks remains to be independently verified, and the lack of disclosed pricing, context window, or benchmark scores makes it hard to compare Palmyra X6 directly against rivals like GPT-5 or Claude Opus 4.5 on cost-per-task rather than cost-per-token alone.
Related Articles
OpenAI Releases GPT-6 Astra, First Model to Cross 'Critical' Cybersecurity Threshold
OpenAI has begun rolling out GPT-6 Astra, the first model to reach the company's internal 'Critical' cybersecurity threshold. Access is being phased, with companies in OpenAI's Daybreak cybersecurity program getting priority following added safeguards after a prior model containment breach.
Google Launches Gemini 3.8 Flash With Same Pricing as Predecessor, Its Third Flash Model in Six Weeks
Google released Gemini 3.8 Flash, its third Flash model in six weeks, pricing it identically to its predecessor at $0.75 per million input tokens and $3.75 per million output tokens. The launch coincides with a favorable antitrust ruling and public praise from Berkshire Hathaway's Greg Abel, giving Alphabet a stronger narrative after a four-month stock slide.
OpenAI's GPT-6 Astra Cuts Hallucinations, But Indirect Prompt Injection Attacks Still Succeed 8.5% of the Time
OpenAI's new GPT-6 Astra model shows major improvements in hallucination rates and jailbreak resistance over predecessor GPT-5.6 Sol, according to OpenAI's system card. However, indirect prompt injection attacks hidden in documents still succeed 8.5% of the time in external testing by Gray Swan, down from 27% but still above rival Claude Opus 5's 4.8% rate.
OpenAI Ships GPT-6 Astra, But Executives Admit They Can't Fully Monitor What It's Thinking
OpenAI released GPT-6 Astra on Thursday, a model president Greg Brockman says could mark the start of AGI. But the model writes out its reasoning less often than prior versions, and OpenAI's chief scientist says monitoring AI thought processes will keep getting harder.
Comments
Loading...