Writer targets token costs with Palmyra X6 and harness overhaul

LLMsAI Agents
Editorial image for Writer targets token costs with Palmyra X6 and harness overhaul

The Core · TL;DR

  • Writer launched Palmyra X6, a new flagship model built as a post-trained variant of Z.ai's open source GLM-5.2, alongside a major agentic harness upgrade on August 13, 2026.
  • Writer estimates the model plus harness changes together could cut costs for basic tasks by up to 50%, though harness efficiency alone drove an average 40% cost reduction across multiple tested models.
  • CEO May Habib says enterprise buyers now prioritize flattening AI costs over chasing benchmark performance.
  • Palmyra X6 can run alongside other Writer models or third-party models imported via Azure or Amazon Bedrock.

Writer isn't chasing another leaderboard spot with its newest model. It's chasing your invoice.

The enterprise AI company released Palmyra X6, its new flagship model, on Thursday alongside a substantial upgrade to its agentic harness, the infrastructure layer that governs how its models plan, call tools, and execute multi-step tasks. Both went live for customers the same day.

Rather than pitching Palmyra X6 as a benchmark leader, Writer is framing the release around a more mundane but pressing enterprise concern: the runaway token costs that come with scaling agentic workloads. CEO May Habib said enterprise buyers are now more focused on flattening those costs than on chasing raw performance gains.

"Enterprise buyers are focused on flattening costs, not chasing benchmarks," Habib said of the shift driving the release.

Writer says the combination of the new model and the retooled harness could cut costs for basic tasks by as much as 50%. That figure comes from the company's own estimate, and it hasn't been independently verified against competing agent stacks.

The harness improvements alone appear to carry most of the weight. Writer published internal research showing that efficiency changes to the harness, independent of which underlying model it ran, reduced costs by an average of 40% across multiple models tested. That suggests the savings are less about Palmyra X6's raw efficiency and more about smarter orchestration: fewer wasted tool calls, tighter context management, and reduced redundant token consumption during agent execution.

Notably, Palmyra X6 isn't built from scratch. It's a post-trained variant of GLM-5.2, the open source model from Chinese AI lab Z.ai, a detail Writer disclosed rather than obscured. Building on an open weights base lets Writer skip the enormous pretraining costs of foundation model development and instead focus its resources on fine-tuning and the harness layer, where it argues the real enterprise value now sits.

Writer is also positioning the release as agnostic rather than closed. Palmyra X6 can run alongside the company's other proprietary models or alongside third-party models imported via Azure or Amazon Bedrock, letting enterprise customers mix models within the same harness rather than committing exclusively to Writer's stack.

The bet underlying the launch is that as agentic AI moves from pilot projects into production, the cost of running these systems at scale, not their benchmark scores, will decide which vendors enterprises keep using.

Original reporting and research used to synthesize this article.

  1. 1Writer introduces new AI model and upgraded harness to contain token coststechcrunch.com
WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram

More from Tools

View all in Tools