Writer introduces new AI model and upgraded harness to contain token costs
热度趋势
百分比基于当前可用热度信号,而非评论数或独立用户人数。
这条记录涉及编程工具或代码能力更新,适合开发者评估工作流变化和可复用价值。
为营销人员提供人工智能工具的Writer公司推出了一款名为Palmyra X6的全新旗舰AI模型。该模型是Z.ai开源模型GLM-5.2的后训练变体,旨在为用户显著降低部署成本。Writer公司估计,结合其对基础设施的改进,Palmyra X6模型可以将基本任务的成本降低多达50%,以应对AI行业日益增长的部署成本问题。
Across the AI industry, users are becoming more conscious of just how expensive their deployments can be —and feeling a new urgency to cut costs. But while open source models offer significantly lower per-token costs, it can be difficult to find the right model for a given job.
On Thursday, Writer, which offers AI tools and agents for marketers, launched a new flagship model called Palmyra X6, aimed at solving that problem for its users. Built as a post-training variation on Z.ai’s open source model GLM-5.2, Writer says the new system should provide deployment-ready capabilities at a much lower price. The company estimates the new model, combined with changes to the companies harness infrastructure, will cut costs for its customers by as much as 50% for basic tasks.
Together with the new model, the company also released significant upgrades to its standard agentic harness. Both features will be available to Writer clients starting Thursday.
“I think the enterprise is absolutely sick of chasing the next benchmark,” CEO May Habib told TechCrunch. “They want flattening cost, and it seems like nobody can deliver that.”
The new approach puts particular emphasis on complex, multi-step tasks, executed faster and with fewer tokens. And Writer sees harness optimization as a crucial lever toward making that happen.
A recent paper from Writer researchers lends credence to this approach, testing small changes in harness efficiency across multiple different models. The research found that, in many cases, changes in the harness were a more reliable way to reduce costs than model choice, with costs falling an average of 40% across their testing.
“The harness is the one component whose efficiency multiplies across every model an organization runs—present and future,” the researchers wrote.
For Writer’s clients, the experience is still model-agnostic: Palmyra X6 will sit alongside other Writer models or outside models imported through Azure or Amazon Bedrock. But Habib also sees the push to cut costs as driving a broader distrust toward major AI labs, which have a financial incentive to drive up token use.
“The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs,” Habib told TechCrunch, adding that the AI labs “don’t deeply understand right how to help an enterprise get benefit from AI.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489.
View Bio