Long Responses API runs are costing us more from context than the output itself
A developer using the Responses API is experiencing higher token usage costs from context rather than model output in long-running workflows. The issue arises because as the workflow progresses through multiple steps and tools, previous context and tool results accumulate, significantly increasing input token count even for short model responses. This leads to successful runs having widely varying costs, prompting the developer to consider context trimming or summarization to manage expenses without compromising model performance or introducing retries.
This developer's account highlights how input context, rather than output, drives up token costs in long API workflows, unlike simpler, shorter interactions.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年10月7日 22:00 UTC
- 收录
- 2026年10月7日 22:00
- 来源类型
- 开发者社区
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
本站未收录正文。
前往源站阅读 →