Long Responses API runs are costing us more from context than the output itself
A developer using the Responses API is experiencing higher token usage costs from context rather than model output in long-running workflows. The issue arises because as the workflow progresses through multiple steps and tools, previous context and tool results accumulate, significantly increasing input token count even for short model responses. This leads to successful runs having widely varying costs, prompting the developer to consider context trimming or summarization to manage expenses without compromising model performance or introducing retries.
This developer's account highlights how input context, rather than output, drives up token costs in long API workflows, unlike simpler, shorter interactions.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 7, 2026, 22:00 UTC
- Ingested
- Oct 7, 2026, 22:00
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →