Skip to content
RCreddit.com·
Not on the current live radar

Long Responses API runs are costing us more from context than the output itself

AI summary

A developer using the Responses API is experiencing higher token usage costs from context rather than model output in long-running workflows. The issue arises because as the workflow progresses through multiple steps and tools, previous context and tool results accumulate, significantly increasing input token count even for short model responses. This leads to successful runs having widely varying costs, prompting the developer to consider context trimming or summarization to manage expenses without compromising model performance or introducing retries.

Why this one

This developer's account highlights how input context, rather than output, drives up token costs in long API workflows, unlike simpler, shorter interactions.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 7, 2026, 22:00 UTC

Ingested
Oct 7, 2026, 22:00
Source type
Dev community

Discussion trend

→ Steady
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Full text isn't available here.

Read at source →
Source·reddit.com