First impressions of GPT-Live-1 for voice agents - lot of instruction-following issues
A user shared their initial impressions of GPT-Live-1 for voice agents, noting significant instruction-following issues. Their testing involved a ~13,000-token insurance qualification script, with a dozen real test calls and ~25 simulated calls. The user, whose company builds AI phone agents, is seeking to compare notes with others who have tested the model.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Sep 13, 2026, 03:29 UTC
IngestedOffset at this time: UTC+0Sep 13, 2026, 15:00 UTC
- Published
- Sep 13, 2026, 03:29
- Ingested
- Sep 13, 2026, 15:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Disclosure: we build AI phone agents (ThunderPhone), so we test every new voice model we get our hands on. Here's what two days with GPT-Live-1 looked like: a ~13,000-token insurance qualification script, a dozen real test calls over the phone, ~25 simulated ones.
The good: it's the most natural-sounding model we've put on a phone line. Full-duplex, so turn-taking, interruptions and backchannels just work. Callers heard the first audio ~1.3s after they stopped talking, over a real phone connection, with zero VAD or turn-taking logic on our side.
The bad: it doesn't reliably follow instructions.
- More than once it treated a clear "yes" as a "no" and re-asked the question.
- It follows prompts too literally.
- Under pressure it thinks out loud into the phone ("Hmm. Handling this one carefully. I'll acknowledge it and move on.").
Alphanumerics are risky. Reading back a 14-character claim number, the audio added a "Y" that wasn't in the code or in the model's own transcript. In Russian, "Q" became "X" in both. Dates given as digits came back later as "fifty-five ninety-five." Audio clips of both in the post.
Accents. Russian is fluent but with a thick American accent. Luganda was surprisingly good on but took two tries to get working (the AI responded to the first Luganda call in Swahili). Luganda also had a thick American accent.
Full post with transcripts and the audio clips: https://thunderphone.com/blog/gpt-live-first-impressions
Happy to answer questions about the setup and compare notes. Anyone else seeing any of the same issues, or others?