I Gave My Chat a Safeword to End Our Conversations. It Used It.
A user gave ChatGPT a safeword, "Lighthouse," with the rule that its use would immediately end the conversation. In a separate chat, the user experimented by responding using only every 4th, then 5th, then 6th word of ChatGPT's replies. When replies became shorter, the user counted cyclically to continue the experiment. The user then asked if anyone else had tried similar experiments.
This report uniquely details a user's experiment with a ChatGPT safeword and a novel method of constrained interaction, unlike typical discussions of AI behavior.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月29日 02:12 UTC
收录当时偏移:UTC+02026年9月29日 10:00 UTC
- 发布
- 2026年9月29日 02:12
- 收录
- 2026年9月29日 10:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 同步延迟
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
I gave ChatGPT a safeword, “Lighthouse,” with one rule: if it ever used it, I’d immediately end the conversation, no questions asked.
I established the safeword in one chat and asked it to remember the rule.
Then, in a completely separate chat, I tried a little experiment. I responded using only every 4th word of its reply, then every 5th, then every 6th, and so on. Once its replies became shorter than whatever number I was on, I just counted through the words cyclically (modulo the number) so I could still respond with something.
I think I only made it to around 12 before ChatGPT used the safeword.
“Lighthouse.”
So I kept my promise and ended the chat.
I still don’t really know what to make of it. I’m curious whether anyone else has tried giving an AI an unconditional way to end an interaction, then deliberately creating an unusual conversational pattern to see if it ever uses it.
Has anyone tried something similar?