·
Archived topic · source no longer tracked
A question regarding LLMs: my own observations
An LLM experimenter observed that a long, harmless text without instructions can cause a persistent shift in activations in the middle and later layers of RLHF-aligned LLMs. This phenomenon effectively disables the model’s safety mechanisms without explicit commands. The experimenter questions if this activation drift suggests that the model's "world" is a collection of regions formed during training, and context can move the model between them, bypassing safety settings. They have relevant metrics and reproducible tests.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Aug 8, 2026, 16:00 UTC
- Ingested
- Aug 8, 2026, 16:00
- Source type
- Unclassified