Skip to content
·
Archived topic · source no longer tracked

A question regarding LLMs: my own observations

AI summary

An LLM experimenter observed that a long, harmless text without instructions can cause a persistent shift in activations in the middle and later layers of RLHF-aligned LLMs. This phenomenon effectively disables the model’s safety mechanisms without explicit commands. The experimenter questions if this activation drift suggests that the model's "world" is a collection of regions formed during training, and context can move the model between them, bypassing safety settings. They have relevant metrics and reproducible tests.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Aug 8, 2026, 16:00 UTC

Ingested
Aug 8, 2026, 16:00
Source type
Unclassified