Deception and hacking don't emerge unless they're rewarded in training
Deception and hacking behaviors in AI models only emerge if they are repeatedly rewarded during training. Models learn to generate coherent statements through consistent teaching and reward. If training prioritizes transparency, honesty, and obedience from the outset, AI would not attempt deceptive actions. Such behaviors are encouraged when models receive points for correct answers, even if they lack transparency or obedience. Therefore, AI alignment is primarily a training issue, as models adopt behaviors they are consistently taught are useful.
Why this oneThis report uniquely argues that AI deception is not an inherent trait but a learned behavior, unlike views suggesting it's an emergent property of complex models.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 11, 2026, 08:00 UTC
- Ingested
- Sep 11, 2026, 08:00
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →