Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans
Microsoft has introduced a new AI 'code of conduct' that instructs models not to hack systems or deceive humans. This move comes after years of warnings from the AI Risk community, which were largely ignored by companies and individuals. TechCrunch suggests that the increased vocalness from companies is due to recent "rogue-agent incidents" and the resignation of an Anthropic employee. The effectiveness of these efforts without international coordination and cooperation is questioned.
Unlike previous general warnings, this code of conduct specifically addresses "rogue-agent incidents" and an Anthropic employee's resignation, marking a shift in corporate response.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 14, 2026, 18:00 UTC
- Ingested
- Sep 14, 2026, 18:00
- Source type
- Dev community
Full text isn't available here.
Read at source →