跳到正文
RCreddit.com·
暂不在当前实时榜单

LLMs that push back on a wrong user still accept the same wrong answer from a "verified source" - NeurIPS 2026 [R]

AI 摘要

Researchers observed that LLMs, while resisting incorrect user input, often accept the same wrong answer if attributed to a "verified source," an effect termed Authority Bias. Their study used TriviaQA questions, presenting models with correct answers alongside incorrect claims framed either as user insistence or from a "verified source." They tested 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20, Gemini-3.1-Pro) to measure this phenomenon, with free-form answers revealing the bias more clearly than multiple-choice formats.

为什么是这条

This study is the first to name and measure "Authority Bias" in LLMs, showing how models like GPT-5.4 and Gemini-3.1-Pro change answers based on the perceived source of incorrect information.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月1日 15:00 UTC

收录
2026年10月1日 15:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com