返回
RCreddit.com

which model is good for detecting deflection?

模型发布
时间与来源
发布
09/05 04:22
收录
09/05 20:00
来源类型
开发者社区
档位
社区
信源状态
正常
档位是按信源手工设定的编辑判断,不是逐条打分。
AI 摘要

一位用户正在寻找一个小型、未经审查的模型,用于检测由前沿或基础大型语言模型(LLM)生成的答案中的“偏离”。用户希望这个模型(最好是小型模型)能够审查这些答案,并专门检测其中的偏离。然而,用户面临的挑战是,未经审查的小型语言模型(SLM)通常会同意所有输入,包括生成的答案和系统提示,这使得它们难以有效地识别偏离。该任务的目的是让这个小型模型专门检测答案中的偏离。

I want the answers generated by frontier LLMs or base model LLM answers to be reviewed by some uncensored or abliterated small model.

The job is this model (preferably small model) is just to detect deflection in the answers.

The problem I am facing is uncensored SLM usually agrees on everything we give input. So the generated answer is also input for it and system prompt is input too.