RCreddit.com
which model is good for detecting deflection?
Model release
- Published
- 09/05, 04:22
- Ingested
- 09/05, 20:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
I want the answers generated by frontier LLMs or base model LLM answers to be reviewed by some uncensored or abliterated small model.
The job is this model (preferably small model) is just to detect deflection in the answers.
The problem I am facing is uncensored SLM usually agrees on everything we give input. So the generated answer is also input for it and system prompt is input too.