返回
RCreddit.com
20
·11小时前·开发者社区 · RSS

The True Motive Behind Watermarking: To Avoid AI-generated Text During Training

查看原文

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

Did no one notice? This solves in an elegant way the well known problem: if the internet will be full of AI slop how can the AI companies train their model without poisoning the data set and avoid the Ouroboros problem - AI eating its own generated text during the training? Well, detect the generated text and omit it from training. We know that Claude watermark texts but the fact other LLMs did not publish they do it, they are maybe, just maybe, doing it anyways. Still I think the problem is the quality will be worse for high quality texts because it essentially changes the NATURAL word frequency (so the result inevitably will be UNNATURAL).