On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs
Research into data poisoning reveals that even a small amount of poisoned data can significantly impact large AI models. For instance, just 250 poisoned documents, making up 0.00016% of training tokens, reliably implanted a backdoor in a 13B-parameter model. Similar success was seen in fine-tuning experiments, where 50–90 poisoned examples achieved over 80% attack success. This raises questions about intellectual property and credit in a future where AI synthesizes human ideas into breakthroughs.
Why this oneThis report uniquely highlights how a minuscule 0.00016% of poisoned data can reliably implant backdoors in large AI models, unlike previous discussions focusing on general data integrity.
Time & source
- Ingested
- 09/09, 05:00 UTC+0
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →