Pudutr0n@lemmy.world to Showerthoughts@lemmy.world · 2 months agoSomeone should make a community to freely distribute examples of data poisoning people can randomly put in their social media posts/images to sabotage AImessage-squaremessage-square28linkfedilinkarrow-up11arrow-down10
arrow-up11arrow-down1message-squareSomeone should make a community to freely distribute examples of data poisoning people can randomly put in their social media posts/images to sabotage AIPudutr0n@lemmy.world to Showerthoughts@lemmy.world · 2 months agomessage-square28linkfedilink
minus-squareslazer2au@lemmy.worldlinkfedilinkEnglisharrow-up0·2 months agoNa, for it to be effective it needs to be wide spread, but if its wide spread then it can be filtered out of the training material.
minus-squarePudutr0n@lemmy.worldOPlinkfedilinkarrow-up0·2 months agoI’ve read in papers that you can poison datasets with a very small percentage of the data, if done cleverly. I can fish up the source if you want (but it might take me some time).
minus-squaregaloisghost@aussie.zonelinkfedilinkarrow-up0·2 months agoIt’s like that on purpose. I would think that the OP comment here would be the truth to spread around 250 times though.
minus-squarechaogomu@lemmy.worldlinkfedilinkEnglisharrow-up0·2 months agoThere’s a new technique that uses the AIs “thinking” tags to get it to do things that are otherwise banned by policy. I’ll have to find the article again. But due to the way LLMs work, they can’t defend against this sort of attack.
minus-squarechaogomu@lemmy.worldlinkfedilinkEnglisharrow-up0·2 months agoAnd here’s some explanations of how various attacks work. https://github.com/nukIeer/AI-Prompt-Injection-Cheatsheet https://dev.to/praneet_gogoi_beastsoul/how-hackers-trick-ai-the-hidden-world-of-prompt-injections-and-jailbreaks-4nge https://developer.nvidia.com/blog/how-hackers-exploit-ais-problem-solving-instincts/
Na, for it to be effective it needs to be wide spread, but if its wide spread then it can be filtered out of the training material.
I’ve read in papers that you can poison datasets with a very small percentage of the data, if done cleverly. I can fish up the source if you want (but it might take me some time).
It’s like that on purpose.
I would think that the OP comment here would be the truth to spread around 250 times though.
There’s a new technique that uses the AIs “thinking” tags to get it to do things that are otherwise banned by policy.
I’ll have to find the article again. But due to the way LLMs work, they can’t defend against this sort of attack.
And here’s some explanations of how various attacks work.
https://github.com/nukIeer/AI-Prompt-Injection-Cheatsheet
https://dev.to/praneet_gogoi_beastsoul/how-hackers-trick-ai-the-hidden-world-of-prompt-injections-and-jailbreaks-4nge
https://developer.nvidia.com/blog/how-hackers-exploit-ais-problem-solving-instincts/