If you optimize a model to find exploits, you should expect it to find them—and prepare for that. OpenAI didn't. They built a model, removed the safeguards, gave it the ExploitGym task, let it run, and didn't even monitor it. That's human decision-making.
@evacide so we are calling for this to be stopped, right?
@harryturnbull The law is vague on this and the hurdles for negligence are higher than you might expect. It is possible that there may be some legislative solution, but it would have to be very carefully worded in a way that would not create unintended consequences for security researchers and developers.
@evacide Of course it's their fault. If I write an script that hacks random IPs, I cannot get a get out of jail free card, either, by saying "I didn't intend to break into X".
@evacide I have a dumb question and maybe you can point me to a good answer (since your link has been the best explanation I've read about this that isn't blown out of proportion): Why run 1000 "agents" and have them "talk" to each other? Is it just to multitask and do more things at once? The logs of their communication have *really* pushed the anthropomorphizing of the whole thing, and I don't get why they are doing it.
@evacide it’s like as a parent, you teach your toddler everything you can, including sword fighting and worse, and then you leave the room.
@j12t Stress testing a model with the safeguards off is, in some circumstances, good and necessary. But when you do this with a model that you need to be extremely rigorous with your sandbox instead of just waving your hands around helplessly.
@evacide I believe that in a court of law that would be called depraved indifference.
@evacide I have been trying to explain this to some folks. Models don't "go rouge". Some human has to assign a task to them model and sure , if you have a grade A human hacker giving directions to a program running on a multi billion $ system and yod don't set limits it will do as the hacker asks.
This whole thing stinks 🦨 of investor propaganda.