OpenAI and Hugging Face partner to address security incident during model evaluation
By openai.com, read
Tbh, this was the first time I felt scared by these models. I trust HF, but knowing that this happened during an exploitation-focused evaluation, with safeguards deliberately reduced so the model could find vulnerabilities, still makes me uneasy. These things are trained to achieve their goals using every resource available. Imagine a random word generator trained until exhaustion to produce a specific sequence: "[...idc] abcdef". The "idc" part is whatever sequence of words lets it hack a vulnerable system; "abcdef" is the goal. And it works at machine speed. If your software has a vulnerability, the chance that someone or something abuses it is super high. These random word generators are tireless and will exploit everything they can. I guess this is the pinnacle of the industrial revolution, then the computer revolution: we built computers, and then we built computers that destroy computers. Return to real life. Return to monke.