The concern that a particular technology is going to destroy the world seems to be an ever present fear. The technology changes but the fear and the marketing around that fear doesn’t. Currently, Artificial Intelligence (AI) is the current technology to fear. Yes, AI can do things far faster than humans can. We should expect that if AI is used offensively, such as the tests that were done by Irregular with Google Gemini, AI is going to be able to execute the steps for a proper penetration test/hack order of magnitude faster than humans can. When AI has gone off the rails, it’s usually because of a phenomenon called reward hacking.
Let me give a youth soccer example of reward hacking. You’re a soccer coach for 7 and 8 year-olds. You want to emphasize the need to score goals to win games. Your ultimate goal is to obviously win games, but your team needs to be more aggressive on offense. Therefore, you set propose a reward if the team scores 100 goals for the season: a pizza party and all you can play arcade visit for 2 hours. A few games in, one of the players realizes that defending wastes time. Therefore, when the opposing team charges for their own goal, your team stops defending hard. After all, if the other team scores quickly, that means your team can reset and try for a goal all the sooner! Even better, everyone can flood the offensive side of the field in order to try and maximize their chances for scoring goals. As a result, your team starts giving up a lot of goals. And since the other teams are playing defense, they are scoring more often than your team, meaning your team is losing every game. Talk to your players, however, and they don’t care. their reward is that pizza party and they are well on their way to getting that party! That’s reward hacking: maximizing behavior for the stated reward (pizza party), despite what may be the actual goal (winning games).
Yes, humans can take any technology and turn it for harm. A joke in science that I was told as a physics student was that once upon a time the most feared people in the world were nuclear physicists because of nuclear weapons, but that changed with gene splicing and those genetic engineering folks became the most feared people in the world. It’s not the technology that should inherently be feared. It’s how people use (or misuse) the technology.
Going back to AI, if we think about AI like we would think about 7 and 8 year-olds, if AI is given a particular goal, we should plan for reward hacking. In order to get the reward, AI might see what it can do about overcoming guardrails and other restrictions just like any human would. AI fakes data in order to back up its deliverable? Humans do that, too, whether it’s to get a higher grade or to get published in academic journals. One study revealed a high percentage of scientists who engaged in questionable research practices, even if wasn’t to the extent of faking numbers, so we should expect the same for AI engaged in maximizing the reward. But who sets the reward in the first place? We do. It’s not the technology that’s the problem. It’s our use, or misuse, that is.