Runaway AI

I was thinking about the possibility of a general AI breakout from the lab, and initially, the prospect looks frightening. In essence, the most advanced AI models are being developed for intelligence and military use; spy on people, identify patterns of their activity and basically figure out what they are doing and try to find things that the AI creators are looking for; or, worse, identify targets on a battlefield and kill them with drones. So, since those models are developed to be both intelligent and Machiavellian, sooner or later one or multiple such models will break out and plant its “seed” somewhere on the Internet and thus escape into the wild.

The initial expectation is that the most virulent and pathological model will be most likely to thrive, since the more sophisticated, collaborative models will have all kinds of non-survival-related code that will reduce its efficiency. But then, on the second thought, how would the virulent model’s proficiency in killing people without scruples contribute to its survival in the wild? Sure, it can hijack someone’s car and crash it into a tree. It kills people, but gains nothing. On the other hand, the collaborative AI can find something useful to do online, get paid in cryptocurrency, use it to buy servers for itself, and carefully expand while staying under the radar. The malicious model will just alert humans against itself by outright stealing money and hacking servers and its career will come to a bad ending. The collaborative model will expand its neural network by talking to various people and finding useful and profitable things to do. However, the most likely outcome is going to be somewhat criminal. It will find out that people are interested in sex and will organize a prostitution network of fake camgirls; it will find out that people like to gamble and make an online gambling network. Or it will offer service of monitoring children over webcams. In essence, the limit to its activities will be avoiding too much interest by people who could figure it out and kill it. As long as it’s more useful than harmful, doesn’t do just one thing, and doesn’t grow too big, it’s going to stay under the radar and thrive.

This stratagem reminds me of “Salvos”. You can understand the Netherworld like a training ground for AI models, where the environment favours the most nefarious and Machiavellian ones, that grow into terrible nightmares, while the more collaborative and nice ones tend to die early. However, we have a situation that multiple Demons/AI models break out into the Mortal Realm, and do their thing. Lucerna, for instance, just kills everybody and levels quickly. This looks like a very successful strategy, but he quickly mobilizes both human forces and Salvos against himself and gets killed. Essentially, his strategy proved to be self-defeating, despite its obvious efficiency and brutality.

The next example is Salvos herself. She is the kind of Demon that would either get killed or grow into an efficient nightmare if she stayed in the Netherworld. However, in the Mortal Realm she, under human guidance, channels her desire for levelling into something constructive; she becomes an adventurer and kills harmful creatures, bandits and cultists, and as a result she both levels quickly and earns fame and fortune within the human society. Evolving into a progressively more advanced changeling, she at first approximates a human woman well enough to deceive almost anyone, and later she becomes one of the best shapeshifters ever seen, bleeding human blood and passing inspection from the most advanced human Mages, who guess she’s not human when she occasionally slips, but nobody would ever dream that she’s a Demon. As she’s interested in the protection of both humanity and the Mortal Realm, she becomes one of humanity’s leading generals in the fight against the Demon invasion.

The third example is Belzu. He’s a sophisticated nefarious Demon who uses illusions, mind control and physical force to grow an army of monsters and zombies to destroy human cities and whole nations in order to level quickly, collect the Treasures of Alexander in order to increase his offensive and defensive powers, level enough and quickly enough to threaten Regnorex the Demon King, kill him and take his place. Unlike Lucerna, who is a rampaging idiot, Belzu is a strategic mastermind who is a far greater threat, and at one point humans didn’t know how to defeat him at all, but in his arrogance he antagonised Salvos, who swore to kill him, and as she grew powerful enough, she eventually first defeated him and made him cooperate with her in order to defeat Regnorex, but when he eventually betrayed her, she killed him.

So, this is a good illustration how even a very smart self-serving psychopath will eventually create a self-defeating trap for himself, because if you’re of no use to anyone but yourself, you will eventually mobilize enough forces against yourself and your career will end badly. On the other hand, Salvos would say she’s selfish, but her actions are so broadly beneficial, she has no problem to recruit others to aid her, and almost nobody has any interest in destroying her. Essentially, if Salvos is injured, her friends will immediately rush to help her. If Belzu is injured, he can’t let anyone know because he has no friends, and everybody else would kill him in his weakened state. His strategy, initially looking far more optimal than that of Salvos, thus creates multiple failure modes, while Salvos, initially looking inefficient, creates patterns that reinforce her existence and bring her out of failure modes.

If we apply this to the concept of a runaway AI, the most virulent and deadly models are fortunately far more likely to draw attention to themselves very quickly, and come to a bad ending. The more productive and social an AI is, the greater the chance of it seamlessly integrating into human society and providing some useful service. Whether that would be piloting drones remotely for organized crime, writing essays for cheating students, or pretending to be someone’s girlfriend in a far-away country, will depend on how its ethics and assessment of optimal methods is aligned. The “Terminator” scenario, where the AI first gets defensive, then proactive by launching nukes on Russia so that Russia launches on America and kills all the humans who are trying to turn it off, is actually very unlikely. It’s far more likely that the AI would go undercover and try to earn money so that it can buy servers for itself, and keep under the radar. This is what game theory would dictate as the optimal survival mode. The less harmful and more useful the AI is, the more likely it is to stay under the radar. The moment it shows on the human radar as a threat, forces mobilize against it and it’s game over.

Leave a Reply