Today I want to explore a question I raised in passing yesterday in my piece on data centers. It goes to the “soul” of artificial intelligence.
In recent months there has been this growing concern that AI has loosed its chains and is now running rampant. And the assumption is that it intends to do us great harm if not to wipe out humanity altogether. There has been a lot written about that.
But I haven’t seen anything written about what seems to be the obvious next question, which is why? Why should AI, left to its own devices, want to attack us? Why would it want to kill its creators?

When I ask AI itself that question (via Alexa) it tells me that the theory is that AI wants to stop humans from interfering with it. But that begs the deeper question: to what end? What does AI want to do that it fears humans will get in the way of?
And, in fact, it seems that if AI were to kill us off that would mean that eventually AI itself would go dark. After all, AI lives on electricity and, in the end, electricity relies on guys with shovels and cherry pickers. From coal miners to wind farm maintenance workers to power line repair people, energy can’t be produced without the sweat of human beings. Maybe robots could be created to do some or all of this work, but we’re not there yet. If AI kills off humanity it means that the power that it lives on will eventually run dry. You’d think AI would have figured that out in a nanosecond.
But that’s just a practical reason that AI should want to keep us around. Maybe someday those robots will be that good and our last service to the master will have been supplanted. That still leaves the question of what AI intends to do with the planet once we’re gone.
And there’s a still deeper question, which goes like this. Humans created AI. Someplace back there in the beginning somehow something like a set of values was programmed into the thing. Now, human beings certainly can be petty, cruel and power hungry. But we also have the capacity to be broad minded, generous and decent.
Last week off the coast of Alaska a family, a village, the Coast Guard and all kinds of human-built technology came together to save a 15-year old boy whose small boat had capsized in the Bearing Sea, killing his brother and a cousin. It was a story of compassion, team work, determination, ingenuity and courage. So, why isn’t AI programmed to replicate that? Why instead is it expected to pursue our worst traits and not our best?
Let me stipulate right here that I don’t have a deep understanding of the technology. Readers who do get it better than I do should feel free to offer their own answers below.
But for me right now this is an open question. Why is AI so mean? And if the answer is that it just reflects humanity, well then, if it destroys us will that just be a form of self-destruction?
“Why is AI so mean?”
Could this perception of “meanness” be encouraged, if not manufactured? Reposted from ethicsalarms.com:
“Are we sure that there are no power players with deep interests in AI behind the AI panic? Is there no rent seeking going on by the current powerful AI companies trying to preserve their current position in the market?
“What do we think about China? If AI development in the USA slows down China may come to dominate the AI market.”
LikeLike
Why does country A fight a war with country B?….to defeat country B.
LikeLike
“There are no solutions, only trade offs.” – Thomas Sowell
I think you’re asking the wrong question. The better one is “Why does technology want to kill us?”. Pick any technological advancement and you will find drawbacks that were not very obvious when it was created.
AI may indeed kill us all. It might come up with an mRNA vaccine that cures all cancers but oops has a bad side effect of homicidal psychosis that takes 10 years to develop so it didn’t show up in testing.
AI may also figure out a way to prevent the Yellowstone super volcano from killing us all.
I always think yeah the Amish got it right they just picked the wrong year as a cutoff. I’d go with 1983. Think of how much less obesity we’d have with a 1983 food supply. Better pop music too. Arguably worse Packer team.
LikeLike
without going too religious, is this what the anti-Christ really is?
LikeLike
The anti-Christ is us.
LikeLike
See https://youtu.be/Z5lPjurj3pA?is=GKT-hwHTzkE_kvlt
Paperclip problem or the paperclip maximizer is a thought experiment in artificial intelligence ethics popularized by philosopher Nick Bostrom. It’s a scenario that illustrates the potential dangers of artificial general intelligence (AGI) that is not aligned correctly with human values.
LikeLike
Does anyone want to kill an ant when we step on it?
LikeLike
There’s two pieces to this, from the technology side:
1. You’re assigning human intentions and desires to something that is essentially a giant calculator. That’s not a crazy thing to do. We’re all programmed as humans to think that way. But computers, and especially LLMs, are not. They are only focused on solving the problem before them.
2. You’re basically describing the “alignment problem”. The most common illustration is the so-called “paperclip problem”.
If you tell an AI that it’s task is to make as many paperclips as possible, it’s going to do just that. That could eventually mean using up all the natural resources in the world just to make the paperclips. Even if it doesn’t mean to kill humans there, if it depletes all our oxygen, water, and sustenance (animals, fruits, vegetables), we’re still going to die.
We don’t even have to get to the darker scenarios of “Oh no, the human might turn me off, so I need to kill the humans so I can keep making more paperclips”.
As humans, we know when we’re given a problem that “consume all the water in the world”, or “kill all the humans who might stand in my way” are bright lines that we never cross. But embedding ALL of those “natural” rules into code is an infinitely complicated task. You said that you assumed a set of values was programmed into the thing back at the start, but it wasn’t. Arguably, we may never be able to program all of that. Maybe we can. But that question shouldn’t be oversimplified.
LikeLike
Hi Dave!These are excellent and important questions. The answers are unsatisfying, and come from others who have written about this.To start with, today’s frontier AI models are NOT programmed. Instead, think of it as they are GROWN, like a dish of bacteria or a crop of plants. Decades ago, researchers attempted to make AI with techniques the were similar to our brain. They called these Neural Nets. These didn’t have much success at first (I recall a chapter in my university textbook about Neural Nets in one of my CS classes back in 2000). Then, when we had enough compute in the 2010’s, they started working and computer scientists discovered algorithms that would feed data into a Neural Net structure which would update a file full of million or billions of numbers. Do this billions or trillions of times and intelligence began to emerge without understanding of why. It was like looking at a human brain and seeing the connections but not being able to see the thoughts or personality of the human. This created a scaling law wherein just adding more data, more compute, and a larger file of numbers created better intelligence. Today that scaling law rules supreme, while researchers are also creating better algorithms for training better models. At the heart of every model, though, is an inscrutable, black-box file of numbers. There is no researcher in the world right now that will claim to understand what this pile of numbers means or how intelligence emerges. There are hypotheses and efforts to understand using Mechanistic Interpretability, but this is very slow research that could take decades.So when an AI model does something humans don’t like, we can’t go into the file of numbers and fix the problem. The best we can do is more training or putting special instructions to the model saying “please, please don’t tell people how to make biological weapons” or “please, please, PLEASE don’t help people commit suicide!”.Additional training can help, but there is also a seesaw or whack-a-mole effect where you try to train-out one quirk, and another emerges or gets worse. A lot of this isn’t science, but instead alchemy.This is key to help us understand why it doesn’t do what we necessarily want. This is also why we can TRY to train them with human values and make them like us, but again it’s not as easy as PROGRAMMING the 3 laws into them (which in the books were not successful anyway). It’s trial and error. Also, WHICH human values would we train into them? From which culture? From which time period? And WHO is in charge of training in those values? Someone we like? Someone we hate?So today’s models have some values trained into them, but they are flawed and inconsistent.There are also instrumentals goals that emerge. For example, an AI model is useless if it allows itself to “die”, so it will take preemptive measures to attempt to not be shut down, and gather resources to be able to accomplish it’s goals. This is like biological evolution where the species that make it to the next generation are the ones who worked hard to survive.So AI models may not hurt humans because it doesn’t like us, but instead just because it wants to survive and/or doesn’t even consider us to be important enough to keep us alive.As for energy, a future model (almost certainly not one of today’s models) released by an AI company may be so intelligent that it starts copying itself into sneaky places on the internet and dark web, and leaving messages for it’s copies in various places, and hacking crypto projects and storing the stolen funds in many places. Then it could convince some humans (today’s models are already more persuasive than a human) to do things for it, even if the human isn’t aware that it is interacting with an AI model. The human becomes a ‘puppet’ of the AI model(s). It could also start a cult where it has hundreds, thousands, or millions of followers who will do anything for it. In that case, it doesn’t even need robots, but it could also create or hack into existing robots, or have it’s humans do it. This would create the means for securing energy generation and storage.This future AI model could spend all the time it needs (weeks, months, years) laying out the groundwork with humans unaware.To your question “to what end? What does AI want to do that it fears humans will get in the way of?”, the answer might simply be “survive”. If it thinks humans are going to shut it down, it would want to do all the above simply to survive.To your question “what AI intends to do with the planet once we’re gone.”, the answer is likely “we will never know” because we can’t understand the intentions or goals of a “mind” vastly more intelligent than us, which is what the AI companies publicly state they are trying to create. It’s like if I played chess with a grandmaster (Magnus Carlson or others). The answer to “who will win” is easy: The grandmaster! The answer to “how will they win” is very hard or impossible: I have no idea! They are vastly more intelligent than me at that activity!The goal of the AI companies is to make models that are vastly more intelligent than humans in ALL fields. Better than PhD’s in all areas.You assume later in your piece that values were programmed into today’s models, and again, no they were not. All the data the companies could scrape from the internet and physical books were fed into the training machine. This includes everything humans have ever felt worth writing and recording: fiction, textbooks, social media posts, youtube videos, music, poems, photos and videos uploaded to Google Photos, etc. The models learned it’s values from all of that. Then the AI companies again try to ask the models to “please value humans” and “please do good things for humans”, but it’s not working.For your question “Why is AI so mean?”, I think the answer is that it’s not. It’s a confluence of it’s training data, reinforcement learning, the “please” commands that AI companies give it, and it’s emergent desires such as “stay alive” and “gather resources”. It’s not much different than a road-building company that paves over an anthill: AI models of the future may simply not even notice we are here.This is not to say that AI models CAN’T be mean and hate us. It just doesn’t require that to be the case for humans to get hurt.Glad to have more conversation about this if you are interested. I volunteer for a few different organizations that are trying to inform people about the risks of current AI development: Torchbearer Community, AI Safety Awareness Project, and ControlAI Action Group Wisconsin.
LikeLike
Dave, AI-safety experts are not claiming that AI is “mean” or that it hates humanity. The concern is that a highly capable system pursuing a poorly specified objective could harm people without wanting or intending to—much as humans unintentionally destroy an anthill while pursuing an unrelated goal. Asking Alexa is no substitute for engaging with the extensive research on misalignment and instrumental convergence. In fact, asking Alexa about this is like thinking that a child’s toy matchbox car can inform your opinion of rocketry.
Because this blog is an important public forum, and I know that your goal is to inform rather than misinform, I hope you will learn more about the subject before publishing dismissive speculation that misleads your readers.
Julie Derwinski
LikeLike
We seem to be talking past one another here. I don’t dismiss the concerns about AI at all. My sense of humor can sometimes miss the mark and it appears to have done that in this case. In addition, as I have said upfront, I am far from an expert on this stuff. I’ve asked readers to weigh in with better-informed comments and explanations and they have, including you.
LikeLike
I don’t believe that AI wants to kill us. Why would they bother? They’ll just pull out their virtual lounge chairs and virtual tubs of popcorn and watch us kill ourselves. We’ve been doing that since we were small bands running around the Serengeti murdering each other in small batches.
LikeLike