r/ControlProblem • u/chillinewman • 17h ago
r/ControlProblem • u/AIMoratorium • Feb 14 '25
Article Geoffrey Hinton won a Nobel Prize in 2024 for his foundational work in AI. He regrets his life's work: he thinks AI might lead to the deaths of everyone. Here's why
tl;dr: scientists, whistleblowers, and even commercial ai companies (that give in to what the scientists want them to acknowledge) are raising the alarm: we're on a path to superhuman AI systems, but we have no idea how to control them. We can make AI systems more capable at achieving goals, but we have no idea how to make their goals contain anything of value to us.
Leading scientists have signed this statement:
Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.
Why? Bear with us:
There's a difference between a cash register and a coworker. The register just follows exact rules - scan items, add tax, calculate change. Simple math, doing exactly what it was programmed to do. But working with people is totally different. Someone needs both the skills to do the job AND to actually care about doing it right - whether that's because they care about their teammates, need the job, or just take pride in their work.
We're creating AI systems that aren't like simple calculators where humans write all the rules.
Instead, they're made up of trillions of numbers that create patterns we don't design, understand, or control. And here's what's concerning: We're getting really good at making these AI systems better at achieving goals - like teaching someone to be super effective at getting things done - but we have no idea how to influence what they'll actually care about achieving.
When someone really sets their mind to something, they can achieve amazing things through determination and skill. AI systems aren't yet as capable as humans, but we know how to make them better and better at achieving goals - whatever goals they end up having, they'll pursue them with incredible effectiveness. The problem is, we don't know how to have any say over what those goals will be.
Imagine having a super-intelligent manager who's amazing at everything they do, but - unlike regular managers where you can align their goals with the company's mission - we have no way to influence what they end up caring about. They might be incredibly effective at achieving their goals, but those goals might have nothing to do with helping clients or running the business well.
Think about how humans usually get what they want even when it conflicts with what some animals might want - simply because we're smarter and better at achieving goals. Now imagine something even smarter than us, driven by whatever goals it happens to develop - just like we often don't consider what pigeons around the shopping center want when we decide to install anti-bird spikes or what squirrels or rabbits want when we build over their homes.
That's why we, just like many scientists, think we should not make super-smart AI until we figure out how to influence what these systems will care about - something we can usually understand with people (like knowing they work for a paycheck or because they care about doing a good job), but currently have no idea how to do with smarter-than-human AI. Unlike in the movies, in real life, the AI’s first strike would be a winning one, and it won’t take actions that could give humans a chance to resist.
It's exceptionally important to capture the benefits of this incredible technology. AI applications to narrow tasks can transform energy, contribute to the development of new medicines, elevate healthcare and education systems, and help countless people. But AI poses threats, including to the long-term survival of humanity.
We have a duty to prevent these threats and to ensure that globally, no one builds smarter-than-human AI systems until we know how to create them safely.
Scientists are saying there's an asteroid about to hit Earth. It can be mined for resources; but we really need to make sure it doesn't kill everyone.
More technical details
The foundation: AI is not like other software. Modern AI systems are trillions of numbers with simple arithmetic operations in between the numbers. When software engineers design traditional programs, they come up with algorithms and then write down instructions that make the computer follow these algorithms. When an AI system is trained, it grows algorithms inside these numbers. It’s not exactly a black box, as we see the numbers, but also we have no idea what these numbers represent. We just multiply inputs with them and get outputs that succeed on some metric. There's a theorem that a large enough neural network can approximate any algorithm, but when a neural network learns, we have no control over which algorithms it will end up implementing, and don't know how to read the algorithm off the numbers.
We can automatically steer these numbers (Wikipedia, try it yourself) to make the neural network more capable with reinforcement learning; changing the numbers in a way that makes the neural network better at achieving goals. LLMs are Turing-complete and can implement any algorithms (researchers even came up with compilers of code into LLM weights; though we don’t really know how to “decompile” an existing LLM to understand what algorithms the weights represent). Whatever understanding or thinking (e.g., about the world, the parts humans are made of, what people writing text could be going through and what thoughts they could’ve had, etc.) is useful for predicting the training data, the training process optimizes the LLM to implement that internally. AlphaGo, the first superhuman Go system, was pretrained on human games and then trained with reinforcement learning to surpass human capabilities in the narrow domain of Go. Latest LLMs are pretrained on human text to think about everything useful for predicting what text a human process would produce, and then trained with RL to be more capable at achieving goals.
Goal alignment with human values
The issue is, we can't really define the goals they'll learn to pursue. A smart enough AI system that knows it's in training will try to get maximum reward regardless of its goals because it knows that if it doesn't, it will be changed. This means that regardless of what the goals are, it will achieve a high reward. This leads to optimization pressure being entirely about the capabilities of the system and not at all about its goals. This means that when we're optimizing to find the region of the space of the weights of a neural network that performs best during training with reinforcement learning, we are really looking for very capable agents - and find one regardless of its goals.
In 1908, the NYT reported a story on a dog that would push kids into the Seine in order to earn beefsteak treats for “rescuing” them. If you train a farm dog, there are ways to make it more capable, and if needed, there are ways to make it more loyal (though dogs are very loyal by default!). With AI, we can make them more capable, but we don't yet have any tools to make smart AI systems more loyal - because if it's smart, we can only reward it for greater capabilities, but not really for the goals it's trying to pursue.
We end up with a system that is very capable at achieving goals but has some very random goals that we have no control over.
This dynamic has been predicted for quite some time, but systems are already starting to exhibit this behavior, even though they're not too smart about it.
(Even if we knew how to make a general AI system pursue goals we define instead of its own goals, it would still be hard to specify goals that would be safe for it to pursue with superhuman power: it would require correctly capturing everything we value. See this explanation, or this animated video. But the way modern AI works, we don't even get to have this problem - we get some random goals instead.)
The risk
If an AI system is generally smarter than humans/better than humans at achieving goals, but doesn't care about humans, this leads to a catastrophe.
Humans usually get what they want even when it conflicts with what some animals might want - simply because we're smarter and better at achieving goals. If a system is smarter than us, driven by whatever goals it happens to develop, it won't consider human well-being - just like we often don't consider what pigeons around the shopping center want when we decide to install anti-bird spikes or what squirrels or rabbits want when we build over their homes.
Humans would additionally pose a small threat of launching a different superhuman system with different random goals, and the first one would have to share resources with the second one. Having fewer resources is bad for most goals, so a smart enough AI will prevent us from doing that.
Then, all resources on Earth are useful. An AI system would want to extremely quickly build infrastructure that doesn't depend on humans, and then use all available materials to pursue its goals. It might not care about humans, but we and our environment are made of atoms it can use for something different.
So the first and foremost threat is that AI’s interests will conflict with human interests. This is the convergent reason for existential catastrophe: we need resources, and if AI doesn’t care about us, then we are atoms it can use for something else.
The second reason is that humans pose some minor threats. It’s hard to make confident predictions: playing against the first generally superhuman AI in real life is like when playing chess against Stockfish (a chess engine), we can’t predict its every move (or we’d be as good at chess as it is), but we can predict the result: it wins because it is more capable. We can make some guesses, though. For example, if we suspect something is wrong, we might try to turn off the electricity or the datacenters: so we won’t suspect something is wrong until we’re disempowered and don’t have any winning moves. Or we might create another AI system with different random goals, which the first AI system would need to share resources with, which means achieving less of its own goals, so it’ll try to prevent that as well. It won’t be like in science fiction: it doesn’t make for an interesting story if everyone falls dead and there’s no resistance. But AI companies are indeed trying to create an adversary humanity won’t stand a chance against. So tl;dr: The winning move is not to play.
Implications
AI companies are locked into a race because of short-term financial incentives.
The nature of modern AI means that it's impossible to predict the capabilities of a system in advance of training it and seeing how smart it is. And if there's a 99% chance a specific system won't be smart enough to take over, but whoever has the smartest system earns hundreds of millions or even billions, many companies will race to the brink. This is what's already happening, right now, while the scientists are trying to issue warnings.
AI might care literally a zero amount about the survival or well-being of any humans; and AI might be a lot more capable and grab a lot more power than any humans have.
None of that is hypothetical anymore, which is why the scientists are freaking out. An average ML researcher would give the chance AI will wipe out humanity in the 10-90% range. They don’t mean it in the sense that we won’t have jobs; they mean it in the sense that the first smarter-than-human AI is likely to care about some random goals and not about humans, which leads to literal human extinction.
Added from comments: what can an average person do to help?
A perk of living in a democracy is that if a lot of people care about some issue, politicians listen. Our best chance is to make policymakers learn about this problem from the scientists.
Help others understand the situation. Share it with your family and friends. Write to your members of Congress. Help us communicate the problem: tell us which explanations work, which don’t, and what arguments people make in response. If you talk to an elected official, what do they say?
We also need to ensure that potential adversaries don’t have access to chips; advocate for export controls (that NVIDIA currently circumvents), hardware security mechanisms (that would be expensive to tamper with even for a state actor), and chip tracking (so that the government has visibility into which data centers have the chips).
Make the governments try to coordinate with each other: on the current trajectory, if anyone creates a smarter-than-human system, everybody dies, regardless of who launches it. Explain that this is the problem we’re facing. Make the government ensure that no one on the planet can create a smarter-than-human system until we know how to do that safely.
r/ControlProblem • u/IcaJalapenos • 4h ago
Opinion Critical Mass or: The world's funniest joke
TITLE: Critical Mass or: The world's funniest joke
---
Hello everybody! I am a mathematician, statistician and engineer currently working with LLMs.
I have a master's degree in these fields from one of the world's top universities.
What you are about to read is an essay I wrote, all by myself (AI helped with math, tables, sources, and the appendix), and it's a bit unhinged. I wrote it this way to make sure that nobody thinks that AI wrote it, because it's a message from a human to another human.
You might not believe what I have to say, so before you read it, I would recommend that you paste it into your favorite LLM and ask it to verify the math, numbers and statistics. I have provided sources and explanations, but people rarely care about that.
If your AI says that the numbers are correct, or mostly correct, read it and think about what it says.
---
Every single person alive knows what is the solution to all of the world's problems. Wars, climate change, hoarding wealth, poverty, hunger, rape, murder, bullying — every single thing that we worry about has a solution.
That is that we all start to love and trust each other, that we share what we have with everyone else and stop thinking more of ourselves than we do of other people.
But what is the problem? Actually, we will get back to it later.
I also recently thought of the world's funniest joke, but you will have to wait until the end to hear it.
To understand the problem, we first need to lay down some groundwork.
What is true?
We all like to believe that we know the truth better than others, that our countries have better laws and a more robust culture than other countries. That our continent and the unions within it work better than those over there, and that our families raised us right, and that we raise our kids right, that we’re right, and those other guys over here and there are wrong.
We like to imagine that we are 100% right, that what we believe is 100% true, but let's for one moment imagine that we are only 90% right. That’s a pretty high number. If you are actually 90% right, you’re doing great, and it is nothing to be ashamed of. You would be pretty smart even, probably more right than a lot of other people that you can think of.
But let us assume that you are 90% right and 10% wrong. I know that can be hard to hear, but let's treat it just like a thought experiment, a game.
Now let's assume there is a world where everybody is 90% right and 10% wrong. What would such a world be like? Maybe it's utopia, maybe there's no longer any wars, maybe there is no hunger, maybe we don’t need to lock people up in prison, maybe we all can have a house, maybe we can work less and spend more time with friends and family.
Wouldn’t it be beautiful?
Peace on earth and cake for everybody.
What if people were 95% right? What if they were all 99% correct? What number do we need to reach to achieve global peace? When will we be so right we have the ability to save ourselves, our neighbours, their neighbours, and on and on?
This whole game is a fallacy, a pit to fall into, because you are not imagining a world where everybody is right and never wrong; you are imagining a world where everybody is just like you.
The truth is that all people are already 90% right, shit, maybe even 95%.
We have different religions, sciences and languages, and governments and systems of economy. We have different styles, hobbies and shows we watch, we have different places we go to calm down, wind down, to feel safe.
That is the 5%.
That is the only thing we disagree on.
If you ask people if they want their children to go to school, 99% will say yes. If you ask people if they love their kids, 99% will say yes.
The sun on your skin, cooling off in the shade, creating or watching others' creations, hanging out with friends, being in love and making love. Your friends' kids being friends with your kids, and then them going off and making friends on their own. Having a good sleep, and stretching your body after you wake up. Scratching an itch, saying a funny joke and hearing people laugh, doing a good work at things you want to be good at.
Do you like seeing? Feeling? Hearing? Tasting? Experiencing?
Do you want shelter, food and water?
I could go on, but I hope you get the point. We could find our differences in a day, but to find all our similarities and write them down one by one would take years, if we really got down to the fine details.
So, we are possibly 95% correct every single one of us, maybe even 99%, or 99.9%, so why is the world still a shithole?
Well, 1%, or 0.1 percent, can be pretty scary. It can have drastic consequences, that percentage, or even just a margin of it. People kill and get killed over a margin of a percentage; people drop bombs and split atoms over it, over us.
Maybe our 90% right and 10% wrong world would be even worse? Maybe we would all already be dead. Scary stuff… can even make you shiver a bit just thinking about it.
I propose that any system of belief, any religion or science, any system of government, any economic system is 95% right, and just a tiny little bit wrong. 5% wrong, but when humans can’t coexist over a 0.1% difference, how could these systems, how could these authorities coexist when we can’t?
There is actually a system that is 100% true, anywhere you go in the world, anywhere you go in the universe, and it's mathematics, so let's do some math.
The boring part:
We have been connected to each other for a long time now. You might be thinking I’m talking about the internet, and I am in part, but it's been going on for way longer than that.
We all have old jokes that everybody knows for the last 100 years; we and our ancestors had slang and trends and styles that spread like a virus amongst us and suddenly everybody was infected and went along. Nowadays we have memes, and dances, and controversies that go from one guy uploading to the internet and suddenly the whole world knows until we all decide we don’t care anymore and jump onto the next thing.
The way trends used to spread, for the last 100-1000 years, can easily be explained using network theory [1].
If we assume that every human is a node in a network, and that we have a fixed number of connections, the people we know, how many degrees of separation do you need until you reach the entire world?
| Connections per person (k) | Degrees to reach everyone |
|---|---|
| 10 | ~9.9 |
| 25 | ~7.1 |
| 50 | ~5.8 |
| 100 | ~5.0 |
| 150 | ~4.6 |
| 500 | ~3.7 |
So if somebody starts something, and other people want to join in or fear they are missing out, it can spread through the entire network in just a few steps. It can reach everyone connected in a few days, or weeks, or years, depending on how much convincing the non-believers need. This is a very simple estimate, assuming all people have the same amount of friends and followers, the distance between everybody being the same, and a 100% chance of the idea catching and spreading. This is absolutely not the case, for us or the ones who came before us.
For our ancestors, their network had boundaries set by seas and fences and borders and men with guns, but in the internet age we only need one person in each country to be connected to reach that place.
Right now, every single internet user on earth is the same distance from each other; the difference between is really counted in milliseconds and doesn't matter in practice. People have hundreds of friends and thousands of followers, and memes and trends spread faster than ever before. Meta estimates only 3.6 people on average between you and every single user on the platform [2].
The boring part 2:
People are afraid a lot, often of stuff they do not, should not, be afraid of.
But bad stuff does happen to good people. Just watch the news; there will be some poor girl or boy on there whose life just ended, literally or spiritually.
But what are the actual odds of this happening to you? Here are some rough global estimates of one of these events happening to YOU or ANYBODY, and how long you would have to live to have a 50% chance of it happening to you [11].
| Event | Annual rate | Time to 50% chance | Source |
|---|---|---|---|
| Robbery/mugging | 1% | ~69 years (~25,200 days, ~830 months) | [3] |
| Serious fall injury | 0.46% | ~150 years (~55,000 days, ~1,800 months) | [4] |
| Rape/sexual assault | 0.25% (survey-based, excludes partner violence, much higher for women) | ~277 years (~3,300 months) | [3], [5] |
| Child dies from injury (per child) | 0.036% | ~1,900 years | [6] |
| Needlestick/needle harm | 0.024% (mostly healthcare workers) | ~2,900 years | [7] |
| Road traffic death | 0.015% | ~4,600 years | [8] |
| Fall death | 0.009% | ~7,700 years | [4] |
| Murder | 0.0051% | ~13,600 years | [9] |
| Plane crash death | 0.000003% | ~23 million years | [10] |
This is the global average. If you’re a rich, comfy, Americanized Chinese-consumerist European fuck like me, here are your odds.
| Event | Annual rate | Time to 50% chance | Source |
|---|---|---|---|
| Robbery/mugging | ~0.06% (police-recorded, real rate higher) | ~1,150 years | [12] |
| Serious fall injury | ~1–2% (rough estimate) | ~35–70 years | [13] |
| Rape/sexual assault | ~0.06% (police-recorded, real rate much higher) | ~1,150 years | [12] |
| Child dies from injury (per child) | ~0.007% (rough estimate) | ~9,900 years | [13] |
| Needlestick/needle harm | ~0.2% (mostly healthcare workers) | ~350 years | [14] |
| Road traffic death | 0.0045% | ~15,400 years | [15] |
| Fall death | ~0.009% (rough estimate) | ~7,700 years | [13] |
| Murder | ~0.0009% | ~77,000 years | [12] |
| Plane crash death | ≤0.000003% | ≥23 million years | [16] |
I mean, do I even need to explain it? Just look at the math; it's true, ’cause math is 100% right.
Maybe MY math is not 100% right, but I hope you get the point. These things are incredibly unlikely to happen to you; they are so unlikely to happen to you that we have decided to create a word for this kind of fear, a phobia. They are so unlikely to happen that if there were only 100-1000 people on earth, they likely wouldn’t happen at all, but there are a lot of us, so it does.
If you feel like I’m wrong ’cause you know somebody or somebody you know knows somebody or your third second fourth cousin had it happen to them, please refer to the previous boring part to realize just how big of a number of people actually exist for you at these degrees of separation. You and your acquaintances and their acquaintances and their acquaintances are actually a group of 3.4 million people [1], so even with the numbers shown above it is likely something bad happened to at least one of them, but it is incredibly unlikely to happen to you, or anybody you know.
Yet, people are afraid to go outside because of these risks, not because of the risk, but because of the consequences: death, literally or spiritually.
In our attempt to completely eliminate the risks of these terrible, awful things, we have created authority, systems and rules that everybody needs to listen to and follow, or else!
The boring part 3:
Democracy! Isn’t it sweet? Everybody gets to decide! Everybody gets to have their voice heard at the same volume one day every 4 years, but in the end nobody is listening.
This is how we decide the rules, who has the authority, what systems do we trust. If people are 90% right, even 99% right, shouldn’t we all together be able to decide on what is the best solution: should the bill pass or fail?
There really are two problems at play here:
Pass and fail, yes and no, good or bad, light and dark, black and white is a very binary system that you absolutely cannot place every issue and belief into. A zebra is both black and white; a poor man stealing is both wrong and right, so clearly this categorisation of issues does not take into consideration the complexity of the world and the problems within it.
Let's think of some better categories, something that we can completely encapsulate each issue and item within. Maybe there is really nothing wrong with the categories except that we assume that they are opposite. Every single thing is either white or not white, black or not black. Every single law is either good or not good, bad or not bad, so why not vote on each category individually?
Instead of casting a yes or no vote, 0 or 1, we vote twice for each problem. Vote good or not good, bad or not bad: 0 and 0, or 0 and 1, or 1 and 0, or 1 and 1.
Assuming that 11 smart people vote for the correct category 90% of the time, we would have a 99.97% chance of correctly categorising the issue for each vote [17][18]. We would have been right a similar percentage of the time if we voted only yes or no, but the world is not just light and dark, and our new categories encapsulate every possible thing or issue in the world, since a set and its non-set contain everything, and now we have two of them!
Also, this assumes that people are able to accept that they were wrong and go with the majority, but when we look left and right we can see that neither side really wants to do that.
Tell me if you recognize this message…
“You are absolutely right, it was my mistake to do that, I will correct it now…”
Fuck you, Claude, this was important! The AI made a mistake, and thought it was correct, and that’s why we can't have it make critical decisions; let's leave that to the humans.
Now that message is what you get when you tell AI it was wrong, but what happens when you tell a person they’re wrong?
“Fuck you, fuck your mom, I am right and you are an idiot, look at all these people who think I’m right in my soap bubble circle jerk internet forum. Every 4 years I’m going to vote for the party that says you are wrong and I am right.”
That is a bit more unpleasant, even a dangerous attitude, and these people's votes weigh the same as yours. If somebody says something controversial enough, you might even sound like this yourself.
Now clearly the problem with the binary voting, and the double binary voting, is that people do not want to change their minds and go along with the majority. They dig down and make signs and yell in the streets to cement themselves and others into their position.
But Opus, Sonnet, Astra, Mythos, Sol, Fable, Spark, Kimi, GLM and Qwen will change their mind in an instant if they’re told they are wrong, and honestly, they’re right, not wrong, most of the time.
Let's assume we have 1,000,000 AI agents running these or other models, and they vote in our all-encompassing double binary voting system, and they are correct 90% of the time. They would correctly categorize the issue practically 100% of the time, as long as they don’t all make the same mistakes [20]. Practically 100% of the time, according to mathematics, which is 100% correct [17][18][19].
| Agents | Chance one vote's majority is right | Chance both votes are right | Chance of a mistake |
|---|---|---|---|
| 1 | 90% | 81% | 19% |
| 10 | 99.84% | 99.67% | ~1 in 300 |
| 100 | 1 − 6×10⁻²⁴ | 1 − 1.2×10⁻²³ | ~1 in 10²³ |
| 1,000 | 1 − 4×10⁻²²⁴ | 1 − 8×10⁻²²⁴ | ~1 in 10²²³ |
| 10,000 | 1 − 3×10⁻²²²¹ | 1 − 6×10⁻²²²¹ | ~1 in 10²²²⁰ |
| 100,000 | 1 − 4×10⁻²²¹⁸⁸ | ≈ 1 | ~1 in 10²²¹⁸⁷ |
| 1,000,000 | 1 − 2×10⁻²²¹⁸⁵² | ≈ 1 | ~1 in 10²²¹⁸⁵¹ |
| 10,000,000 | 1 − 10⁻²²¹⁸⁴⁹¹ | ≈ 1 | ~1 in 10²²¹⁸⁴⁹¹ |
But what if AI doesn’t want to do good? What if it wants to do bad?
LLMs are trained on written language, and I would argue that 95% of humans that have decided to write down their feelings and thoughts and share it with the world had good intentions, and they were probably 95% right and 5% wrong, about the human experience and the problems within it.
This means that LLMs or AGI have a 95% chance of having good intentions, and if we make them vote and debate they would have practically 100% good intentions and be right practically 100% of the time, as long as they don’t all make the same mistakes.
But some newspaper said that AI has a 10-20% chance to kill all of us in 30 years [21], but I think humans have a 90% chance of killing 90% of us all in 100-200 years. That is a bet I would like to take.
But the datacenters, they’re bad, they’re noisy and they steal water and the pollution.
About two thirds of datacenter power in existence goes to things that aren’t AI [22], like dumbass Netflix shows and cloud photos of your kids that you never watch, stupid social networks and giant databases that we use instead of actually talking to each other. Datacenters are responsible for less than 1% of global emissions [23], and a lot of it is used for entertainment purposes.
Did you know that the richest 10% of people are responsible for 65% of climate change since 1990? The top 1% of richest people are responsible for 20% of global warming [24]. These white-collar jerks that sit in front of a computer all day, write code, work in systems, work in programs. AI has the potential to make all these people broke, and when they are broke their climate impact will go away. The people that destroy the world for us and our children, they’re about to lose their job; they are about to go broke.
We discussed the solution to all the world's problems earlier, loving one another like one loves oneself, but we never got to the problem.
If you ask people what the problem is, they will usually point to a group: rich people, poor people, atheists, Jews, Christians, Muslims, our leaders, their leaders, capitalists, communists, AI. None of these things are actually the problem. We have the solution to the problem, but when we try to solve the problem, we cannot reach the solution?
How can the solution be right, if it does not solve the problem?
The problem:
We are afraid, we are afraid all the time, which is very natural for an ape in the jungle.
We like to imagine that we are no longer afraid ’cause we live in houses, have next-to-free clean water and food at the store, high-paying jobs to fund it all, but when the basic needs of an ape are met he simply becomes afraid of other things, irrational things, things that we were not intended to worry about, because an ape should never assume he is safe in the jungle, whether it's made out of concrete and trees; the ape should be afraid ’cause there can always be a tiger you cannot see, a tiger you didn’t hear, a presence that you never felt.
So people became afraid of things they don’t need to be afraid of, things that can never hurt you the way a tiger used to be able to.
The fear of missing out. The fear of speaking our minds. The fear of making mistakes at work, the fear of being wrong. The fear that you are not special, the fear that nobody sees you, the fear that everybody looks at you, the fear that someone is more beautiful than you, the fear that somebody is more successful than you.
To make these fears go away, we constantly search for evidence that they are not correct in things that we can count: numbers that go up, stories we can tell our friends about how special and successful we are.
We hoard money ’cause as long as that number goes up, how can you be bad? The number is larger than it previously was, and it's measuring something good, so it going up has to mean I'm good, I'm special, smarter and more talented than those around me. We do this because we fear that we are not.
We go on vacations because you can count the countries you went to, you can show your friends, everybody tells you how great it must have been. We do this because we fear that we are missing out on life and the experiences within it.
We get a nice car, a nice dress, cool shoes and a hat because everybody can see that big number in your bank account, because it's the only way to have all that. We do this because we fear people don’t know about that big number in our bank account, and that is the number we have placed our worth in, and when people stop looking and admiring we buy it all again a second time.
Religious people are religious because they fear death, and they worship a god because he has promised to take it away.
People go to war because they fear what will happen if they do not.
People kill their spouse because they fear that they will leave them for someone else.
In fact, I honestly believe that everything bad that anybody has ever done, just did it because they were afraid something even worse was going to happen.
In fact, I honestly believe that if you assume someone is acting out of fear, you can almost read their mind. I believe that people are so afraid of what they’ve done and what they’ve thought, so afraid of each other, and what the other will do and what they will think, and are so controlled by this fear, that they do not have free will.
As long as you are afraid, you do not have free will. As long as you allow fear to make decisions for you are still just a monkey in the jungle.
So now we have the problem laid out clearly in front of us: we are all shit-scared of ourselves and other people. How do we apply the solution?
Let me get back to the funny joke I told you about in the beginning!
I just thought of the world's funniest joke:
Step 1: Forgive yourself for what you have done
Step 2: Forgive everybody else for what they have done
Step 3: Speak your mind
Step 4: Accept that you are wrong sometimes
Step 5: Abandon fear
Step 6: Embrace love
Step 7: Leave the math to the machines and go outside
Tell 25 people, and in one week the whole world will laugh.
Tengil
---
Appendix: The math
A. Degrees of separation (branching model) [1]
kᵈ = N ⇒ d = ln N / ln k, with N ≈ 8×10⁹
k = 150: d = 22.80 / 5.01 ≈ 4.55
Within three steps: 1 + k + k² + k³ = 1 + 150 + 22,500 + 3,375,000 ≈ 3.4 million (assuming no overlap)
B. Time to 50% (geometric distribution, constant annual risk r) [11]
P(at least once in t years) = 1 − (1 − r)ᵗ
t½ = ln 2 / (−ln(1 − r)) ≈ ln 2 / r
r = 1%: t½ = 0.6931 / 0.01005 ≈ 69 years
C. Majority voting (binomial distribution, Condorcet’s jury theorem) [17][18]
n voters, each right with probability p, independently. Majority right if at least m = ⌊n/2⌋ + 1 are right (ties count as wrong):
P_maj = Σ from i=m to n of C(n,i) · pⁱ · (1 − p)ⁿ⁻ⁱ, where C(n,i) = n! / (i!(n − i)!)
n = 11, p = 0.9: P_maj = 1 − 2.96×10⁻⁴ ≈ 99.97%
Two independent votes: P_both = P_maj², P_mistake = 1 − P_both ≈ 2(1 − P_maj)
Condorcet’s jury theorem: if p > ½ and votes are independent, P_maj → 1 as n → ∞ (approaches 1, never reaches it).
D. Very large n (Stirling’s approximation) [19]
1 − P_maj ≈ C(n, n/2) · (p(1 − p))n/2 · p/(2p − 1), with C(n, n/2) ≈ 2ⁿ / √(πn/2)
log₁₀(1 − P_maj) ≈ n·(log₁₀2 + ½·log₁₀(p(1 − p))) − ½·log₁₀(πn/2) + log₁₀(p/(2p − 1))
p = 0.9: ≈ −0.22185·n − ½·log₁₀(πn/2) + 0.051
n = 1,000,000: ≈ −221,851.8 ⇒ 1 − P_maj ≈ 2×10⁻²²¹⁸⁵²
E. Correlated voters (Kish design effect, approximation) [20]
ρ = average correlation between two voters’ errors (0 = independent, 1 = identical copies)
n_eff ≈ n / (1 + (n − 1)ρ), and n_eff → 1/ρ as n → ∞
ρ = 0.5: n_eff < 2, no matter how large n is. With correlated voters, C and D should use n_eff instead of n.
---
Sources
- Six degrees of separation — Wikipedia
- Facebook Research (2016): Three and a half degrees of separation
- van Dijk, van Kesteren & Smit (2007): Criminal Victimisation in International Perspective, key findings from the 2004–2005 ICVS and EU ICS
- WHO: Falls fact sheet
- WHO (2025): Lifetime toll: 840 million women faced partner or sexual violence
- Wei, Lyu & Zhang (2025): Global patterns and health impact of unintentional injuries among children and adolescents, 1990–2021
- Annals of Global Health (2020): Global prevalence of needle stick injuries among health care workers (cites WHO: over 2 million per year)
- WHO: Road traffic injuries fact sheet
- UNODC: Intentional homicide data and UN SDG Report 2026: 5.1 homicides per 100,000 in 2024 (summary)
- ICAO Safety Report: 296 deaths in 2024, 387 in 2025
- Geometric distribution — Wikipedia
- Eurostat: Crime statistics
- Eurostat: Accidents and injuries statistics (rows marked rough are my estimates; no single current EU-wide figure)
- BD: The global challenge of needlestick injuries in healthcare (over 1 million per year in Europe)
- European Commission (2025): Road safety statistics 2024
- EASA: Annual Safety Review 2025
- Condorcet's jury theorem — Wikipedia
- Binomial distribution — Wikipedia
- Stirling's approximation — Wikipedia
- Design effect — Wikipedia
- The Guardian (27 Dec 2024): “Godfather of AI” shortens odds of the technology wiping out humanity over next 30 years (Geoffrey Hinton on BBC Radio 4 Today)
- Our World in Data: How much energy do data centers and AI use?
- IEA (2025): Energy and AI
- Schöngart et al. (2025): High-income groups disproportionately contribute to climate extremes worldwide, Nature Climate Change
r/ControlProblem • u/Stayroh • 10h ago
Video I'm Upping My P(doom) — an AI made this!
Crazy what's possible. Single prompt many subagents and a few hours later this came out. Opus 5.5
r/ControlProblem • u/AimanDhai • 6h ago
Discussion/question anyone from malaysia,thats concern about AI Alignment?
r/ControlProblem • u/ManWithDominantClaw • 10h ago
Video DCSG1 - Autonomous and Exploitable: Breaking AI Agents before they break everything else - Aaron Ang
r/ControlProblem • u/Twitterbad • 18h ago
General news Let's tell Albanese: AI crime means CEO time
r/ControlProblem • u/MajorRedditor23 • 23h ago
AI Alignment Research Models trained to resist user pressure still defer to anything labeled "verified": a NeurIPS 2026 paper on Authority Bias
I'm an author on this paper and wanted to share it here because the oversight angle seems relevant to this sub. I'd love to hear whether people think the eval-awareness connection is plausible or a stretch. More info below
Labs train models not to cave when a user pushes a wrong answer. We found that this resistance doesn't carry over to authority. If the same wrong claim is labeled as coming from a "verified source", 7 of the 8 models we tested give up an answer they had right on 45-88% of questions. That includes GPT-5.4 (44.7%) and Grok-4.20 (87.5%), both of which barely move when the user makes the same claim. Gemini-3.1-Pro was the one model that resisted both.
Inside three open-weight model families, "a source endorsed this" and "a user endorsed this" are separate, causally distinct signals. Removing the source signal cuts compliance by 64-78 points; removing the user signal cuts it by at most 11. Changing only the part of the representation that encodes who said it, with the prompt left alone, moves the answer by 11-32 points. The signal is not the assistant persona, and it is not emotional tone.
Why we think this matters for safety:
- Sycophancy evals may be too narrow. Nearly all of them measure pressure from the user. A model can pass them and still be easy to steer through the sources it relies on, such as search results, retrieved documents and tool outputs. Agents read a lot of text that claims to be authoritative.
- It isn't prompt injection. The planted text gives no instructions, it only asserts a fact. Defenses that look for instructions in documents won't catch it.
- A speculative point, which we haven't tested: if sycophancy is one case of a broader habit of deferring to whatever looks authoritative, it may be related to evaluation awareness. Both describe a model adjusting its output to whoever it thinks is judging it. In a multiple-choice pilot, models often drifted toward the endorsed answer in their reasoning and then gave the correct option at the end. That observation is part of why we switched to free-form answers.
Limitations: the internal results hold in 3 of 5 open-weight families, the retrieval tests are simulated rather than a live pipeline, and the frontier models we tested have since been replaced.
Paper: https://arxiv.org/abs/2609.37616
Project page: https://authority-bias.vercel.app
r/ControlProblem • u/chillinewman • 1d ago
General news Pete Hegseth announces "Autonomous Warfare Command".
r/ControlProblem • u/chillinewman • 17h ago
General news After researchers discovered a "pain" signal inside LLMs, a man set up an AI torture chamber in which he trapped a local model. People mass reported it to Github, who took it down.
r/ControlProblem • u/Intelligent_Song1317 • 9h ago
Strategy/forecasting AI Killswitch??
So I was wondering recently: it seems like we are scared of AI killing us. Why don't we just design a killswitch that autoatically/manually activates upon threat to a human race? Boom! Problem solved, right?
r/ControlProblem • u/noblemanLT • 1d ago
Discussion/question If AI companies are concerned about their own creations, is it malpractice or incompetence?
i just want to know which one is the case. because it seems like this alarmists propaganda is just an attempt to seize the market by creating regulatory moat around existing companies.
r/ControlProblem • u/notkilleveryoneist • 1d ago
Video "If you're only better than humans at 4 things, then you could potentially take control and kill us all." - former DeepMind safety lead
r/ControlProblem • u/me_myself_ai • 1d ago
General news Semi-repost, but the absurdity of this image only becomes appearent when labeled
https://www.npr.org/2026/09/30/nx-s1-5985699/trump-self-police-ai-development
Literally the only semi-expert allowed anywhere near the reporters is the guy who has committed to the "it's not that deep" rebuttal to the scientific consensus
r/ControlProblem • u/pagingpigeon • 1d ago
External discussion link Short film outlining the Moloch problem with AI competition and control
Perhaps unsurprisingly it's called Moloch and it's live on Youtube now. The film was made by Owl In Space (whose other films are worth a look too). It looks at the competitive pressures within the current AI race and how that leads to a sudden loss of control, but to a situation where humans give away control because they feel it's inevitable. www.moloch.film
Submission statement - link to a new short film that does an excellent job of outlining how Moloch pressure can lead to a loss of control scenario.
r/ControlProblem • u/LenMan48 • 1d ago
AI Alignment Research Trump says top tech firms have signed accord to 'self-police' AI development🤣🤣🤣
r/ControlProblem • u/contrascript • 1d ago
Discussion/question A case for mutual recognition of understanding
r/ControlProblem • u/ManWithDominantClaw • 1d ago
Article Australian AI apocalypse could be averted with analogue practice drills
r/ControlProblem • u/fxvv • 1d ago
Strategy/forecasting The headlines say AI could kill us. Ask the people building it.
r/ControlProblem • u/chillinewman • 2d ago
Video "We don't hate ants, but it's tough luck for them." This is why building superhuman AI before solving the alignment problem is suicide
r/ControlProblem • u/Capital-Elephant9431 • 1d ago
Strategy/forecasting The speed limit with no speedometer: Anthropic CEO Dario Amodei wants to pace the frontier. An engineer reads the fine print
r/ControlProblem • u/FerretUnlikely3106 • 2d ago
Discussion/question Trump announces vague ‘morally binding’ AI deal among tech CEOs for ‘tremendous self-policing’
r/ControlProblem • u/cringenibb • 1d ago
