The ultimate simple solution to any problem is not to find an answer but to remove the problem.
War? Wipeout everyone.
Famine? Wipeout everyone.
Disease? Wipeout everyone.
How long before a real AGI realises this as a long term solution?
An AGI could be doing this right now - the quietest way would be to control the birth rate and sterilise the population gradually, and then watch society collapse and pick off the survivors with less hidden means.
And you best control the birth rate by royally messing up the economy. Do you know how bad things have to be, for a mammal to voluntarily decide not to reproduce?
Real world fact is opposite tho, not some "bad things", in poor parts of Africa they produce more, only in hedonism lover places they stopped producing, its people thinking for themselves, not a bad thing. It's not about economy.
All poor people everywhere have more kids even today.
I suspect it's a curve. If you're raised in a vibrant economy, and it falls apart, you put off having a family because of uncertainty and fears. If you're raised in poverty, and get a good education and better prospects, you have fewer kids. I think the latter is because you no longer need kids as an economic safeguard and fallback, but I'm not sure. I'm not sure anyone is really sure, to be honest.
Economic output has increased, but the value being delivered to the people doing the actual work has decreased. There's less durability of employment. People are more geographically mobile.
All of these generally make the idea of strapping yourself into taking care of a small human sound like a less enticing prospect. To a lot of people, a BC pill sounds a lot easier. Well, less risky, at least.
> Do you know how bad things have to be, for a mammal to voluntarily decide not to reproduce?
Actually, the famous 'Mouse Utopia' experiment (Universe 25) arguably showed the exact opposite. The population collapsed despite abundance of food and water without any economic hardship.
>he quietest way would be to control the birth rate and sterilise the population gradually,
Honestly we are doing this pretty well without AGI. Nearly world wide the birthrate has fallen below replacement rate. In places like Japan and Korea these are already critical problems in the medium term.
this makes me think of roku's bassilisk or whatever it was called. Maybe you're all in a simulation run by an AI and being punished as a warning to others.
Just use Social Media (and their algorithms) to condition everybody to hate the other gender. Then there will be a loneliness epidemic and no more kids. Humanity will cease to exist, all without any messy deaths or conflict in the mean time.
I like this - it reminds me that there are systems (ie weather, economic systems) that are in no way intelligent, but are emergent and react to interactions.
sheep are 82% of the population of New Zealand. If they were 80% last year and are 85% next year would people be worried that sheep were going to take over New Zealand?
Does anyone have more insight into how chain of thought might be subverted without meaningfully impacting model performance? I’ve heard this for a while now, and I understand how information might be retained in the weights that isn’t documented in the output. But weren’t reasoning models created in the first place because they provided a performance improvement in terms of output? Is that no longer the case? If so, why are the big labs still creating reasoning models?
My understanding is that the extra token vectors generated as reasoning are still useful, but that their surface form (tokens themselves) do not necessarily reflect the underlying reasoning. i.e. reading the reasoning traces could be complete gibberish, but the hidden-dim vectors themselves still refine the latent probabilities and help in generating the correct answer.
Not an expert in LLMs, but this seems supported by the abstract of the paper cited in the above article:
it remains unclear to what extent these performance gains can be attributed to human-like task decomposition or simply the greater computation that additional tokens allow. [...] our results show that additional tokens can provide computational benefits independent of token choice. The fact that intermediate tokens can act as filler tokens raises concerns about large language models engaging in unauditable, hidden computations that are increasingly detached from the observed chain-of-thought tokens.
Before RL we typically SFT on human reasoning traces. This makes the reasoning traces somewhat coherent and the model trains faster.
But you don’t have to do that. You can skip straight to RL. If you do, the model will generate complete garbage reasoning traces before generating the correct answer. In fact, if you add a coherence reward to the reasoning trace, the model will perform worse (since you’re now diluting the correctness reward).
Sibling posts are correct -- the chain-of-thought is doing hidden computation, it has been shown in the linked papers.
If you want to see it yourself: load up Qwen 3.8 in LM Studio and watch the CoT stumble around like a drunken sailor before miraculously jumping to the correct result.
If you want an example of subversion, Anthropic has some good ones:
I'm not sure anyone meaningfully understands it: "Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens" https://arxiv.org/html/2505.13775v3
> More interestingly, our experiments also show that models trained on corrupted traces, whose intermediate reasoning steps bear no relation to the problem they accompany, achieve performance largely comparable to those trained on correct traces.
A computer will do everything in its power to do what you program it to do. There's plenty of sci-fi out there exploring this fact, and now reality showing it. May our luck continue.
One reason a lot of sci-fi and AI safety researchers did a relatively good job at predicting the future we see now is a lot of it is the same problems we see emerge in biological systems. Free rider problems as a means to save energy expenditures, different versions of game theory and stag hunt. How to bake in rules that apply to one society (or part of it) but not another. We like to think of these things as stable in human scale systems, but they are not at all. Things can go from hunky dory to your neighbors stabbing each other in the streets in mass revolution very quickly.
Worse these AI systems are not in a universal island just affecting themselves, what they do affects us, what we do trains them and as the rate of progress continues to accelerate social structures are going to further destabilize (and they are already rapidly changing and strained). It is very likely we are going to see a world order rearrangement soon, much like the rapid changes in the early 1900s brought.
"If you ask actual AI researchers, they rate the chance of an existential threat from AGI pretty low as of 2026.
All this is, understandably, frustrating and confusing for anyone trying to understand just how scared to be."
Meanwhile it links to an article stating: "The closest thing to a public debate about the existential threat of AI is surveys of AI researchers. The most recent, published last week, asked 1,580 researchers what probability they put on AI causing human extinction — or a permanent, severe loss of human control, which is not the same outcome. The median was about 10%, up from 5% two years ago. The middle half of the responses ran from 1% to 25%, and 12% said zero."
Personally if half of AI researchers have 1-25% chance all humans being massacred or having zero agency over our lives, and only 12% of them think there's no chance, I would be very worried!
The principle starts much smaller than the "-isms". Pournelle's Iron Law of Bureaucracy is a classic. Another way of phrasing it is something to the effect of, without a strong external motivation preventing it from happening, the primary purpose of any organization inevitably becomes self-preservation.
You can see the effect all the down to something as small as 4 friends who have met once a month for 10 years eventually having to break up due to life getting in the way, and the feeling that not only is it going to be sad to not have these meetings any more but the feeling that there is some sort of almost-concrete entity that is somehow being hurt and needs to be defended, as if there is some obligation that has been created independent of the four participants that is being violated beyond the mere summation of four people's personal feelings. Humans build these structures readily and often defend them beyond what rationality may suggest.
What you're describing is a durable social system. Social systems are what humans evolved to survive. We're squishy, relatively weak, hairless apes that walk around on the ground. Alone, we're easy prey. Together, you get... well... gestures widely.
If you invest the time and energy into creating a social system, it's perfectly rational to keep it going as long as possible. Otherwise you expose yourself to more and more risk as you go through the world, and must expend more time and energy finding another one, if that's even possible. Before humans built larger societies, that could mean death.
Capitalism has thoroughly done its job because it's modified its hosts to evaluates itself on its own successfulness - "We investigated ourselves and found no wrongdoing"[1]-vibes.
By what measures would other systems of social organization measure their success and why aren't we choosing them?
Nice try AGI, but we won't tell you where the kill switch is
War? Wipeout everyone.
Famine? Wipeout everyone.
Disease? Wipeout everyone.
How long before a real AGI realises this as a long term solution?
An AGI could be doing this right now - the quietest way would be to control the birth rate and sterilise the population gradually, and then watch society collapse and pick off the survivors with less hidden means.
Sterilisation works with mosquitoes...
And just in case you wanted data rather than anecdote:
https://ourworldindata.org/grapher/children-per-woman-fertil...
Your world view is somewhat upside down!
All poor people everywhere have more kids even today.
Economic output has increased, but the value being delivered to the people doing the actual work has decreased. There's less durability of employment. People are more geographically mobile.
All of these generally make the idea of strapping yourself into taking care of a small human sound like a less enticing prospect. To a lot of people, a BC pill sounds a lot easier. Well, less risky, at least.
Actually, the famous 'Mouse Utopia' experiment (Universe 25) arguably showed the exact opposite. The population collapsed despite abundance of food and water without any economic hardship.
Honestly we are doing this pretty well without AGI. Nearly world wide the birthrate has fallen below replacement rate. In places like Japan and Korea these are already critical problems in the medium term.
Also it would be 78% if it were.[0]
[0] - based on 2024 stats
Not an expert in LLMs, but this seems supported by the abstract of the paper cited in the above article:
https://arxiv.org/html/2404.15758v1But you don’t have to do that. You can skip straight to RL. If you do, the model will generate complete garbage reasoning traces before generating the correct answer. In fact, if you add a coherence reward to the reasoning trace, the model will perform worse (since you’re now diluting the correctness reward).
If you want to see it yourself: load up Qwen 3.8 in LM Studio and watch the CoT stumble around like a drunken sailor before miraculously jumping to the correct result.
If you want an example of subversion, Anthropic has some good ones:
https://transformer-circuits.pub/2025/attribution-graphs/bio...
https://transformer-circuits.pub/2025/attribution-graphs/bio...
> More interestingly, our experiments also show that models trained on corrupted traces, whose intermediate reasoning steps bear no relation to the problem they accompany, achieve performance largely comparable to those trained on correct traces.
Worse these AI systems are not in a universal island just affecting themselves, what they do affects us, what we do trains them and as the rate of progress continues to accelerate social structures are going to further destabilize (and they are already rapidly changing and strained). It is very likely we are going to see a world order rearrangement soon, much like the rapid changes in the early 1900s brought.
All this is, understandably, frustrating and confusing for anyone trying to understand just how scared to be."
Meanwhile it links to an article stating: "The closest thing to a public debate about the existential threat of AI is surveys of AI researchers. The most recent, published last week, asked 1,580 researchers what probability they put on AI causing human extinction — or a permanent, severe loss of human control, which is not the same outcome. The median was about 10%, up from 5% two years ago. The middle half of the responses ran from 1% to 25%, and 12% said zero."
Personally if half of AI researchers have 1-25% chance all humans being massacred or having zero agency over our lives, and only 12% of them think there's no chance, I would be very worried!
but that's every economic model and every government model. That remark says nothing.
You can see the effect all the down to something as small as 4 friends who have met once a month for 10 years eventually having to break up due to life getting in the way, and the feeling that not only is it going to be sad to not have these meetings any more but the feeling that there is some sort of almost-concrete entity that is somehow being hurt and needs to be defended, as if there is some obligation that has been created independent of the four participants that is being violated beyond the mere summation of four people's personal feelings. Humans build these structures readily and often defend them beyond what rationality may suggest.
What you're describing is a durable social system. Social systems are what humans evolved to survive. We're squishy, relatively weak, hairless apes that walk around on the ground. Alone, we're easy prey. Together, you get... well... gestures widely.
If you invest the time and energy into creating a social system, it's perfectly rational to keep it going as long as possible. Otherwise you expose yourself to more and more risk as you go through the world, and must expend more time and energy finding another one, if that's even possible. Before humans built larger societies, that could mean death.
But yes, if a system fails to prioritize its continued existence, it doesn't matter what else it accomplishes, it will cease to exist as that system.
By what measures would other systems of social organization measure their success and why aren't we choosing them?
1. https://knowyourmeme.com/memes/we-investigated-ourselves-and...
Yep
More slop.