// HACKER NEWS — CYBERSECURITY
The Education of a Doomer
If you follow me on Twitter, or read this blog, you have noticed that I went
from being generally optimistic and excited about AI to being extremely
concerned. And I thought, I should explain why I changed my mind. Each section
of this post is about some aspect of AI where my views shifted. I begin each
section by explaining my previous beliefs, and why I held them, and then explain
why those beliefs changed.
Automation has been good. We’ve automated 99% of the jobs people did in 1790,
and the result is not mass unemployment, rather, we are wealthier, healthier,
more educated, we have more leisure, etc. I had this vague, inductive idea that,
while I can’t predict what jobs will exist after AGI, there will be demand for
me to do something. If nothing else, the much higher economic growth of the
post-AGI world means that the human niche, while small in absolute terms, might
be much larger than today’s economy.
And this may yet be true. Or it may show a lack of imagination on my part. If AI
develops such that we have enduring complementarity between humans and AIs, then
we might still have jobs in the post-AGI future. But if AI becomes truly
general, the G in AGI, and on top of that it is vastly smarter, faster, and
cheaper than humans, then there might be nothing for us to do, except live off
UBI.
When people talk about UBI, they typically worry about the problem of meaning in
a world without work. I’ve never had this worry. When I was funemployed last
year, I spent my time reading books and writing code and hanging out with
friends. If the future is an infinite UBI-funded vacation, I know what I’ll
do. “Before the Singularity, read books and throw house parties; after the
singularity, etc.”
But then I started thinking about the political consequences of AGI, and started
writing about it:
I had hope that alignment would turn out to be a normal engineering problem,
that we solve through empirical experimentation and investigation, like
everything else. It helps that the early LLMs were not incomprehensible alien
minds, like something from a Stanisław Lem novel, but rather immensely
human. It’s hard not to anthropomorphize them. The huge core of unsupervised
learning in an LLM understands human morality just fine: you can talk to them
about it, they will explain, eloquently, in detail, why something is “good” or
“bad” according to some moral system. And, because capabilities were weaker, the
failures were very small. What’s the worst ChatGPT in 2022 could do?
Since like 2024, reinforcement learning has been the main technique to push the
frontier forward. And reinforcement learning agents work exactly like Yudkowsky
says. Consequently, capabilities have increased markedly but the models are
harder to understand (literally: their prose is increasingly incomprehensible)
and are increasingly misaligned as RL scrambles their brains in pursuit of
reward. Incidents of serious misalignment are more common and more
consequential. It’s clear that AI capabilities are growing far, far faster than
our ability to control or even understand them.
I thought—or, rather, I implicitly believed—that people would want to
remain in control of the AIs. And if we want to retain control, and solve
alignment, then we will stay in control. Simple enough. But recently I started
to think: no, we will probably hand over control to the AIs.
The weak version of the disempowerment thesis is something like the prisoner’s
dilemma: people/companies/polities that hand more power to AI outcompete those
that don’t, so there’s competitive pressure towards disempowerment. This is easy
to believe.
The strong version of disempowerment is: the AIs will be so smart,
knowledgeable, personable, moral etc. that we will willingly, voluntarily hand
power to them. We’ll think, “they can do a better job than us”, and we’ll be
right. Believing this requires you to be somewhat cynical about humanity’s
desire for autonomy vs. material considerations. But, over the past few m