← blog

Why I Doubted AI Risk

Analyzing My Biases

7 min read

Until a few weeks ago, I considered myself fairly skeptical of AI safety concerns. I was aware of them intellectually, and although I did not grow up reading The Sequences, it had been years since I’d first heard of “P(doom).” Nonetheless, it was usually not something which had occupied a significant share of my mind.

I’d like to take a moment here to think back on that version of Joel, and consider where he disagreed with my current self. What led him to believe that these issues were not worth considering? What cognitive fallacies was he falling for?

Genetic Fallacy

It has been my experience with humanity, that people are quite likely to believe what is convenient or desirable for them to believe. One particularly prominent instance of this in my early career has been in people’s selection of jobs and the moral weight which they assign this.

In quantitative finance, you will find many who solemnly swear by the value of providing liquidity to the market. Their classmates, who instead stumbled into big tech, will discover that in fact the most important engine of the economy is advertising.1 The most devout communists I knew are now happy bankers and landlords.

So, when my friends in AI told me that it would be the most important thing and change the world, perhaps even that the survival of humanity would depend on their work, I didn’t believe their story that this was why they were attracted to the field. Instead, I believed they had been lured to it by the normal reasons: the prestige, the high salaries, the intellectually stimulating nature, and more. Then, their proximity to it had led to their belief in its importance.

I may not even have been wrong about the causality here, nor is it crazy for this to have made me more skeptical of their claims. If they had been making immense personal sacrifice to do this work, that would have been a stronger signal, since there would have been fewer alternative explanations. Still, the fact that their reasoning may not have been pure and unbiased was not incriminating on its own. To dismiss concerns about the importance and risk of AI technology based on this was to fail to meaningfully engage with the ideas.

Guilt By Association

Early thinking about AI risk largely grew out of a few related online subcultures, in particular the effective altruists and the rationalists. Often, this was to such an extent that it could be hard to untangle from their other ideas. The effective altruist view in particular was salient to me, given my milieu.

Effective altruism was a movement which thought about how to more efficiently allocate money to charitable causes, often by expanding one’s domain of empathy, and giving for example in poorer countries where your money would go further. Rather than contribute to a GoFundMe for a child in your state needing an incredibly expensive surgery to save their life, you might instead spend it on bed-nets to protect against mosquitoes carrying malaria, and save countless lives. This, to me, feels noble and just.

Soon, however, that instinct to expand the domain of empathy was taken further. Not only should you care about people geographically distant from you, but also those temporally distant: future humans who might come to exist. In its most extreme versions, by assuming the possibility of an infinite number of future humans, who all had the same moral worth as current humans, you could justify a belief that any infinitesimal reduction in the probability of human extinction (multiplied by the value of infinite future humans) was worth more than any concrete gain today. This belief was known as “longtermism”, and its adherents were among those most concerned about AI safety, which they saw as a possible long-term cause of human extinction.

I am not, and was not, a supporter of longtermism. I feel that if you are willing to give your contribution to extinction risk credit for the compounding returns of infinite future humans, then you must also permit the same analysis to apply to concrete help now. Perhaps your bed-net keeps someone healthy, who has a child, who mentors a student, who loans money to a friend, who founds a company, which employs a new grad, who goes on to reduce extinction risk. Choosing to prioritize concrete help now is not then because you believe this person has more moral worth than future humans, but because they have the potential to compound and help those future humans as well.2 Furthermore, I find it incredibly hubristic to imagine that you can estimate your tiny fraction of a percent impact on far future extinction risk with any degree of precision: surely that future human who is facing the risk more directly is better positioned to combat it, and your best move is to help them.3

As I saw the world, these two beliefs, longtermism and concern about AI extinction risk, were coupled pretty tightly, and people believed both or neither. Since I was not a longtermist, I therefore must not be concerned about AI leading to extinction some time in the long-term future. I did not consider the possibility that in fact, catastrophic harm caused by AI could occur not only within my lifetime, but within only a few years’ time. If that were the case, my generation would indeed be the right people to focus on this issue.

Rational Irrationality

Blaise Pascal famously argued that you should believe in god, not because of blind faith, but because of cold calculating rationality. If god exists, then believing brings you enormous benefit (eternal paradise), but even if god does not exist, then believing costs you nothing. In game theoretic terms, belief in god is a “dominant strategy”, which always performs at least as well as not believing.

Something similar, perhaps, could be said about belief in a high probability of doom brought about by recursive super-intelligence. If doomsday does not come, and you have been planning your life around doomsday (no need for savings beyond that point and similar), then you wind up in a bad position. In contrast, if doomsday does come, then all you get is to say “I told you so” before you go up in flames (or engineered virus) along with all the others. Most doomsday cults at least offer the assurance that the chosen believers will somehow be whisked to safety during The Rapture.

Sometimes I even added an extra row to this matrix, claiming that if beneficial aligned super-intelligence appeared in a non-disruptive way, then I along with my friends and family were already well positioned to benefit, and so it would provide diminishing returns to optimize my life towards an even better result in that case.4 In any scenario, planning around AI risk seemed like a bad strategy for me personally.

It similarly did not feel as if I was likely to have particular insight that would prove the difference between aligned and unaligned AI. In my view the sort of work done by AI safety researchers seemed woefully inadequate to the situation they described us as in, and perhaps inevitably doomed. After so many years of failing to gain more than trivial insight into the workings of our own minds, did they really expect alien minds to be somehow more comprehensible? Even if the doomsday scenario were real, it wouldn’t necessarily make the interpretability work it inspired valuable.

Of course, Pascal’s wager doesn’t imply that god exists, and neither does mine show that AI is safe. It only tells me strategically how I might respond to the world, and what I “should” believe, but not what is ultimately correct. As someone who generally prefers to have an accurate view of the world, I should have cared more about answering that second question and not allowed myself to stop at the first point.

Today

Although I have come around to the risks, I see some of these same patterns today. In people who see AI risk as something pushed by tech titans who have lied to them before, or as something they can have no influence over when their government seems ossified and concerned only with the interests of money and lobbying groups. I hope to treat these objections with grace, since just like my own, they may be rational responses, if not necessarily indicative of the underlying reality.

Footnotes

  1. As you can perhaps tell, I allowed myself to be fairly cloistered among CS majors in college, but I imagine the same thing can happen in other fields too. ↩

  2. Really, I find this belief almost necessary for any charitable giving, otherwise it would always seem preferable to let your money compound in the stock market for more years on the theory that it could save more people later. ↩

  3. A longer version of a similar argument can be found here: https://windowsontheory.org/2022/05/23/why-i-am-not-a-longtermist/. ↩

  4. Not to mention that it’s hard to optimize for, considering e.g. OpenAI’s famous disclosure that “it may be difficult to know what role money will play in a post-AGI world.” ↩

Comments

Loading comments…