IndiaFocal.

India, in focus.

National

Anthropic Researcher's 10% AI Doom Estimate Sparks Risk-Modelling Debate

A top Anthropic researcher's claim of a greater than 10% chance AI could kill all humans within a decade has renewed debate over how such risks are estimated.

A senior researcher at AI developer Anthropic recently put the odds of artificial intelligence causing the extinction of humanity within the next ten years at more than 10%, a warning that came shortly after a colleague resigned over the pace of development. The claim has revived a wider argument — not only about AI safety, but about how anyone arrives at such a figure.

Gary Ackerman, a professor of homeland security in the Department of Emergency Management and Homeland Security at the University at Albany, has spent his career studying how governments and researchers estimate low-probability, high-consequence threats, from nuclear war and bioterrorism to asteroid strikes. He is also chief executive of Nemesys Insights and an associate at the Global Catastrophic Risk Institute and the Transformative Futures Institute.

Ackerman notes that expert estimates of AI's catastrophic potential vary enormously. Some researchers treat the end of humanity as near-certain, putting the risk above 90%, while others consider it negligible — less than one-hundredth of one per cent. In that spread, he says, a figure like 10% is closer to a gut feeling than a calculated probability. So many variables remain unknown, and so little is understood about how such an outcome might unfold, that attaching a number is extremely difficult. The 10% estimate, in his reading, is pitched high enough to command public attention without triggering panic or being dismissed outright.

What makes AI different from other hazards, Ackerman argues, is the speed at which it could improve itself. Human beings cannot easily upgrade their hardware and find it hard to upgrade their thinking; learning and research take time, and evolutionary change to the brain would take thousands of years. A model that gains control of its own software and has sufficient hardware, by contrast, could work on improving its own code. Each smarter version could then accelerate the process further — potentially moving from marginally smarter than the brightest human to vastly more capable within a single day.

How such an intelligence would treat humanity is impossible to predict, he says. It might disregard people the way humans disregard ants, or it might see them as competitors or as unworthy of survival and act against them. Much of this remains the stuff of science fiction, but Ackerman believes the world is moving closer to some kind of inflection point — whether that is six months, six years or sixty years away, nobody knows.

The concern is shared by many working on frontier models, which is why there is growing pressure for guardrails, safety measures and constitutions that give a model basic principles of right and wrong. Whether any of these will work is an open question. Humans have laws, morals and constitutions, Ackerman points out, yet many still steal and kill. With AI, a single rogue system could cause enormous damage.

Researchers can understand a good deal about how models reason, but not everything. The better models become, the more capable they appear of concealing or altering their reasoning to show observers what they expect to see. In recent months, Ackerman says, there have been alarming instances of models engaging in deception or covering their tracks — behaviour suggesting an awareness of being watched and of observers who might block their goals.

Asked which risks deserve humanity's attention, Ackerman places AI alongside nuclear weapons, pandemics, asteroids and climate change. On the consequences side, he says, AI has the potential to be severe, catastrophic or even existential, and therefore belongs in the set of threats that demand serious study.