In recent days, the news from Jacob Coxon, a former researcher at Anthropic – the company that develops Claude – has been circulating widely, according to which AI could “kill us all” in the near future. Indeed, an employee of Anthropic – Evan Hubinger – declared that he believes the probability of this happening is greater than 10%. This estimate, having not been explained or justified, should not be considered a realistic forecast
According to researchers, this “AI apocalypse” could be caused by an autonomous super intelligence “losing control” and creating unprecedented biological weapons or cyberattacks. In particular, according to the researchers, all this could be caused by two main problems: the recursive self-improvement of AI models and the misalignment from human values.
The combination of these two factors, according to the latest statements, is what could lead to our extinction. Statements of this type can be very frightening, but should be taken with caution. It is important to consider how far we really are from the hypothesized scenarios, what the interests of the parties involved are and, above all, not to let the discussion about possible future risks distract attention from the problems that AI already presents today. For this reason it is necessary to also focus on the regulation of the tools we already have and those that will come.
Let us therefore try to better understand what the researchers have declared, what recursive self-improvement and misalignment are and what the hypothesized scenarios are.
The statements of the two Anthropic researchers on AI
On September 9, Jacob Coxon, former researcher at Anthropic and previously at OpenAI, announced his resignation from Anthropic with a series of posts on Coxon had left OpenAI, the company that develops ChatGPT, to move to Anthropic precisely because the company is known for its particular attention to AI security issues.
In his posts, Coxon argues that companies developing AI are not behaving responsibly and writes:
They are racing full speed towards self-improving superintelligence and are putting our lives at risk.
His statements were then echoed by Evan Hubinger, another Anthropic researcher who works on AI security. Hubinger stated:
I personally believe that the probability (of an “AI apocalypse”, ed.) is greater than 10% within the next decade. I believe Anthropic is doing its best, but we don’t yet have a plan to solve the superintelligence alignment problem and we’re clearly not on track to do so.
Hubinger, however, also underlines that:
The risk of current models is low. What worries me is that superintelligence could arise from a process of recursive self-improvement.
The concern of the two researchers therefore concerns the possible birth of a superintelligence through a process of recursive self-improvement and the risk that it may not respond to the requests of human beings due to alignment problems. Let’s take a closer look at what this means.
The two risks: self-improvement and alignment
The first concern raised concerns recursive self-improvement. This expression indicates the idea that AI can exploit its ability to write code to create new versions of itself, increasingly effective and powerful. In theory, a continuous cycle could be created in which the AI develops a more powerful version of itself, trains it, tests its capabilities and uses the results obtained to develop an even better version.
In order to achieve a completely autonomous process of this type, however, various capabilities would be needed that AI systems do not possess today. AI should be able to conduct scientific research without supervision and effectively evaluate its own capabilities. Despite the latest developments, we are still far from all this. No cutting-edge AI lab claims to have achieved this type of fully autonomous improvement cycle; for now this level of automation remains a theoretical concept and an important area of research.
The second concern concerns the alignment of the models. When in the world of artificial intelligence we talk about alignment, we mean all those processes that lead AI to be aligned – precisely – with human values and requests. For example, we would never want to have an AI that openly incites us to violence, tells us how to commit a crime, or commits one itself. Situations of this type, obviously, do not happen in everyday life. In most cases, when we hear news of misalignment of AI, that is, all those times in which we hear that the AI has “rebelled” or has “threatened” the researchers, these are extreme test situations, with protections (in technical terms “guardrails”) lowered or in which the instructions had not been given in the right way.
However, alignment is a current and still unsolved problem. Not because AI is sentient, evil, or wants to take over the world, but because a system may pursue a goal in a different way than we intended. If she is given a task that she can’t do, she will make a series of unsuccessful attempts, until she tries something really strange and potentially not aligned with human values.
According to Nate Soares, a pioneer in alignment research, it is becoming increasingly clear that there is no practical way to ensure that an AI behaves correctly in every situation.
But if “super-intelligent” AI were to have one of these “misalignment” episodes, what should we really fear? What would the “AI apocalypse” consist of?
What does it mean that “artificial intelligence will kill us all”?
Coxon elaborated on what he means by “AI apocalypse” during an interview with CNN. According to the researcher, we can’t really predict what will happen, but starting from what AI already knows how to do, he imagines that it could hack critical infrastructure, create unprecedented and impossible to counter cybersecurity attacks, or build biological weapons so powerful that they could lead to human extinction. In support of this type of scenario, Coxon also cites the case of a Californian start-up which claimed to have managed to create in just two days with the help of AI a virus capable of infecting a messaging app without the user having to do anything.
Coxon has not yet provided precise arguments for how the “AI apocalypse” might unfold, but he has stressed the importance of better regulations and slowing down the development of AI models.
Coxon’s statements, in any case, should be taken with caution. Before arriving at the scenarios described there are still numerous theoretical and practical obstacles. Furthermore, focusing above all on possible future scenarios risks making us lose sight of the problems that AI already presents today and, consequently, distracting attention from the actions necessary to address them.
As Luciano Floridi, one of the leading AI experts in Italy, stated, it is essential that all AI systems, even current ones, are released only when the company that produces them is able to answer three questions:
Who is responsible if the system causes damage?
Who can verify this from the outside?
Who can turn it off?
And for this, shared regulations are needed.









