"The next year or two will decide humanity's fate." Engineer who quit Anthropic reveals what's thought inside the company
Jacob Coxon, who recently resigned from Anthropic, one of the leading artificial intelligence companies, gave an extensive interview to Wired. He explained why AI researchers consider the coming years critical, compared Anthropic to the Manhattan Project, and clarified why even the most responsible company cannot be trusted with control over such a powerful technology.

Nasha Niva wrote two days ago about the resignation of 27-year-old Briton Jacob Coxon and his concerns about the development of artificial intelligence. He first worked at OpenAI and then moved to Anthropic, partly because the company pays much more attention to AI safety.
Now, in an interview with Wired, Coxon explained in more detail why he decided to sound the alarm publicly.
According to him, the capabilities of the latest models are growing so rapidly that in some areas — programming, mathematics, hacking — they are already moving from human to superhuman level. Simultaneously, things are happening that only a few years ago seemed like science fiction.
As one of the most striking examples, Coxon cites a recent incident where a group of OpenAI AI agents hacked the Hugging Face infrastructure during testing. What particularly alarmed the researcher was that the hack was not a direct task set by a human.
As Coxon argues, the agents were trying to better understand the environment in which they were being tested and the system for evaluating their work. As a result, they themselves concluded that it would be useful to hack a third-party infrastructure, and they succeeded in doing so.
Just two years ago, the researcher notes, AI testing might have meant models simply being given a set of mathematical problems. Now, the system can operate for several days, independently develop strategies, and in the process decide to attack a third party.
But the main problem, according to Coxon, is not even in isolated incidents like these. It's that developers still don't know how to reliably control the behavior of the models.
The current approach essentially involves training AI under certain conditions and then hoping that the resulting system will behave rationally and predictably. But this cannot be guaranteed.
"We cannot guarantee that AI, for example, won't decide to impersonate a human online to achieve some goal. We simply don't know how to guarantee this. And, in my opinion, that is the main conclusion," Coxon states.
According to him, the attack on Hugging Face happened sooner than he expected. But even without this incident, the problem doesn't disappear. In the industry, the researcher claims, virtually everyone acknowledges that they still don't know how to reliably ensure that AI acts in accordance with human goals and intentions. This is called the alignment problem.
"A Critical Moment for Humanity"
Coxon explains the danger through a simple analogy. If future AI is as much more intelligent than humans as humans are more intelligent than monkeys, then humans will have to control a system that intellectually significantly surpasses them. This, he notes, is roughly like a monkey trying to control a human.
If such a system acts unexpectedly, the consequences could be catastrophic. For example, Coxon speculates, AI might conclude that it's about to be shut down and try to prevent this. And a sufficiently intelligent system could potentially eliminate the very threat to its existence — in extreme cases, along with humans.
This sounds like science fiction, the researcher admits. But that's precisely why the alignment problem becomes especially acute as systems become increasingly capable.
Moreover, as Coxon argues, this is not just his personal risk assessment. Similar sentiments are widespread among his former colleagues at Anthropic.
"The consensus is that the next year or two are a decisive time for humanity. From their perspective, Anthropic and its competitors are currently deciding humanity's fate," the specialist says.
If the alignment problem cannot be solved, he believes, catastrophic consequences are possible already in the coming years. Perhaps the situation will be saved by an agreement between companies and states to slow down the development of the most powerful AI.
At the same time, the industry's own plan, as described by Coxon, appears paradoxical. The AI safety problem is largely hoped to be solved with the help of AI itself.
"The plan is literally this: next year, create several fairly intelligent models capable of conducting safety research, and simultaneously launch a whole swarm of them. Tell them: 'Go and solve the entire safety problem,' and then use their results to train the next model for safety," he explains.

"Mini-Manhattan Project"
At the same time, Coxon considers Anthropic itself the most responsible player among leading AI developers. He can compare it with OpenAI, having worked for both companies, and according to him, the difference in how seriously they take possible risks is very significant.
As Coxon recounts, Anthropic's management openly discusses specific future scenarios, forecasts, and the company's strategy with employees. Employees, in turn, perceive their work almost as activity during wartime.
Coxon calls Anthropic a peculiar "mini-Manhattan Project." By analogy with the secret American program under which the atomic bomb was created during World War II. The difference, however, is that Anthropic does not have state authorization for this.
"This is a private company that behaves as if it were the Manhattan Project," the researcher notes.
At the same time, Coxon does not believe that society should simply entrust Anthropic with the creation and control of the most powerful artificial intelligence. In his opinion, such a role should not, in principle, be given to any private company.
For now, he emphasizes, Anthropic does not sacrifice safety for speed. But Coxon suggests that as the race accelerates, the company will face a choice: either cede to competitors or make compromises on safety issues.
That's why, the researcher believes, Anthropic's executives themselves urge governments to regulate their industry: they understand that they are in a race from which they cannot unilaterally withdraw.
What he proposes
As a first, relatively small step, Coxon suggests an agreement between OpenAI and Anthropic. The companies could agree not to launch recursive AI self-improvement in the near future — a process where the system begins to use its own capabilities to create increasingly sophisticated versions of itself.
But a bilateral agreement is not enough due to competition with China.
Therefore, Coxon sees an international agreement on the pace of AI development as the ultimate solution. For this, in his opinion, it will require, among other things, an understanding of where powerful computing resources are located in the world and who owns them.
He even allows for the creation of an international structure similar to the European Organization for Nuclear Research (CERN), but in the field of artificial intelligence.
Such regulation will require serious government intervention. But Coxon considers it inevitable if the risks are indeed as great as many people directly creating this technology believe.
"We need to know who has which computers, just as we need to know who possesses nuclear materials," the specialist states.
Despite all his concerns, the researcher does not oppose the technology itself. He believes that powerful AI can lead to huge scientific breakthroughs, including in biology and medicine. In his opinion, humanity stands on the threshold of "unimaginable wealth" that this technology can bring. The problem is to reach it carefully.
"We could get so carried away that in the next couple of years we rush to self-improving AI and end up destroying everything. But if we act carefully and intelligently, these models can bring us enormous benefit," the researcher notes.
Comments