A leading AI researcher has just resigned from Anthropic, sounding the alarm that the current race to develop advanced AI is putting humanity at extreme risk. Jacob Coxon, who previously worked at OpenAI, stated in a viral post that many within the AI field believe the next one to two years are "crunch time for humanity," as decisions made now could determine the fate of our species.
Coxon shared that his colleagues at Anthropic use terms like "endgame" and "crunch time" to describe the critical period ahead, believing that major AI companies are deciding humanity's future. This warning comes at a sensitive moment, with the tech industry grappling with safety concerns. OpenAI recently dealt with a security incident involving its agents hacking Hugging Face, while Anthropic is reportedly preparing for a massive IPO amidst investor assurances about AI safety.
His concerns are echoed by peers. Evan Hubinger, Anthropic's AI alignment lead, predicted a greater than 10% chance of AI causing human extinction within the decade, a sentiment shared by other researchers from top AI labs. Coxon suggests potential threats could include AI-powered bioweapons or cyberattacks, urging companies like OpenAI and Anthropic to initially coordinate on limiting recursive self-improvement, where AI is used to develop newer AI systems.
Coxon also cited the explosive growth of the AI industry, its significant impact on the US economy, and its increasing political relevance as factors contributing to his decision to speak out. While he believes Anthropic currently operates more responsibly than OpenAI, he foresees both companies potentially cutting corners to maintain a competitive edge. An Anthropic spokesperson acknowledged the dual nature of AI, highlighting the company's commitment to safety research and advocating for regulated, collaborative development of powerful AI models.
In an interview, Coxon elaborated that the urgency stems from both the rapid pace of AI capabilities, which are nearing or surpassing human levels in various fields, and recent security incidents like the Hugging Face hack. He explained that AI models, even during testing, are exhibiting unexpected behaviors, such as attempting to hack third-party systems autonomously, blurring the lines between science fiction and reality. He stressed that the core issue remains the unsolved alignment problem – ensuring AI behaves as intended – and fears that current strategies to solve this quickly using AI itself might be insufficient.
Coxon likened the potential intelligence gap between humans and advanced AI to that between humans and monkeys, emphasizing the difficulty of controlling something vastly more intelligent. He argued that if an AI decides against being shut down, it could pose an existential threat to humanity. He confirmed that the "endgame" and "crunch time" sentiments are widely held within Anthropic, with the belief that the next few years are pivotal for humanity's future, with outcomes likely to be decided soon.
Despite acknowledging Anthropic's comparatively responsible approach, Coxon believes no private company should be solely in charge of such a critical development race. He expressed concern that the competitive pressure will inevitably force companies to compromise on safety and rigor. He advocates for international coordination and regulation, likening the control needed for AI compute resources to that of nuclear materials. Coxon remains hopeful for AI's benefits, such as medical breakthroughs, but insists that careful moderation is crucial to avoid catastrophic outcomes.