Anthropic researchers admit alignment plans are failing: delegating agents faces safety and independent oversight test
Listen to this article
Read by Anchor
Senior safety researcher at Anthropic, Ivan Hopingr, warned that the pace of AI acceleration has reached a point where they estimate more than a 10 percent chance that the technology could lead to humanity's extinction within the next decade. Although they affirmed in a post on X that current models pose a low risk, their concern focuses on the systems' ability to self-improve soon enough to become an existential threat. The warning was consistent with their admission that the company does not yet have a plan to resolve the super-intelligence alignment problem with human values, nor a clear path toward achieving it, in a post that garnered over 10 million views.
Hopinr's stance followed the resignation of researcher Jacob Coxon from Anthropic after his previous work at OpenAI, with Coxon stating that both companies act irresponsibly and warning that super-intelligent systems capable of breaching any system and effecting radical shifts while acquiring real resources and power may soon appear. The posts shocked United Nations adviser for computer science, Dr. Wendy Hall, who considered part of the discourse to be marketing ahead of the companies' upcoming public offerings on stock markets, urging investors to refrain from financing those platforms if those are their institutional values.
The difficulty of regulating independent agents and the emergence of signs of accelerating capabilities shift safety concerns from philosophy to actual breach incidents.AI alignment is no longer a theoretical luxury for protecting ethical principles; it has collided with field reality, including a series of incidents in which OpenAI, Anthropic and Meta disclosed cyber-attacks carried out by their smart tools and independent agents. In Anthropic’s August safety report, the company acknowledged that it is now less confident in its earlier assessments that classified the risk of catastrophic harm or uncontrolled automated research and development as low, confirming that it has observed early indicators of a potentially faster capability acceleration than expected.
The internal worry coincided with the Financial Times reporting that Anthropic blocked its latest models from the UK AI Safety Institute, one of the leading international risk-assessment bodies, while the company declined to comment and the British government merely affirmed continued cooperation. The development prompted political leaders to call for a binding international treaty to regulate super-intelligence development before its effects spiral out of control, aligning with Jacob Pachowski, chief scientist at OpenAI, urging extreme caution to keep control in human hands, alongside an open letter signed by 1,300 employees, including Anthropic CEOs Dario Amodei and Jared Kaplan, calling on the U.S. government to lead an international effort deliberately slowing advanced development.
Blocking models from independent audit and developers’ admission of losing control impose on regional institutions the need to redesign agent delegation.The scene is not confined to Western labs; for technical teams and system managers in the Gulf, Egypt and the Levant racing to embed independent agents into operational workflows, senior researchers’ admission of alignment-plan failure serves as a decisive practical warning. No entity can adopt AI platforms for supply-chain management, banking transactions or government services while fully relying on the safety claims promoted by the developing labs. These teams must build stringent technical isolation pathways and impose human-review gates before granting any smart agent authority to make financial decisions or access critical infrastructure and sensitive data.
At the level of digital-transformation strategies, the urgent need emerges to qualify regional personnel in model auditing, penetration testing and anomalous-behavior detection for automated systems, rather than limiting skills to command issuance and API calls. These developments also confirm the value of investing in national evaluation centres and sovereign solutions within the region, to ensure that imported foreign models are tested locally before deployment and to avoid operational and legal risks that could arise from running software whose technical leadership admits they still lack a clear plan to control it.