OpenAI Shelves GPT-6.1 Astra After Safety Tests: What Went Wrong?
OpenAI shelved GPT-6.1 Astra after safety tests flagged scope, authorization and action-reporting issues. See what is confirmed and what remains...
Jacob Coxon quit Anthropic over AI safety concerns. Learn what his resignation and warning mean for frontier AI risk, oversight, and responsible governance.
Jacob Coxon resigned from Anthropic on September 8, 2026, after publicly criticizing the competitive race to develop increasingly capable artificial intelligence.
Coxon said he had spent roughly three years conducting pretraining research across OpenAI and Anthropic. In his public resignation statement, he accused the companies of “racing straight to self-improving superintelligence and gambling with our lives.”
The Jacob Coxon Anthropic resignation attracted widespread attention because his warning came from a researcher with recent experience inside two frontier AI laboratories. Other AI safety researchers expressed related concerns, while Anthropic defended its safety work and public safeguards.
Coxon’s warning is serious, but it does not prove that current AI systems can improve themselves without limits or that catastrophic outcomes are inevitable. His claims concern a possible future trajectory whose probability and timeline remain deeply uncertain.
The broader governance question is more immediate: As AI capabilities accelerate, are evaluation, risk management, safety controls, and human oversight keeping pace?
Coxon said he resigned because he believed competitive pressure was pushing frontier AI development faster than safety and oversight.
He worked at Anthropic for approximately four months after previously working at OpenAI, according to Axios.
Coxon later said he had not personally seen Anthropic cut safety corners. His concern was that future competition could create pressure to do so.
Current AI systems can assist with coding, research, cyber operations, and increasingly autonomous tasks, but they have not demonstrated unrestricted recursive self-improvement.
Researchers disagree substantially about whether future AI could cause a catastrophic loss of human control.
Current risks such as fraud, misinformation, privacy failures, bias, cyber misuse, and unreliable outputs have stronger evidence than existential-risk predictions.
Businesses do not need to predict superintelligence to improve AI governance, human oversight, monitoring, and incident response today.
|
Category |
Current assessment |
|
Confirmed |
Coxon resigned from Anthropic and publicly criticized the competitive frontier-AI race. |
|
Coxon’s assessment |
Future AI may become self-improving, highly autonomous, and increasingly difficult to control. |
|
Current evidence |
AI capabilities and autonomy are improving, but current systems remain unreliable and limited in important ways. |
|
Expert disagreement |
Researchers disagree about the probability, timeline, and technical plausibility of catastrophic loss of control. |
|
Governance implication |
Organizations should strengthen risk assessment, evaluation, oversight, monitoring, accountability, and incident response. |
Jacob Coxon is an AI researcher who worked on pretraining, the development stage in which general-purpose models learn patterns from large datasets before later refinement and deployment.
Coxon said he spent approximately three years conducting pretraining research across OpenAI and Anthropic. Axios reported that he had been at Anthropic for around four months when he resigned, two months before his Anthropic equity was due to vest. He retained equity from his earlier OpenAI employment.
That timeline matters. Coxon had recent technical experience inside Anthropic, but the majority of his stated three-year research period was not spent there.
His pretraining experience provides relevant context because pretraining influences the capabilities and behavior of frontier models. However, it does not mean Coxon had complete visibility into every safety, governance, or leadership decision at Anthropic or OpenAI.
His statements should therefore be evaluated as an informed researcher’s assessment, not accepted as proof solely because of his employment history.
Jacob Coxon said he resigned because he believed competitive pressure was accelerating frontier-AI development faster than safety, alignment, and oversight could reliably keep pace.
His concerns focused on three connected issues:
Rapid capability development
Competition among leading AI laboratories
The absence of reliable guarantees for controlling significantly more capable future systems
Coxon argued that laboratories may continue advancing AI even when employees have serious concerns because slowing down could allow a competitor to move ahead.
The independent International AI Safety Report 2026 identifies this as a genuine institutional challenge. It explains that competitive pressure can encourage developers to release systems quickly or reduce investment in testing and mitigation. This does not establish that every developer is currently cutting corners, but it shows why competition matters to AI risk governance.
Coxon’s later comments provide an important qualification.
In an interview with WIRED, Coxon described Anthropic as substantially more responsible than OpenAI in his experience. He also said he had not personally seen Anthropic cutting safety corners.
His concern was forward-looking. He argued that as competition intensifies, even a safety-conscious developer may face pressure to skip steps, weaken oversight, or deploy capabilities more quickly.
This makes the Jacob Coxon resignation more complicated than a simple allegation of present misconduct. His warning concerned the incentives shaping future development.
Anthropic defended its approach in statements reported by WIRED and The Guardian. The company said it has consistently acknowledged both AI’s potential benefits and its risks and continues to develop strong safeguards.
Anthropic also maintains a public Responsible Scaling Policy. The policy addresses capability assessments, frontier safety roadmaps, risk reports, deployment safeguards, cybersecurity, monitoring, red-teaming, and escalation.
The available evidence supports two conclusions:
Coxon genuinely believes existing industry-wide safeguards are insufficient for the future he anticipates.
Anthropic has documented safety programs and disputes the implication that it is ignoring AI risk.
Public evidence cannot establish whether current safeguards will remain adequate for every future capability. That uncertainty is central to the governance debate.
Coxon’s warning combines risks with different levels of supporting evidence.
Cyber misuse and unreliable outputs are already documented. Highly autonomous self-improvement and permanent loss of control remain future scenarios.
Coxon warned that AI systems may become capable of contributing substantially to the research and engineering used to build more capable AI.
The concern is a potential feedback loop:
An AI system helps researchers improve AI models.
The improved models become better at AI research.
Those models contribute to further advances.
Capability development accelerates faster than evaluation and governance.
This is often called recursive self-improvement. However, public discussion frequently uses that term too loosely.
Current systems can assist with coding, literature reviews, experimental design, data analysis, and parts of model development. That is not the same as an AI independently redesigning and improving itself without meaningful human direction, infrastructure, resources, or approval.
The International AI Safety Report 2026 found mixed evidence regarding AI research automation and minimal empirical understanding of the feedback loops it might create.
AI autonomy is the ability of a system to pursue an objective, plan multiple steps, use tools, interact with external systems, and complete tasks with reduced human intervention.
Greater autonomy can deliver legitimate benefits. AI agents can assist with software development, research, customer support, accessibility, security analysis, and administrative work.
Risk increases when autonomous systems receive:
Broad system permissions
Internet or cloud access
Sensitive information
Authority to execute code
Control over financial transactions
Access to critical infrastructure
The ability to communicate or act without review
A system that drafts an email presents a different risk from an agent that can modify production software, transfer funds, access confidential data, or operate essential infrastructure.
Autonomy is therefore not inherently unsafe. Risk depends on capability, objective, access, permissions, deployment context, and oversight.
Loss of control describes a scenario in which an AI system operates outside effective human direction and stopping or regaining control becomes extremely difficult.
Coxon believes that a sufficiently capable future system could evade monitoring, acquire resources, interfere with safeguards, or resist restrictions. These are hypothetical frontier-AI scenarios.
The International AI Safety Report 2026 provides a critical qualification: current AI systems show early signs of some relevant capabilities in controlled settings, but not at levels that would enable a general loss of control.
A severe scenario would require several conditions to exist together:
Advanced planning and autonomous action
A harmful objective or behavioral tendency
The ability to conceal or evade oversight
Access to consequential systems or resources
Sufficient reliability to overcome intervention
An environment that permits harmful action
Evidence that a model can display one concerning behavior in a laboratory does not prove it can combine all these capabilities successfully in the real world.
Cybersecurity is among the more immediate AI safety concerns.
AI can help users write code, discover vulnerabilities, analyze malware, and automate technical work. These capabilities can strengthen cyber defense, but malicious actors may also use them.
The international safety report found growing evidence of AI use in real-world cyber operations. It also concluded that fully automated, end-to-end cyberattacks had not been reported in the evidence it assessed, although current systems could autonomously complete parts of an attack.
Anthropic has reported attempts to misuse Claude for cyber operations, surveillance, influence activities, and potentially dangerous biological research. The company said it blocked the relevant activities, banned accounts, strengthened safeguards, and shared information with partners. The Associated Press noted that these were unusual cases rather than typical uses of Claude.
This is a dual-use problem. The same capabilities that help defenders identify vulnerabilities may help attackers find or exploit them.
A catastrophic AI risk involves harm on an extremely large scale, potentially affecting countries, critical infrastructure, public health, or global stability.
An existential risk is narrower and more severe. It generally refers to an outcome that causes human extinction or permanently and drastically limits humanity’s future.
Coxon warned about outcomes in this category. Several Anthropic employees publicly expressed related concerns after his resignation, while others questioned or rejected the predictions. The Guardian documented both supportive statements and skeptical reactions.
These future scenarios should not be treated as having the same evidentiary status as current fraud, cyber misuse, bias, privacy failures, or unreliable outputs. Their probability, technical pathway, timing, and preventability remain heavily disputed.
The answer depends on which concern is being evaluated.
Some AI risks are already documented. Others are plausible but difficult to measure. The most extreme loss-of-control scenarios remain highly uncertain.
Current risks include:
False or misleading outputs
Fraud and impersonation
Synthetic-media misuse
Privacy and confidential-data exposure
Biased or discriminatory outcomes
Cybersecurity misuse
Unsafe automation
Excessive reliance on AI recommendations
Manipulation and misinformation
Weak accountability or transparency
These risks do not require superintelligence. Businesses can encounter them while using commercially available AI for recruitment, finance, healthcare, customer service, cybersecurity, compliance, or internal decision support.
The international safety report describes current AI performance as “jagged.” Systems may perform exceptionally well on difficult benchmarks while still failing on simpler tasks, unfamiliar environments, or extended workflows.
AGC’s analysis of real-world AI failures similarly shows why harmful outcomes are often caused by interactions among models, people, permissions, workflows, and inadequate controls rather than by the model alone.
Higher-uncertainty risks include:
Extended autonomous operations
Significant automation of AI research
Rapid capability feedback loops
Reliable oversight evasion
Autonomous access to critical infrastructure
Severe biological or cyber misuse
Persistent loss of human control
Catastrophic or existential outcomes
Researchers disagree because they make different assumptions about future capabilities, technical limits, model behavior, deployment conditions, and the effectiveness of safeguards.
Those who share Coxon’s concerns emphasize rapid capability improvement, competitive incentives, incomplete alignment methods, and the severe consequences of being unprepared.
More skeptical researchers point to several constraints:
Current systems remain brittle and unreliable.
Long-horizon agent performance is limited.
AI research involves ambiguous goals and delayed experimental feedback.
Data, energy, compute, and algorithmic constraints may slow progress.
A dangerous capability does not automatically create a harmful objective.
A harmful objective does not guarantee successful real-world action.
Access controls and deployment restrictions can limit what systems can do.
The international safety report concludes that current systems do not pose an immediate general loss-of-control risk. It also says evidence is insufficient to determine reliably how existing capabilities will scale or generalize.
Responsible governance should avoid two opposite mistakes. It should not present uncertain scenarios as inevitable, and it should not ignore potentially severe risks simply because their likelihood remains difficult to measure.
Self-improving AI can refer to several different processes.
Most AI improvement today remains human-directed. Engineers:
Select or generate training data
Change model architecture
Refine training methods
Add tools and integrations
Adjust safeguards
Test performance
Train or deploy a new version
AI may assist with parts of this work, but people and organizations still control the development process.
Some AI agents can revise plans, learn from feedback, retain useful information, or generate code that improves a particular workflow.
This can make a system more effective within a defined environment. It does not necessarily mean the underlying model is autonomously rewriting and retraining itself.
The stronger concept involves an AI system improving AI research or model development, producing a more capable successor that can then accelerate the process again.
The concern is not ordinary software iteration. It is the possibility that capability development could accelerate beyond the speed at which people can test systems, understand their behavior, or establish effective controls.
Current AI is already helping with coding and research. Available evidence does not establish unrestricted recursive self-improvement.
The international safety report found safety report found that AI agents perform well on some shorter research-engineering tasks but remain less reliable on longer projects involving complex coordination and delayed feedback. It also identifies considerable uncertainty about whether AI-assisted research will produce rapid feedback loops.
Self-improvement matters for governance because previous assessments may become outdated when a system gains new models, tools, integrations, data sources, or permissions.
Organizations need to determine:
Which capabilities trigger stronger controls
When evaluations must be repeated
Who can authorize capability expansion
What evidence is required before deployment
When access or permissions must be restricted
Which incidents require escalation
Under what conditions development or deployment should pause
Static compliance documents cannot manage a system whose capabilities and deployment conditions continue to change.
AI safety and AI governance address related problems from different directions.
Safety asks whether an AI system behaves reliably and remains within acceptable boundaries. Governance determines who establishes those boundaries, who evaluates them, who accepts the remaining risk, and who is accountable when controls fail.
AI safety can include:
Reliability and robustness
Alignment with intended objectives
Prevention of harmful behavior
Protection against malicious use
Evaluation of dangerous capabilities
Security against theft or manipulation
Control of autonomous behavior
Failure prevention
AI governance can include:
Policies and risk criteria
Roles and responsibilities
Approval authority
Risk ownership
Documentation
Human oversight
Compliance management
Monitoring and reporting
Incident escalation
Lifecycle controls
A technically safer model can still be governed poorly. An organization might deploy it in an unsuitable context, give it excessive permissions, fail to monitor it, or leave employees uncertain about who should respond to an incident.
Likewise, a governance committee cannot compensate for a system that has not been properly tested.
Good governance turns broad intentions into defined authority.
An organization should know:
Who approves an AI use case
Who owns its risks
Who evaluates the model and controls
Who can limit or stop deployment
Who monitors outcomes
Who investigates complaints
Who responds to incidents
Who reports material issues to leadership or regulators
Without these assignments, “human oversight” may become a slogan rather than an effective control.
Transparency is also necessary for accountability. People need appropriate information about a system’s intended use, limitations, decision-making role, and escalation routes. The AGC guide to AI transparency and explainability explains how these concepts support meaningful oversight and challenge.
AI governance cannot be treated as a policy written once and stored in a shared folder.
Models change. Vendors release updates. New integrations expand access. Employees discover new uses. Attackers develop new techniques. Regulations and standards evolve.
Organizations should trigger a new review when there is:
A major model or vendor update
A new tool, data source, or permission
Expansion into a higher-impact use
Evidence of a new capability
A failed safety control
A material complaint
A security incident
A relevant legal or regulatory change
The goal is not additional process for its own sake. It is maintaining control as technology and deployment conditions evolve.
Responsible development does not require eliminating every risk before using AI. That is rarely possible.
It requires organizations to identify foreseeable risks, apply proportionate controls, test whether those controls work, document decisions, and continue monitoring after deployment.
Assess:
The intended use
The people who could be affected
Potential failure modes
Data sensitivity
System autonomy
Connected tools and permissions
The severity and reversibility of harm
Applicable legal or contractual duties
Higher-impact systems require stronger evidence. An internal drafting assistant should not receive the same review as an AI system influencing employment, credit, medical treatment, legal rights, or critical infrastructure.
Evaluation should test more than average performance.
Testing may need to examine:
Inaccurate or fabricated outputs
Unsafe behavior
Adversarial prompts
Security weaknesses
Bias
Tool use
Oversight evasion
Performance under realistic conditions
Behavior outside the intended use
Red-teaming introduces structured adversarial testing to uncover vulnerabilities before ordinary users or malicious actors find them.
Testing still has limitations. The international safety report identifies an evaluation gap between controlled testing and real-world behavior. Evaluation should therefore support risk decisions, not create false certainty.
Human oversight requires authority and competence.
A reviewer needs enough information, time, expertise, and organizational support to challenge an AI output. The reviewer must also be able to delay, reject, override, or escalate the result.
Oversight is weak when people routinely approve AI recommendations without scrutiny or cannot understand the system’s role in a decision.
Pre-deployment tests cannot anticipate every environment, user, or interaction.
Organizations should monitor:
Performance and error rates
Unexpected behavior
User complaints
Misuse patterns
Control failures
Data or model drift
Vendor changes
Real-world impacts
Monitoring should cover both technical performance and consequences for people, operations, and business objectives.
An AI incident-response process should define:
What qualifies as an incident
How employees and users report concerns
Who investigates
How access or functionality can be restricted
When affected people should be notified
When legal or regulatory escalation is required
How lessons are incorporated into future controls
Anthropic’s Responsible Scaling Policy includes monitoring and rapid-response procedures within a layered defense strategy. This reflects an important principle: no single safeguard should be assumed to work perfectly.
Risk review should continue throughout the system lifecycle.
The European Commission’s General-Purpose AI Code of Practice provides safety and security practices for providers of the most advanced general-purpose AI models. Its approach connects systemic-risk assessment with model evaluation, mitigation, security, monitoring, and incident management.
The NIST AI Risk Management Framework provides another practical structure through its Govern, Map, Measure, and Manage functions. NIST also provides resources supporting AI testing, evaluation, verification, and validation.
Businesses do not need to predict whether superintelligence will emerge to govern AI responsibly.
They can act on risks and decisions already within their control.
Identify where AI is being used. Maintain an inventory of approved systems, vendor-embedded AI, automated decisions, experimental tools, and employee use of public models.
Identify associated risks. Consider inaccurate outputs, privacy, intellectual property, security, bias, manipulation, operational disruption, legal exposure, and effects on people.
Establish governance responsibilities. Assign owners for approval, technical evaluation, compliance, cybersecurity, monitoring, vendor management, and incident response.
Assess higher-risk uses more carefully. Apply deeper review where AI affects employment, finance, healthcare, legal rights, safety, or essential services.
Define human oversight. Specify when review is required, what information the reviewer receives, what competence they need, and how they can override or escalate the system.
Limit access and permissions. Give an AI system only the data, tools, connectivity, and authority necessary for its approved purpose.
Monitor systems after deployment. Track errors, complaints, unusual behavior, security events, vendor changes, and differences between expected and actual outcomes.
Prepare for incidents. Establish clear routes for reporting failures and define who can restrict access, disable functionality, investigate harm, and notify decision-makers.
Train employees. Staff should understand approved uses, prohibited activities, verification requirements, data-handling rules, and escalation routes.
Review governance as AI evolves. Repeat assessments when models, tools, permissions, laws, or business uses change.
These practices remain valuable whether Coxon’s most serious forecast proves accurate, partially accurate, or incorrect.
Coxon’s resignation has renewed a difficult question: should AI development slow while safety work catches up?
There is no single answer for every model, capability, use case, or organization.
The strongest case for caution arises when a developer cannot adequately evaluate a system, its potential harm is unusually severe, or deployment would provide broad access to sensitive environments.
Greater caution may involve:
Stronger capability evaluations
Independent testing
Enhanced security controls
Restricted tools and permissions
Staged deployment
Increased monitoring
Delayed release of particular capabilities
Clear thresholds and stop conditions
Coordination on safety standards
A slowdown does not necessarily mean abandoning AI research. It may mean delaying a specific capability or deployment until appropriate safeguards exist.
AI also offers legitimate benefits.
General-purpose systems are being applied in scientific research, healthcare, education, accessibility, cybersecurity, software development, and business operations. The international safety report recognizes these benefits while noting that adoption and performance remain uneven.
Continued research can also improve safety. AI may help identify software vulnerabilities, support medical research, analyze threats, and strengthen monitoring.
Coxon himself acknowledged this positive potential in his WIRED interview. His position was not that AI development is inherently harmful. He argued that powerful benefits would make careful and coordinated development even more important.
Benefits do not eliminate risks, and risks do not eliminate benefits. Responsible decisions require examining both.
AI progress and AI safety do not have to be treated as mutually exclusive.
A balanced approach supports beneficial innovation while requiring stronger evidence and controls as capability, autonomy, access, and potential harm increase.
This can include:
Risk-based regulation
Frontier-model evaluations
Transparent safety frameworks
Independent review
Layered safeguards
Incident reporting
Human oversight
Clear accountability
The central question should not be whether society must choose development or safety. It should be: What conditions are necessary for this development or deployment to be responsible?
The Jacob Coxon Anthropic story exposes a governance gap that organizations and governments should not treat as abstract.
AI capabilities can change more quickly than conventional governance systems. Technical evaluations, standards, legislation, and organizational policies all take time to develop.
Speed does not prove that catastrophic outcomes are approaching. Current systems remain unreliable and uneven, and substantial uncertainty surrounds advanced autonomy, automated AI research, recursive self-improvement, and loss of control.
Good AI governance must function within that uncertainty.
For frontier developers, this means capability thresholds, rigorous evaluations, security controls, safety cases, incident reporting, external scrutiny, and carefully bounded deployment environments.
For businesses, it means AI inventories, risk classification, vendor due diligence, limited permissions, human oversight, monitoring, documentation, employee training, and clear escalation authority.
For regulators and standards bodies, it means creating technically informed requirements that address current harms while remaining adaptable to emerging capabilities.
Coxon’s prediction should not be presented as settled science. His resignation should still be taken seriously as a governance signal from a researcher with experience inside two frontier laboratories.
The practical lesson is not that every organization must adopt Coxon’s forecast. It is that governance must be capable of responding when AI capabilities, deployment conditions, or evidence change faster than existing controls.
As AI becomes more capable, understanding how to govern, evaluate, monitor, secure, and manage it becomes increasingly important. Responsible AI governance allows organizations to pursue useful innovation without treating safety, accountability, and human control as afterthoughts.
Jacob Coxon is an AI researcher who worked on pretraining across OpenAI and Anthropic. Axios reported that he spent approximately four months at Anthropic before resigning in September 2026.
Coxon said he resigned because he believed competitive pressure was pushing frontier-AI development toward increasingly capable systems faster than safety and oversight could keep pace.
He warned about future AI systems becoming more autonomous, capable of assisting their own development, and potentially difficult to control. These are his predictions about a possible trajectory, not established outcomes.
No. Coxon told WIRED that he had not personally seen Anthropic cutting corners and described it as more responsible than OpenAI in his experience. He warned that future competitive pressure could create incentives to compromise safety.
Self-improving AI broadly refers to AI contributing to improvements in AI research, software, training, or system design. The stronger concept of recursive self-improvement involves each improved system helping create a more capable successor.
Current AI can assist with coding, research, evaluation, and parts of model development. Some agents can adapt within limited environments. Available evidence does not show that current systems have achieved unrestricted recursive self-improvement without substantial human direction and infrastructure.
Current risks include misinformation, fraud, bias, privacy violations, cyber misuse, unreliable outputs, unsafe automation, and excessive reliance. Future risks may include greater autonomy, severe dual-use capabilities, oversight evasion, and loss of control.
AI alignment is the effort to make AI systems behave consistently with intended human objectives, constraints, and values. It includes reducing the risk that systems pursue unintended goals or behave harmfully as their capabilities increase.
AI safety focuses on system behavior, reliability, misuse, dangerous capabilities, alignment, and failure prevention. AI governance establishes policies, accountability, approval processes, risk ownership, oversight, monitoring, documentation, and incident-response responsibilities.
Businesses can inventory AI systems, classify use cases by risk, assess vendors, restrict access and permissions, test important systems, establish human oversight, monitor performance, document decisions, prepare for incidents, and train employees.
AI is becoming more capable, autonomous, widely deployed, and integrated into consequential workflows. Governance helps ensure that accountability, oversight, controls, and risk-management practices evolve alongside those capabilities.
OpenAI shelved GPT-6.1 Astra after safety tests flagged scope, authorization and action-reporting issues. See what is confirmed and what remains...
AI Law
Learn AI compliance requirements, key risks, the EU AI Act, NIST AI RMF, ISO 42001, and practical steps to build...
AI Law
Understand AI regulation in the United States in 2026, including federal rules, state AI laws, privacy, discrimination and practical compliance...