OpenAI Shelves GPT-6.1 Astra After Safety Tests: What Went Wrong?
OpenAI shelved GPT-6.1 Astra after safety tests flagged scope, authorization and action-reporting issues. See what is confirmed and what remains...
Learn how AI risk controls turn identified risks into practical safeguards through clear ownership, testing, monitoring, and residual-risk evaluation, helping organizations strengthen AI governance and manage changing risks throughout the AI lifecycle with greater accountability.
AI risk controls turn identified risks into safeguards that can be assigned, tested, and monitored. Identifying an AI risk does not control it, and deciding to mitigate that risk does not prove the selected safeguard will work when an AI system meets real users, changing data, and operational pressure.
AI risk controls are technical, organizational, procedural, human, or other safeguards designed to prevent, detect, reduce, contain, or respond to identified AI risks. They translate risk-treatment decisions into operational action.
For every material control, an organization needs a traceable connection between the risk, the control objective, the accountable owner, evidence that the safeguard operates, and the residual risk that remains. Controls should also be proportionate to the system and its consequences. Applying every available safeguard to every AI use case creates complexity, not necessarily protection.
AI risk controls translate risk-treatment decisions into practical safeguards.
Every control should address a specific risk, cause, exposure, or consequence.
Material risks often require technical, organizational, and human controls together.
A documented or deployed control is not automatically an effective control.
Controls need clear objectives, owners, evidence, testing, and monitoring.
Residual risk must be evaluated after relevant controls are implemented and validated.
An AI risk control is a safeguard implemented to modify an identified risk or improve prevention, detection, response, or recovery. Controls are part of broader AI risk management, which also covers identification, assessment, treatment, governance, and review.
An AI risk combines an uncertain event, condition, or scenario with its potential consequences. A mitigation strategy decides what to do. A control is the safeguard that implements that decision.
Consider an AI-supported decision tool that may produce unreliable recommendations. The treatment objective is to reduce the possibility that unreliable outputs influence consequential decisions. Relevant controls could include mandatory human review, validation rules, restricted automation, and defined escalation criteria.
The NIST AI Risk Management Framework Core calls for internal controls to be documented, their effectiveness assessed regularly, and residual risks recorded. A controls list is therefore insufficient without evidence that each safeguard addresses the intended risk.
The following five categories offer a practical way to organize AI controls. They are not a taxonomy mandated by NIST or ISO. Controls may be preventive, detective, or corrective, and material risks often require a combination.
|
Control category |
Main purpose |
Example |
Evidence to check |
|
Governance and process |
Structure decisions and accountability |
Approval or escalation requirement |
Approval and exception records |
|
Technical |
Restrict or modify system behavior |
Access or automation limits |
Configuration and test evidence |
|
Data and model |
Manage data or model-related risk |
Validation and quality checks |
Evaluation results |
|
Human oversight |
Enable meaningful intervention |
Review before consequential action |
Review and override records |
|
Monitoring and response |
Detect changes, failures or incidents |
Alerts and incident escalation |
Monitoring and incident records |
These controls structure AI decisions. Examples include policies, approvals, assigned ownership, supplier review, deployment restrictions, exception management, escalation, and documentation requirements.
Technical controls constrain system behavior through measures such as access restrictions, permission limits, output filters, automated validation, automation limits, and fail-safe mechanisms. Each control must target a documented risk pathway.
Data-quality checks, input restrictions, model evaluation, performance thresholds, and change controls can address data or model risks. Evidence should reflect deployment conditions, not only laboratory performance.
Human controls include approval before consequential action, output review, escalation, overrides, and defined authority. Effective oversight requires enough information, time, competence and authority to question or stop an output. Merely adding a person to the workflow is insufficient.
Monitoring, alerts, complaint channels, exception tracking, incident reporting, and rollback and suspension procedures support detection and response. The NIST Generative AI Profile offers technology-specific risk-management actions, illustrating why controls should be adapted to the technology and use case.

Control categories describe where a safeguard operates, while control functions describe what it is intended to do. A balanced control design may use several functions across the same risk pathway.
Preventive controls reduce the likelihood that an unwanted event will occur. Examples include access restrictions, prohibited-use rules, input validation, deployment approval, data-quality gates and automation limits.
Detective controls identify failures, deviations, or emerging risks. Examples include performance monitoring, output sampling, anomaly alerts, complaint analysis, audit logs and exception reporting.
Corrective controls contain harm, restore safer operation, or prevent recurrence after a problem is detected. Examples include rollback, model withdrawal, access revocation, incident response, retraining, and corrective action.
Compensating controls provide alternative protection when the preferred safeguard is unavailable or insufficient. For example, enhanced human review and narrower deployment may temporarily reduce exposure while a technical control is being improved.
A safeguard can serve more than one function. Human review may prevent an unsuitable decision, detect a recurring output problem, and initiate corrective action. The organization should still state the primary control objective so that effectiveness can be tested against the intended result.
Start with the risk, not a generic checklist. Move from a precise scenario to a treatment objective, a control that interrupts the risk pathway, and a proportionality decision.
A workable risk statement explains what could happen, under which conditions, who could be affected, and why it matters. Labels such as “bias,” “privacy,” or “AI error” are too vague.
For example: “When recruiters use model-generated rankings, unsuitable training data may produce unreliable recommendations that disadvantage qualified applicants.” This identifies the event, conditions, affected people, and consequences.
The objective may be to remove a risk source, reduce likelihood, exposure, or severity, detect failure sooner, or improve recovery. Organizations choose AI risk mitigation strategies first, then select controls that implement the treatment.
Ask whether the control addresses the cause, event, exposure, consequence, detection, or recovery. For confidential information disclosure, an accuracy check misses the pathway. Access limits, input restrictions, data-loss controls, and incident response may be relevant.
A low-impact internal tool and a system influencing employment, healthcare, finance, or safety should not receive identical controls automatically. Greater consequences, uncertainty, or exposure may justify stronger safeguards, independent testing, and higher approval authority.
Control selection belongs within the AI risk management process, after assessment and before monitoring. ISO/IEC 23894:2023 provides guidance for integrating risk management into AI-related activities.
Every material control needs enough detail to operate consistently, test objectively, and be reviewed when circumstances change.
State which risk and treatment objective the control addresses and the expected change. A measurable objective supports credible testing.
Name who is accountable for keeping the control effective. This control owner may differ from the risk owner, who owns the overall risk decision. Identify the operator and exception authority too.
Document what happens, when, by whom, upon which trigger, and with which exceptions. Replace “human review required” with clear review criteria, authority, and escalation.
Identify evidence such as approvals, tests, review logs, configurations, monitoring, exceptions, and incidents. Specify the response to control failure, missing evidence, or excessive risk.
The NIST AI RMF Playbook offers voluntary suggested actions for operationalizing Govern, Map, Measure, and Manage outcomes. It can inform implementation but is not a mandatory checklist or fixed sequence.
A control record should provide enough information for another qualified person to understand, operate, and test the safeguard. For each material control, document:
The related risk scenario
The treatment objective
The control description and function
The risk owner and the control owner
The operator and exception authority
The operating frequency or trigger
The required implementation evidence
The effectiveness-testing method
The failure or escalation threshold
The expected effect on residual risk
The monitoring and review conditions
The approval and next review date
Documentation should remain proportionate. A low-impact internal tool may require less formal evidence than a system influencing employment, healthcare, finance, safety, or access to essential services. The objective is traceability and reliable operation, not paperwork for its own sake.
Before approving a control, ask whether the assigned owner has the authority and resources to operate it, whether evidence will be produced consistently, and whether failure leads to a defined response. If any answer is unclear, the safeguard is not yet operationally ready.
For organizations building a wider control structure, this record can become the foundation of a reusable AI control library. Each approved control can include its intended risks, implementation conditions, evidence, testing method, limitations, and review triggers. Reuse should never replace checking whether the control fits the current system and context.
A control can exist on paper or in software and still fail to reduce risk. Testing should answer three questions.
Design effectiveness asks whether the safeguard can address the identified risk if operated as intended. Examine its relationship to the cause, event, exposure, or consequence, including limitations and dependencies.
Operating effectiveness asks whether the safeguard works consistently. Evidence may include tests, monitoring, samples, reviews, exceptions, incidents, and audit findings. Use realistic deployment conditions and check whether users bypass or weaken the control.
Suppose human approval is required before an AI recommendation affects a customer. Sampling shows reviewers approve almost every recommendation without scrutiny. The control exists, but its risk reduction is weaker than intended.
After validating controls, reassess the remaining likelihood, exposure, and consequences. Decide whether residual risk meets defined criteria, needs more treatment, requires changed deployment or must be escalated for authorized acceptance. Record the decision-maker, rationale, conditions, and review trigger.
The OECD's work on advancing accountability in AI connects risk treatment with governance, monitoring, and lifecycle review. Control-test results should inform decisions, not remain isolated assurance records.

Consider an AI-supported recruitment tool that ranks applicants. A complete risk-to-control chain might work as follows.
The risk scenario is that unsuitable historical data or model behavior produces unreliable rankings that disadvantage qualified applicants. The treatment objective is to reduce the likelihood that unreliable rankings influence consequential hiring decisions.
Preventive controls could restrict the tool to decision support, prohibit automatic rejection, and require validation before deployment. Detective controls could compare selection patterns, sample ranking quality, track overrides, and monitor complaints. A human control could require trained recruiters to review relevant evidence before acting on a recommendation. Corrective controls could suspend the tool, investigate the cause, and revise data, configuration, or operating procedures.
Evidence might include configuration records, validation results, recruiter review logs, monitoring reports, complaints, and corrective-action records. Testing would examine whether reviewers genuinely challenge outputs, whether restrictions prevent automatic rejection, and whether monitored outcomes reveal disproportionate errors.
The organization would then reassess residual risk and record the authorized decision, conditions of use, and reassessment triggers. A new model version, a different applicant population, or an expanded hiring context would reopen the relevant control and risk decisions.
This is a hypothetical example. The appropriate controls depend on the system, organization, jurisdiction, and applicable employment and data-protection requirements.
Controls that worked at deployment may weaken as the system or its environment changes. Model updates, retraining, new data, integrations, expanded uses, new user groups, greater automation, emerging threats, incidents, and altered business processes can all change the risk-control relationship.
Monitoring should ask whether the risk is changing, whether the control still operates, and whether it remains sufficient. A functioning control may no longer provide adequate protection if exposure or consequences have increased.
Retest after material model or data changes, new use cases, major deployment expansion, significant incidents, or control failures. Routine review frequency should reflect risk and context. These checks continue throughout the AI risk management lifecycle, where monitoring evidence can trigger reassessment, stronger controls, restricted use, or suspension.
Buying or accessing an AI system does not remove the need for controls. The organization may have less visibility into training data, model changes, testing, and incidents, so supplier and contractual controls become part of the risk response.
Depending on the context, relevant measures may include vendor due diligence, defined data-use restrictions, model-change notification, incident notification, documentation access, testing rights, service-level expectations, fallback arrangements, and an exit plan. Technical controls may also restrict the data, permissions, tools, or decisions available to the external system.
Third-party controls should reflect actual leverage and evidence. A contract term is not effective simply because it exists. The organization should know how compliance will be checked, what happens after a supplier failure, and whether operations can continue safely if the service changes or ends.
A familiar safeguard may not interrupt the current risk pathway. Control selection should begin with the assessed scenario and treatment objective.
A control owner who cannot obtain evidence, enforce the process, or escalate failure cannot keep the safeguard effective.
Human oversight becomes weak when reviewers lack time, information, competence, or authority, or when approval becomes routine.
Operational conditions, users, data, and dependencies change. Material changes and monitoring evidence should trigger proportionate retesting.
Frequent overrides, informal workarounds, and missing evidence may show that the control is impractical, misunderstood, or routinely bypassed.
An alert does not reduce risk unless someone reviews it and has authority to investigate, restrict, correct, or suspend the system.
Implementing controls does not automatically make the remaining risk acceptable. Residual risk requires an explicit, authorized decision.
NIST organizes AI risk-management outcomes through Govern, Map, Measure, and Manage. Controls connect across these functions through accountability, risk mapping, testing, treatment, and monitoring. The Playbook suggests actions but is not a control catalog.
ISO/IEC 42001:2023 specifies requirements for establishing and continually improving an AI management system. It connects risk assessment and treatment with controls, responsibilities, evaluation, and corrective action. ISO 42001 risk management covers the detailed AIMS workflow, while ISO/IEC 23894 provides complementary guidance.
Article 9 of the EU AI Act provides a scoped example. For covered high-risk AI systems, risks should be eliminated or reduced through design and development where technically feasible, with mitigation and control measures for remaining risks. This is not a universal requirement for every AI system.
Organizations can align shared evidence, roles, and monitoring without forcing a one-to-one mapping. The guide to using NIST and ISO 42001 together explains the integration in depth.
Selecting safeguards is only one part of responsible AI risk management. Professionals also need to define credible risk scenarios, evaluate potential impacts, assign accountability, test control effectiveness, and monitor residual risk throughout the AI lifecycle.
If you work in AI governance, compliance, risk, audit, data, technology, or business leadership, explore AI risk management training to strengthen your ability to design, evaluate, and document practical controls that support responsible AI decisions.
Controlling AI risk requires more than collecting safeguards in a checklist. Effective AI risk controls connect to specific risks, have clear objectives and owners, produce evidence, undergo design and operating-effectiveness testing, support residual-risk decisions and remain subject to monitoring.
For every material control, an organization should be able to answer: Which risk does this control address, who owns it, how do we know it works, and what happens if it fails? If those answers are unclear, the control is not yet operationally defensible. The next action is to strengthen the risk-to-control record, test the safeguard, and assign an explicit response to failure.
AI risk controls are technical, organizational, procedural, or human safeguards used to prevent, detect, reduce, contain, or respond to identified AI risks. Each needs an objective, owner, operating method, and evidence requirement.
Risk mitigation is the strategy for addressing a risk. AI risk controls are the safeguards that implement it. For example, limiting exposure is an objective; restricting access and automation are controls.
Examples include access restrictions, data validation, model testing, human approval, monitoring, automation limits, incident reporting, and rollback. The right combination depends on the risk and context.
Preventive controls reduce the likelihood of an unwanted event. Detective controls identify failures or emerging risks. Corrective controls contain harm, restore safer operation, or prevent recurrence. A control may perform more than one function.
A named control owner should be accountable for keeping the safeguard operational and effective. The control owner may differ from the risk owner, who remains accountable for the overall risk decision and residual-risk acceptance.
It should identify the related risk, treatment objective, control description, owner, operator, frequency or trigger, required evidence, testing method, failure threshold, escalation response, residual-risk effect, and review date.
Test whether the control is designed to address the risk, then use records, samples, monitoring, and incidents to determine whether it operates consistently. Finally, reassess residual risk.
Review frequency should reflect risk, uncertainty, and context. Retest after material model or data changes, new uses, expanded deployment, significant incidents, control failures, or unreliable assessment assumptions.
The organization should assess the failure, determine whether exposure or residual risk has increased, and follow a defined escalation process. Appropriate action may include remediation, additional safeguards, restricted use, rollback, suspension, or renewed risk acceptance.
Requirements depend on the jurisdiction, sector, organization, role, and AI system. Some laws impose specific risk-management or control duties for covered systems, while frameworks such as NIST AI RMF are voluntary. Organizations should assess the rules that apply to their circumstances.
OpenAI shelved GPT-6.1 Astra after safety tests flagged scope, authorization and action-reporting issues. See what is confirmed and what remains...
AI Law
Learn AI compliance requirements, key risks, the EU AI Act, NIST AI RMF, ISO 42001, and practical steps to build...
AI Law
Understand AI regulation in the United States in 2026, including federal rules, state AI laws, privacy, discrimination and practical compliance...