← All guidesISO/IEC 23894 — AI Risk Management

ISO/IEC 23894 AI Risk Management Guidance: A Reference for Boards and Compliance Teams

Source
Khullani M. Abdullahi, JD / Techné AI
URL
https://techne.ai/insights/iso-iec-23894-reference/
Type
practitioner guide
Retrieved
2026-08-17
License note
Summary and analysis by On The Ground (OTG). Original article © source_organization. This is an original summary, not a reproduction of the source text — see source_url for the complete original.

Overview

ISO/IEC 23894:2023, "Information technology — Artificial intelligence — Guidance on risk management," is a joint ISO/IEC standard published in February 2023 by ISO/IEC JTC 1/SC 42, providing AI-specific risk management methodology. It is a methodological document: it describes how to identify, analyze, evaluate, treat, monitor, and report on the risks arising from developing, deploying, and operating AI systems.

Crucially, this article emphasizes that ISO/IEC 23894 is guidance, not a certifiable management-system standard. It contains no mandatory ("shall") requirements and no clause structure built for third-party audit, so there is no certification pathway against it directly. Organizations seeking a certifiable AI standard instead pursue ISO/IEC 42001, the AI management-system standard, which imports ISO/IEC 23894's methodology as the substance behind its own risk-management clauses. The two are meant to be used together — 42001 supplies the auditable structure, 23894 supplies the risk-management method inside it.

How it relates to ISO 31000

ISO/IEC 23894 is structured as an AI-specific extension of ISO 31000:2018, inheriting that standard's principles, framework, and process essentially unchanged and populating each with AI-specific content. ISO 31000's eight principles (integrated, structured and comprehensive, customized, inclusive, dynamic, best-available information, human and cultural factors, continual improvement) describe what good risk management looks like; its framework covers the organizational arrangements that embed risk management into governance; and its process is the operational sequence — communication and consultation, scoping, risk assessment, treatment, monitoring, and recording — by which specific risks actually get managed.

Because ISO/IEC 23894 keeps this architecture rather than replacing it, an organization cannot meaningfully adopt 23894 in isolation. It needs an underlying ISO 31000 (or ISO 31000-aligned, e.g. COSO ERM) risk methodology already in place, onto which the AI-specific layer is grafted.

The risk management process, stage by stage

  • Communication and consultation. For AI, the relevant stakeholder group is unusually broad — it can include people directly affected by a system's decisions (job applicants, patients, residents of a policed neighborhood), the human reviewers in the loop, the developers building the model, procurement/vendor-management for purchased AI, and the executives accountable for governance. The article frames this as continuous engagement across the AI life cycle rather than a single pre-launch consultation.
  • Scope, context, and criteria. Because an AI system is a composite of training data, model architecture, deployment environment, human-oversight arrangements, system integrations, and operational context, scoping an AI risk assessment takes in more moving parts than scoping a conventional IT risk assessment. External context covers the legal, regulatory, market, and societal environment; internal context covers the organization's own policies, risk appetite, culture, and governance maturity; criteria define what consequences and likelihoods are considered tolerable.
  • Risk assessment — identification. Uses two parallel registers: risk sources (the underlying characteristics of the system, its data, its context, or its development process that could produce harm) and risk events (the concrete ways those sources materialize into failure — biased or incorrect outputs, adversarial-input exploitation, drift from the training distribution, automation bias in human reviewers, cascading failures across linked AI systems, loss of oversight, or leakage of private/IP-protected training data). The standard's informative Annexes A and B supply catalogs of AI-related objectives and risk sources meant to make this identification systematic rather than ad hoc.
  • Risk assessment — analysis. Both consequence and likelihood are harder to pin down for AI than for conventional risks: the same failure can produce very different harm depending on deployment context (a flawed recommendation in a marketing tool versus in a clinical triage tool), and likelihood often depends on things that aren't directly observable, such as data drift or adversarial probing. Where quantification is incomplete, the article notes the standard supports qualitative methods, scenario analysis, and red-teaming instead.
  • Risk assessment — evaluation. Analyzed risks are weighed against the criteria set during scoping to decide what needs treatment, what can be accepted, and what needs escalation. The article stresses that for AI this is where the organization's risk appetite becomes a real governance decision — one boards should expect to be consulted on, not a technical calculation left entirely to staff.
  • Risk treatment. Selecting and implementing mitigations appropriate to the evaluated risks, done iteratively: treat, reassess residual risk, decide whether further treatment is needed.
  • Monitoring and review. The article identifies this as the stage where AI risk management diverges most from conventional enterprise risk management, because AI systems are non-stationary — inputs drift, environments change, performance degrades, and new failure modes surface as users find capabilities the developers never anticipated. Ongoing monitoring is treated as a first-order obligation rather than a periodic afterthought, covering things like distributional drift detection, emergent post-deployment behavior, and whether human reviewers are still meaningfully exercising oversight or have started deferring to the system by default.
  • Recording and reporting. Documents the assessment, treatment decisions, residual-risk acceptance, monitoring results, and incident lessons through governance channels. For a system subject to a regulatory filing obligation (EU AI Act technical documentation, a state safety-and-security protocol, a CCPA-style risk assessment), these records become the substance of that filing.

AI-specific risk sources (per the standard's Annex B, as summarized in the article)

The article groups AI risk sources into: data quality and bias (representativeness gaps, historical bias in labels, sampling bias, poisoning, provenance defects); training and tuning choices (objective function, optimization, hyperparameters, reward specification, fine-tuning); algorithmic transparency and explainability (treated as a spectrum calibrated to use case, stakeholder, and regulatory regime rather than a binary property); model robustness and reliability (brittleness under distribution shift, sensitivity to perturbation); adversarial-input and security vulnerabilities (adversarial examples, prompt injection, model extraction, training-data extraction, model inversion, data poisoning — distinct from, and additional to, conventional cybersecurity risk); deployment-context risk (scope creep, population shift, downstream integration amplifying errors); human oversight and automation bias (treating oversight itself as something requiring ongoing calibration, since reviewers of a high-accuracy system tend to defer to it over time); societal and environmental impact (labor market effects, effects on affected communities, information-environment effects, and the energy/water footprint of training and inference); and lifecycle risk generally, which the standard's Annex C maps across life-cycle stages including the comparatively neglected area of decommissioning (data retention/deletion, dependency risk for downstream consumers).

AI-specific risk treatments

Correspondingly, the article summarizes the standard's treatment categories as: data curation and governance (provenance documentation, representativeness analysis, bias auditing, synthetic data to fill representation gaps); model design choices (choosing more interpretable model classes where explainability matters, fairness-aware training, and the inherent trade-offs between performance, robustness, interpretability, and privacy); validation and testing (representative and adversarial/out-of-distribution test sets, subgroup fairness testing, red-teaming, documented validation evidence); output filtering and validation at runtime (business-rule checks, content filtering, confidence thresholds, guardrails); human-in-the-loop controls (full review, sampled review, or exception-based escalation — with oversight itself needing calibration and monitoring, not treated as a one-time control that discharges further obligation); monitoring and observability post-deployment (logging, drift detection, fairness monitoring, incident-reporting channels reaching people with authority to act); and decommissioning processes (orderly retirement, transition planning for downstream consumers, data/artifact retention and deletion, and a governance check that winding the system down doesn't itself create new risk).

Relationship to ISO/IEC 42001 and the NIST AI RMF

The article positions three documents as the current international AI-governance methodology stack: ISO/IEC 42001 (the certifiable AI management-system standard, structured like ISO/IEC 27001 or ISO 9001, specifying what an AI management system must contain — policy, leadership commitment, roles, risk processes, controls, audit cadence, management review); ISO/IEC 23894 (supplying the substantive risk-management method that 42001's risk clauses point to); and the NIST AI Risk Management Framework (the US voluntary framework, organized around four functions — Govern, Map, Measure, Manage — rather than ISO 31000's process stages, but covering substantially overlapping ground). The article's view is that these are complementary rather than competing: ISO 23894 tends to be preferred where an organization already runs on ISO 31000 or is pursuing ISO 42001 certification, or operates in jurisdictions (EU/EMEA) where ISO standards carry more weight; NIST AI RMF tends to be preferred in a US-only context or when documentation needs to be legible to US regulators and procurement counterparties. Many multinational organizations cite both.

How the article ties this to specific regulatory regimes

  • EU AI Act, Article 9. Requires providers of high-risk AI systems to establish, implement, document, and maintain a risk-management system across the system's life cycle, specifying the required components (identification, estimation, evaluation, post-market monitoring, and mitigation measures) without prescribing a specific methodology — which is the gap the article says ISO/IEC 23894 is commonly used to fill. The article also notes the EU's Digital Omnibus deferral pushed the compliance deadline for high-risk-system obligations to December 2, 2027, and that once the AI Act's harmonized standards (several of which are expected to reference ISO/IEC 23894) are formally adopted, conformity with them would create a presumption of conformity with Article 9 — but that citing ISO/IEC 23894 today is a defensible reference, not yet a safe harbor.
  • New York RAISE Act. Requires large frontier-model developers to publish a safety and security protocol addressing critical-harm risk, again without prescribing the underlying risk-assessment method; the article describes ISO/IEC 23894, used alongside ISO/IEC 42001 and NIST AI RMF, as a common methodological basis for building that protocol, ahead of the Act's January 1, 2027 effective date.
  • State impact-assessment regimes. The article notes this area moved quickly: the original 2024 Colorado AI Act's algorithmic impact-assessment obligation was repealed in May 2026 before it took effect, and Colorado's replacement (SB 26-189) requires disclosures rather than formal assessments. Where an assessment obligation does apply elsewhere (e.g., CCPA-style risk assessments for automated decision-making, or EU AI Act documentation), the article's point is that the underlying methodology — not the specific statute — is what persists, and ISO/IEC 23894 is offered as a defensible basis for that methodology regardless of which particular law is in force at a given time.
  • NYC Local Law 144. Requires bias audits of automated employment decision tools; the audit methodology itself is set by NYC rules, but the article notes the broader risk-management process around the audited tool is left to the deployer, which is where it sees ISO/IEC 23894's data-quality, bias, and monitoring content being used.
  • Sector regulators and procurement. The article notes that sector regulators and enterprise procurement due-diligence increasingly reference NIST AI RMF, ISO/IEC 23894, and ISO/IEC 42001 together as recognized methodologies for documenting AI risk work.

Limitations the article flags

  • It is guidance, not a requirement — no certifiable clause structure, and it can't stand alone as a compliance artifact; it needs to be embedded in a management system or cited within an actual regulatory filing to carry weight.
  • It is a methodology, not a control specification — it tells an organization how to identify, evaluate, and treat AI risk, not which specific control to apply to which specific risk, so two organizations can apply it faithfully and still land on different treatment decisions.
  • It is not a substitute for legal analysis — it doesn't interpret which regulatory obligations apply to a given system in a given jurisdiction; that has to be done separately, with ISO/IEC 23894 then used as the methodology for operationalizing whatever obligations that legal analysis identifies.

Practical steps the article recommends

Confirm the organization already has an ISO 31000-aligned enterprise risk baseline; name an accountable owner for AI risk (typically the CRO or CCO, with CTO/CISO as technical lead); build and maintain an inventory of AI systems in scope (developed, deployed, procured, and embedded in third-party tools); run the full scope-identify-analyze-evaluate-treat-record process against each in-scope system, scaling depth to the system's risk level; cite ISO/IEC 23894 explicitly in the resulting risk artifacts so auditors and counterparties can tell the work was done against a recognized method rather than improvised; feed the same underlying work into the AIMS, the regulatory filings, and board reporting rather than redoing it for each audience; and refresh assessments at least annually, plus whenever the system, its operating context, or an incident materially changes.

Note on generative AI and foundation models

The article observes that ISO/IEC 23894 was finalized shortly after the public release of ChatGPT, and was written generally enough to apply across AI system types — including generative and agentic systems — without singling any of them out, though practical application to those systems has developed considerably since publication. Organizations applying it to frontier or generative systems are described as typically supplementing it with more model-specific resources (e.g., NIST's Generative AI Profile, red-teaming guidance) within the same overall process structure.

Suggested citation

Abdullahi, Khullani M. "ISO/IEC 23894 AI Risk Management Guidance: A Reference for Boards and Compliance Teams." Techné AI, May 12, 2026. https://techne.ai/insights/iso-iec-23894-reference