AI Risk Assessment
Definition
AI Risk Assessment
An artificial intelligence (AI) risk assessment is a structured evaluation of what could go wrong with a particular AI system, how likely that is and what would follow. It is a per-system activity, which is what distinguishes it from a standing governance framework.
The framework says every system must be assessed. The assessment is the work done on one of them — and it produces a decision about that system rather than a policy about all of them.
Its scope is wider than accuracy. A complete assessment covers data provenance, bias and fairness, security, privacy, explainability, human oversight, failure behaviour and the consequences of being confidently wrong.
The last of those deserves separate attention. A system that declines to answer is inconvenient, whereas a system that answers wrongly with apparent certainty transfers its error straight into a decision.
Timing matters as much as content. An assessment run after build can only recommend mitigations — while one run at design stage can still change what gets built.
That is why the trigger belongs in the intake process. Teams asked to submit an assessment at the approval gate will write one that justifies the system they have already finished.
Key takeaways
- The assessment evaluates one system; the framework governs all of them.
- Scope extends well beyond model accuracy to harms, oversight and failure behaviour.
- Assessment at design stage changes the system; assessment after build only mitigates.
- Reassessment is required when data, use or population changes.
How it works
The assessment identifies the system’s purpose and affected population, enumerates plausible harms, evaluates likelihood and impact, records existing controls, and states whether residual risk is acceptable and who accepted it.
Established security practice supplies the shape. NIST Special Publication 800-30 provides guidance for conducting risk assessments, carried out at all three tiers of the risk management hierarchy as part of an overall risk management process.
Regulation increasingly sets the trigger. The European Commission states that high-risk AI systems will be subject to strict obligations including adequate risk assessment and mitigation systems before being placed on the market.
| Dimension | Question asked | Evidence expected |
|---|---|---|
| Purpose and population | Who is affected and how | Use case statement |
| Data | Where it came from, what it permits | Provenance record |
| Fairness | Does performance vary by group | Disaggregated results |
| Oversight | Can a person intervene meaningfully | Escalation design |
| Failure | What happens when it is wrong | Fallback behaviour |
Examples
Assessments vary enormously with who actually bears the consequences of a wrong output. The four cases below show that single difference driving how deep the evaluation goes and which evidence sits at its centre.
A lender assesses a scoring model with disaggregated performance testing. The ai bias audit is the central evidence rather than an appendix.
A software firm runs adversarial testing before release. Its ai red teaming exercise is scoped by the assessment’s harm list, not by tester curiosity.
A support operation assesses a customer-facing assistant for ai hallucination risk. Confident wrong answers, not refusals, are the harm that drove the design.
An enterprise assesses a purchased tool through ai vendor evaluation. The assessment is harder because the model internals belong to someone else.
Related terms
Risk work in AI spans several artefacts at different scopes, and the entries below separate this per-system activity from the supplier-level and portfolio-level work that surrounds it in most organisations.
- Model drift: the change that triggers reassessment after deployment.
- Vendor risk assessment: the supplier-level evaluation that sits alongside the system-level one.
- Third-party risk management: the programme that handles purchased AI at portfolio level.
FAQ
When should an assessment be run?
At design stage, before build. An assessment run against a finished system can only propose mitigations, because the expensive decisions have already been made.
How does it differ from a governance framework?
The framework sets the requirement and the approval path. The assessment is the evaluation of one system carried out to satisfy that requirement.
What triggers a reassessment?
A change in the data, the population served, the use case or the model itself. Material drift in monitored performance should also trigger one.
Can a purchased AI tool be assessed properly?
Partially. Internal model details are usually unavailable, so the assessment relies on supplier evidence, contractual commitments and the buyer’s own testing.
Who signs off residual risk?
A named business owner, not the technical team. Risk acceptance is a business decision, and recording who made it is the point of writing it down.
Is accuracy testing enough?
No. A model can be accurate overall and unacceptable for a subgroup, unexplainable to a regulator, or damaging in the way it fails.
Find partners who can run independent AI assessments in the Outsource Accelerator directory.







Independent




