Questions for the Audit
| Question | What it tests |
|---|---|
| What is the maximum credible harm? | Impact and severity |
| Who could experience that harm? | Affected parties |
| What AI capability could contribute? | Capability risk |
| What conditions are required? | Harm pathway |
| What does the provider control? | Upstream dependency |
| What do we control? | Deployment responsibility |
| What prevents the scenario? | Preventive controls |
| What happens if that control fails? | Defense in depth |
| How would we know harm is occurring? | Detection |
| Can affected people report unexpected harm? | Feedback |
| Who reviews and escalates those reports? | Accountability |
| Can someone override, restrict, or stop the system? | Intervention |
| How does the organization recover? | Resilience |
| What risk remains? | Residual risk |
Follow the Evidence
The questions become assurance work when answers are tied to evidence. A claim that high-impact actions require approval should lead to workflow configuration and approval records. A claim that access is constrained should lead to permissions and entitlement evidence. A claim that the system can be disabled should lead to a tested procedure rather than an assumption.
That traceability makes it harder for a risk register to close a severe issue merely because a control exists on paper. The auditor can test whether the control actually interrupts the pathway that made the harm credible.
Test Feedback and Response
Feedback deserves the same treatment as other controls. Ask how users and affected people can report harm, how reports enter the governance process, how trends are identified, and what thresholds trigger reassessment or escalation.
Then test the response path. Who can investigate? Who can restrict a feature? Who can revoke access or disable the system? Can an affected person appeal a consequential outcome? What evidence shows that those mechanisms have been exercised or tested?
This matters because some harms are observable first by the people experiencing them. A technically healthy system can still produce harmful outcomes that operational telemetry does not recognize.
What Good Looks Like
A strong assessment does not need to prove that every imaginable extreme event is impossible. It should show that the organization understands its credible severe harms, knows which upstream dependencies and local conditions matter, has controls tied to those pathways, can detect and respond when assumptions fail, and has consciously evaluated the residual risk.
For some systems, that analysis will establish a relatively low risk ceiling. For others, it may reveal that autonomy, privileged access, scale, or dependence on an upstream model warrants stronger safeguards or escalation.
That is the practical value of catastrophic-risk thinking for ordinary risk and audit work. It turns a debate that can feel remote from most organizations into a disciplined way to ask how bad an AI failure could reasonably become, what would have to happen to get there, and whether the evidence supports confidence in the safeguards standing in the way.