
We Built an ISO 42001 Evidence Checklist for AI Companies — Here’s What Auditors Actually Look For
Controls are rarely why organisations fail an ISO 42001 audit. Evidence is…A clause-by-clause evidence checklist, the seven patterns that separate a pass from a finding, and the questions auditors actually ask.
ISO 42001 evidence checklist, ISO 42001 audit, AIMS certification, Stage 2 audit, AISIA, Statement of Applicability, AI system register
Across this series I’ve made the same claim four times, in four different contexts, and it’s time to give it a post of its own:
Controls are almost never why organisations fail. Evidence is.
When I led ShareVault through ISO 42001 Stage 2 certification on the first audit attempt and later served as their internal auditor, the pattern held throughout. The difference between a clean pass and a nonconformity was rarely whether a control existed. It was whether we could put a dated artifact on the table showing that the control operated, on a specific date, under a specific policy version, owned by a named person.
So this post is the checklist I wish more teams had before Stage 1 — organised the way an auditor actually works through it, rather than the way the standard is numbered.
One framing that will save you time. Every control an auditor examines gets tested against four questions:
- Does it exist? — the document
- Is it operating? — the record showing it ran
- Who owns it? — a name, not a team
- Show me a specific instance. — one dated example, produced now
Most programs can answer 1. Certification requires all four.
Stage 1 and Stage 2 test different things
This trips people up more than any technical requirement, so it’s worth being explicit.
Stage 1 is a documentation review. The auditor is checking whether your AIMS is designed adequately: does the mandatory documented information exist, is the scope coherent, is the SoA complete, does the risk methodology make sense. You can pass Stage 1 with a management system that has never actually run.
Stage 2 tests whether it operates. Records, not documents. The auditor samples: show me the impact assessment for this AI system, the approval that let it deploy, the monitoring output from last quarter, the internal audit that covered this control, the corrective action that closed that finding.
The single most common failure mode I see is a Stage-1-ready programme presented at Stage 2. Beautiful policies, signed and versioned, with no operating history behind them. If your AIMS was built in the eight weeks before the audit, Stage 2 will find that out — not because the auditor is suspicious, but because records have dates.
Practical implication: your AIMS needs an operating history before Stage 2. Three months is thin. Six is comfortable.
The mandatory documented information
Start here, because these are non-negotiable under Clause 7.5 and their absence is an automatic finding. Nine items:
| # | Artifact | Clause | The evidence that makes it real |
|---|---|---|---|
| 1 | AIMS scope document | 4.3 | Inclusions, exclusions, and justification for exclusions |
| 2 | AI policy | 5.2 / A.2.2 | Signature of top management, date, version, evidence of communication |
| 3 | AI risk assessment records | 6.1.2 / 8.2 | Executed assessments per AI system, with dates and treatment decisions |
| 4 | AI system impact assessment (AISIA) records | 6.1.2 | One per in-scope AI system, updated on material change |
| 5 | Statement of Applicability | 6.1.3 | All 38 Annex A controls, applicability decision, justification, status |
| 6 | AI objectives | 6.2 | Measurable, with the monitoring record showing measurement happened |
| 7 | Internal audit reports | 9.2 | Audit plan, findings, auditor independence evidence |
| 8 | Management review records | 9.3 | Minutes with decisions and action items, not attendance lists |
| 9 | Nonconformity and corrective action records | 10.2 | Root cause, action, and effectiveness review |
Three of these fail more often than the rest.
The SoA (#5) fails when exclusions are justified with something like “not currently a priority.” That is not a justification. A valid exclusion explains why the control is not applicable to your context — typically because you’re an AI user rather than a provider, so provider-specific controls in A.4, A.6.2, and A.7 may genuinely not apply. Write the reason, not the intention.
AI objectives (#6) fail when they’re aspirational. “Improve responsible AI practices” is not measurable. “Complete AISIA for all in-scope AI systems by 30 June,” “100% AI awareness training completion,” “AI incident MTTR under X hours” are. And Clause 9.1 then requires you to show you measured them — the objective without the measurement record is half a finding.
Corrective actions (#9) fail on the last step. Teams log the nonconformity, log the action, and stop. Clause 10.2 requires a review of whether the action was effective. That effectiveness review is the single most commonly missing artifact I encounter, in both 42001 and 27001.
Clause-by-clause: what to have on the table
| Clause | What the auditor asks for | Common nonconformity |
|---|---|---|
| 4.1 Context | Contextual analysis covering AI regulation, public trust, internal AI maturity | Generic corporate context with no AI dimension |
| 4.2 Interested parties | Stakeholder register including individuals affected by AI decisions, regulators, model vendors | Register lists customers and investors only — omits affected individuals |
| 4.3 Scope | AIMS scope document with justified exclusions | AI tools used in HR screening excluded without justification |
| 5.1 Leadership | Management meeting minutes discussing AI governance; resource allocation | Auditor interviews an executive who cannot describe the AIMS scope |
| 5.3 Roles | RACI or roles document naming AIMS owner, AI risk owner, system owners, data governance lead, incident manager | “The security team owns it” — no named individuals |
| 6.1.2 Risk + AISIA | Executed risk assessments and impact assessments per system | AISIA done once at implementation, never revisited |
| 6.1.3 Treatment | Risk treatment plan with owners, timelines, residual risk acceptance | Residual risk not formally accepted by anyone |
| 7.2 Competence | Competence matrix by role, training records, effectiveness evaluation | Training records exist; effectiveness never evaluated |
| 7.3 Awareness | Awareness programme evidence with attendance covering all staff | Attendance list covers a fraction of headcount, no follow-up |
| 8.1 Operation | Change management showing risk/impact reassessment when AI systems changed | Model version changed; no reassessment triggered |
| 9.1 Monitoring | Metrics register or dashboard with actual readings over time | Metrics defined, never populated |
| 9.2 Internal audit | Audit programme, plan covering all clauses over the cycle, reports, independence evidence | Internal auditor audited their own work |
| 9.3 Management review | Minutes covering the full required agenda | Review held, but agenda missed risk assessment results or audit findings |
| 10.2 Improvement | Nonconformity log with root cause and effectiveness review | Effectiveness review absent |
A note on 9.2 independence: the auditor cannot audit their own work. In a small company this is a real constraint, and the usual resolutions are to have a different function audit the AIMS, bring in an external internal auditor, or split the audit so no one reviews the area they built. Plan for it early — it’s a structural problem, not a documentation one.
Annex A hot spots
Thirty-eight controls across nine domains, and the failures cluster predictably. The ones I’d stress-test first:
A.2.3 — Alignment with other policies. The AI policy exists, but HR, procurement, IT, and data governance policies were never updated to reflect AI. Auditors check this because it’s a fast test of whether the AIMS is real or bolted on.
A.3.3 — Reporting of concerns. A channel for staff to raise ethical concerns, bias observations, or unexpected AI outputs without reprisal. Most organisations have a security incident channel and assume it covers this. It doesn’t — and the auditor will ask an employee whether they know where to report an AI concern.
A.4.6 / 7.2 — Competence and awareness. AI-specific competency requirements per role, not general security awareness with an AI slide.
A.5.2–A.5.5 — Impact assessment. The process document, the records, and specifically the societal impact dimension (A.5.5), which teams routinely skip because it feels abstract. Environmental cost of compute, systemic bias at scale, labour effects — write something considered, even if brief.
A.6.2.4 / A.6.2.5 — Verification and deployment. Bias and fairness testing across demographic groups, adversarial testing, and a documented go/no-go authorisation before deployment. “We tested it” without a record of the authorisation decision is a finding.
A.6.2.8 — Event logs. AI system logs sufficient for incident investigation and accountability, with defined retention and access controls. See the agent section below — this control has quietly become much harder.
A.7.3 / A.7.5 — Data acquisition and provenance. Legal basis for training data, provenance documentation, chain of custody. If you’re an AI user rather than provider, these may be excluded — but then your SoA justification needs to say so, and A.10.3 supplier evidence has to carry the weight instead.
A.9.2 / A.9.4 — Responsible use and intended use. Acceptable use processes, human oversight of outputs, escalation and override procedures, and enforcement of intended purpose. Use beyond documented intended purpose must be identified and controlled — which is the control that catches shadow AI.
A.10.2 / A.10.3 — Allocation and suppliers. Responsibilities allocated across the AI value chain, supplier tiering, AI-specific due diligence, contractual clauses. Your model provider’s terms are evidence here; go read them before the auditor does.
The seven patterns that separate a pass from a finding
This is the part I’d put on the wall. Any artifact you plan to present should satisfy all seven.
- Dated and versioned. An undated document proves nothing about when a control operated. If your version numbering doesn’t track chronology — v1.3 dated before v1.2 — expect a document control finding regardless of content quality.
- Signed by the right person. Not just signed. The AI policy needs top management under Clause 5.2. Residual risk acceptance needs the risk owner. An approval signed by whoever was available is a finding waiting to be written.
- Shows a decision, not just a document. Auditors distinguish artifacts that record a judgement from artifacts that describe a process. “Our deployment process requires impact assessment” is a document. “Impact assessment for System X, classified Medium, approved for deployment by [name] on [date], with these conditions” is evidence.
- Shows the loop closed. Finding → root cause → action → effectiveness review. Three out of four is a nonconformity. This applies to internal audit findings, incidents, and supplier issues alike.
- Covers the population, not a convenient sample. If you have eleven AI systems and eight AISIAs, the auditor will find the three. Completeness against the register is the test, which is why the register itself has to be accurate.
- Independent where independence is required. Clause 9.2 auditor independence, and — increasingly relevant — separation between whoever operates a control and whoever reviews it.
- Producible on request, during the audit. This is the practical one. If retrieving an artifact takes a week of searching shared drives, you have a records problem that will read to the auditor as a control problem. My rule of thumb: any mandatory artifact should be retrievable in under ten minutes by someone who isn’t the person who wrote it.
And the underlying principle, borrowed from control-effectiveness rating practice: a control is not effective because a policy exists. Rate honestly — 0 not implemented, 1 ad hoc, 2 partial, 3 defined, 4 implemented and evidenced, 5 measured and continuously improved. Most organisations sit at 3 and report 4. Stage 2 is where that gap surfaces.
Ten questions to rehearse
Auditors vary, but these come up in some form nearly every time. If you can’t answer one in two sentences with an artifact, that’s your gap list.
- Show me your AI system register. Is it complete, and when was it last updated?
- Which AI systems are excluded from scope, and why?
- Walk me through the impact assessment for this system. Who approved it?
- What changed about this system in the last six months, and did that trigger a reassessment?
- Who is accountable for this AI system? (The auditor may then go ask that person.)
- How does a member of staff raise a concern about an AI system?
- Show me your last internal audit report and how the findings were closed.
- What AI incidents have you had, and how were they handled?
- How do you assess your AI suppliers, and what’s in the contract?
- Show me the management review where AI risk was discussed.
Question 5 is the one that most often unravels a program, because auditors follow it up by interviewing the named person. If the RACI says someone owns a system and that person doesn’t know it, the document is evidence against you.
The new gap: agentic systems
This is where the evidence bar has risen fastest, and where checklists written even a year ago fall short. If you run agents, add these:
Agents belong in the AI system register (Clause 4.3, A.6.2.7). Including evaluation, test, and CI agents. As the OpenAI–Hugging Face incident demonstrated, non-production is not low-risk — an evaluation harness with code execution and reachability into shared infrastructure is a high-impact AI system whatever environment it nominally sits in.
A.6.2.8 event logs now have to answer four questions: who authorised this action, what context did the system have, what did it decide, and was that consistent with policy — with the policy version recorded. Agent action volume makes retrofitting this impractical; design it in.
A.9.2 human oversight needs to be demonstrable, not declared. Under EU AI Act Article 14 the standard is a demonstrated capability to intervene, interrupt, and disregard. The corresponding evidence is a tested kill switch with a record of the test and how long it took. An untested kill switch is an assumption, and auditors have started asking.
A.10.3 extends to model providers and MCP servers. Tiering, due diligence, contractual terms including data handling and no-training clauses, and change notification — because a silent model swap underneath you is a change you’re accountable for.
A.6.2.4 verification should include adversarial testing. Prompt injection, tool misuse, privilege escalation, memory poisoning, and approval bypass, with expected denials, version-controlled and re-run on material change.
Five nonconformities I’d bet on finding
If I walked into a first-time AI company audit tomorrow, these are where I’d look first, in order:
- Effectiveness reviews missing from corrective actions (Clause 10.2)
- AISIA completed once, never updated after material change (Clause 6.1.2)
- Objectives defined but never measured (Clause 6.2 into 9.1)
- Adjacent policies not updated for AI — HR, procurement, IT (A.2.3)
- AI system register incomplete — embedded vendor AI and internal agents missing (Clause 4.3)
None of these require sophisticated controls to fix. All of them require having actually run the management system for a couple of quarters.
Test yourself this afternoon
A genuinely useful exercise that takes about an hour:
Pick three controls at random from your SoA. For each, ask someone who did not build it to produce, within ten minutes: the governing document with its version and date, one dated record showing the control operated in the last quarter, and the name of the person accountable.
Count how many of the nine you get. That number is a better predictor of your Stage 2 outcome than any maturity assessment, and it costs you an hour instead of a certification cycle.
Work with DISC InfoSec
DISC InfoSec helps B2B SaaS and financial services organisations build AI management systems that survive an audit rather than describe one — AIMS scoping, AI system and agent inventories, AISIA methodology, Statement of Applicability, evidence architecture, internal audit, and Stage 1 / Stage 2 readiness.
I led Virtual Data Room (VDR) through ISO 42001 Stage 2 certification on the first audit attempt as the internal practitioner, served as their internal auditor, and authored their MCP Governance Standard. I’ve sat on both sides of the table, which is why this checklist is organised around what gets asked rather than how the standard is numbered.
Readiness path:
- Free 15–20 minute readiness call
- ISO 42001 gap assessment — | ISO 27001 gap assessment — clause-level, with a prioritised remediation roadmap
- Quick-Start engagement, 7–10 days — the core artifact set built and handed over
- Full implementation and certification support, including internal audit
DiscInfoSec — Principal Consultant, DISC InfoSec (Deura Information Security Consulting LLC), Petaluma, AICP, CISSP, CISM | ISO/IEC 42001 & ISO/IEC 27001 Lead Implementer | PECB Authorized Training Partner
📅 calendly.com/hd-deurainfosec 📧 info@deurainfosec.com 📞 (707) 998-5164 🌐 deurainfosec.com
This checklist reflects general practice and my own experience as an implementer and internal auditor. Certification bodies and individual auditors vary in emphasis; nothing here substitutes for your own certification body’s guidance or your accredited auditor’s judgement.
References
- ISO/IEC 42001:2023 — Clauses 4–10; Annex A (38 controls across A.2–A.10); Annex B implementation guidance
- ISO/IEC 42005 (AI system impact assessment guidance); ISO/IEC 23894 (AI risk management)
- NIST AI RMF 1.0 (NIST AI 100-1)
- Regulation (EU) 2024/1689 (EU AI Act), Arts. 14, 26
- OWASP Agentic Security Initiative; OWASP AI Agent Security Cheat Sheet
MachineLearning & Artificial Intelligence
AI Vulnerability Scorecard: Discover Your AI Attack Surface Before Attackers Do
Your Shadow AI Problem Has a Name-And Now It Has a Score
Most AI Security Tools Won’t Pass an Audit. Here’s a 15-Minute Way to Find Out.

InfoSec services | InfoSec books | Follow our blog | DISC llc is listed on The vCISO Directory | ISO 27k Chat bot | Comprehensive vCISO Services | ISMS Services | AIMS Services | Security Risk Assessment Services | Mergers and Acquisition Securit
DISC InfoSec blog | DISC InfoSec Site
- ISO 42001 Evidence Checklist: What Auditors Actually Look For (2026)
- Agents don’t produce wrong answers anymore They take wrong actions – A practitioner’s guide to agent security
- AI Governance for Bay Area Startups: What to Put in Place Before Enterprise Customers Ask
- How Much of Your Job Can Become an AI-Executable Workflow — and What’s Left Standing When It Does
- The AIMS/ISMS readiness ladder: seven steps from curious to certified




















