IT Audit of AI Systems
Last updated:
IT audit of an AI system is the independent examination of the controls around a system that learns or acts, rather than of the model itself. Most of the control set does not change. What changes is evidence: a conventional control produces a record that stays true, while an AI control produces a record that describes behaviour at one moment, from a system whose behaviour can move without anyone changing the code.
A practitioner reference on auditing artificial intelligence, written for internal audit, IT audit, and assurance functions rather than for model builders. It sets out what genuinely changes in an audit programme when a system learns or acts, what does not change and should not be rebuilt, and how to test a control whose behaviour is probabilistic.
What auditing an AI system actually means
An AI audit is the independent examination of the controls around a system that learns or acts. It is not an inspection of the model. Auditors rarely have the access, and almost never the mandate, to open a vendor's weights, and an opinion that depended on doing so would be unusable in most organisations. The examinable surface is the process around the system: who approved it, on what evidence, with what authority, under whose ownership, and with what condition for switching it off.
That framing matters because it decides what the function has to build. If the audit target is the model, an audit function needs data scientists it cannot hire. If the audit target is the control environment, the function needs its existing discipline applied to a system whose behaviour is not fixed at deployment. The second is achievable now.
What changes, and what does not
Most of the control catalogue survives unchanged. Access control, change management, segregation of duties, logging, vendor management, and business continuity all apply to an AI system exactly as they apply to any other production system, and rebuilding them under a new name is the most common waste in an AI audit programme.
Four things do change. Authority: a system that acts needs a written boundary on what it may do without a human, and irreversible actions held behind a checkpoint. Evidence: an output cannot be reproduced by re-running the same input, so the log has to capture the decision and its inputs at the time. Drift: the system can move without a change ticket, which breaks the assumption that change management catches change. Provenance: the data a system learned from is part of its control environment, and is usually undocumented.
An audit programme that adds those four and leaves the rest alone will cover more real risk than one that starts from a hundred-control AI framework.
Evidence is the hard part
Conventional IT audit rests on a walkthrough: observe the control once, confirm it operated as described, and rely on change management to tell you if it moved. That logic assumes a control that is stable between changes. An AI control is not stable between changes, because retraining, a vendor model update, or a shift in the input population can move its behaviour without anyone touching the system.
The practical consequences are set out in the walkthrough stops working: design effectiveness moves upstream into the approval record, sampling loses its meaning against a population that can be tested in full, and an audit opinion acquires a shelf life. The last point is the one that most often surprises an audit committee. An opinion on an AI control describes a moment, and the reasonable question is not whether the control worked but when it was last confirmed to work.
Auditing agents, not only models
An agent plans, calls tools, and takes multi-step actions rather than returning text. That moves the risk from output quality to authority, and it creates two problems most organisations have not addressed. The first is evidence: an agent's reasoning is not reliably reconstructable after the fact, so unless the system logs its decisions as it makes them, the audit trail does not exist. The second is identity: an agent authenticates, holds permissions, and acts, which makes it an identity in the estate that no joiner-mover-leaver process governs.
Both are examined in Who Audits the Agent?, together with what belongs in the audit programme. The short version for an audit plan: treat every agent as an account, require a named owner and a documented authority boundary, and test whether the stopping condition was written before deployment rather than after the first incident.
Which frameworks actually apply
Three are worth the effort. The NIST AI Risk Management Framework gives a defensible structure for govern, map, measure, and manage, and is the easiest to map onto an existing risk taxonomy. ISO/IEC 42001 provides a certifiable management system, which matters when an organisation needs to demonstrate governance externally. ISACA's guidance is the closest to audit practice, because it is written for the assurance function rather than for the builder.
None of them is a control catalogue an auditor can test directly, and treating any of them as one produces a programme of policy checks. They are useful for structuring coverage and for showing an audit committee that the programme rests on something recognised. The testable controls still have to be written for the specific system, which is what the AI control checklist and the AI model register are for.
Where an audit function should start
Start with an inventory, because most organisations cannot yet answer what AI they are running. That is a finding in itself, and it is the one an audit committee understands immediately. The use case intake form and the model register are enough to begin; a complete inventory is not the goal of the first pass.
Then take one consequential system and audit it properly rather than surveying twenty. Test the approval record, the authority boundary, the logging, the ownership, and the stopping condition. A single well-evidenced audit of a system that matters establishes the method and gives the function something to generalise from. Surveying everything produces a maturity score and no assurance.
The governance side of the same problem, written for the people who own the systems rather than the people examining them, is on the public sector AI governance hub and in the AI governance playbook. The security exposure that an audit will keep meeting is set out under AI security and assurance.
Where this perspective comes from
Shahzad Asghar is Head of Data and Digital Solutions at UNESCWA, the United Nations Economic and Social Commission for Western Asia, and holds CISA, CISSP, and CISM, together with ISACA membership. The combination is the reason for this cluster: CISA is the audit qualification, CISSP and CISM are the security and governance ones, and AI assurance sits where the three meet.
The writing here is drawn from building and running systems inside a United Nations organisation subject to real audit, rather than from advising on frameworks. That is also its limit, and it is worth stating: this is the perspective of an audit-qualified practitioner who builds, not of an audit partner signing opinions across a client portfolio.
Common questions
What is an AI audit?
An AI audit is the independent examination of the controls around a system that learns or acts, covering how it was approved, what authority it holds, who owns it, what it logs, and under what condition it is switched off. It is not an inspection of the model itself, which auditors rarely have the access or the mandate to open.
Who is responsible for auditing AI in an organisation?
Internal audit or IT audit provides the independent assurance, but they are the third line. The system owner is accountable for the control operating, and a risk or governance function usually owns the framework. The most common failure is an organisation treating an AI governance committee as the assurance, when the committee is part of what should be audited.
How often should an AI system be re-audited?
More often than a conventional system, because an AI control can move without a change ticket. Rather than fixing an interval, tie re-examination to events: retraining, a vendor model update, a material shift in the input population, or an expansion of the system's authority. An opinion issued without such a trigger should carry an explicit as-at date.
What evidence should an AI system produce for audit?
At minimum: the approval record showing what was decided and on what basis, a log capturing each consequential decision with its inputs at the time, the authority boundary in writing, the named owner, the documented stopping condition, and the provenance of training or retrieval data. Evidence that has to be reconstructed after the fact is usually not evidence.
Does auditing AI require new controls or the existing ones?
Mostly the existing ones. Access control, change management, segregation of duties, logging, vendor management, and continuity apply unchanged. Four areas are genuinely new: the authority an autonomous system holds, evidence that cannot be reproduced by re-running an input, drift that bypasses change management, and the provenance of the data the system learned from.
What qualifications does an IT auditor need to audit AI systems?
No new certification is required to begin. CISA remains the relevant audit qualification, and CISM or CISSP cover the governance and security side. What an auditor does need is the ability to read an evaluation result, an approval record, and a system log critically, and to tell the difference between a model claim and a control that can be tested.
Get the next essay by email
One practical essay a month on AI governance, agentic AI, and digital delivery in the UN system. No marketing, no forwarding of your address.