What Internal Audit Can and Cannot Assure About AI

Internal audit’s scope is changing as AI becomes embedded in business operations. Your audit function has core responsibilities in this new landscape: you must verify whether governance and control frameworks are in place and working, identify gaps where business appetite and actual risk do not align, and provide early warning before a material AI failure reaches the board. What internal audit cannot do is evaluate the statistical performance of a machine learning model in the way a data scientist can, or make business trade-offs between competing demands on an AI system. Those decisions belong to the model owners and compliance teams.
This distinction matters because it shapes what evidence you collect. When you audit hiring AI systems used by HR, you are not testing whether the model’s false-negative rate is 2 percent or 3 percent. You are verifying that management has defined acceptable false-negative rates, documented what testing was done, and put controls in place to catch harmful outcomes before they harm candidates. You are checking that HR explained the use of AI to candidates and that people have a right to request human review. You are sampling actual hiring decisions to see whether the system’s outputs were used properly, logged, and challenged when they seemed wrong. For example, a financial services firm once discovered that its AI-driven credit scoring model was not being applied consistently across regions. Although the model performed well statistically, internal audit flagged that the lack of standardisation led to inconsistent lending decisions and potential regulatory exposure. This was not a technical failure but a governance one.
Internal audit’s value in AI governance comes from asking whether the organisation has done what it said it would do. This means building audit programmes that test the end-to-end control environment, from the moment someone proposes an AI use case to the point where the system is retired or superseded. You will need to understand enough of the model life cycle to ask the right questions and follow the answers. You do not need to be a machine learning engineer, but you must learn what each stage involves, what could go wrong at each point, and what constitutes good evidence. A healthcare provider’s internal audit team reviewed an AI system used for diagnosing medical images. They focused not on whether the model was accurate in its predictions, but on whether the process had been properly documented, whether staff had been trained on its limitations, and whether there was a clear escalation path when the system flagged a case. This approach led to improved transparency and better integration of human oversight into the decision-making process.
The Three Lines Model Applied to AI
First line controls come from model owners and the teams that build or deploy AI. They decide what data to use and set thresholds. Second line controls are built by compliance, data governance and risk teams who design the frameworks and policies that constrain first line decisions. Third line oversight is your role: you verify whether first and second line controls work in practice. An e-commerce company’s AI fraud detection system was flagged for inconsistent performance. Internal audit reviewed the division between first and second line responsibilities. It found that although the fraud team had implemented thresholds, the compliance team had not established clear policies on how those thresholds should be applied. This led to a re-evaluation of responsibilities and improved alignment between the teams.

The Governance Skip and Why It Fails
Many organisations start AI audit work by diving straight into model testing. They ask for code repositories and evaluation metrics. This approach surfaces technical problems but misses governance failures entirely. If no one documented what the model should do or tested it against that specification, then technical excellence is irrelevant. Audit governance first, then model controls. A UK retail chain once conducted an audit of its customer segmentation AI. The audit team focused on performance metrics and found the model was highly accurate. However, upon deeper investigation, they discovered that the model had been built using data that was not compliant with GDPR. This led to a major compliance issue that had been overlooked because the audit had focused on performance rather than governance. This example illustrates why governance must be the starting point of any AI audit.
Scoping Your First AI Audits
Start with high-risk systems: those that make decisions affecting people or operate at scale. Hiring systems, credit decisioning and fraud detection models are natural starting points. Shadow AI and unauthorised tool adoption run close because they sit outside all formal controls. A telecommunications company began its AI audit by reviewing its customer churn prediction model. Although the model was in production, it had never been formally approved or tested for bias. Internal audit flagged this gap and led to a process review that included bias testing and stakeholder communication. This was a critical step in ensuring that the model was not only technically sound but also ethically and legally compliant. Another example is a logistics firm that used an AI system to predict delivery times. Although the system was performing well, it had not been tested for fairness or transparency. Internal audit recommended that the model be audited for unintended bias and that its decision-making process be made explainable to customers. This led to a more dependable and trustworthy system that improved customer satisfaction and regulatory compliance.
