Sampling Evidence: Model Cards, Logs, Datasets and Approvals
An internal auditor cannot examine every piece of evidence. An organisation might have 50 deployed AI systems, 200 team members, 12 months of logs, and thousands of data records. Sampling is how auditors are efficient without sacrificing credibility. Sampling means selecting a subset of evidence that represents the whole population. When auditors sample credibly, their findings are generalised to the whole population.
Model cards are foundational evidence. A model card documents what an AI system is, what data it was trained on, how it performs, what its limitations are, and what it is intended for. ISO/IEC 42001 Annex A.4 requires documented information about AI systems. Model cards are the practical implementation. When you audit, ask to see model cards for all deployed systems. If you cannot see a model card, the system lacks required documentation. If you can see 10 model cards and only 7 meet required standards, that is a finding. Sampling might mean: “Requested model cards for all 5 high-risk systems deployed in production. Examined all 5 in detail.” When high-risk systems are sampled, the sample is representative.
Logs are evidence of operational control. Logs record what happened: when a model was deployed, when data was updated, when changes were approved, when alerts fired. Logs reveal whether procedures written in the policy are actually followed. An approval log might show that changes are supposed to be signed off within 24 hours. Audit the logs and find that three of the last ten changes lacked any approval record. That is a major nonconformity. Sampling logs might mean: “Examined the model change log for the past three months. Verified that all 23 logged changes had documented approval records.” Or sampling might mean: “Examined logs for three representative systems over the past six months.” Logs are objective evidence; they show real behaviour.
Datasets are evidence of data governance. You cannot read every row of a training dataset, but you can ask questions: How was the dataset constructed? What is its source? How was it validated? Are there known limitations or biases? Were ethical review protocols followed? A dataset inventory might show that a historical employment data set was used to train a resume screening AI. That raises a red flag. Sampling datasets means: “Requested inventory of all training datasets used in production AI systems. Interviewed the data engineer responsible for preparing each dataset. Examined three representative datasets in detail for completeness, accuracy and potential bias indicators.” Understanding data provenance is essential to understanding system risk.
Approvals and change control records show whether governance processes are followed. An approval record states that a specific person, on a specific date, reviewed an AI system or change and approved it. Approvals must be recorded so that responsibility is clear and decisions can be traced. Audit the approval records and verify that required signatories actually signed. Look for blank approval lines, unsigned documents or approvals by people without authority. Sampling approvals might mean: “Examined 100 percent of system deployment approvals for the past year. Verified that each approval bore the signature of the required reviewer.” Audit trails prevent systems from evolving without oversight.

Interview evidence is often overlooked. Data scientists can explain procedures that documents do not capture. Ask them: “Walk me through how you train a model. Who approves the training data? Who approves the model before deployment? What happens if performance degrades?” Their answers reveal whether procedures are followed in practice or just documented on paper. Interviews surface actual workflows, not just ideal procedures.
Documentation does not always exist. Do not assume absence is acceptable. If a control requires documented information and you cannot find the document, that is a nonconformity. Record it precisely: “Clause 8 requires documented information for operational controls. The data governance procedure does not exist as a documented control.” Absence of required documentation is a specific, defensible finding.
Combining multiple evidence types often gives the clearest picture. You might request model documentation (document evidence), examine deployment logs (record evidence), interview the model owner (interview evidence), and observe the system in production (observational evidence). A finding that rests on only one piece of evidence is weaker than a finding that rests on multiple converging pieces. If both interview and documents confirm that approvals are missing, that is stronger than if only documents show missing approvals.
Finally, keep detailed notes during your sampling. Record what you examined, who you interviewed, what they said, and what documents you reviewed. These notes become your audit trail. They allow you to reconstruct your findings later and explain your reasoning to management. If a finding is contested, your notes are your defence.
