05 / 06

Test and monitor

Test AI systems before they go live. Monitor them continuously after. AI is not "set and forget" — models drift, data changes, and operational conditions evolve. Governance requires ongoing vigilance, not a one-time approval.

High priority Technical integrity — prevents silent system failures Accountable parties: Data Science, QA, InfoSec, Operations
Pillar explained

What this pillar requires from your organisation.

An AI system that worked correctly at deployment is not guaranteed to work correctly six months later. Training data becomes stale. Operational conditions change. The model's assumptions diverge from reality. This is called algorithmic drift — and it is silent. No alarm sounds when an AI model starts producing unreliable outputs. You need a monitoring system to detect it before it causes harm.

Pre-deployment testing is the first requirement. Every AI system must be tested against documented acceptance criteria — for accuracy, robustness, bias, security, and usability — before it goes live on a real project. A deployment authorisation sign-off from the accountable person is required. Verbal approval is not sufficient.

For high-risk or generative AI systems, additional evaluations apply: red teaming to expose system vulnerabilities, prompt manipulation testing for generative tools, and independent auditing for systems with high autonomy or significant consequence.

An AI system that is not monitored is not governed. If you cannot detect when a model starts failing, your governance framework is theoretical.
Contractor scenario — Document control

When an untested AI flags the wrong documents as non-compliant

An AI-powered document management system is deployed to check project documentation for compliance with contract specifications. The system was not tested against the specific contract template in use on this project. It begins flagging compliant documents as non-compliant and approving non-compliant ones — because the training data did not include the relevant contract format. The errors go undetected for two months because there is no monitoring framework. By the time the issue is identified, significant rework has been required and a delay claim has been lodged.

Contractor scenario — HSEQ

When a safety monitoring AI drifts without detection

A site safety monitoring AI platform — used to analyse observation data and identify precursors to incidents — was accurate and reliable at deployment. Over eighteen months of operation on an evolving project, the physical site conditions, workforce composition, and work types have changed substantially. The model's performance has degraded significantly, but no monitoring framework exists to detect this. The system continues to generate weekly safety dashboards that management trusts — based on a model that no longer reflects site reality.


Pre-deployment testing

What every AI system must be tested against

Pre-deployment testing must be documented, conducted against these criteria, and signed off by the accountable executive before the system goes live.

Accuracy

Does the system produce correct outputs against a representative test dataset? What is the acceptable error rate for this use case and risk level?

Robustness

Does the system perform reliably under edge cases, unusual inputs, and conditions that differ from the training environment? How does it fail — gracefully or catastrophically?

Bias

Does the system produce systematically different outputs for identifiable groups — gender, age, cultural background, employment type — that cannot be justified by the purpose of the system?

Security

Is the system resilient to adversarial inputs, prompt injection (for generative AI), and data poisoning? Have known vulnerabilities in the underlying model been assessed?

Usability

Can the operators who will use the system understand its outputs and identify when those outputs may be incorrect? Is the system's interface appropriate for its operational context?

Data governance

Does the system handle personal data in accordance with privacy obligations? Are data usage rights, retention periods, and cross-border transfer restrictions documented and observed?

Accountability mapping

Testing and monitoring RACI — who does what

Pre-deployment testing and ongoing monitoring require clear ownership across technical, quality, security, and operational functions.

AAccountable
RResponsible
CConsulted
IInformed
Function / Activity Executive Accountable Official Data Science & Analytics Quality Assurance Information Security Operations & Business Units
Pre-deployment testing I R A C C
Deployment authorisation sign-off A C R C I
Ongoing performance monitoring I R C C A
High-risk system evaluation (red teaming) I R C A I
Data governance & IP compliance I R C A C
Required audit documentation

The documents an auditor will ask for.

These four artefacts form the minimum documentation set for Pillar 05. They must exist per system, not as generic organisational policies.

Document 01

Pre-Deployment Testing Report

A formal report documenting the testing methodology used, the acceptance criteria applied across each testing dimension (accuracy, robustness, bias, security, usability), the test results, any issues identified, and the remediation actions taken. Must be signed off by the Quality Assurance function and the Executive Accountable Official before deployment proceeds.

Required before deployment Download template — available in full pack
Document 02

Continuous Performance Dashboard Framework

A defined set of metrics and monitoring processes for each AI system in operation — tracking accuracy trends, error rates, user override frequency, and any indicators of algorithmic drift. Must specify the monitoring cadence, the person responsible for reviewing results, and the escalation thresholds that trigger a formal review or system suspension.

Ongoing — post-deployment Download template — available in full pack
Document 03

Data Governance & Cybersecurity Protocols

Documented procedures covering: data provenance and quality assessment for training and operational data, personal data handling in AI systems (aligned to Privacy Act obligations), intellectual property protections for generative AI outputs, data residency and cross-border transfer controls, and cybersecurity protections specific to AI interfaces and model endpoints.

Reviewed annually and after incidents Download template — available in full pack
Document 04

AI System Monitoring Plan

A per-system plan that specifies what is being monitored, by whom, at what frequency, with what tools, and what constitutes an acceptable versus unacceptable result. Must include escalation triggers, the responsible reviewer, and linkage to the Incident Response Plan for when thresholds are breached. This is distinct from the dashboard framework — the monitoring plan is the governance document; the dashboard is the operational tool.

Per system — updated after material changes Download template — available in full pack
Audit-ready checklist

Five questions a compliance auditor will ask.

Your progress is saved automatically. Print this page to include your checklist status in a compliance submission.

Pillar 05 — Testing and monitoring checklist

Tick each item when you have the required documentation in place.

0 of 5 items complete 0%
Audit risk

Common gaps auditors find in contractor submissions.

These are the findings that appear most frequently when contractors are assessed against Pillar 05.

No pre-deployment testing documentation

AI systems are procured and deployed without formal pre-deployment testing — or with informal testing that was never documented. "The vendor tested it" is the most common response. Vendor testing validates the system against the vendor's test environment, not your operational context. The deployer is required to conduct and document their own acceptance testing before use on a live project.

No monitoring after deployment

Organisations test AI systems at deployment and then operate them indefinitely without structured monitoring. Performance dashboards exist for IT systems but not for AI-specific metrics — accuracy trends, override frequency, data drift indicators. Without monitoring, algorithmic drift is invisible until it produces a significant operational failure. By then, the harm is done and the audit trail is absent.

Generative AI used without data governance

Project teams use generative AI tools — ChatGPT, Copilot, Gemini — to draft documents, analyse contracts, and summarise meeting notes. No data governance policy has been established governing what data can be input into these tools. Commercially sensitive project data, personal worker information, and client-confidential material is routinely processed by third-party AI systems under no data processing agreement. This is both an IP and Privacy Act exposure.

No escalation threshold defined

Organisations have monitoring processes but have not defined what constitutes a problem that requires escalation. Monitoring without a trigger is observation without governance. The Monitoring Plan must specify: at what error rate does the system require a formal review? At what point does it require suspension? These thresholds must be documented before deployment, not determined ad hoc after a failure.

Next step

Put testing and monitoring frameworks in place before deployment — not after failure.

The full AI Governance Compliance App includes pre-deployment testing report templates, a monitoring plan framework, and data governance policy templates — all configured for the operational realities of Australian heavy industry projects.

← Previous 04 — Share essential information Next → 06 — Maintain human control