Auditing automated eligibility decisions: what a working audit actually requires
Most audit obligations in current legislation can be satisfied without anyone examining whether a system produces defensible decisions. This paper sets out what an audit would have to include to be worth requiring.
Audit has become the default remedy in technology regulation. Where a legislature is uncertain what to require, it requires an audit. The appeal is obvious: audit is familiar from finance, it implies independence, and it defers the hard questions to a later technical process.
The difficulty is that almost none of the audit obligations now in force specify what the auditor must look at, what would constitute a failure, or what happens when one is found. An organisation can comply by commissioning a review of its own documentation, conducted by a firm it selects and pays, against criteria it supplies, with a report it is not required to publish. That is not an audit. It is a procurement exercise with a reassuring name.
Four things an audit must contain
We propose that any audit requirement worth imposing specifies at least the following. Each is drawn from practice in clinical governance and financial audit, where these questions were settled decades ago.
- A defined subject. Not “the system” but the specific decision, the population it is applied to, and the outcome being assessed. An audit of a triage tool that does not state which patients and which outcome has assessed nothing.
- An external comparator. Performance measured against the decision the institution would otherwise have made, not against the system’s own training objective. Systems routinely optimise for a proxy that diverges from the goal.
- Access sufficient to fail the system. An auditor who cannot query the model, examine held-out cases, or interview the staff overriding it cannot find anything the operator did not already know.
- A published result. Including negative findings. An audit regime in which failures are confidential produces a market for auditors who do not find failures.
The override problem
The most consistent finding in our review of deployed decision-support systems is that the formal right of human override is preserved while the practical ability to exercise it is not. Staff are permitted to disagree with the system, and are measured on throughput that disagreement reduces. Override rates fall to near zero within months, and the institution reports this as a sign of alignment.
A right of override that is never exercised is not a safeguard. It is a record of the pressure not to use it.
Any audit that reports an override rate without reporting the incentives acting on the people who hold that right is reporting a number without its meaning. We recommend that override rates be published alongside the performance measures applied to the staff concerned.
What we are asking for
This paper is written to be drafted from. Sections 4 and 5 are set out as model provisions that a regulator or ministry can adapt directly, with drafting notes explaining which choices are load-bearing and which are stylistic.
We would rather this paper were criticised than agreed with. If the model provisions would fail in your jurisdiction, or would be unworkable in your institution, that is the most useful thing you can tell us.