Model Audit Provisions for Automated Decisions
Model wording a regulator or ministry can draft from directly, with notes on which choices are load-bearing.
The distance between what these systems can do and what any institution is equipped to oversee is the widest in any field we work on, and it is still opening. Rules drafted for products that ship once, in a fixed form, with a manufacturer who can be inspected, do not fit systems that change weekly, behave differently in each deployment, and are frequently assembled from components nobody in the chain fully controls.
The failure mode we see most often is not absent regulation but unenforceable regulation. An obligation is written in terms that cannot be audited — 'appropriate safeguards', 'adequate human oversight' — and is then administered by a body with two technical staff for several hundred deployments. What follows is a compliance industry rather than compliance: documentation improves, assessment reports multiply, and nobody establishes whether the systems concerned produce defensible decisions. A well-documented bad system reads better than a badly documented good one, every time.
We work on governance that survives contact with practice. Obligations expressed with enough technical specificity that an auditor could fail a system against them. Evaluation that measures behaviour in the population actually being treated rather than laboratory performance. Review triggered before deployment rather than after harm, with a stopping rule defined in advance. Supervisory bodies with the staff, the pay scales and the statutory access to exercise the powers they already hold. And, throughout, the question that is asked far less often than it should be: does this mechanism change outcomes, or does it only change what gets written down?
Model wording a regulator or ministry can draft from directly, with notes on which choices are load-bearing.
An eight-page briefing for legislators on why enforcement is limited by staffing rather than statutory powers.
An empirical study of what actually limits enforcement in technology supervision — comparing statutory powers against the staff, skills and time available…
A controlled trial of whether removing throughput pressure restores meaningful human override in a deployed decision-support system — with the result published…
An open call inviting teams anywhere to measure human override of automated decisions using a shared published protocol, so results can be…
Shared technical capacity across supervisory authorities that individually cannot justify the headcount but together comfortably can.
Developing an auditable standard for the review of automated decisions affecting individuals, drafted with the regulators and engineers who would have to…
A working session on the model audit provisions, with supervisory staff, engineers and caseworkers in the same room.
Most audit obligations in current legislation can be satisfied without anyone examining whether a system produces defensible decisions. This paper sets out…
Our policy paper on audit obligations is published, with model provisions a regulator can draft from directly.
Medicine spent a century building the institutional machinery for deploying interventions that can harm people. Much of it transfers directly, and is…
The group is now accepting participants from outside the institute — particularly people who expect the draft standard to fail.