AI Capability Evidence Framework
The capability evidence framework records one row per affected role: the capability required, what counts as evidence of it, where that evidence is recorded, and when it is reassessed. This is what turns a training programme into a durable organisational expectation rather than an event.
Training that is not recorded as an expectation decays within a year and has to be repurchased. Evidence attached to roles survives staff turnover in a way that attendance records do not.
The template
For each affected role, complete this row. This is the document that turns a programme into a durable expectation.
| Role | Capability required | Evidence of capability | Where it is recorded | Reassessed |
|---|---|---|---|---|
| Example: service manager | Can assess whether a proposed system should be adopted in their service | Completed use case assessment reviewed by the CoE | Annual objectives | Yearly |
| Example: data engineer | Can design and run an evaluation against live-like data | System passing internal gate | Role profile and objectives | Yearly |
How to use it
- Complete one row per affected role rather than per course.
- Define evidence as something observable, not attendance.
- Record it in the system that already governs the role, so it survives the programme that created it.
- Set a reassessment interval, because capability claims age as tools change.
Where this goes wrong
The failure modes below are the ones worth checking for first. Each describes a way this artefact stops doing its job while still appearing to be in use.
- Evidence is collected after the claim is made. Assembling proof to support a decision already taken produces selection, not evidence. Fix the measure before the system is built and record what would count as failure.
- A demonstration is accepted as evidence of capability. Demonstrations run on chosen inputs in favourable conditions. Capability is what the system does on the distribution it will actually meet, including the cases nobody prepared for.
- Accuracy is reported without a baseline. A model at eighty-five per cent means nothing until you know what the existing process achieves and what chance would achieve. Many published gains disappear against that comparison.
- The evidence is never tied to a decision. A framework that produces a report rather than a go, no-go, or conditional-approval outcome adds process without adding control.
Common questions
How do you make AI training stick?
Record capability against roles rather than tracking course attendance. For each affected role, state the capability required, what counts as observable evidence of it, where that evidence is recorded, and when it is reassessed. This turns a programme into a durable expectation that survives staff turnover, where attendance records do not.