Depth 10.3.1 · Model operations · Level 4
Evaluation harness and regression set
An evaluation harness is the test suite for an AI feature: a fixed set of inputs with expected behaviour, run automatically on every change to the prompt, model or retrieval, so a regression shows up before release instead of in front of users.
Specification
- Input
- The AI feature, real inputs it has seen, and your team's judgement of what a correct answer looks like.
- Method
- We collect representative and failing cases, write a check for each (exact match, rubric or citation validity), wire the suite into your pipeline and add every new failure found in use to the set.
- Deliverable
- The regression set in your repository, the harness in your pipeline, and a report format that shows which cases changed.
- Measured by
- Cases passing per release, and failures found in use that were not already in the set.
What the file contains
- Case collection
- Checks per case
- Pipeline integration
- Change report per release
Part of Model operations.
Ask about the evaluation harness and regression set
Tell us which systems it should cover. We reply with the access we would need and what the first version would contain.
Ask about this deliverable
A written reply, no call needed.