Depth 10.3 · AI engineering and automation · Level 3

Model operations

Model operations is the routine that keeps an AI system working after launch. Model providers update their models, documents change and the questions people ask drift, and any of these can quietly lower quality. Model operations catches that with repeatable tests, logs and a clear process for changing prompts or models.

Specification

Input
The AI system in question, a sample of real inputs and outputs, and people who can judge what a correct output looks like.
Method
We build a test set from real cases your team has reviewed, run it on every change to a prompt, model or data source, and log production use so new failure types are added to the set.
Deliverable
Evaluation harness and regression set: the test cases, the scoring rules and a script that runs them before any change ships.
Measured by
Scores on the regression set over time, and the number of changes that were blocked before release because a score dropped.

Most often bought inB2B software

The artefact, before and after

Evaluation harness and regression set: the test cases, the scoring rules and a script that runs them before any change ships.

Before: a blank model operations record with a row for each of reviewed evaluation set from real cases, automated scoring with human spot checks, regression run on every prompt or model change, production logging and failure review, version history of prompts and models, and no entries. After: every row carries a finding and a checked status, which is the state the work hands over.

Inside it

  • 01Reviewed evaluation set from real cases
  • 02Automated scoring with human spot checks
  • 03Regression run on every prompt or model change
  • 04Production logging and failure review
  • 05Version history of prompts and models
Deliverable · Level 4→ Evaluation harness and regression setAn evaluation harness is the test suite for an AI feature: a fixed set of inputs with expected behaviour, run automatically on every change to the prompt, model or retrieval, so a regression shows up before release instead of in front of users.

Typical first engagement

The usual order of a first piece of work. Timing and price are set at scoping, once we have seen the stack.

  1. 01InventoryYou get: Every model call in production with its prompt, model version and owner.
  2. 02HarnessYou get: An evaluation harness that runs your regression set on each change.
  3. 03Release gateYou get: A pipeline rule that blocks changes which lower agreed scores.

AI engineering and automation

Next in this line

Ask about this part of the stack

Write one question about model operations in your own systems. We answer it in writing and name what we would need to see to answer it properly.

Send one question

A written reply, no call needed.