Depth 06.1 · Managed services · Level 3

Monitoring and support

Monitoring and support is the work of watching production systems around agreed hours, raising an alert when something breaks or starts to degrade, and having a named person respond. Support covers the human side: someone to report a problem to, a record of each incident, and a follow up so the same cause does not return.

Specification

Input
The list of systems and user journeys that matter most, the hours they must be covered, and who on your side is called for which decisions.
Method
We monitor from the user's side as well as from inside the system, tune alerts until each one needs action, and handle incidents to a written procedure. Response targets are those agreed in your contract; we do not use a generic published figure.
Deliverable
Live monitoring, an incident log, and a short written review after each serious incident.
Measured by
Acknowledgement and resolution times against the contracted targets, alerts that needed no action, and incidents with a repeated cause.

Most often bought inPayments and fintech, B2B software, Professional services, Logistics and supply chain, Ecommerce

The artefact, before and after

Live monitoring, an incident log, and a short written review after each serious incident.

Before: a blank monitoring and support record with a row for each of uptime and synthetic journey checks, error and log alerting, severity levels defined in the contract, incident log and written reviews, monthly report, and no entries. After: every row carries a finding and a checked status, which is the state the work hands over.

Inside it

  • 01Uptime and synthetic journey checks
  • 02Error and log alerting
  • 03Severity levels defined in the contract
  • 04Incident log and written reviews
  • 05Monthly report

Typical first engagement

The usual order of a first piece of work. Timing and price are set at scoping, once we have seen the stack.

  1. 01Coverage mapYou get: What is watched today, what is not, and who is told when it breaks.
  2. 02Alerts and routingYou get: Checks on the paths that matter, with alerts sent to named people.
  3. 03RunbooksYou get: A written first response for each alert, kept in your repository.

Managed services

Next in this line

Ask about this part of the stack

Write one question about monitoring and support in your own systems. We answer it in writing and name what we would need to see to answer it properly.

Send one question

A written reply, no call needed.