MTTR (Mean Time To Resolve / Recover)
Average time between an incident being detected and the service being fully restored. Lower MTTR is a primary goal of incident response.
Source
- Atlassian — MTTR: Mean Time To Repair / Recover
Related terms
Uptime
The percentage of time a service is reachable and responding correctly over a given window.
Downtime
Any period during which a service fails to respond, returns errors, or is so degraded it is unusable.
SLA (Service Level Agreement)
A formal contract between a service provider and a customer that defines minimum service levels and the con…
SLO (Service Level Objective)
An internal reliability target — for example, "99.9% of API requests succeed within 500 ms over a rolling 3…
SLI (Service Level Indicator)
The raw measurement that an SLO is calculated from — request success ratio, latency percentile, or freshnes…
MTTD (Mean Time To Detect)
Average time between the start of an incident and the moment your team is first alerted.
Start monitoring in 2 minutes — 10 free commercial-safe monitors
Start free