Mean Time between Failures
Definition
Mean Time between Failures
Mean time between failures is the average operating time a repairable system delivers between one breakdown and the next, calculated from total uptime divided by failure count. It is the headline reliability number, and it is routinely misread as a lifespan guarantee.
The word repairable is doing real work. The measure applies to systems that get fixed and returned to service, not to components that are discarded once they fail.
It is also an average, not a promise. A server with a five-year mean time between failures can still fail in week one — without contradicting the figure.
Key takeaways
- Mean time between failures divides total operating time by the number of failures.
- The measure applies only to repairable systems, not to single-use components.
- Repair time is excluded; that belongs to mean time to repair.
- A high figure describes an average, never a guaranteed run length.
How it works
Mean time between failures is calculated by dividing the total operating time of a system across a period by the number of failures recorded during that same period, giving an average uptime between breakdowns.
The formula is: total operating time ÷ number of failures.
Reliability terminology is precise here, and the distinctions matter when writing a contract.
| Measure | What it covers | Applies to |
|---|---|---|
| Mean time between failures | Uptime between breakdowns | Repairable systems |
| Mean time to failure | Time until first failure | Non-repairable items |
| Mean time to repair | Time spent restoring service | Repairable systems |
| Availability | Uptime as a share of total time | Both, contractually |
The distinction in row two is not pedantry. The NIST/SEMATECH e-Handbook notes that a repairable system can be restored to satisfactory operation by any action, and that failure rates and hazard rates apply only to first failure times in non-repairable populations.
The underlying property has a formal definition. The American Society for Quality defines reliability as the probability that a product, system, or service performs its intended function adequately.
That same definition covers operating in a defined environment without failure across a specified period.
Define failure before measuring anything. A degraded service that still responds may or may not count — and that single choice moves the figure by an order of magnitude.
Read it beside availability rather than instead of it. A system that fails rarely but takes two days to fix can have worse availability than one failing weekly and recovering in minutes.
The metric belongs inside the wider service level agreement (SLA) rather than being quoted on its own in a sales conversation.
Never extrapolate the figure to an individual unit. Population averages say nothing about which specific machine fails next week.
Examples
Reliability expectations vary by how much redundancy exists and how expensive an outage actually is. Five cases show how the measure gets applied and contracted.
Data centre operators quote uptime rather than mean time between failures. Redundancy means individual component failures never reach the customer, so availability is the honest measure.
Manufacturers publish component figures in hours. Those numbers come from accelerated testing, not from field experience, which is why field results often differ.
Telecom networks measure it per element and per route. A single link failing matters far less than a route with no alternative path.
Software platforms adapted the term for services. Failures there mean incidents rather than hardware faults, so the definition has to be written into the contract.
Outsourced infrastructure teams report it alongside repair time — buyers should ask for both, since a strong figure paired with slow recovery still produces poor availability.
Related terms
Mean time between failures sits inside the reliability and availability family used across IT and infrastructure operations. The terms below cover the roles, the contracts, and the recovery planning around it.
- Site Reliability Engineer: the role that owns availability targets and failure budgets.
- Network Engineer: the role managing redundancy across links and routes.
- Systems Administrator: the role handling day-to-day recovery and patching.
- Technical Support Engineer: the role restoring service when a failure reaches users.
- IT Support Technician: the front line that logs and escalates failures.
- Business Continuity Plan (BCP): the document covering what happens during a long outage.
- Service Level Agreement (SLA): the contract that turns reliability into an obligation.
FAQ
How is mean time between failures calculated?
Divide total operating time across the period by the number of failures recorded in that same period.
Does it include repair time?
No. Time spent restoring service is measured separately as mean time to repair.
What is the difference from mean time to failure?
Mean time between failures applies to repairable systems, while mean time to failure applies to items discarded once they fail.
Does a high figure guarantee a long run?
No. It is a population average, so an individual unit can still fail early without contradicting it.
Why report availability alongside it?
Because rare failures with slow recovery can produce worse availability than frequent failures fixed quickly.
How should failure be defined?
Explicitly and in writing, since counting degraded performance as a failure changes the figure dramatically.
Curious how reliability commitments are structured across outsourced infrastructure teams? Start with Outsource Accelerator.







Independent




