- MTBF is operating time divided by number of failures. MTTR is total repair time divided by number of repairs. Both are averages, and averages hide the failures that hurt.
- Availability = MTBF / (MTBF + MTTR), which is where the two become useful together.
- Measure repair time from the moment the machine stopped, not from when the technician arrived, or you will optimise the wrong half.
If MTTR only counts spanner time, a four-hour wait for a part is invisible and your metric looks healthy while the line stands idle. Measure from stop to running, and MTTR starts telling you about your spares and your response rather than your technicians.
The formulas
- MTBF = total operating time / number of failures. A machine running 500 hours with 5 failures has an MTBF of 100 hours.
- MTTR = total time to restore / number of repairs. Five repairs totalling 10 hours gives an MTTR of 2 hours.
- Availability = MTBF / (MTBF + MTTR). With the figures above: 100 / 102, about 98%.
MTBF applies to repairable equipment. For items you replace rather than repair, the equivalent is MTTF, mean time to failure. The distinction matters when you are comparing components with suppliers.
Where MTBF misleads
It is an average, so it says nothing about distribution. A machine with an MTBF of 100 hours might fail predictably every 100 hours, or run for 480 and then fail four times in a shift. The second is far worse to operate and looks identical in the metric.
Plot the intervals rather than only the mean. Clustered failures usually mean a repair that did not address the cause, and that pattern is invisible in a monthly MTBF figure.
Where MTTR misleads
- Scope. Wrench time versus stop-to-start. Only the second reflects what the business lost.
- Averaging. Twenty ten-minute resets and one eight-hour rebuild produce a comfortable mean that describes neither.
- What it excludes. Waiting for parts, waiting for a technician and waiting for a permit are usually the largest components and the most fixable.
Using them together
MTBF and MTTR pull on different levers. Raising MTBF is reliability work: root cause analysis, better preventive maintenance, addressing repeat failures. Lowering MTTR is response work: spares availability, documented procedures, faster diagnosis, better handover.
When availability is short of target, the split tells you which team owns the problem. Frequent short stops is a reliability problem; rare long stops is a response and spares problem.
Getting data you can trust
Both metrics are only as good as the timestamps behind them. The common failure is that stop and restart times are written up at the end of the shift from memory, which systematically shortens recorded downtime.
Capturing the stop and the restart at the moment they happen, by the person at the machine, is what makes the numbers usable. If that capture takes more than a few seconds it will not happen reliably, which is a good argument for it being a task on the technician's phone rather than a form in the office.