MTTR is the average time it takes to get a broken service working again. MTTA, the shorter one, measures the gap between an alert going off and a person acknowledging it, and MTBF counts the stretch of working time between one failure and the next, averaged over however many failures you had. The first two get counted in minutes and the third in days, and all three move a long way depending on where you let the clock start.
That last part is the bit the definitions leave out, and most of the grief with these numbers comes out of exactly that. The letters themselves are easy enough. Nobody argues about the M or the T. What goes missing is an agreement inside the one company about which moment counts as the start of an incident and which counts as the end of it. Without one, two careful people can report the very same outage with numbers more than a day apart. Both of them will be telling you the truth.
The same Tuesday, reported three different ways
Take one made-up Tuesday, since the thing is easier to see with a single incident laid out in front of you. At 08:02 the sign-in on your app stops working for anybody coming in through SSO from Google Workspace or Microsoft 365, and nobody knows it yet, on account of the health check looking at a different page entirely. The first customer email lands at 08:09, polite, asking whether it is only them. The alert goes off at 08:14. Your on-call engineer is on a train with the phone face down on the table, acknowledges at 08:25, rolls back the change from the night before and watches sign-ins come good at 08:56, and as far as engineering is concerned the incident is over and the write-up can wait for the afternoon.
The desk has a different morning of it. Forty-odd tickets came in while sign-in was broken, and a good share of those people went and tried their password four or five times before they wrote, so when sign-in came back plenty of them were locked out of their accounts anyway and wrote in again, crosser this time. Reset requests run on into the afternoon. A couple of people who wrote at half eight do not hear back until after lunch, because the desk is working the pile from the top down, the way desks do. The last ticket from that incident gets closed on the Wednesday at 16:30.
So on Thursday three numbers go round for the one outage. The engineering review says 42 minutes, alert to restore, which is exactly the stretch their paging tool can see. Somebody who went and read the logs properly would put it at 54, since the thing was broken for 12 minutes before any alarm knew about it. The support lead has 32 hours and a bit written down, first ticket to last ticket closed. Each team counted the part of it that its own tools can see. Nobody fiddled anything, to be honest with you, which is what makes it so awkward to sort out afterwards, since there is no mistake anywhere to go and correct.
What does the R in MTTR actually stand for
Funnily enough, the answer depends on which decade the person asking learned the trade in. The letters started out in hardware maintenance, where the R meant repair and the clock ran only for as long as a technician was actually standing at the failed machine with the tools out. That suited a world of servers in racks and spare parts in a cupboard. It suits a software outage very badly, since the fix its self is often ten minutes and the finding out is the rest of it.
Software teams went and borrowed the letters and quietly changed what they meant. For most people who deploy code now the R is recovery, or restore, and the clock runs from the moment the service fails to the moment people can use it again, whether or not anybody understands yet why it broke. DORA’s research on software delivery did a good deal to make that the default when it put time to restore service among its four key measures.
Then the service desks got hold of it. On an IT ticketing system run the ITIL way, which is how most people in ITSM first learned it, an incident counts as resolved once normal service is back for the person who reported it. The deeper cause gets handed off to problem management as a separate record of its own, so on a desk like that resolve and restore land fairly close together. Plenty of desks that have never opened ITIL use resolve to mean the lot, cause found, fix shipped, follow-up sent, and on those desks MTTR can run on for a good long while, days sometimes. Respond crept in somewhere along the way too. It measures how quickly somebody starts work on the thing and so belongs in the same drawer as MTTA, and a team reporting that one as its MTTR will look very quick indeed next to a team that means recovery, without either of them doing anything sly to get there.
Write one sentence at the top of your own report saying which R you mean, and do it before the numbers underneath go anywhere near a slide, since the industry is not going to settle it on your behalf any time soon.
How do you calculate MTTR, MTTA and MTBF
Stay with the made-up month the Tuesday belongs to, 30 days of it with three incidents in all. The sign-in outage was down for 54 minutes. A week later outbound email sat stuck in a queue for 96 minutes on a Wednesday afternoon, and near the end of the month the API went slow enough, past its own threshold, to count as down for 30 minutes at two in the morning. Add the three together and you have 180 minutes of downtime for the month. Divide by the 3 incidents and your MTTR, meaning recovery, comes out at 60 minutes. That is the whole of it as arithmetic goes, the time spent broken divided by the number of times it broke.
MTTA goes the very same way on a different pair of timestamps. The three acknowledgements took 11 minutes, 4 minutes and 24 minutes, the last one being the page at two in the morning, so 39 minutes over 3 incidents gives an MTTA of 13 minutes. If you track detection as well, the gap between the service breaking and any alarm knowing, the month ran 12, 3 and 9 minutes, which is an MTTD of 8. It takes a minute.
For MTBF you need the calendar out. A 30-day month is 720 hours. Take off the 3 hours the service was down and you are left with 717 hours of working time, and 717 divided by 3 failures gives an MTBF of 239 hours, a little short of 10 days between one outage and the next on average. The same two figures give you availability if anybody upstairs asks for it, MTBF divided by MTBF plus MTTR, so 239 over 240, which is 99.58%, and you get the very same answer the long way round by dividing 717 by 720.
Written out as rules rather than as a month, with the made-up month kept alongside so the sums can be checked:
| Metric | Clock starts | Clock stops | How to work it out | The made-up month |
|---|---|---|---|---|
| MTTD, mean time to detect | The service fails | An alert fires or someone notices | Total time to detect ÷ number of incidents | 24 min ÷ 3 = 8 min |
| MTTA, mean time to acknowledge | The alert fires | A person acknowledges it | Total time to acknowledge ÷ number of incidents | 39 min ÷ 3 = 13 min |
| MTTR, mean time to recover (restore) | The service fails | Users can use it again | Total downtime ÷ number of incidents | 180 min ÷ 3 = 60 min |
| MTTR, mean time to resolve | The incident is reported | The incident record is closed | Total time open ÷ number of incidents | Depends on what your desk counts as closed |
| MTBF, mean time between failures | The service comes back | The service fails again | Total working time ÷ number of failures | 717 h ÷ 3 = 239 h |
| Availability | Start of the period | End of the period | MTBF ÷ (MTBF + MTTR) | 239 ÷ 240 = 99.58% |
One thing about averages, mind you, before you lean on any of these. A single long outage drags a mean a long way, so a month with one 4-hour incident and six 5-minute ones reports an MTTR of about 39 minutes, which describes not one of the seven. The median is the fairer middle for a spread like that, and we went through mean against median at more length on first response time, since the very same trap sits under that number too.
Why is MTTA the easiest of the three to flatter
MTTA is the favourite in a lot of teams, and fair is fair, it is the one you can do the most about by Friday. The acknowledgement its self records less than any of the others. What the timestamp captures is the moment somebody pressed acknowledge on the page, and a person can press acknowledge half asleep and then take twenty minutes to find the laptop, and the MTTA for that incident will still say four. Now and then a team sets its alerting up to acknowledge automatically once it posts into a Slack or Microsoft Teams channel. Nobody switches that on to cheat. It usually goes on to stop the channel filling up with red, and the MTTA drops to near nothing as a side effect while telling you nothing about whether a human being saw it.
The honest use for MTTA is smaller than the dashboards make it look. It tells you whether a page reaches a person at all, and how long that takes at night against during the day, and those two want reporting apart, since an average of both describes a shift nobody actually works. In the made-up month the page at two in the morning took 24 minutes and the two daytime ones averaged seven and a half, so the month’s 13 is hiding the most useful bit of it, which is that the overnight rota needs a second phone or a louder one.
On a help desk the same gap goes by another name. First response time, FRT on plenty of dashboards, is the customer’s MTTA, the stretch between their email arriving and a person answering it. It gets flattered the very same way by an auto-reply, which is why our piece on SLA compliance spends a good while on what should and should not be allowed to stop that clock.
How can MTBF improve while the outages get worse
Picture two quarters side by side. The first has 6 outages of about an hour each. The quarter after, the service only falls over 4 times, which looks like progress on the slide, except every one of those runs for 4 hours. A 90-day quarter is 2,160 hours, so the first quarter leaves 2,154 working hours over 6 failures, an MTBF of 359 hours, and the second leaves 2,144 over 4, an MTBF of 536 hours. So MTBF went up by nearly half, while the time your customers spent unable to use the thing went from 6 hours to 16.
Nothing in that is a trick. MTBF its self counts how often, and it has no opinion at all about how long or how bad, so it will cheerfully report a better quarter while the worse one is going on all round it. It only makes sense sat next to MTTR, which in this case went from one hour to four and would have told you straight away.
The other way it drifts is through what gets counted as a failure in the first place, and that is a decision somebody makes, most of the time without writing it down. If the bar moves from anything a customer notices to anything that breaches the SLA, a month with nine wobbles in it becomes a month with two, and MTBF goes up more than four times over without one line of code changing. Somebody goes and merges two outages into one on account of them sharing a root cause, and that does something similar. So does leaving planned maintenance out of the count, which is fair enough if customers were told in good time, and a good deal less fair if the window was announced in a channel nobody outside engineering reads.
Truth be told, the number was never really built for software. MTBF grew up in reliability engineering for electronics, the handbooks the US military wrote on predicting failure rates, MIL-HDBK-217 being the one people still quote, where parts wear out at rates you can estimate and failures turn up more or less at random. Software mostly breaks around change, a deploy, a config pushed on a Friday, the certificate nobody remembered to renew. So an MTBF on a software service ends up measuring how often you change things and how carefully you go about it, which is worth knowing and is still a different question from the one MIL-HDBK-217 was written to answer.
What is a good MTTR
We are not going to give you an industry figure for this. The benchmarks that get passed around, under an hour for this sort of team and under four for that sort, tend to arrive without a sample or a date or even a clear meaning for the R, and we could not trace one of them back to anything we would stand behind our own selves. So they are not on this page.
What you can do instead is work it backwards from the promise you have already made. A 30-day month has 43,200 minutes in it. If you have told customers 99.9% availability you are allowed 0.1% of that as downtime, which is 43.2 minutes a month, all in. At 99.5% it is 3.6 hours, and at 99% it is 7.2. Now take the number of incidents you usually have. At 2 incidents a month on a 99.9% promise, each one has to be over in a little under 22 minutes on average, detection and acknowledgement and the fix all inside that. The made-up month above, at 180 minutes of downtime, would have missed it by a long way while sitting comfortably inside 99.5%.
That is a good MTTR in the only sense that is any use to you, one that fits inside your own promise with a bit of room left over. Site reliability people call the leftover the error budget, and Google’s Site Reliability Engineering book explains it better than we could, but the heart of it is plain enough. When the budget runs low you stop shipping risky changes for a while, until the month turns over or the service settles down, and the rest of the time you get on with it.
Where does the help desk’s clock start and stop
Go back to the Tuesday. Engineering’s clock stopped at 08:56 and the desk’s ran on into the Wednesday afternoon. It is the desk’s number that comes closest to what the customer actually lived through, since for them the incident started when they could not get in and ended when somebody told them they could, or when the reset they asked for finally came. The engineering tools see none of that, since they only ever watched the servers.
So the desk needs its own clock, and it needs the incident tickets pulled together so they can be counted as the one thing instead of as forty separate complaints. What you actually say to those forty people while it is down, and once it is back, is a subject of its own and we went at it on incident management. The measuring half is simpler. It only needs every email to be a ticket with its own times on it, and something marking which tickets belong to which incident.
Over here at Maxdesk that is how the shared inbox works to begin with, every email that lands on your support address becomes a ticket with the minute it arrived written on it. An automation rule can bump the priority on anything with a word like down in the subject the moment it lands, so the incident mail floats to the top without anybody sorting it by hand. You can tag those tickets as incidents as well, so they can be pulled back out and counted as a set later. The SLA policy holds a target for the first reply and another for resolution against each priority, both counted in your own business hours, with a countdown sitting on every ticket. And the audit trail writes down every assignment, priority change, status change and note, which is how you find out a week afterwards the exact minute the desk first knew something was wrong, and when each of those tickets finally got closed. All of it is billed per workspace and never per agent, so the engineer who wants to go and read through the incident tickets can be given a login without it costing anybody anything.
The half of these numbers a help desk never sees
We do not watch your servers. Nothing inside of Maxdesk monitors anything or pages anybody at two in the morning, and the on-call rota and the status page belong to other tools as well, so MTTD and MTTA and MTBF all live in whatever your engineers use for those jobs. A help desk that goes and claims to report them properly would be guessing from the outside. What we see is the customer’s end of it. That is a real half, no two ways about it, and it is still only the half.
The dashboard we ship puts open tickets, SLA breaches, CSAT and first response time in front of you, with 7 day trends on them. An average resolution time per incident is not one of those tiles today. So the desk’s own MTTR is arithmetic you do your own self from the tickets you tagged. Export sits on the Pro plan, where it comes to an export and a spreadsheet and ten minutes of work. Without it, on the free plan, somebody reads the times off each ticket by hand, which is all right for three incidents and gets old quickly after that.
History is the other limit, and it bites MTBF harder than anything else on this page. The free plan keeps a rolling 3 months of tickets, reports and audit trail, which is enough to see a quarter and not enough to see a year, and MTBF over a single quarter is a thin number when you only fail a handful of times inside it. Pro keeps 12 months and Elite keeps 24. And the desk only counts people who emailed, since email is the one channel we run, so a customer who rang your office about the outage or complained about it on social media is missing from our clock altogether.
What should you write down after the next outage
Before the next one arrives, get engineering and the desk to agree a sentence for each clock, where it starts and where it stops and which R you mean, and put those sentences at the top of whatever report the numbers end up in. It takes a quarter of an hour, and you can draft the desk’s half your own self before anybody books a room for it. That alone would have spared the Thursday review most of its hour.
Then after the next incident, while it is still fresh in everybody’s head, write down six times for it by hand. A shared spreadsheet is plenty.
- When the service actually broke, read back from the logs afterwards
- When the first customer wrote in
- When the alert fired
- When a person acknowledged it
- When users could use the service again
- When the last affected customer was told, or had their follow-up sorted
Do that for three incidents and you have your first honest MTTR, MTTA and MTTD, with the desk’s clock written down beside engineering’s instead of arguing with it. The gap between the first customer email and the alert is worth a look on its own account. If customers keep beating your alerts to it, the monitoring is most likely checking the wrong page, which was the whole story of the made-up Tuesday.
If the desk’s times are the ones you cannot find, the emails scattered here and there across personal Gmail and Outlook inboxes and nobody sure which reply went out first, that is the part a help desk is for. Maxdesk keeps every one of those times on the free plan for as long as its 3 month window runs. The engineering columns stay with engineering, since the alerts and the logs sit with them.
Common questions about MTTR, MTTA and MTBF
Is MTTR the same as mean time to resolve?
Sometimes. MTTR can stand for repair, recovery, resolve or respond, and teams use it for all four. Most engineering teams mean recovery, from failure to the service working again. On an ITIL service desk resolve lands close to that, while other desks count resolve until the underlying cause is fixed too, which can run days longer. Write down which one yours means.
What is the difference between MTBF and MTTF?
MTBF, mean time between failures, is for things you repair and put back into service, so it measures the working time between one failure and the next. MTTF, mean time to failure, is for things you replace rather than repair, a disk or a fan, and measures how long one lasts before it fails the once.
How is MTTA different from first response time?
MTTA runs from an alert firing to a person acknowledging it, and it belongs to whoever is on call. First response time runs from a customer’s email arriving to a person replying, and it belongs to the help desk. They measure the same kind of gap from opposite ends of an incident, and automatic acknowledgements can flatter both.
What is MTTD?
Mean time to detect, the average gap between something breaking and any alarm knowing about it. It is usually worked out after the fact, by reading the logs back to the first error. If customer emails regularly arrive before your alerts do, your MTTD is longer than your monitoring thinks it is.
How do MTBF and MTTR give you availability?
Divide MTBF by MTBF plus MTTR. With an MTBF of 239 hours and an MTTR of 1 hour, availability is 239 ÷ 240, which is 99.58%. As a check the other way, 99.9% availability over a 30-day month allows 43.2 minutes of downtime in total.
Should a small IT help desk track MTBF?
Only once it has enough incidents, and enough history kept, for the number to mean something. Three failures in a quarter give an MTBF that swings wildly with one more or one fewer. For a small desk, MTTR and first response time, with their start and stop points written down, will tell you more for a good while yet.
Definitions and figures reviewed on 29 September 2026. The worked month is invented for the example; the Maxdesk plan details were checked on the live site the same day.
