Mean Time to Detect
Definition
Mean Time to Detect
Mean time to detect is the average interval between an incident actually beginning and the moment the organisation becomes aware of it. It is the blind window, and it is the only incident metric that is entirely about what you did not know.
Detection is not the same as reporting. An alert firing at 3am that nobody reads until 8am gives five hours of undetected exposure — whatever the monitoring dashboard claims.
Start time is the hard part. Establishing when an incident really began usually requires log reconstruction after the fact.
Key takeaways
- Mean time to detect averages the gap between incident start and organisational awareness.
- Incident start time must be reconstructed from logs, not taken from the alert timestamp.
- Customer-reported incidents represent complete detection failure and belong in a separate count.
- Detection time dominates total incident cost in security far more than repair time does.
How it works
Mean time to detect is calculated by summing the intervals between incident start and detection across all incidents in a period, then dividing by the number of incidents to give an average blind window.
The formula is: total detection intervals ÷ number of incidents.
How an incident is found tells you more than the average itself, so split the number by detection source.
| Detection source | What it indicates | Target share |
|---|---|---|
| Automated alert | Monitoring is working | Highest |
| Internal staff notice | Gaps in alert coverage | Moderate |
| Customer report | Detection failed outright | Near zero |
| Third-party notification | Serious blind spot | Zero |
The bottom two rows are the ones that matter for governance. Learning about your own outage from a customer or a regulator is a monitoring failure, not an incident-response one.
Security work treats this gap as central. The U.S. National Institute of Standards and Technology publishes the Cybersecurity Framework, whose CSF 2.0 release helps organisations understand and reduce cybersecurity risk.
National advisories show how fast the threat picture moves. The U.S. Cybersecurity and Infrastructure Security Agency publishes operational advisories on active threats against critical infrastructure, updated continuously.
Monitoring coverage is the main lever. Instrumenting the services customers actually touch beats instrumenting the ones that are easiest to measure.
Alert fatigue lengthens detection more than missing tools do. A team ignoring 400 daily alerts will miss the one that mattered — and the tooling will look perfectly healthy in the audit.
Ownership usually sits with a security operations center (SOC) for threats and with engineering for availability, and the two often report the number differently.
Never average security and availability incidents together. The distributions differ so much that the blended figure describes nothing.
Examples
Detection performance depends on monitoring coverage, on alert discipline, and on whether anyone is actually watching outside normal business hours. Five cases show the practical spread across very different kinds of operation.
Cloud platforms detect availability issues in seconds. Synthetic checks running continuously catch failures before most users notice anything.
Financial institutions detect fraud patterns in minutes. Transaction monitoring is real-time by regulation, so the blind window there is deliberately tiny.
Manufacturers detect equipment faults through sensor thresholds. Detection is fast, but only for the failure modes someone thought to instrument.
Retailers frequently detect checkout failures from customer contacts. Support ticket spikes become the detection mechanism, which is exactly the pattern to design out.
Outsourced monitoring teams report detection by source and by hour — buyers should ask for overnight figures specifically, since coverage gaps show up there first.
Related terms
Mean time to detect opens the incident timeline that response, repair, and resolution metrics complete. The terms below cover the functions, the tooling, and the contracts around it.
- Security Operations Center (SOC): the function watching for security incidents around the clock.
- Information Security Analyst: the role investigating alerts and confirming incidents.
- Live Monitoring: the continuous observation practice detection depends on.
- AI Observability: the emerging discipline extending monitoring to model behaviour.
- Site Reliability Engineer: the role owning availability detection and alerting.
- Business Continuity Plan (BCP): the plan triggered once a major incident is confirmed.
- Service Level Agreement (SLA): the contract that increasingly sets detection targets.
FAQ
How is mean time to detect calculated?
Sum the intervals between incident start and detection across all incidents, then divide by the number of incidents.
How do you establish when an incident started?
By reconstructing it from logs after the fact. The alert timestamp marks detection, not the beginning.
Why separate customer-reported incidents?
Because they represent complete detection failure and should never be averaged in with successful automated detection.
What lowers detection time fastest?
Better monitoring coverage of customer-facing services, followed by reducing alert noise so real signals get read.
Should security and availability incidents be combined?
No. Their distributions differ so much that the blended average is meaningless.
Who owns the metric?
Usually a security operations function for threats and engineering for availability.
Source partners running monitoring and incident detection for clients can compare models via Outsource Accelerator hubs.







Independent




