Maximum delay to answer
Definition
Maximum delay to answer
Maximum delay to answer is the longest a caller waits before an agent picks up — the ceiling, not the average. Contact centers pair it with a service-level target, 80% answered inside 20 seconds. Any call past the cap counts as a breach.
Key takeaways
- Maximum delay to answer caps the WORST acceptable hold time, not the average.
- Standard SLA pairs it with “80% of calls answered inside 20 seconds.”
- Timing starts when the caller exits the IVR and ends when an agent answers.
- Breaches trigger financial penalties inside most BPO service-level agreements.
- Voice cap sits tighter than chat cap; emergency queues run tightest at 10 seconds.
The metric matters because it caps the WORST caller experience, not the typical one. Average speed of answer can read healthy at 28 seconds while 5% of callers wait three minutes and churn, complain, or call back.
It’s both operational and contractual — most BPO contracts write the cap into service-level agreements with financial penalties if breached. Workforce planners staff around the cap, not around the average.
The paired percentage target (the “80” in 80/20) decides how strict the cap is in practice. A 90/20 SLA is far tighter than an 80/20 SLA at the same 20-second cap because it forbids more breaches per hundred calls.
How it works
The call center clock starts the moment a caller exits the IVR menu and stops the second an agent answers. If that gap crosses the cap, the platform flags a breach in real time and logs it against the SLA.
Most workforce management platforms (Genesys, NICE CXone, Five9, Talkdesk) set the cap as a hard threshold and stream breach counts to real-time dashboards. Supervisors see the meter creep, then reroute or add agents before the queue tips over.
Modern platforms tag each call with a queue-timestamp and an answer-timestamp, then subtract to get the wait interval. That timing feeds two reports: a rolling breach count and a monthly SLA breach rate.
Real-time reroute logic reads breach rates the moment they cross a warning threshold. When the daytime peak threatens the 20-second cap, the system spills overflow to a nearshore backup site or promotes idle agents from a lower-priority queue.
Cap targets shift by channel and by urgency. Voice runs tighter than chat, and emergency lines run tightest of all. The table below shows the bands that most 2024 contracts follow:
| Channel | Typical cap | Common SLA target |
|---|---|---|
| Inbound voice (sales) | 20 seconds | 80% answered ≤ 20s |
| Inbound voice (support) | 30 seconds | 80% answered ≤ 30s |
| Emergency / safety lines | 10 seconds | 90% answered ≤ 10s |
| Premium / VIP queues | 15 seconds | 90% answered ≤ 15s |
| Web chat | 45 seconds | 80% answered ≤ 45s |
Breach reporting varies. Some platforms treat the cap as a rolling 15-minute average, others as a hard per-call flag. The industry primer at Wikipedia’s contact-center overview still uses the per-call definition.
Some contracts blend the cap with abandonment metrics. If a caller hangs up during the wait, the cap gets a check because the customer left; if they stay past the cap, the breach counts once against the SLA.
Examples
Teleperformance, the world’s largest BPO, handled more than 8 billion customer interactions across 200,000 agents in 2024, with voice queues held to a 20-second cap and premium accounts pushed to a 15-second ceiling.
Emergency dispatch centers — 911 in the United States and 999 in the United Kingdom — enforce a 10-second cap because a 45-second wait can decide a life. Both regulators publish monthly breach rates.
Healthcare triage lines sit somewhere in the middle. A 20-second cap on nurse hotlines is common in 2024, tightening to 10 seconds if the caller flags chest pain or bleeding through a screening IVR.
Retail banks running premium concierge lines target 90% answered inside 15 seconds. A Harvard Business Review study found that reducing customer effort, including wait time, is a stronger loyalty driver than trying to delight.
Web chat runs looser. Most 2024 SLAs allow 45 seconds because typing gives the agent breathing room, and callers stay engaged. Live sports betting and stock trading push chat back down to 20 seconds during peak.
Airlines apply a two-tier cap during operational disruptions. Elite tier callers stay on a 15-second cap; economy flows to a 60-second cap with a call-back option, protecting the SLA without stranding low-tier passengers on hold.
Utility companies got hit hard during the 2023-2024 winter storms. Many missed their 30-second cap on outage lines and paid out state-mandated credits to customers who waited longer than five minutes.
Vetted providers in Outsource Accelerator’s verified BPO directory publish their cap targets and breach rates alongside quality scores, so clients can compare like-for-like.
Related terms
Terms that sit next to maximum delay to answer in most contact-center scorecards. Each has its own metric definition, its own target, and its own line inside the SLA. They intersect but do not substitute for the cap.
- Average speed of answer: the mean wait time, not the ceiling.
- Service level: the paired % target that defines an acceptable breach rate.
- Abandonment rate: the share of callers who hang up before an agent answers.
- First call resolution: whether the caller’s issue closes on the first contact.
- Average handle time: how long each answered call actually lasts.
- Call center: the operational unit where the cap is enforced.
- Workforce management: the scheduling discipline that staffs to the cap.
FAQ
Common questions about maximum delay to answer, from how contact centers set the cap to how they measure it and how it interacts with the surrounding SLA metrics.
What is a maximum delay to answer in a call center?
Maximum delay to answer is the longest a caller should wait before a live agent picks up. Contact centers set it in seconds paired with a service-level percentage, typically 80% of calls answered inside 20 seconds. Calls past the cap breach the SLA.
How is maximum delay to answer different from average speed of answer?
Maximum delay to answer is the ceiling; average speed of answer is the mean. A center can hit its ASA target while quietly breaching the cap on 5% of calls. Those tail-end callers are the ones who churn.
What is a typical maximum delay to answer for inbound support?
Most 2024 support SLAs cap voice at 30 seconds with an 80/30 target. Sales queues tighten to 20 seconds, and emergency lines cap at 10 seconds. Web chat usually runs looser at 45 seconds.
What happens when the cap is breached?
Most BPO contracts trigger service-level penalties calculated as a percentage of monthly fees. Repeat breaches can escalate to contract review or termination. The provider also loses bonus tiers tied to over-performance.
Who owns the maximum delay to answer metric?
The workforce-management team owns the staffing math and the operations manager owns real-time recovery when the queue tips over. Finance owns the SLA penalty exposure downstream. Everyone reads the same dashboard.
Does chat use the same cap as voice?
No: chat SLAs typically run 45 seconds because typing conceals the delay from the caller. Voice caps sit tighter, usually 20 to 30 seconds.
Explore more OA terms and guidance at Outsource Accelerator.







Independent




