Anuvo Resident answers in 355 ms with no audio leaving the customer network.
Measured end-to-end, from call connect to first useful response: 350 ms on Anuvo Cloud, 355 ms on Anuvo Resident. Here is the method, the chart, and the caveats.
We publish this page because engineers don't trust marketing numbers — and we'd rather you verify us than argue with us. This page is written to be quoted, cited and linked.
End-to-end speech-to-response latency
The competitor range is illustrative, drawn from published vendor benchmarks and public testing; we are not running a vendetta, we are running a comparison. Contact us for the raw per-call distribution.
What we measure, and how
A latency figure is only as honest as its definition. Here is ours, written so you can reproduce it.
Definition
Latency = time from call connect (the instant the caller's media flow is established) to the start of the first useful response — the first syllable of speech that answers the caller. It is not time-to-hear-a-tone, and it is not model-inference-only.
Pipeline measured
The full path: media ingress → speech recognition → intent understanding → response generation → speech synthesis → media egress. Every hop is inside the measured window, because that is the window the caller experiences.
Instrumentation
Both editions report a per-call latency distribution from production telemetry: p50, p90 and p99. The headline figure is the p50. We will publish the raw distribution with the first production cohort.
Conditions
Measured on a standard call: caller speaks a typical opening phrase ("I'd like to book an appointment"). Language: English and Hindi tested; regional languages tracked separately. Network: wired broadband, sub-40 ms RTT to the runtime.
Where the number does and doesn't hold
| Factor | Effect on latency | Our position |
|---|---|---|
| Poor network / satellite links | +50–200 ms | Budget the last mile; Resident is unaffected by public network jitter. |
| Heavy background noise | +p99 variance, not p50 | Recognition re-attempts add to the tail, rarely to the median. |
| Very long or unusual caller phrasing | +100–300 ms | Barge-in recovers mid-sentence; the budget is enforced per-turn. |
| Under-provisioned Resident hardware | Degrades the whole curve | Sizing guidance in the planner prevents this before it happens. |
| Third-party integrations on the response path | +integration time | Booking/HIS calls run post-answer, off the latency critical path. |
The first second of silence is a verdict
Callers judge you in the first second of silence. Above 400 ms, a caller perceives the agent as slow or broken; below it, the response feels like a person picking up. Sub-400 ms is not a vanity metric — it is the difference between a caller staying on the line and hanging up.
And the self-hosted number matters for a different reason: most on-premise voice stacks sit above a second. Resident proves that residency and speed are not a trade-off.