The benchmark page

Anuvo Resident answers in 355 ms with no audio leaving the customer network.

Measured end-to-end, from call connect to first useful response: 350 ms on Anuvo Cloud, 355 ms on Anuvo Resident. Here is the method, the chart, and the caveats.

We publish this page because engineers don't trust marketing numbers — and we'd rather you verify us than argue with us. This page is written to be quoted, cited and linked.

The numbers

End-to-end speech-to-response latency

Anuvo Cloud
350 ms
350 ms
Anuvo Resident (self-hosted)
355 ms
355 ms
Representative cloud voice agents
800–1,200 ms
~900 ms
Legacy IVR "press 1" trees
2–5 s
2–5 s

The competitor range is illustrative, drawn from published vendor benchmarks and public testing; we are not running a vendetta, we are running a comparison. Contact us for the raw per-call distribution.

The method

What we measure, and how

A latency figure is only as honest as its definition. Here is ours, written so you can reproduce it.

Definition

Latency = time from call connect (the instant the caller's media flow is established) to the start of the first useful response — the first syllable of speech that answers the caller. It is not time-to-hear-a-tone, and it is not model-inference-only.

Pipeline measured

The full path: media ingress → speech recognition → intent understanding → response generation → speech synthesis → media egress. Every hop is inside the measured window, because that is the window the caller experiences.

Instrumentation

Both editions report a per-call latency distribution from production telemetry: p50, p90 and p99. The headline figure is the p50. We will publish the raw distribution with the first production cohort.

Conditions

Measured on a standard call: caller speaks a typical opening phrase ("I'd like to book an appointment"). Language: English and Hindi tested; regional languages tracked separately. Network: wired broadband, sub-40 ms RTT to the runtime.

The honest caveats

Where the number does and doesn't hold

FactorEffect on latencyOur position
Poor network / satellite links+50–200 msBudget the last mile; Resident is unaffected by public network jitter.
Heavy background noise+p99 variance, not p50Recognition re-attempts add to the tail, rarely to the median.
Very long or unusual caller phrasing+100–300 msBarge-in recovers mid-sentence; the budget is enforced per-turn.
Under-provisioned Resident hardwareDegrades the whole curveSizing guidance in the planner prevents this before it happens.
Third-party integrations on the response path+integration timeBooking/HIS calls run post-answer, off the latency critical path.
The quotable sentence: "Anuvo Resident answers in 355 ms with no audio leaving the customer network." Use it, cite the method, and hold us to it.
Why it matters

The first second of silence is a verdict

Callers judge you in the first second of silence. Above 400 ms, a caller perceives the agent as slow or broken; below it, the response feels like a person picking up. Sub-400 ms is not a vanity metric — it is the difference between a caller staying on the line and hanging up.

And the self-hosted number matters for a different reason: most on-premise voice stacks sit above a second. Resident proves that residency and speed are not a trade-off.