Back to all posts

AI receptionist performance: the KPIs that actually matter

Uptime and conversation counts feel like progress. They don't tell you whether the thing is actually working. Here's what to track instead. Six KPIs to track.

MN
Meera Nair
Customer Success

Most businesses that deploy an AI receptionist watch exactly one thing afterward: whether anyone complains. If the phones are quiet and nobody's upset, it must be working. That's a real signal, but it's a lagging one — by the time a pattern is bad enough to generate a complaint, it's already cost you weeks of calls handled badly. Here's what we tell customers to actually track, starting in week one.

The numbers that feel like progress but aren't

Total conversations handled and uptime are the two metrics every dashboard shows first, and they're the two that tell you the least. A high conversation count just means the phone rang a lot — it says nothing about whether those calls went well. Uptime tells you the system was reachable, not that it was useful. Neither number would change if every third caller hung up frustrated. Treat them as plumbing checks, not performance indicators — worth glancing at to confirm nothing's broken, not worth building a weekly review around.

Five KPIs worth watching

Resolution rate

The share of conversations that ended with the customer's need met, without a human having to step in. This is the closest thing to a headline number, and it's the one we publish a range for — 82% to 96% across our customers — because it depends heavily on how connected your setup is to your actual booking system, not just how good the underlying voice or chat is. A business with a thin FAQ and no live calendar connection will sit at the low end of that range no matter how capable the underlying agent is; a business that's configured its exceptions and edge cases carefully will sit at the high end. This is why resolution rate is worth watching as a trend over your first month, not judging on day one.

Response time

How long a caller or chatter waits before getting a reply, and how natural the back-and-forth feels once the conversation is moving. On voice specifically, this is measurable in milliseconds — AIVA averages around 198ms, faster than the roughly 300-millisecond gap humans typically leave each other in conversation — and it's worth checking, because a system that pauses for a second or two mid-sentence reads as "thinking," not "listening," even when it eventually gives a correct answer. This one rarely needs active monitoring once it's confirmed healthy at launch; it's more of a one-time check than an ongoing metric, unless call volume grows sharply and you want to confirm it's still holding.

Escalation rate, and why

Some escalation is correct and expected. What matters is the reason code attached to each one. A rising share of escalations tagged "customer asked for a human" is healthy. A rising share tagged "system couldn't handle the request" is a configuration problem, not a customer preference. The fix for the second kind is usually specific and fast — an FAQ gap, a phrasing the agent wasn't prepared for, an exception nobody wrote down — and it's the kind of thing worth checking weekly rather than letting accumulate for a month.

Booking conversion

Of the customers who asked about availability, how many actually left with a confirmed slot. This is the number that ties most directly to revenue, and it's easy to miss if you're only watching resolution rate — a conversation can be "resolved" by correctly telling someone you're fully booked. If booking conversion sits noticeably lower than resolution rate, that gap is worth investigating on its own: it can mean real availability, or it can mean the booking flow itself is losing people who wanted to say yes. How the booking flow is structured has a direct effect on this number specifically.

After-hours and off-peak capture

How much of your total volume is happening outside the hours your team is normally working the phones. This is usually the most surprising number to a new customer, and the clearest evidence of what an AI receptionist is actually adding, rather than just replacing. Businesses that assumed their call volume was mostly a 9-to-6 phenomenon are often surprised to see a third or more of it landing in the evening or on a day they're closed — after-hours coverage is often the single easiest KPI to point to when explaining the value to a co-owner or partner who wasn't part of the decision to try it.

Uptime tells you the system was reachable. It doesn't tell you whether anyone was actually helped.

Read them together, not one at a time

A single high number can hide a problem the others would catch. High resolution rate with a climbing "system couldn't handle it" escalation reason means the AI is quietly narrowing what it attempts, not actually getting better. High booking conversion with a shrinking after-hours share might just mean your daytime team got faster — not a problem, but worth knowing which one moved. A genuinely healthy deployment tends to show all five moving in a consistent direction, or holding steady together; when one moves sharply while the others don't, that's usually the more interesting story than the number itself.

Common mistakes reading these numbers

The most common one is checking only on a bad day — a single frustrating call prompts an owner to pull up the dashboard, see one escalation, and read it as a trend when it's a single data point. The second is comparing resolution rate across fundamentally different types of calls, like a simple hours question against a multi-step booking with an exception, and treating them as the same difficulty when they aren't. The third is watching the dashboard obsessively in week one and then never again — the numbers that matter in month three are different from the ones that matter on day one, and a KPI worth setting up once is worth glancing at periodically afterward, not abandoning once the novelty wears off.

Setting a cadence that doesn't become a chore

Week one, check daily — this is when configuration gaps surface fastest and are cheapest to fix. Weeks two through four, check weekly, comparing against the prior week rather than against day one. After the first month, once resolution rate and escalation reasons have settled into a stable pattern, monthly is usually enough, with an exception for any period where you've changed your FAQs, added a location, or launched a new channel — treat those like a mini pilot and watch more closely again for the first couple of weeks, the same way you would during an initial rollout.

Tying the numbers back to the business case

These five KPIs are also the raw inputs for the question most owners actually care about, which is whether the whole thing is worth what it costs. Booking conversion and after-hours capture feed directly into an ROI calculation — captured bookings and freed-up staff time are the benefit side of that math, and they're sitting in the same dashboard you're already checking for performance. There's no need to run a separate exercise to build the business case; it's mostly a matter of pulling numbers you're already tracking and putting them next to what the usage actually costs.

Where to actually look

All five of these live in the same place: a real-time dashboard with a 90-day retention window and an auto-categorized breakdown of what customers are actually asking about, so you're not left guessing at the "why" behind a number that moved. The data is exportable too, if you want to fold it into a broader monthly review alongside bookings or revenue. Check it weekly for the first month, then monthly once the pattern's stable. See what it looks like on AIVA's analytics, or read how we think about what "resolution rate" actually measures in more depth before you start comparing numbers across weeks.

Share
MN
Written by
Meera Nair
Customer Success

FAQ

Common questions.

Across AIVA's customers, resolution rate typically runs 82% to 96%, depending mostly on how well the system is connected to your real booking calendar and how complete your FAQs are — not just how good the underlying voice or chat is.

Weekly for the first month, while configuration is still being calibrated against real conversations, then monthly once the pattern is stable. Checking daily in week one is reasonable too — daily after month one usually isn't a good use of time.

No — it depends on the reason code. A rising share of escalations tagged as the customer asking for a human is healthy and expected. A rising share tagged as the system being unable to handle the request is a configuration problem worth fixing.

A conversation can be fully resolved without producing a booking — for example, correctly telling a caller you're booked up. Resolution rate alone won't show you that; booking conversion tracks the number that ties most directly to revenue.

All five live in AIVA's analytics dashboard — a real-time view with 90-day retention and an auto-categorized breakdown of what customers are actually asking about, so a number that moves comes with a reason attached.

Yes, the analytics data is exportable, which is useful if you want to review it alongside other business numbers or share a specific stretch with a partner or accountant.

Uptime. It confirms the system was reachable, not that any given caller was actually helped — a business could have perfect uptime and a mediocre customer experience at the same time.

Only loosely. Resolution rate depends heavily on how connected your setup is to your booking system and how complete your FAQs are, so a lower number than someone else's usually points to a configuration gap on your end worth checking, not an apples-to-apples ranking.

Like this? Get more.

One email a month. Engineering deep-dives, product launches, customer stories. No fluff.

4,200+ subscribers. Unsubscribe anytime.