Back to all posts

What 4 million conversations taught us about customers

Not what our customers told us in interviews — what their customers actually say, at 4 AM, in three languages at once, when nobody's coaching them on phrasing.

AP
Arjun Patel
Co-founder

We wrote before about what we learned from our first 100 customers — the businesses themselves, how they used AIVA, why some succeeded and some churned. This is a different dataset: not the businesses, but the people calling them. Four million conversations a month gives us a genuinely large window into how small business customers actually behave, as distinct from how we assumed they'd behave when we started. Some of it confirmed what we expected. A good amount of it didn't.

When people actually call

The single clearest pattern: a large share of conversations — close to a third, by our internal count — happen outside the business's stated opening hours. Weekend mornings, weekday evenings, the occasional 2 AM booking attempt from someone who clearly just remembered they needed an appointment. This is the exact gap we built the product to fill, so it's less a surprise than a confirmation, but the scale of it still struck us. A meaningful fraction of a small business's customer demand simply doesn't happen during business hours, and until recently, all of that demand went to voicemail or a closed sign.

What people actually ask

The shape of the questions matches what we found in our earlier customer research — the large majority is logistics, not complexity. Are you open, how much does it cost, do you have anything tomorrow. What the conversation-level data adds that the customer-level research couldn't is texture: the actual phrasing is far more informal and far more mixed-language than a business's own FAQ page would suggest. A significant share of our Hindi and Gujarati voice conversations code-switch mid-sentence into English for specific words — prices, appointment types, technical terms — in a way that would break a system built around one language at a time. People don't ask questions in the tidy, monolingual way businesses write their own websites. They ask however they actually talk.

The other thing the aggregate view surfaced: the specific mix of questions varies more by industry than we expected going in. A driving school's most common question isn't a salon's most common question, even though both fall under the same broad bucket of "logistics, not complexity." We've written a more detailed breakdown of how the same underlying FAQ shape gets asked completely differently across industries, which is a more useful read than this piece if you're configuring AIVA for a specific one.

Is this dataset representative, or just of who already adopted it?

I want to name the honest limitation here before going further, because it's the obvious objection to treating any of this as universal small-business behavior. Our four million conversations a month come exclusively from businesses that already chose to deploy an AI receptionist — which is not a random sample of small businesses generally. It likely skews toward businesses with higher call volume than average (the pain of missed calls is what motivates adoption in the first place), toward businesses with more after-hours demand (the exact problem AIVA solves), and possibly toward business owners who were already more comfortable with newer technology than a typical owner in their category.

That means a finding like "a third of calls happen after hours" might be somewhat inflated relative to small businesses broadly — we may be disproportionately serving the businesses for which that was already true, rather than discovering it's true everywhere. I don't have a clean way to correct for this selection effect from inside our own data; we'd need visibility into businesses that haven't adopted AIVA to know for sure. I still think the patterns are informative — the mechanisms behind them (people have things to book outside a 9-to-5, people phrase questions informally, repetition is annoying regardless of who or what you're talking to) don't obviously depend on the selection effect. But I'd rather flag the limitation than let a large number imply more certainty than the sampling actually supports.

Patience is higher than the industry assumes, until it isn't

The most useful thing we've learned: callers are considerably more patient with a competent automated voice than the conventional wisdom about AI skepticism would suggest — as long as the call resolves quickly and correctly. Our average voice response sits around 198 milliseconds and resolves about 96% of calls without a human, and at that speed, most callers don't seem to consciously register that they're not talking to a person, or don't especially mind if they do.

Patience isn't about whether someone's talking to a machine. It's about whether the machine makes them say anything twice.

That patience collapses fast under one specific condition: being asked to repeat something. Across every dataset we've cut this by, the single strongest predictor of a bad rating isn't "sounded robotic" or "couldn't help me" — it's repetition, which lines up with Gartner's own research on customer effort finding that effort predicts disloyalty more reliably than satisfaction scores do. A caller who has to restate their name, their request, or their answer to a question the system should already know produces frustration that's disproportionate to how minor the inconvenience actually is. We've reorganized a fair amount of engineering priority around this single finding.

What the repetition finding actually changed

It's worth being concrete about what "reorganized engineering priority" meant in practice, rather than leaving it as an abstract claim. Once repetition surfaced as the dominant driver of bad ratings — well ahead of raw accuracy or how natural the voice sounded — it changed what we measured, not just what we built. A resolution number alone doesn't tell you whether a caller had to say their name twice to get there, which is part of why we've been careful about what "resolution rate" actually captures and where it runs out as a single headline metric. We started tracking repeat-request rate as its own number, separate from resolution, specifically because this finding showed us a call could count as "resolved" and still have annoyed the person on the other end of it.

Small business customers aren't enterprise customers

The other pattern worth naming: expectations differ meaningfully by context, in ways that aren't really about AI at all. A caller to a neighborhood clinic or a local salon expects something closer to a relationship than a transaction — "hi, is this Priya's salon?" is a common opening in a way it wouldn't be calling a national chain's support line. Formality and menu-style navigation, which feel normal on an enterprise support call, read as cold and bureaucratic in a small business context. People calling a small business want to feel like the business knows who they're talking to, even when, mechanically, it's the first time.

This shows up especially clearly in the language data. It's not just code-switching within a sentence — it's that the register people use with a small, local business is warmer and less formal than the register the same person would use calling a large company, in any language. A system trained on generic customer-service phrasing, rather than on how people actually talk to businesses they consider "theirs," reads as noticeably stiffer than it should. Getting this right has been a bigger part of handling Indian languages accurately than getting the accent recognition right on its own — tone and register turned out to matter nearly as much as words.

What this adds up to

None of this was visible from our first hundred customers alone — patterns like code-switching frequency or the exact repetition penalty only show up at real volume. Four million conversations a month stopped being a number we cite in a pitch and started being an actual, humbling picture of how people talk to the businesses they rely on, with the honest caveat that it's a picture of people who called a business that had already adopted AIVA, not a picture of every small business customer everywhere. We're still finding things in it. If you want to see what the shape of this data looks like for a specific deployment, AIVA's analytics break it down the same way, per business, not just in aggregate.

Share
AP
Written by
Arjun Patel
Co-founder

FAQ

Common questions.

That callers are far more patient with a competent automated voice than conventional wisdom about AI skepticism suggests — right up until they're asked to repeat something, which is the single strongest predictor of a bad rating we've found.

Close to a third, by our internal count — weekend mornings, weekday evenings, and the occasional very-late booking attempt. It's the exact gap AIVA was built to fill, so it's more confirmation than surprise, but the scale of it still struck us.

Mostly not, as long as the call resolves quickly and correctly. At AIVA's average response time of around 198ms, most callers don't seem to consciously register — or particularly mind — that they're not talking to a person.

Repetition — a caller having to restate their name, their request, or an answer the system should already know. It outranks "sounded robotic" and "couldn't help me" by a wide margin in every dataset we've cut this by.

Not perfectly — it's drawn from businesses that already chose to adopt an AI receptionist, which likely skews toward higher call volume and more after-hours demand than the average small business. We think the patterns still generalize, but we hold that claim more loosely than the raw numbers might suggest.

Yes, and more informally than a business's own FAQ page would suggest. A significant share of Hindi and Gujarati voice conversations code-switch mid-sentence into English for prices, appointment types, and technical terms.

That earlier piece looked at the businesses themselves — who they were, why they succeeded or churned. This dataset is the people calling those businesses: patterns like code-switching frequency and the exact repetition penalty only became visible at much higher volume.

Like this? Get more.

One email a month. Engineering deep-dives, product launches, customer stories. No fluff.

4,200+ subscribers. Unsubscribe anytime.