Somewhere around call thirty of a Tuesday, Rohan walked past my desk, looked at my headset, and asked if I'd finally lost it. I was calling our own AIVA line, hanging up, changing one variable — a pause length, a word order, a greeting — and calling again. I did this, on and off, for most of a week. I do a version of it most weeks, actually. It's the least glamorous part of my job and probably the most useful.
The thing a transcript doesn't show you
As a designer, my instinct is to review things visually. A transcript of a call reads fine almost every time — the words are correct, the information is right, the grammar holds up. But a transcript is not a phone call. It strips out everything that actually makes a call feel like talking to a person, or feel like talking to a machine: the timing between when you stop speaking and when the reply starts, the tiny acknowledgment that tells you you're being listened to, the rhythm of how a sentence gets broken up.
None of that shows up if you're reading. All of it shows up if you're the one on the call, phone against your ear, exactly like a customer would be. I've written elsewhere about why voice design has to work this way — there's no canvas to review, no mockup to point at, only the call itself.
What the tenth call teaches you that the first one doesn't
The first call always sounds fine. That's the trap. A single cold call is the worst way to evaluate a voice experience, because you're paying full attention and forgiving of anything slightly off. Call ten is where you start noticing the pattern you missed the first nine times — a phrase that gets reused too often, a pause that's a beat too long after a specific kind of question, a greeting that's technically warm but somehow lands flat on repetition.
Call fifty is where you stop noticing the content of what's being said at all and start hearing only the shape of it — which, it turns out, is roughly the state a real repeat customer eventually reaches too. If the shape is wrong, fifty calls in is where it starts to grate, even if every individual answer was correct.
"Isn't this exactly what automated testing is for?"
I get asked this a lot, usually by people outside design, and it's a fair question — we're an eight-person team, and spending a week of a designer's time dialing a phone number by hand sounds like a strange use of scarce hours when synthetic call generation exists and scales infinitely. We do use automated checks for a real chunk of this. Whether a fact is correct, whether a booking actually lands in the calendar, whether a specific intent gets recognised — that's all covered by tests that run constantly, without anyone picking up a phone, and it should stay that way. Nobody's proposing we manually re-verify pricing accuracy every week.
What automated testing can't do, at least not yet in any way I'd trust, is tell you whether a correct answer felt considerate or felt cold. That's not a measurement problem with an obvious fix — it's closer to asking a script to grade its own bedside manner. A synthetic caller doesn't get quietly irritated by a confirmation that arrives a half-beat too fast after bad news, because a synthetic caller doesn't feel anything at all. It can confirm the words were right. It cannot tell you the call was slightly unpleasant to be on, and that gap is exactly the part of the product I'm responsible for.
Calling in languages I don't speak fluently
The part of this job that surprised me most, when we started supporting more than English and Hindi, is that I couldn't rely on my own ear the way I used to. Rhythm and pacing that feel natural to me in English don't automatically transfer — a pause that reads as thoughtful in one language can read as hesitant or slow in another, and politeness markers that make a Tamil sentence sound respectful don't map cleanly onto a Bengali one. W3C's internationalization guidance makes a version of this same point about text and layout — that conventions which feel neutral in one language are often specific to it — and it holds just as true for the sound of a spoken conversation as it does for how a page reads. For the languages AIVA speaks where I'm not a confident native listener, I call alongside a teammate who is, and I watch their face more than I trust my own read of the audio. It's a slower process than testing in English, and it's the only honest way to do it — deciding whether a Marathi greeting "feels right" is not a call a non-Marathi speaker gets to make alone, no matter how many calls they've sat through.
The call that made me change something small
A few months ago I was testing a rescheduling flow, back to back, mostly on autopilot. Somewhere in there I noticed AIVA was answering my questions perfectly and I was still faintly irritated. It took a few more calls to figure out why: it never paused before confirming a change. It moved straight into confirmation, at the same clipped pace, whether I'd asked to reschedule a haircut or cancel something more inconvenient. Correct, fast, and slightly relentless.
We added a beat — a genuinely tiny pause, well under a second — before any confirmation involving a change to an existing booking. Nothing about the words changed. The information was identical. But to an actual human ear, it reads as the difference between a system executing a command and something that briefly considered what you just asked. I would never have found that from a transcript. I only found it because I was still on call forty-something, mildly annoyed, and curious enough to ask myself why.
Every design flaw that actually matters in a voice product is invisible on a transcript and obvious on the fortieth call. That gap is basically my whole job.
The one that fooled me first
Not every finding comes from something going wrong. One of the more useful calls I've made was one where nothing felt off at all, and that was the problem — I'd gotten so used to the sound of our own voice product that I'd stopped being able to tell if it was actually good, or if I'd just gone numb to its flaws the way you stop smelling your own kitchen. I've since started deliberately mixing in calls after a few days away, specifically to catch what fatigue was letting slide. It's a strange kind of QA — testing your own ability to notice things, not just testing the product — but it's caught real issues that back-to-back testing sessions missed entirely, including ones related to how AIVA handles being talked over mid-sentence, which is much easier to miss when you're not fully present on the call yourself.
Why I still do this
I could hand this work to a testing team, and eventually parts of it should live there. But right now, the person deciding what "feels right" needs to be the person who sat through the calls where it didn't. Reading about a bad pause is not the same as being the one waiting through it, four calls after the last time it happened, wondering if it's you or the phone. It's also, frankly, a job that fits a team our size — we've stayed at eight people on purpose, and one consequence of that is nobody gets to be purely a manager of the work. I still do the work myself, which is uncomfortable some Tuesdays and is exactly why the calls still get made.
So I keep calling. Rohan keeps asking if I'm okay. I'm not entirely sure yet, but the product's better for it. If you want to hear the result rather than take my word for it, the voice line is live — call it as many times as you want. We do.