Most of my design career happened on screens, and most of what I learned there turned out to be actively misleading the day I started designing for a phone call. A screen forgives you: show five options at once, let a gaze skip past what it doesn't need, lay out a hierarchy and trust people to find their way through it. None of that exists on a call. There's no layout. The only interface is time — what gets said, in what order, how long each part takes. Most voice products, including early versions of our own, still get designed like a screen with the pictures removed, and it shows.
Designing without a canvas
The thing nobody warns you about, moving from visual design to voice, is that you lose the artifact. A screen design is something you can put in front of someone — a file, a frame, a link. A conversation isn't a thing until someone actually has it. You can write a script, but a script is not the experience any more than a shot list is a film. The real design work is in how a system responds to what wasn't in the script, and there's no reviewing that from a document. You have to hear it, in something close to real conditions, before you know if it's any good.
That changed what I think a "design deliverable" even is here. Increasingly, it isn't a spec. It's a set of real recorded calls, annotated, that the team argues about the way a design team would argue about a mockup. I've written before about what that actually looks like week to week — it's less elegant than a Figma file and more useful for this particular job.
Information architecture for the ear
Screen design has a trick voice doesn't get to use: show several things at once and let the eye do the sorting. Say three appointment options out loud, one after another, and by the third, most people have lost the first — not from inattention, but because holding a spoken list in working memory is a genuinely harder task than scanning a written one. It's remarkable how often voice products get designed as if that weren't true, and it's a large part of why usability research into voice interaction keeps landing on the same conclusion screen-first teams tend to relearn the hard way.
The fix isn't a shorter list. It's usually not presenting a list at all — asking what the caller actually wants and offering the fewest options that get them there, the way a person who knew the schedule would just say "Thursday afternoon's open, does that work?" instead of reciting the whole week. Concretely, the difference looks something like this: a screen-shaped version of the call says "We have Tuesday at 2, Wednesday at 11, Thursday at 3, or Friday at 10 — which works for you?" and expects the caller to hold four options in their head and pick. The voice-shaped version says "What day usually works better for you, earlier or later in the week?" and narrows from there, one branch at a time, the way an actual conversation narrows. Both get to the same booking. Only one of them feels like talking to someone who's actually listening.
"Doesn't offering fewer options limit the customer?"
This is the objection I hear most from people encountering the approach for the first time, usually framed as a fairness concern — who are we to decide the caller only needs to hear one option, when they might genuinely want to compare all four? It's a legitimate worry, and the answer isn't to ignore it, it's to make the narrower path the default without making it the only path. AIVA is built to lead with the option most likely to match what was just asked, and to expand the moment a caller signals they want more — "actually, what else do you have" gets a real answer, not a redirect back to the narrow flow. The design bet isn't that nobody wants options. It's that most callers, most of the time, want the fastest correct path to done, and the ones who don't will ask, plainly, and get what they asked for. Optimizing for the common case while staying genuinely open to the exception is different from just cutting choices to save time, and the distinction matters enough that we test for it specifically — deliberately trying to sound like we're boxing a caller in, and fixing it when a call session drifts that way.
Tone is a spec you can't point at
The hardest thing I do isn't deciding what AIVA should say. It's specifying how — and there's no shared vocabulary for that the way there is for a screen. I can tell an engineer a corner radius is 12 pixels and we'll agree instantly on what that looks like. I cannot tell someone "warm but efficient, reassuring without being saccharine" with anything close to that precision, and yet that's the actual brief for almost every conversational moment we design.
What's worked, imperfectly, is triangulating with examples instead of adjectives — this recorded exchange, not that one — and accepting the spec is a direction you keep correcting toward, not a value you set once. It's also why the more useful conversation-design references, like Google's own conversation design guidelines, lean so heavily on persona and worked sample dialogue rather than adjectives — the format itself is an admission that tone doesn't survive translation into a spec sheet.
I can specify a corner radius to the pixel. I cannot specify "reassuring" to anyone's satisfaction. Most of my actual job lives in that gap.
The same words, a different language, a different design
The gap gets wider once you're designing across the languages AIVA speaks rather than just one. A pause that reads as attentive in English can read as hesitant in a language where responses are expected more immediately. A level of directness that sounds efficient in one language can sound curt or under-polite in another, where a softer opening phrase is doing real social work before the actual content of the sentence even starts. I don't speak all twelve fluently, which means a meaningful part of this design work isn't mine to finish alone — I lean on teammates who are native speakers to tell me when something is technically correct and quietly wrong, the same way a translated screen can pass every spell-check and still feel foreign. Treating "voice design" as a single deliverable that gets localized afterward is, I think, the single most common mistake in this category, ours included in the earliest version of the product.
Where design ends and engineering begins
How fast AIVA responds, how it handles being talked over mid-sentence, whether a pause reads as thoughtful or broken — that's latency and turn-taking work, and it's Rohan's craft, not mine, even though it shapes the experience as much as anything I do. Interruptions specifically are their own design problem, because a real caller doesn't wait politely for a sentence to finish — they backtrack, correct themselves, talk over the response when they've already heard enough. My job sits upstream and downstream of the engineering work: what gets said, how much of it, in what order, in what tone. The two disciplines meet in the same few seconds of a call, and neither covers for the other — a beautifully timed response to the wrong information is still a bad call, and a perfectly worded response delivered a beat too slow feels broken regardless of how good the words were.
We've gotten this wrong in both directions at different points. Wording so careful it arrived a fraction too slow to feel natural, and pacing so fast the words hadn't been trimmed enough to fit it, so a rushed sentence came out sounding clipped instead of efficient. Neither team could fix either problem alone. The fix, both times, was sitting in the same review together, listening to the same call, rather than handing a finished spec across a wall.
What we actually optimise for
Not a voice that sounds impressively human — that's a parlour trick, and a shallow goal on its own. We optimise for a conversation that doesn't ask more of a caller's attention than it has to: nothing to hold in memory that didn't need holding, nothing said that could have been shorter, no moment where the caller has to work out what's being asked of them. Almost none of that shows up in a feature list. Almost all of it shows up in whether someone hangs up feeling understood, or like they talked to a very polite machine. That gap is the entire job, and it's a design job as much as it's anything else.
If you want to hear whether any of this actually holds up rather than take a designer's word for it, the voice line is live — the fastest way to judge a conversation is still to have one.