Northwind Traders runs one of the largest B2B logistics networks in western India — 200,000 shipments a month, 40,000 active business accounts, a 180-person customer support operation.
When they came to us in January, their average hold time was 8.2 minutes and their CSAT was 3.2 out of 5. Their NPS was negative.
Three weeks after going live with AIVA Voice, their average resolution time was 78 seconds and CSAT was 4.6. Nobody was let go.
Illustrative example. Northwind Traders is a composite scenario built to show how AIVA works for this kind of business — not a verified customer account.
The problem
Northwind's support team was drowning in repetitive queries. "Where's my shipment?" "When does it arrive?" "I need to reschedule a delivery."
Their own analysis showed that 74% of all incoming contacts were some variation of these three questions — and every one of them required a human agent to log into three different systems, find the right record, and read it back over the phone.
The agents weren't adding value. They were doing data entry out loud. The customers were frustrated because waiting eight minutes to hear information that already existed in a database felt absurd. The agents were frustrated because they'd joined to help people, not to be a slow API.
This is the shape of nearly every support operation we've seen in courier and logistics businesses: the volume isn't hard, it's just relentless, and it crowds out the work that actually needs a person.
Why hold time was the wrong metric to fix
Northwind had tried to fix this before, twice, and both attempts targeted hold time directly.
The first attempt was more agents. It worked, briefly, and then volume grew into the new capacity. The second was an IVR menu deep enough to pre-sort callers by query type. It made hold time slightly worse, because callers now spent 90 seconds navigating a tree before joining the same queue.
The insight that changed the outcome was that hold time is a symptom. The disease is that a human is required at all for a question the database can answer. You don't fix that by making the queue move faster — you fix it by taking the query out of the queue.
The agents weren't replaced. They were finally doing the job they were hired for — handling the hard stuff, not reading shipment numbers out loud.
The implementation
We deployed AIVA Voice into Northwind's existing Twilio setup and connected it via API to their shipment tracking system, their delivery partner feeds, and their calendar and rescheduling backend. The full list of what connects out of the box is on the integrations page.
Setup took three days:
- Day 1–2 — integration and tuning. Connecting data sources, configuring escalation rules, and teaching the agent Northwind's specific shipment vocabulary. This was the bulk of the work and almost all of it was terminology rather than capability: docket numbers, LR numbers, hub codes, and the half-dozen ways a customer might describe the same delivery problem.
- Day 3 — parallel run. AIVA handled inbound calls with human agents shadowing every conversation and able to take over instantly. This is the step most teams want to skip, and the one we insist on — it's where you find the gap between what you configured and what customers actually say.
- Day 4 — live. AIVA handling all tier-one queries without human backup. Agents freed for escalations, complaints, and edge cases.
If you're planning something similar, our guide to piloting an AI receptionist covers the same sequence in a form you can run yourself.
Getting escalation right was the deciding factor
The single configuration decision that determined this outcome was where the handoff line sits.
Northwind's instinct — and this is near-universal — was to escalate whenever the agent was uncertain. We pushed back hard on that, for reasons we've written up separately in our escalation rules post.
Escalating on uncertainty produces an expensive FAQ lookup: the agent handles only what's trivially routine and hands a human everything else, which is exactly the workload the humans already couldn't cope with.
What we configured instead escalates on who the caller needs, not what they're asking: emotional distress, billing disputes over a threshold, account security questions, and any explicit request for a person.
A caller asking four related questions about a delayed shipment is complex but doesn't need a human — it needs system access and patience. The buyer's explanation of how handoff works covers the same logic without the internals.
The results
The numbers after three weeks:
| Metric | Before | After |
|---|---|---|
| Average resolution time | 8.2 minutes | 78 seconds |
| CSAT | 3.2 | 4.6 |
| First-contact resolution | 41% | 94% |
| NPS | −12 | +34 |
| Agent focus | 100% tier-one | 100% escalations and complex cases |
First-contact resolution is the number we'd point at first — it's the one that best predicts whether a deployment sticks, and it's worth understanding what resolution rate actually measures before comparing it across vendors.
The surprise was NPS, which moved from −12 to +34. Customers don't fill out surveys to say "the bot answered my question correctly."
They responded positively because getting a fast, accurate answer — in Hindi or Gujarati, at any hour, across the twelve languages we support — felt genuinely different from what they expected of a logistics support line.
What we'd do differently
Two things, in hindsight.
We under-invested in the terminology pass. Two days felt generous at the time and was barely enough. Logistics vocabulary is dense and locally variable — the same document is a docket, an LR, or a "bilty" depending on who's calling and from where.
The first week of live traffic surfaced maybe forty phrasings we hadn't configured. None of them broke anything, because the agent asked a clarifying question rather than guessing, but each one was a caller doing work the configuration should have done.
If we ran it again, we'd spend day one reading actual call transcripts rather than interviewing the support team. What agents report customers saying and what customers say diverge more than anyone expects — agents summarise, and the summary is already normalised into internal vocabulary.
We should have staged the rollout by query type, not by time. We went from parallel run to full tier-one coverage in a single step. It worked, but it meant that when something needed adjusting in week one, we were adjusting it under full volume.
Starting with tracking queries only, then adding rescheduling a few days later, would have given us the same end state with a much smaller blast radius on each change.
Neither of these changed the outcome at Northwind. Both are now standard in how we run large deployments.
The outcome nobody forecast
Northwind's support director told us something that stuck: the best outcome wasn't the CSAT improvement. It was that agent attrition dropped to zero in the first quarter.
When people aren't doing mindless work, they don't quit. That's not a metric we'd thought to promise, and it's now the first thing we ask about when a large support team evaluates us.
If your team is spending its day reading database records out loud, the fix is available in an afternoon. Start free with ₹500 of credit, or watch what your own call mix looks like in the analytics dashboard before you change anything.