Skip to main content
Artificial Intelligence · Business & Lab

What to Test Before Investing in a Voice Agent

← Artificial Intelligence

Written by Evren BalPublished  · 7 min read

A single blue call token reaches a completed task after several calls stop earlier.
Discuss this article with your AI

A voice-agent demo can make an investment look compelling. The conversation is fluent, questions are understood, and the exchange feels natural. For a business struggling with capacity in sales or customer service, it is easy to move from that demonstration to an expectation of reaching more people and completing more work.

But by the time a demo begins, contact with the customer has already been made. In an outbound scenario, that contact is the first uncertainty that determines the investment's outcome. Will the intended audience answer the phone, listen to why they are being called, and want to continue the conversation? A technically capable agent cannot recover opportunities lost before those stages.

That is why deciding whether to use a voice agent requires more than assessing its conversational ability. You also need to test whether the channel fits the audience and whether the conversation leads to the business outcome the system was meant to produce.

The limit revealed by more than 2,000 calls

I saw this distinction in an experiment of our own. We directed prospective customers who left contact details through several channels to WhatsApp or live chat. When they did not write first, we followed up by message and reached some of them that way. Calling the group we could not reach looked like a possible way to add sales to the existing flow.

We used a voice agent to make more than 2,000 calls to a selected group of leads from the United States who had left US phone numbers. Most calls were unanswered or went to voicemail. Some people who answered ended the call shortly after hearing the agent. Very few conversations developed into a meaningful exchange. As I remember it, there were one or two additional sales.

These were leads we had not reached through earlier contact channels. The result therefore does not represent all US customers or voice-agent use in general. We also do not know, one by one, why the calls were not answered. The experiment did show that the ability to reach people and the ability to hold conversations need to be evaluated separately. Improving the agent's conversational quality would have addressed a problem we had not yet encountered in most calls.

The audience and the reason for the call belong together

In an inbound scenario, the customer has chosen the phone channel at that moment. The business may be trying to reduce waiting time, answer after hours, or handle more requests than its team can cover. A voice-agent investment can then be assessed by whether it serves existing demand better.

In an outbound scenario, starting the conversation is part of the work. Leaving contact details is not the same as being ready for an unexpected call. Returning a requested call, following up on an old application, and calling to confirm an appointment each create different expectations.

A Pew Research Center survey conducted in the US in July 2020 also found a pattern of people not answering unknown numbers. About eight in ten adults said they generally did not answer these calls on their mobile phones. This dated, self-reported finding does not give us an answer rate for a sales campaign today. It provides context for why having a phone number and being able to reach someone by phone are separate conditions.

The country, language, time of call, how the caller ID appears, and the person's prior relationship with the business should be part of a pilot's scope. Consent to call and call-recording conditions also need to be clear before the system goes live. An approach that works in one market or customer group cannot be assumed to produce the same outcome in another.

What should happen by the end of the conversation?

“Talk to prospective customers” is not a success measure on its own. The expected result needs to be defined before the call begins. It might be booking an appointment, confirming information, or transferring the customer to the right person.

A fluent conversation is only one part of that work. Giving a clear answer to a customer's question can make for a good experience. But if an appointment is booked incorrectly or a transfer never happens, the conversation is not complete from the business's perspective.

Artificial Analysis's Speech Agent Arena evaluates conversational preference and task completion separately. A system participants find more natural does not necessarily perform equally well at completing the requested task. The evaluation uses paid, screened participants and specified scenarios, most of which are not public. Its results therefore do not show the answer rate of unexpected calls or real customers' purchasing behaviour.

A business needs to verify success in its own pilot from its records. Was the appointment actually made? Did the customer reach the right person? Was the request resolved? The distinction between output, completed work, and business effect that I examined in Is Your AI System Actually Delivering Business Results? applies here too.

A pilot should show where the loss begins

The purpose of a pilot is not only to find conversations that went well. It is to see at which stage people are being lost. Keeping the first test to one customer group and one calling purpose makes that possible. Instead of looking only at a single conversion rate, work through these questions in order:

StageQuestion
CallsHow many distinct people were called, and how many attempts were made per person?
ConnectionHow many people answered, and how many calls went to voicemail?
ConversationHow many people continued the conversation, and how many ended it early?
TaskDid the requested appointment, confirmation, or human handoff happen?
OutcomeDid the business gain an additional sale, resolve a request, or free up capacity?

If people cannot be reached before contact is made, reconsider the audience, timing, and expectation for the call. If they answer but end the conversation immediately, the reason for calling or the way it is presented may be the problem. If the conversation continues but the task is not completed, examine the call content and the handoff to a person.

Repeated attempts to reach the same person should be kept separate from the number of people reached. A success rate calculated only among people who continued the conversation does not describe the whole pilot. For an inbound scenario, waiting time, resolved requests, handoffs to people, and repeat calls about the same issue are more useful measures.

The additional result should determine the investment

These measures reduce the investment decision to one question: does voice calling create an additional outcome that justifies the resources spent?

The cost is not limited to the agent's conversations. Setup, calls, follow-up, handoffs to staff, and the work needed to correct errors all belong in the calculation. People who cannot be reached and calls that produce no outcome must be included too.

Where possible, leave a comparable group in the existing communication flow. That comparison makes it easier to see whether voice calling is producing an additional result. Without it, it is difficult to attribute every sale that occurs to the agent.

There is no universal answer rate or conversion rate that applies to every business. A few additional sales may matter in a high-value transaction; the same result may not cover the cost of follow-up in a low-value one. The decision to scale should depend on whether the added result justifies the total cost, not on how human the agent sounds.

If the pilot does not show a sufficient contribution, building a more sophisticated agent is not automatically the right next step. Changing the calling purpose, testing a different group, or keeping the existing messaging flow are all valid options. A new investment needs evidence that voice calling can produce a measurable result for this audience.

If this article was useful

Linking to it from a relevant page on your website or sharing it on social media genuinely helps it reach more people. Thank you for your support.

Linking and brand guidelines →

About this article

Use of artificial intelligence
AI-assisted — The experience, central idea, and scope of this article were supplied by Evren Bal. AI-assisted tools supported source research and drafting.