Skip to content
by Vacademy

Voice AI

How to Measure AI Call Quality: Metrics That Actually Matter

The metrics that show whether AI calls work: connect, qualified and booking rates, cost per qualified lead, dead air, talk-over and a weekly QA routine.

By the Telleo team 7 min read

Key takeaways

  • Measure two things: the funnel (connect, conversation, qualified, booked) and the health of each conversation (silences, talk-over, repeats, missed answers).
  • Write down what counts as a connect, a qualified lead and a booking before you count anything. Most bad numbers come from loose definitions.
  • Cost per qualified lead, not cost per minute, is the number that tells you whether AI calling is paying off.
  • Listen to a structured sample every week: random calls, flagged calls and the shortest and longest calls, scored against one rubric.
  • Use the same rubric on your human team’s calls, so you compare like with like.

Two numbers get quoted most about AI calling, and both mislead. “Calls made” measures activity, not results. “Accuracy” usually means how well the speech recognition did on a test set, which is not your callers on your phone lines.

This post covers the numbers that matter: a short funnel of operational metrics, a set of conversation-health signals that show why the funnel looks the way it does, and a weekly routine that turns both into fixes. It applies to any AI calling setup, and most of it applies to human calling too. We give formulas, not benchmarks, because your baseline is the only one that counts.

Define every outcome before you count it

Most arguments about AI calling results are really arguments about definitions. Write these down first:

  • Connected: a person answered and spoke. Voicemail, IVR menus and calls where nobody said anything are not connects.
  • Conversation: the person engaged, for example by answering at least one of the agent’s questions.
  • Qualified: your exact criteria, such as budget in range, course of interest, city and timeline. Not the agent’s impression.
  • Booked: a meeting, visit or demo with a specific day and time agreed. “Call me next week” is not a booking.
  • Incomplete: a call where the person said nothing. Keep these apart from “not interested”, or you will blame the script for what is really a dialling or voicemail problem.

The funnel: seven operational metrics

MetricFormulaWhat it tells you
Connect rateConnected calls ÷ dial attempts (also track connected leads ÷ leads dialled)Lead quality, freshness, timing and caller ID. Rarely the script.
Conversation rateConversations ÷ connected callsWhether the opening line works. A low number means people hang up early.
Talk timeMedian and spread of connected call length, per outcomeLong qualified calls are fine. Long “not interested” calls mean the agent isn’t letting go.
Qualified rateQualified leads ÷ conversationsWhether the agent asks the right questions and the lead source fits
Booking rateBookings with a set day and time ÷ qualified leadsWhether the agent asks for the next step clearly
Transfer successTransfers answered by a person ÷ transfers attemptedWhether your team is actually available for hot leads
Cost per qualified leadTotal spend in the period ÷ qualified leads in the periodWhether the whole thing pays off

Use the median talk time, not the average. A handful of very long calls can hide a pattern where most calls end in the first few seconds.

For cost per qualified lead, include everything: call minutes, monthly minimums, telephony, number rental, DLT, analysis fees and GST. Also count the retries. A lead that takes four attempts to connect costs four attempts.

Conversation-health signals

The funnel tells you where leads drop out. These signals tell you why. Most can be measured from call timings and the transcript.

  • Dead air. Silence where the caller expects a reply. Measure the gap between the end of the caller’s speech and the start of the agent’s reply. Track the longest gap per call. Callers on a phone line assume a long silence means the call dropped, and start saying “hello?”.
  • Talk-over and barge-in handling. Look for both voices at once. There are two failures: the agent keeps talking after a real interruption, or it stops mid-sentence for every “haan”. Both show up when you listen.
  • Repeated questions and repeated lines. The agent asks something it already asked, or repeats a sentence. Search transcripts for duplicate agent sentences in the same call.
  • Missed answers. The caller answered but the field is empty, or it holds the wrong value. Compare the transcript with the captured fields on a sample. Short answers are the riskiest. In the Voice of India speech-recognition benchmark, word error rates for clips under two seconds were far higher than for clips over five seconds (for example 18.74% against 10.45% for Amazon’s system).
  • Voicemail and IVR mistakes. Calls marked “not interested” where only a machine spoke, or the agent talking to a voicemail greeting. These inflate your failure rate and waste minutes.
  • Early hang-ups. The share of connected calls that end within the first few seconds after the opening line. This is the clearest test of your opening.
  • Transfer failures. A transfer attempted and not answered, or answered with no context. A qualified lead left on hold is worse than no transfer at all.

A weekly listening routine

Dashboards tell you something is wrong. Listening tells you what. A routine that fits in about two hours a week:

  1. 1

    Pull a structured sample

    Take 10 random connected calls, every call flagged as a problem, the five shortest and five longest connected calls, and two or three calls from each outcome. Aim for 20 to 40 calls.

  2. 2

    Listen with the transcript open

    Listen to the audio, not just the transcript. Silences, talk-over, tone and mispronunciations don’t show up in text.

  3. 3

    Score against the rubric

    Use the same rubric every week (below), and write one line on the biggest problem in each call.

  4. 4

    Pick one or two fixes

    Group the problems and fix the most common one or two: the opening line, a missing question, an objection the agent handles badly, a word it mispronounces.

  5. 5

    Change one thing at a time

    If you change the script and the lead source in the same week, you won’t know which one moved the numbers.

  6. 6

    Re-measure next week

    Check the funnel metric the fix was meant to move, and listen to calls where the problem used to happen.

Build a QA rubric

Keep it short enough to score a call in two minutes. Score each item 0 (missed), 1 (partly) or 2 (done well), and write down what each score means for your business so two people give the same call the same score.

CriterionQuestion to ask
OpeningDid it say who is calling and why, clearly, within the first sentence or two?
DiscoveryDid it ask every required question, and follow up on vague answers?
AccuracyDid it state only correct facts, prices and dates, and admit when it didn’t know?
ListeningDid it handle interruptions and short answers, without repeating itself or talking over the caller?
ObjectionsDid it address the objection the caller raised, rather than repeat the pitch?
Next stepDid it secure a specific next step: a booking, a transfer or a callback time?
ComplianceWas the call within your calling window, and did it respect “don’t call me” and any required disclosures?
Data capturedDo the captured fields match what the caller actually said?

Use the same tools on your human calls

Call-intelligence tools that transcribe and score calls are not only for AI agents. If you run the same rubric on your telecallers’ calls, you get three things:

  • A fair comparison. AI and human calls scored on the same criteria, on similar leads, at similar times. Humans often get the hotter leads, so compare like with like.
  • Coaching material. Talk ratio, objections raised and whether they were handled, and whether the call ended with a next step. These point to specific coaching topics for each rep.
  • Better scripts. The best human calls show the phrasing and objection handling your AI agent should use. The worst AI calls show where a human handoff should happen sooner.

Mistakes that make the numbers lie

  • Counting voicemails as connects. This inflates connect rate and deflates conversation rate.
  • Averages instead of medians. Averages hide the many very short calls.
  • Treating “not measured” as zero. A signal that wasn’t captured is unknown, not perfect.
  • Small samples. A week with 30 conversations can swing either way. Look at trends over several weeks.
  • Changing several things at once. Script, lead source, calling hours and voice all move results.
  • Judging by the dashboard alone. A call can score “qualified” and still sound bad enough to lose the lead.

Where Telleo fits

Every Telleo AI call gets a green, amber or red health verdict with a plain headline. For example: the agent could not hear the caller, a long silence, caller answers were discarded, the agent kept restarting a reply, voice synthesis stalled, probably an answering machine, a failed transfer or slow responses. A signal that could not be measured shows as “not measured”, never as zero. Every call has a transcript and timings, and recording can be switched on for every call. The call log has search, filters and export. See call health and QA.

After each call you get a disposition from your own list, a summary, a lead rating and the answers captured. Call intelligence is free on AI calls. It also works on your team’s human calls or uploaded recordings, charged per minute, and scores reps against a rubric you can edit.

Book a demo

Sources

  1. PIB — TRAI Strengthens Consumer Protection with Amendments to TCCCPR, 2018 (12 Feb 2025) — checked 25 September 2026
  2. Bhogale et al. — Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India, arXiv 2604.19151 (v4, July 2026) — checked 25 September 2026

Frequently asked questions

What is a good connect rate for AI calls?+

We don’t quote a benchmark, because connect rates depend mostly on the lead source, how fresh the leads are, the time of day and the caller ID. Measure your own baseline over two to four weeks, and compare the AI agent with your human team on the same kind of leads at the same times.

How many AI calls should I listen to each week?+

Enough to see patterns. For most teams that is 20 to 40 calls a week: a random sample, every call flagged as a problem, and the shortest and longest calls. Score each against the same rubric, and fix the one or two most common problems before listening again.

What is dead air on an AI call and how do I measure it?+

Dead air is silence where the caller expects the agent to speak, usually after they finish a sentence. Measure the gap between the end of the caller’s speech and the start of the agent’s reply. Track the longest gap per call and the number of gaps above a threshold you set after listening to some calls.

How do I calculate cost per qualified lead for AI calling?+

Add up everything you spent in the period (call minutes, platform minimums, telephony, number rental, DLT, analysis fees and GST) and divide by the number of leads that met your written definition of qualified in the same period. Use the same formula for your human team, including salaries and tools.

Can I use the same QA rubric for human and AI calls?+

Yes, and you should. Score opening, discovery, accuracy, listening, objection handling, next step and compliance for both. Some items will matter more for one than the other, but a shared rubric is the only fair way to compare them.

Hear it on a real phone call.

Book a 20-minute demo and we'll build a first agent around your script, or ring our test line and talk to one right now.