Research & Benchmarks · 2026-07-13 · 12 min read

2026 Dealership AI Voice Benchmark: 191,685 Production Calls

A transparent, anonymized 90-day analysis of production dealership AI voice calls—with definitions, limitations, and operational lessons for dealers.

Most automotive AI benchmark claims begin with a percentage and hide the denominator. This report begins with the data dictionary.

TrafficDriver analyzed a rolling 90-day production cohort extracted on July 13, 2026. The dataset contained 191,685 call records across 50 dealership IDs and more than 45,000 unique customer records. The published results are aggregate and anonymized: no dealership, employee, or customer identity is included.

Executive summary

  • 191,685 production calls were recorded in the 90-day cohort.
  • 50 dealerships were represented.
  • completed was the largest carrier/application status: 71,766 calls (37.4%).
  • No-answer, voicemail, and left-voicemail statuses together represented a large share of outbound attempts, demonstrating why carrier status cannot be reported as customer intent.
  • The inbound-agent subset contained 1,870 calls across 39 dealers.
  • Among 1,093 inbound calls with structured resolution summaries, 420 (38.4%) ended in an appointment or department-transfer code, 225 (20.6%) had a general question answered, and 444 (40.6%) were disconnected or unresolved.
  • The report is descriptive, not causal. It does not claim that every appointment, transfer, or sale was incremental.

Methodology

Cohort window

The extraction selected call records whose created_at timestamp fell within the prior 90 days at query time on July 13, 2026. Because production continued during analysis, small count differences may appear in a later rerun of the same rolling query.

Included workflows

The cohort included records tagged as internet, inbound, equity, missed-appointment, null, and unclassified agent types. Agent-type spelling and casing were normalized for analysis where possible; two legacy equity labels remain distinct in the source data.

De-identification

The published aggregation uses counts of distinct internal dealer and customer identifiers but does not expose those identifiers. It does not publish names, phone numbers, emails, VINs, transcripts, private recordings, dealership-level rankings, or customer-level outcomes.

Status hierarchy

This report separates three layers:

  1. Call status: carrier/application result such as completed, no-answer, voicemail, busy, failed, or transferred.
  2. Resolution summary: post-call business disposition such as appointment booked, transferred to service, question answered, not interested, or disconnected/no resolution.
  3. Downstream outcome: CRM/DMS appointment, show, sale, or repair order. Those outcomes require separate reconciliation and are not inferred from carrier status.

90-day call-status distribution

Call result Calls Share of cohort What it does—and does not—mean
Completed 71,766 37.4% The call connected and ended at the carrier/application layer; not automatically a contact or appointment
No answer 38,208 19.9% The attempt did not produce an answered connection
Voicemail 34,195 17.8% A voicemail system was detected or reached
Left voicemail 32,594 17.0% The workflow recorded that a voicemail was left
Busy 8,491 4.4% The called line returned a busy condition
Failed 4,671 2.4% The call failed at the telephony/application layer
Null/legacy 1,170 0.6% No normalized call result was stored
Transferred 590 0.3% The call result recorded a transfer; transfer completion still needs business validation

The practical lesson is simple: a vendor cannot divide appointments by completed and call the result universal conversion. A completed call may still be a rejection, wrong person, opt-out, partial conversation, or unanswered business need.

Workflow distribution

The 90-day cohort was dominated by internet-lead outreach, with smaller inbound, equity, missed-appointment, and unclassified populations. That mix matters. Outbound attempts naturally produce more no-answer and voicemail statuses than inbound calls, while inbound calls need a resolution and routing taxonomy.

Normalized workflow label Calls
Internet agent 171,444
Null/legacy 9,453
Equity-agent legacy label 8,016
Inbound agent 1,870
Equity agent 451
Missed-appointment agent 432
Unclassified 16

These groups should not be blended into one appointment benchmark. Intent, eligibility, call direction, customer history, and expected next step differ.

Inbound benchmark: resolution matters more than answer status

The inbound subset included 1,870 calls, 39 dealers, and 694 unique customer records. Its call-result field contained 1,099 completed, 300 transferred, 470 null/legacy, and one voicemail status. That is still not enough to judge customer value.

A more useful subset is the 1,093 inbound calls with a structured resolution summary:

Structured inbound outcome Calls Share of summarized inbound calls
Appointment or department transfer 420 38.4%
General question answered 225 20.6%
Disconnected/no resolution 444 40.6%
Other coded outcomes 4 0.4%

Appointment or department transfer includes appointment booked, sales appointment booked, service appointment booked, transfer to sales/follow-up, transfer to service, and transfer to parts. These are grouped as valid next-step outcomes, not treated as equivalent revenue.

The result exposes the operational value of a recovery queue. A disconnected or unresolved call is not a final customer disposition. It should create a time-bound recovery task with context and an accountable owner, subject to consent and suppression rules. See the dealership missed-call recovery guide.

Linked lead and appointment subset

A smaller 90-day subset contained 1,106 call-linked lead rows, of which 391 had a non-null appointment timestamp. That is not a network appointment rate. It applies only to calls that had linked lead records in the database, and the presence of an appointment timestamp does not independently prove show, sale, or incrementality.

This subset is published to show why lineage matters: a defensible appointment analysis needs an explicit join path, one appointment definition, duplicate handling, an attribution window, and downstream reconciliation.

What dealers should demand from every AI voice benchmark

1. A fixed cohort

Require dates, departments, stores, direction, campaign eligibility, exclusions, and deployment stage. A pilot containing only answered inbound calls is not comparable to a database-wide outbound campaign.

2. A status dictionary

Ask for exact definitions of attempt, connected, contact, conversation, transfer attempted, transfer completed, appointment, confirmed appointment, show, sale, repair order, opt-out, and duplicate.

3. Denominators

Every rate should name its denominator. Appointment per eligible lead, per attempt, per connected call, and per two-way contact are different metrics.

4. Source-system reconciliation

Verify appointments, shows, sales, and repair orders against the CRM or DMS. Do not rely exclusively on the vendor dashboard.

5. Call evidence

Review recordings or transcripts across wins, rejections, opt-outs, transfers, wrong numbers, background noise, interruptions, and unresolved calls. TrafficDriver publishes selected production call examples in the missed-call recovery guide.

6. Limitations

A useful report says what it cannot prove. This analysis does not establish a randomized counterfactual, incremental gross profit, or a universal conversion target.

Data quality findings

The cohort surfaced three issues any production voice program should monitor:

  • Legacy labels: workflow and result values can drift in spelling, casing, or null state. Normalize them before reporting trends.
  • Summary coverage: only a subset of total calls had structured resolution summaries. Carrier status should remain separate until post-call enrichment is complete.
  • Duration coverage: recording-duration data was sparse for some workflows, so this report does not publish a network duration benchmark.

Publishing missingness is better than filling gaps with a precise-looking estimate. The next benchmark release should fix the cohort at extraction time, publish summary-coverage rates by workflow, and expand reconciled show and outcome data.

Operational recommendations

  1. Normalize agent type and call result at write time.
  2. Require a structured resolution summary for every connected call.
  3. Separate transfer attempted from transfer connected.
  4. Create recovery tasks for disconnected and unresolved inbound calls.
  5. Join appointments through one documented key and deduplicate revisions.
  6. Reconcile shows, sales, and repair orders on a fixed schedule.
  7. Sample calls by outcome and store instead of reviewing only successes.
  8. Publish cohort definitions alongside every dashboard rate.

Research context

The 2025 Cox Automotive Car Buyer Journey Study reported that efficiency, connected experiences, and useful digital tools were associated with strong buyer satisfaction. NADA Data documents the scale of franchised dealership sales and service operations. Those sources explain why communication operations matter; they do not replace dealership-level production measurement.

For compliance and data governance, review the FTC Telemarketing Sales Rule guide, FTC automobile-dealer Safeguards Rule FAQs, and applicable FCC and state requirements.

Frequently asked questions

How large is the 2026 TrafficDriver production-call benchmark?

The rolling 90-day extraction on July 13, 2026 contained 191,685 production call records across 50 dealership IDs and more than 45,000 unique customer records. The cohort included inbound, internet-lead, equity, missed-appointment, and unclassified workflows. Dealer and customer identities were removed from the published analysis.

Does a completed call status mean the customer booked an appointment?

No. Completed is a carrier or application call status, not a business outcome. A completed call can end with interest, an answered question, a rejection, an opt-out, or no useful resolution. Business reporting should use structured post-call outcomes and then reconcile appointments to the CRM or DMS.

What happened in the inbound-agent cohort?

The 90-day inbound cohort included 1,870 calls across 39 dealers. Among 1,093 calls with structured resolution summaries, 420 (38.4%) were coded as an appointment or department transfer, 225 (20.6%) as a general question answered, and 444 (40.6%) as disconnected or unresolved.

Are these benchmark results a performance guarantee?

No. They describe a mixed production cohort across different dealerships, hours, workflows, data quality, deployment stages, and caller intents. They are not a controlled experiment and do not isolate incremental sales, repair orders, or calls that would otherwise have gone unanswered.

Why publish call-status and resolution definitions?

Because conversion claims become misleading when completed calls, contacts, transfers, appointments, and shows are treated as interchangeable. Publishing the denominator and status taxonomy allows dealers and researchers to reproduce calculations, identify missing data, and compare vendors using the same business-outcome definitions.

How should a dealership use this benchmark?

Use it to design a local audit, not to set an automatic target. Export your own call-detail records, establish department and time-window cohorts, define every status, listen to sampled calls, reconcile appointments and transfers, and compare a limited pilot against the dealership's baseline.

Can journalists or integration partners request a methodology review?

Yes. TrafficDriver can provide the aggregate query definitions, cohort rules, status mapping, and a walkthrough of how dealer and customer identifiers were excluded. Customer-level records, private transcripts, dealership identities, credentials, and protected operational data are not part of the public dataset.

Media and partner notes

TrafficDriver invites automotive retailers, publications, and CRM/DMS or telephony partners to scrutinize the methodology. Suggested editorial angle: why dealership AI voice benchmarks fail when carrier status is mistaken for customer outcome—and what a 191,685-call production cohort reveals about transfers, unresolved calls, and recovery workflows.

  • DealerRefresh: practitioner discussion and methodology critique
  • CBT News contributor page: original automotive-retail analysis
  • AutoSuccess: dealer operations and technology audience
  • Integration partners: co-author a technical note on status normalization, transfer completion, or CRM/DMS outcome reconciliation

For methodology requests, use the TrafficDriver contact page. For implementation, review AI voice agents for dealership sales, CRM integration architecture, and the missed-call recovery pillar.