Request a callbackBook a call
← All posts

AI Voice Agent Cost Per Minute (2026): Every Line Item, Managed vs Custom, 20K to 1M Minutes

TL;DR
  • Managed voice platforms land between about 8 and 14 cents a minute once telephony is counted; Retell's default is $0.11 plus $0.015 for its telephony. I moved a production voice stack from about 10 cents a minute on a managed platform to about 2.5 cents on a custom LiveKit stack.
  • The 2.5 cents is seven lines: carrier $0.0032, LiveKit's SIP fee $0.004, speech-to-text $0.0033, the model $0.0004, the voice $0.009 and about $0.0051 of workers, state and logs, with no platform fee. The voice and the prompt size swing the total more than any other choice.
  • Stay managed below roughly 20,000 minutes a month, where fixed infrastructure and maintenance eat the saving. At 100,000 minutes a custom stack saves about $5,700 a month and repays a six-week build in about three months; at 1 million it saves about $73,000 a month.
The seven lines in a voice AI minute
1 · Carrier

The phone minute itself, on a SIP trunk. Telnyx lists inbound local at $0.0032, outbound local from $0.005 and toll-free inbound from $0.015.

$0.0032
2 · SIP bridge and media

Getting the call into a LiveKit room. LiveKit Cloud charges $0.004 a minute on Ship for a trunk you bring, $0.003 on Scale, $0.01 for its own number. Self-hosted LiveKit has no per-minute fee.

$0.0040
3 · Speech to text

Streaming transcription of the whole call, silence included. xAI lists $0.20 an hour; Deepgram Flux is $0.0065 a minute on LiveKit's rate card.

$0.0033
4 · Language model

Every turn resends the prompt and the history. Gemini 2.5 Flash-Lite is $0.0004 a minute on LiveKit's assumptions; Claude Haiku 4.5 with a 6,800-token turn is $0.0072.

$0.0004
5 · Text to speech

Billed per character the agent speaks. At 600 characters a minute xAI's voice is $0.009 and Cartesia Sonic 3 is $0.030.

$0.0090
6 · Workers, state and logs

Agent processes, Redis, log storage and observability. Nobody sells this per minute; it is the measured gap between list prices and my production bill.

$0.0051
7 · Platform fee

The orchestration margin: Retell's voice infrastructure line is $0.055, Vapi's hosting fee $0.05, LiveKit Cloud's agent session $0.01. On a custom stack it is zero.

$0.0000
The badges add up to $0.0250 a minute, the figure I run in production. Lines 1 to 5 come from published rate cards checked on 23 September 2026, line 6 is measured, and line 7 is the one you are deciding whether to keep paying.

How much does an AI voice agent cost per minute in 2026?

Between about 2.5 and 14 cents a minute, depending on who orchestrates the call. Managed platforms land at roughly 8 to 14 cents once telephony is counted, LiveKit Cloud's own calculator defaults to 4.8 cents, and a custom LiveKit stack runs 2 to 3 cents. My production stack went from about 10 cents to about 2.5.

The published rates, checked on 23 September 2026. Retell prices its default configuration at $0.11 a minute ($0.055 voice infrastructure, $0.04 model, $0.015 voice) and adds $0.015 if you use its telephony. Vapi charges a $0.05 hosting fee and passes models and voices through at cost; its own example prices 1,000 minutes at $82 to $129. Bland lists $0.14 a minute on Start and $0.12 on Build, which carries a $299 monthly fee, with telephony billed separately. ElevenLabs lists its Speech Engine at $0.08 a minute, and xAI lists speech-to-speech at $0.08 a minute of audio.

My own figures are the baseline I trust most, because they come from real invoices. I integrated Retell into an AI interview pipeline in 48 hours, and that integration cut false-positive assessments from 50% to 15%, a result Retell published as a case study. Later, at volume, I rebuilt the voice layer on LiveKit and took the cost from about 10 cents a minute to about 2.5 cents, a 75% cut with assessment quality unchanged.

Name the baseline whenever you quote a multiple. Retell's default costs 4.4 times as much as 2.5 cents, LiveKit Cloud's default of $0.0479 costs 1.9 times as much, and my old bill was 4 times as much. The market moves quickly: in August the same LiveKit calculator defaulted to $0.0672, mostly because its default voice line has since fallen from $0.030 to $0.009 a minute.

What the market charges per minute
$ per minute, published list rates on 23 September 2026 plus two measured figureslower is better
Bland Starttelephony billed separately$0.140
Retell default plus Retell telephony$0.11 + $0.015$0.125
Bland Build$299 a month, telephony extra$0.120
Vapi, its own 1,000-minute example$0.05 hosting plus pass-through$0.082 to $0.129
My managed platform bill, beforemeasured$0.100
ElevenLabs Speech Engine$0.080
xAI speech-to-speechaudio only$0.080
LiveKit Cloud calculator defaulthosted agents$0.0479
Custom LiveKit stack, aftermeasured$0.025
Managed rates leave out some telephony and every add-on, so a real invoice usually lands above its bar, not below it. Speech-to-speech and cascaded pipelines are different products; compare them on latency and control as well as price.

What are the line items in an AI voice agent minute?

Seven: the carrier, the SIP bridge into your media server, speech-to-text, the language model, text-to-speech, the workers and logs that run the call, and a platform fee. Managed platforms fold most of them into one rate and bill telephony beside it. A custom stack pays each vendor at list price and replaces the platform fee with fixed costs.

The carrier and the bridge are the phone call. Telnyx lists inbound local at $0.0032 a minute, outbound local from $0.005 and toll-free inbound from $0.015; Twilio lists Elastic SIP inbound local at $0.0034. Bringing that trunk into LiveKit Cloud costs a third-party SIP fee of $0.004 a minute on Ship after 5,000 included minutes, or $0.003 on Scale after 50,000 (LiveKit pricing). LiveKit's own US number is $0.01 a minute. Media inside the room is covered by plan allowances, 150,000 participant-minutes on Ship, and past them costs at most about a tenth of a cent a minute by my arithmetic for two participants at $0.0005 each. Run the open-source LiveKit server and SIP service yourself and the bridge has no per-minute fee, only servers. The carrier comparison is in LiveKit SIP trunking with Twilio vs Telnyx.

Speech in both directions is where the market moved most. xAI lists streaming speech-to-text at $0.20 an hour, which is $0.0033 a minute, and text-to-speech at $15 per million characters. LiveKit's rate card converts voices to a per-minute price at 600 characters a minute, so xAI's voice is $0.009 and Cartesia Sonic 3 is $0.030 on Ship. Transcription listens to the whole call, silence included, so it bills on call minutes; the voice bills only for what the agent says.

The model line depends on how much text you resend each turn, not on the price per token alone. On LiveKit's assumption of 3,000 input and 175 output tokens a minute, Gemini 2.5 Flash-Lite costs $0.0004. A support agent with a 6,000-token cached prompt, 800 new tokens and an 80-token reply, four turns a minute, costs $0.0072 on Claude Haiku 4.5 at $1 and $5 per million tokens with cache reads at $0.10 (Anthropic pricing). A different model and a bigger prompt: eighteen times the model bill.

LineWhat you pay forPublished ratePer minute, my stackRange you will see
CarrierInbound local SIP minutesTelnyx $0.0032 in, from $0.005 out, from $0.015 toll-free$0.0032$0.0032 to $0.015
SIP bridgeYour trunk into LiveKit Cloud$0.004 Ship, $0.003 Scale; LiveKit number $0.01$0.0040$0 self-hosted to $0.01
MediaRooms, participant minutes, bandwidth150,000 participant-minutes and 250 GB included on Ship$0.0000$0 to about $0.001 past the allowance
Speech to textStreaming transcription of the whole callxAI $0.20 an hour; Deepgram Flux $0.0065 a minute$0.0033$0.0033 to $0.0065
Language modelPrompt, history and reply, every turnGemini 2.5 Flash-Lite $0.0004; Claude Haiku 4.5 profile $0.0072$0.0004$0.0004 to $0.0144
Text to speechCharacters the agent speaksxAI $15 per million characters; Cartesia Sonic 3 $0.030 at 600 a minute$0.0090$0.0045 to $0.030
Workers, state, logsAgent processes, Redis, logs, observabilityNo rate card; a c7i.xlarge is $0.1785 an hour$0.0051$0.0002 compute alone to $0.02 hosted on LiveKit Cloud
Platform feeThe orchestration marginRetell $0.055; Vapi $0.05; LiveKit Cloud agent session $0.01$0.0000$0 to $0.055
TotalOne inbound minute, my configurationList-price lines sum to $0.0199$0.0250Managed platforms: $0.08 to $0.14

How does a 2.5-cent voice AI minute add up?

Start from LiveKit Cloud's own default of $0.0479 and make four changes: host the agent workers yourself, bring your own SIP trunk, switch speech-to-text to xAI, and use a Flash-Lite model. The list-price lines fall to $0.0199. Workers, Redis, logs and observability add about $0.0051, and the measured total is about 2.5 cents.

The arithmetic, line by line. LiveKit's default minute is $0.0100 agent session, $0.0100 observability, $0.0100 telephony, $0.0090 voice, $0.0075 transcription and $0.0014 model. Self-hosted workers remove the session and observability lines, because LiveKit bills agent session minutes for agents deployed on its cloud. A Telnyx trunk plus the $0.004 SIP fee replaces $0.0100 of telephony with $0.0072. xAI transcription replaces $0.0075 with $0.0033, and Gemini 2.5 Flash-Lite replaces $0.0014 with $0.0004. That leaves $0.0199.

The last half cent is what no rate card sells. On AWS a c7i.xlarge is $0.1785 an hour, and LiveKit sizes a 4-core, 8 GB server at 10 to 25 concurrent calls (LiveKit deployments), so raw compute at 15 calls a server is $0.0002 a minute. The rest of the $0.0051 is workers sized for peaks rather than averages, Redis, log storage and observability. That is why I quote 2.5 cents and not 2.

Four variations cover most real deployments. Outbound calls pay $0.005 at Telnyx instead of $0.0032, which is why I plan outbound at about 2.7 cents. A toll-free number adds about 1.2 cents. A Haiku-class model with a long prompt adds about 0.7 cents. And if the agent talks for half the call rather than all of it, the voice line halves to $0.0045.

From LiveKit's $0.0479 default to a 2.5-cent minute
  1. 1
    LiveKit Cloud default$0.0479

    Agent session $0.0100, observability $0.0100, LiveKit number $0.0100, voice $0.0090, transcription $0.0075, model $0.0014.

  2. 2
    Host the agent workers yourselfto $0.0279

    Agent session minutes apply to agents deployed on LiveKit Cloud, and your own logs replace its observability line. You now pay for servers instead.

  3. 3
    Bring a Telnyx trunkto $0.0251

    $0.0032 of carrier plus the $0.004 SIP fee on Ship replaces the $0.0100 bundled number.

  4. 4
    Switch transcription to xAIto $0.0209

    $0.20 an hour is $0.0033 a minute, against the default line of $0.0075.

  5. 5
    Use a Flash-Lite modelto $0.0199

    Gemini 2.5 Flash-Lite at $0.0004 a minute on LiveKit's token assumptions, against $0.0014.

  6. 6
    Add what nobody sells per minuteabout $0.0250

    Workers sized for peaks, Redis, log storage and observability: about $0.0051 a minute, measured on my own bill.

Every step before the last uses a published price. The last is my own bill, and it is the step most per-minute comparisons leave out.

Which choices move the per-minute cost the most?

Three choices: who takes the platform fee, which voice you use, and how much text the model reads each turn. The platform fee is worth up to 5.5 cents, the voice about 2 cents between budget and mid-priced options, and the model about 0.7 cents in a typical support profile. Transcription and carrier choice move fractions of a cent.

The platform fee is the biggest line on a managed invoice and the one a custom stack removes. Half of Retell's $0.11 default is its $0.055 voice infrastructure line, and Vapi's $0.05 hosting fee is charged before any model or voice. That margin buys real things: orchestration, a dashboard, telephony plumbing and a support contact. At 10,000 minutes a month it is $500 to $550, less than any engineer. At a million minutes it is $50,000 or more a month.

The voice is the widest spread among the vendor lines. At 600 characters a minute, xAI's voice is $0.009, Cartesia Sonic 3 is $0.030 on LiveKit's Ship rate card and $0.0225 on Scale, and ElevenLabs lists Flash and Turbo at $0.05 per 1,000 characters on its own API, which is also $0.030. Choose the voice by ear first, but know that a mid-priced voice costs more per minute than the model and transcription of my stack combined. The trade-offs are in the best TTS for voice agents.

The model is cheap until the prompt grows. Every turn resends the system prompt, tool schemas, retrieved passages and history, so the 6,800-token profile above costs $0.0072 a minute on Claude Haiku 4.5 and $0.0144 on Claude Sonnet 5 at $2 and $10 per million tokens with cache reads at $0.20, more than the voice. Cache the stable prefix, keep retrieval to four to six passages, and route only hard turns to a bigger model. Gemini 3.1 Flash-Lite, at $0.25 and $1.50 per million tokens with cache reads at $0.025 (Gemini pricing), runs the same profile for $0.0019 a minute.

One minute, one change at a time
$ per minute: my measured minute with a single line changedlower is better
My stack as measured$0.0250
With Claude Haiku 4.5 and 6,800-token turnsmodel line $0.0072$0.0318
With a toll-free numbercarrier from $0.015$0.0368
With Claude Sonnet 5 and the same turnsmodel line $0.0144$0.0390
With agents and observability hosted on LiveKit Cloud$0.02 instead of $0.0051$0.0399
With Cartesia Sonic 3 as the voicevoice line $0.030$0.0460
Each bar changes one line and keeps the rest. The voice and the prompt are decisions made in an afternoon that set the bill for years, which is why they deserve a spreadsheet before a demo decides them.

How do managed platforms, LiveKit Cloud and a custom stack compare on cost?

They sell the same pipeline at three levels of ownership. A managed platform charges 8 to 14 cents and runs everything. LiveKit Cloud's agent hosting charges about 4.8 cents by default and runs the media and the workers. A custom stack pays vendors at list, runs its own workers, and lands near 2.5 cents plus fixed costs.

What you give up moving right is what the platform fee pays for. Retell includes 20 concurrent calls and sells more at $8 each a month, along with post-call analysis, a dashboard and one support contact. Vapi includes 4 concurrent lines without a support package, sells more at $10 a line a month, and charges $2,000 a month for HIPAA. On your own stack each of those becomes a ticket in your backlog, and the pager is yours.

LiveKit Cloud sits in the middle and is a sensible first step off a platform. You keep hosted agents, turn detection and Agent Insights, which shows transcripts, traces, logs and recordings on one timeline for each session, and you pay $0.01 a minute for agent sessions after the included minutes. The Scale plan adds a SOC 2 Type II report and a signed BAA for HIPAA.

The custom column is the one this post is about, and its catch is fixed cost. You pay for servers whether or not calls arrive, and someone has to own upgrades, vendor changes and the 2am page. I model that as $600 a month of infrastructure plus two engineer-days of maintenance at a $150,000 loaded salary, about $1,754 a month, the same inputs as the voice AI build vs buy break-even.

Three levels of ownership
Managed platform
Fastest to live, highest minute
  • 8 to 14 cents a minute on published rates
  • Platform fee of $0.05 to $0.055 before any model
  • Concurrency, dashboard and support included or sold as add-ons
  • The right answer below about 20,000 minutes a month
LiveKit Cloud agents
Hosted pipeline, vendors you choose
  • $0.0479 a minute on the calculator default
  • Agent session and observability at $0.01 a minute each
  • Transcription, model and voice picked from a rate card
  • A step off a platform without running servers
pick
Custom LiveKit stack
Lowest minute, you own the pager
  • About 2.5 cents a minute, measured
  • No platform fee; about $1,754 a month fixed in my model
  • Every vendor replaceable the week a better one ships
  • Needs someone who owns incidents and upgrades
Moving right lowers the minute and raises the fixed cost. Where you should sit is set by your volume and your staffing, not by preference.

What does a voice AI agent cost at 20,000, 100,000 and 1 million minutes a month?

At 20,000 minutes a managed platform at 10 cents costs $2,000 and a custom stack $2,254 with fixed costs, so stay managed. At 100,000 it is $10,000 against $4,254, a saving of $5,746 a month. At 1 million it is $100,000 against $26,754, and the build repays itself in about a week.

The model behind those numbers: a managed rate of 10 cents, which is what my old platform cost; a custom rate of 2.5 cents; $1,754 a month of fixed infrastructure and maintenance; and a six-week build, about $17,300 at a $150,000 loaded salary. The break-even is $1,754 divided by the 7.5-cent gap, about 23,400 minutes a month. Against Retell's list price with its telephony, 12.5 cents, it falls to about 17,500. That is where the rule of thumb of roughly 20,000 minutes comes from.

Payback is the second number to check, because break-even ignores the build. At 100,000 minutes the $5,746 monthly saving repays $17,300 in about three months, or about two against Retell's list price. At 20,000 minutes the saving against Retell's list price is $246 a month, a 70-month payback, which is another way of saying never.

Run your own inputs in the voice AI cost calculator: minutes, vendor rate and engineering cost. If the case only works at a salary well below what you pay, or only with credits applied, it does not work.

Monthly minutesManaged at 10 centsRetell list with telephony (12.5 cents)Custom: 2.5 cents plus $1,754 fixedSaving against 10 centsBuild payback
20,000$2,000$2,500$2,254loses $254never
50,000$5,000$6,250$3,004$1,9968.7 months
100,000$10,000$12,500$4,254$5,7463.0 months
250,000$25,000$31,250$8,004$16,9961.0 month
1,000,000$100,000$125,000$26,754$73,246about a week
Monthly cost by volume
140,000105,00070,00035,000010K20K50K100K250K500K1M$ per monthMinutes per month
about 23,400 min/mo
Managed at 10 cents a minuteCustom at 2.5 cents plus $1,754 fixedRetell list with its telephony, 12.5 cents
Below the crossover the fixed line dominates and a platform is cheaper; above it the gap widens by 7.5 cents every minute. The x-axis is not to scale, so read the values rather than the slopes.

What changes as you grow from 20,000 to 1 million voice minutes a month?

The question moves from whether to build to which line to squeeze. At 20,000 minutes the fixed costs decide and a platform wins. At 100,000 the plan and the voice start to matter. At 1 million every tenth of a cent is $1,000 a month, vendor concurrency caps bind, and running your own SIP bridge is worth pricing.

At 20,000 minutes a month, about 670 a day, the busy hour carries about 1.3 simultaneous calls if 12% of a day's calls land in it, which is my assumption, and five channels keep blocking under 1%. Nothing in the stack is near a ceiling. Stay on a platform and use the time to log cost and latency per call, so a later move has a baseline.

At 100,000 minutes the plan matters. LiveKit's base fee plus third-party SIP minutes comes to $430 a month on Ship against $650 on Scale; Scale only wins on those two lines above about 320,000 minutes, so move for its SOC 2 report or its limits, not for price. The voice is now worth $2,100 a month between xAI and Cartesia on Ship. The busy hour carries about 6.7 simultaneous calls, or 14 channels, well inside Deepgram's pay-as-you-go limits and LiveKit Ship's 20 inference connections.

At 1 million minutes, the smallest line is worth money. Scale's SIP fee is $3,350 a month, enough to price running the Apache-licensed LiveKit server and SIP service yourself. The busy hour needs about 82 channels: Telnyx sells inbound channels at $12, $11 and $9 a month by tier, so 82 cost $848 against $3,200 of per-minute inbound. And 82 calls exceed Deepgram's published 45 text-to-speech streams on pay-as-you-go and, if you buy models through LiveKit, Scale's 50 inference connections, so plan for an enterprise tier or a second vendor. Bland, Vapi and LiveKit all list custom enterprise tiers, so treat list rates as a ceiling.

What to change at your volume
What should you do at your monthly volume?
Under 20,000 minutes a month
Stay managed and instrument

Fixed costs exceed the saving. Log cost and latency per call now so a later move has a baseline.

20,000 to 100,000 minutes, two engineers who can own it
Build when the saving covers the pager

The monthly saving rises from nothing to about $5,746 against a 10-cent platform, and payback falls to about three months.

100,000 to 500,000 minutes
Build, then tune the voice and the prompt

The voice alone is worth thousands a month, and LiveKit Scale beats Ship on SIP price above about 320,000 minutes.

Over 500,000 minutes
Negotiate every line and price self-hosting

Inbound channels instead of per-minute billing, enterprise concurrency, and your own LiveKit SIP service once its fee passes an engineer's time.

A compliance agreement needed inside 90 days
Buy from a vendor that signs

Every vendor in a custom chain needs its own agreement, and that is procurement time, not engineering time.

Two of five branches say keep paying a platform. The dividing lines are volume, staffing and paperwork, not ambition.

When should you stay on a managed voice AI platform?

When you run under about 20,000 minutes a month, when nobody on the team can own a voice pipeline and its pager, when a compliance agreement is needed this quarter, or when time to market matters more than margin. Those are common situations, and paying a platform through them is a sound decision.

I say that as someone who sells builds. Retell went into production in my interview pipeline in 48 hours, and the custom stack came later, when the volume made 7.5 cents a minute worth an engineer's attention. Both decisions were right when I made them. The mistake is not starting on a platform; it is staying on one long after the invoice has passed the cost of owning the pipeline.

If you do move, move gradually: run both stacks behind a routing layer, compare cost per completed call as well as cost per minute, and ramp traffic in stages. The full sequence is in migrating from Retell or Vapi to LiveKit, the architecture you are moving to is in how to build your own voice AI platform on LiveKit, and for a support line at thousands of calls a day the product version is an AI call center.

If you want the stack built rather than modelled, that is voice AI development, with voice agent builds from $12,000. It is built at $0: the work is split into checkpoints with acceptance criteria agreed before work starts, and each is invoiced only after you have seen it and accepted it.

AI voice agent cost per minute: common questions

→How much does an AI voice agent cost per minute?

Managed platforms land at about 8 to 14 cents a minute once telephony is included: Retell's default is $0.11 plus $0.015 for its telephony, and Bland lists $0.12 to $0.14 before telephony. A custom LiveKit stack costs about 2.5 cents a minute in my production system, down from about 10 cents on a managed platform.

→What is the most expensive part of a voice AI minute?

On a managed platform it is the platform fee: $0.055 of Retell's $0.11 default, or Vapi's $0.05 hosting fee. On a custom stack it is usually the voice, $0.009 a minute for xAI and $0.030 for Cartesia Sonic 3 at 600 characters a minute, unless the model resends a long prompt every turn.

→At what volume should I build my own voice AI stack?

Around 20,000 minutes a month. With $1,754 a month of fixed infrastructure and maintenance, a custom stack at 2.5 cents breaks even at about 23,400 minutes against a 10-cent platform and about 17,500 against Retell's 12.5 cents with telephony. At 100,000 minutes it saves about $5,746 a month and repays a six-week build in three months.

→Do per-minute voice AI prices include telephony?

Often not. Retell's $0.11 default assumes your own carrier, and its telephony adds $0.015 a minute. Bland bills telephony separately. On a custom stack the carrier is its own line: Telnyx lists inbound local at $0.0032 a minute and toll-free from $0.015, plus LiveKit Cloud's $0.004 SIP fee on the Ship plan when you bring your own trunk.

→Is a speech-to-speech model cheaper than a cascaded pipeline?

Not at list price. ElevenLabs lists its Speech Engine at $0.08 a minute and xAI lists speech-to-speech at $0.08 a minute of audio, about three times a tuned cascade at 2.5 cents. Google's Gemini 3.8 Live lists audio at $0.005 a minute in and $0.018 out, which is closer, before text tokens and telephony. Compare latency and control as well.

→Why do per-minute voice AI quotes differ so much?

Because three assumptions hide in every quote: how many characters a minute the agent speaks, how much of the call it talks, and how many tokens each turn resends. LiveKit assumes 600 characters and 3,000 input tokens a minute. Halving the talk share halves the voice line, and 6,800-token turns on Claude Haiku 4.5 cost eighteen times LiveKit's Flash-Lite default.

Take this into your own chat

Open the article in your assistant with one click and ask it how this applies to your product.

Keep reading

See it in production: voice AI at 2.5¢ a minute, the case study

Ready to talk numbers?

Twenty minutes, straight to the engineer. No sales rep, no deck.