AI Voice Agent Cost Per Minute (2026): Every Line Item, Managed vs Custom, 20K to 1M Minutes
- Managed voice platforms land between about 8 and 14 cents a minute once telephony is counted; Retell's default is $0.11 plus $0.015 for its telephony. I moved a production voice stack from about 10 cents a minute on a managed platform to about 2.5 cents on a custom LiveKit stack.
- The 2.5 cents is seven lines: carrier $0.0032, LiveKit's SIP fee $0.004, speech-to-text $0.0033, the model $0.0004, the voice $0.009 and about $0.0051 of workers, state and logs, with no platform fee. The voice and the prompt size swing the total more than any other choice.
- Stay managed below roughly 20,000 minutes a month, where fixed infrastructure and maintenance eat the saving. At 100,000 minutes a custom stack saves about $5,700 a month and repays a six-week build in about three months; at 1 million it saves about $73,000 a month.
The phone minute itself, on a SIP trunk. Telnyx lists inbound local at $0.0032, outbound local from $0.005 and toll-free inbound from $0.015.
$0.0032Getting the call into a LiveKit room. LiveKit Cloud charges $0.004 a minute on Ship for a trunk you bring, $0.003 on Scale, $0.01 for its own number. Self-hosted LiveKit has no per-minute fee.
$0.0040Streaming transcription of the whole call, silence included. xAI lists $0.20 an hour; Deepgram Flux is $0.0065 a minute on LiveKit's rate card.
$0.0033Every turn resends the prompt and the history. Gemini 2.5 Flash-Lite is $0.0004 a minute on LiveKit's assumptions; Claude Haiku 4.5 with a 6,800-token turn is $0.0072.
$0.0004Billed per character the agent speaks. At 600 characters a minute xAI's voice is $0.009 and Cartesia Sonic 3 is $0.030.
$0.0090Agent processes, Redis, log storage and observability. Nobody sells this per minute; it is the measured gap between list prices and my production bill.
$0.0051The orchestration margin: Retell's voice infrastructure line is $0.055, Vapi's hosting fee $0.05, LiveKit Cloud's agent session $0.01. On a custom stack it is zero.
$0.0000How much does an AI voice agent cost per minute in 2026?
Between about 2.5 and 14 cents a minute, depending on who orchestrates the call. Managed platforms land at roughly 8 to 14 cents once telephony is counted, LiveKit Cloud's own calculator defaults to 4.8 cents, and a custom LiveKit stack runs 2 to 3 cents. My production stack went from about 10 cents to about 2.5.
The published rates, checked on 23 September 2026. Retell prices its default configuration at $0.11 a minute ($0.055 voice infrastructure, $0.04 model, $0.015 voice) and adds $0.015 if you use its telephony. Vapi charges a $0.05 hosting fee and passes models and voices through at cost; its own example prices 1,000 minutes at $82 to $129. Bland lists $0.14 a minute on Start and $0.12 on Build, which carries a $299 monthly fee, with telephony billed separately. ElevenLabs lists its Speech Engine at $0.08 a minute, and xAI lists speech-to-speech at $0.08 a minute of audio.
My own figures are the baseline I trust most, because they come from real invoices. I integrated Retell into an AI interview pipeline in 48 hours, and that integration cut false-positive assessments from 50% to 15%, a result Retell published as a case study. Later, at volume, I rebuilt the voice layer on LiveKit and took the cost from about 10 cents a minute to about 2.5 cents, a 75% cut with assessment quality unchanged.
Name the baseline whenever you quote a multiple. Retell's default costs 4.4 times as much as 2.5 cents, LiveKit Cloud's default of $0.0479 costs 1.9 times as much, and my old bill was 4 times as much. The market moves quickly: in August the same LiveKit calculator defaulted to $0.0672, mostly because its default voice line has since fallen from $0.030 to $0.009 a minute.
What are the line items in an AI voice agent minute?
Seven: the carrier, the SIP bridge into your media server, speech-to-text, the language model, text-to-speech, the workers and logs that run the call, and a platform fee. Managed platforms fold most of them into one rate and bill telephony beside it. A custom stack pays each vendor at list price and replaces the platform fee with fixed costs.
The carrier and the bridge are the phone call. Telnyx lists inbound local at $0.0032 a minute, outbound local from $0.005 and toll-free inbound from $0.015; Twilio lists Elastic SIP inbound local at $0.0034. Bringing that trunk into LiveKit Cloud costs a third-party SIP fee of $0.004 a minute on Ship after 5,000 included minutes, or $0.003 on Scale after 50,000 (LiveKit pricing). LiveKit's own US number is $0.01 a minute. Media inside the room is covered by plan allowances, 150,000 participant-minutes on Ship, and past them costs at most about a tenth of a cent a minute by my arithmetic for two participants at $0.0005 each. Run the open-source LiveKit server and SIP service yourself and the bridge has no per-minute fee, only servers. The carrier comparison is in LiveKit SIP trunking with Twilio vs Telnyx.
Speech in both directions is where the market moved most. xAI lists streaming speech-to-text at $0.20 an hour, which is $0.0033 a minute, and text-to-speech at $15 per million characters. LiveKit's rate card converts voices to a per-minute price at 600 characters a minute, so xAI's voice is $0.009 and Cartesia Sonic 3 is $0.030 on Ship. Transcription listens to the whole call, silence included, so it bills on call minutes; the voice bills only for what the agent says.
The model line depends on how much text you resend each turn, not on the price per token alone. On LiveKit's assumption of 3,000 input and 175 output tokens a minute, Gemini 2.5 Flash-Lite costs $0.0004. A support agent with a 6,000-token cached prompt, 800 new tokens and an 80-token reply, four turns a minute, costs $0.0072 on Claude Haiku 4.5 at $1 and $5 per million tokens with cache reads at $0.10 (Anthropic pricing). A different model and a bigger prompt: eighteen times the model bill.
| Line | What you pay for | Published rate | Per minute, my stack | Range you will see |
|---|---|---|---|---|
| Carrier | Inbound local SIP minutes | Telnyx $0.0032 in, from $0.005 out, from $0.015 toll-free | $0.0032 | $0.0032 to $0.015 |
| SIP bridge | Your trunk into LiveKit Cloud | $0.004 Ship, $0.003 Scale; LiveKit number $0.01 | $0.0040 | $0 self-hosted to $0.01 |
| Media | Rooms, participant minutes, bandwidth | 150,000 participant-minutes and 250 GB included on Ship | $0.0000 | $0 to about $0.001 past the allowance |
| Speech to text | Streaming transcription of the whole call | xAI $0.20 an hour; Deepgram Flux $0.0065 a minute | $0.0033 | $0.0033 to $0.0065 |
| Language model | Prompt, history and reply, every turn | Gemini 2.5 Flash-Lite $0.0004; Claude Haiku 4.5 profile $0.0072 | $0.0004 | $0.0004 to $0.0144 |
| Text to speech | Characters the agent speaks | xAI $15 per million characters; Cartesia Sonic 3 $0.030 at 600 a minute | $0.0090 | $0.0045 to $0.030 |
| Workers, state, logs | Agent processes, Redis, logs, observability | No rate card; a c7i.xlarge is $0.1785 an hour | $0.0051 | $0.0002 compute alone to $0.02 hosted on LiveKit Cloud |
| Platform fee | The orchestration margin | Retell $0.055; Vapi $0.05; LiveKit Cloud agent session $0.01 | $0.0000 | $0 to $0.055 |
| Total | One inbound minute, my configuration | List-price lines sum to $0.0199 | $0.0250 | Managed platforms: $0.08 to $0.14 |
How does a 2.5-cent voice AI minute add up?
Start from LiveKit Cloud's own default of $0.0479 and make four changes: host the agent workers yourself, bring your own SIP trunk, switch speech-to-text to xAI, and use a Flash-Lite model. The list-price lines fall to $0.0199. Workers, Redis, logs and observability add about $0.0051, and the measured total is about 2.5 cents.
The arithmetic, line by line. LiveKit's default minute is $0.0100 agent session, $0.0100 observability, $0.0100 telephony, $0.0090 voice, $0.0075 transcription and $0.0014 model. Self-hosted workers remove the session and observability lines, because LiveKit bills agent session minutes for agents deployed on its cloud. A Telnyx trunk plus the $0.004 SIP fee replaces $0.0100 of telephony with $0.0072. xAI transcription replaces $0.0075 with $0.0033, and Gemini 2.5 Flash-Lite replaces $0.0014 with $0.0004. That leaves $0.0199.
The last half cent is what no rate card sells. On AWS a c7i.xlarge is $0.1785 an hour, and LiveKit sizes a 4-core, 8 GB server at 10 to 25 concurrent calls (LiveKit deployments), so raw compute at 15 calls a server is $0.0002 a minute. The rest of the $0.0051 is workers sized for peaks rather than averages, Redis, log storage and observability. That is why I quote 2.5 cents and not 2.
Four variations cover most real deployments. Outbound calls pay $0.005 at Telnyx instead of $0.0032, which is why I plan outbound at about 2.7 cents. A toll-free number adds about 1.2 cents. A Haiku-class model with a long prompt adds about 0.7 cents. And if the agent talks for half the call rather than all of it, the voice line halves to $0.0045.
- 1LiveKit Cloud default$0.0479
Agent session $0.0100, observability $0.0100, LiveKit number $0.0100, voice $0.0090, transcription $0.0075, model $0.0014.
- 2Host the agent workers yourselfto $0.0279
Agent session minutes apply to agents deployed on LiveKit Cloud, and your own logs replace its observability line. You now pay for servers instead.
- 3Bring a Telnyx trunkto $0.0251
$0.0032 of carrier plus the $0.004 SIP fee on Ship replaces the $0.0100 bundled number.
- 4Switch transcription to xAIto $0.0209
$0.20 an hour is $0.0033 a minute, against the default line of $0.0075.
- 5Use a Flash-Lite modelto $0.0199
Gemini 2.5 Flash-Lite at $0.0004 a minute on LiveKit's token assumptions, against $0.0014.
- 6Add what nobody sells per minuteabout $0.0250
Workers sized for peaks, Redis, log storage and observability: about $0.0051 a minute, measured on my own bill.
Which choices move the per-minute cost the most?
Three choices: who takes the platform fee, which voice you use, and how much text the model reads each turn. The platform fee is worth up to 5.5 cents, the voice about 2 cents between budget and mid-priced options, and the model about 0.7 cents in a typical support profile. Transcription and carrier choice move fractions of a cent.
The platform fee is the biggest line on a managed invoice and the one a custom stack removes. Half of Retell's $0.11 default is its $0.055 voice infrastructure line, and Vapi's $0.05 hosting fee is charged before any model or voice. That margin buys real things: orchestration, a dashboard, telephony plumbing and a support contact. At 10,000 minutes a month it is $500 to $550, less than any engineer. At a million minutes it is $50,000 or more a month.
The voice is the widest spread among the vendor lines. At 600 characters a minute, xAI's voice is $0.009, Cartesia Sonic 3 is $0.030 on LiveKit's Ship rate card and $0.0225 on Scale, and ElevenLabs lists Flash and Turbo at $0.05 per 1,000 characters on its own API, which is also $0.030. Choose the voice by ear first, but know that a mid-priced voice costs more per minute than the model and transcription of my stack combined. The trade-offs are in the best TTS for voice agents.
The model is cheap until the prompt grows. Every turn resends the system prompt, tool schemas, retrieved passages and history, so the 6,800-token profile above costs $0.0072 a minute on Claude Haiku 4.5 and $0.0144 on Claude Sonnet 5 at $2 and $10 per million tokens with cache reads at $0.20, more than the voice. Cache the stable prefix, keep retrieval to four to six passages, and route only hard turns to a bigger model. Gemini 3.1 Flash-Lite, at $0.25 and $1.50 per million tokens with cache reads at $0.025 (Gemini pricing), runs the same profile for $0.0019 a minute.
How do managed platforms, LiveKit Cloud and a custom stack compare on cost?
They sell the same pipeline at three levels of ownership. A managed platform charges 8 to 14 cents and runs everything. LiveKit Cloud's agent hosting charges about 4.8 cents by default and runs the media and the workers. A custom stack pays vendors at list, runs its own workers, and lands near 2.5 cents plus fixed costs.
What you give up moving right is what the platform fee pays for. Retell includes 20 concurrent calls and sells more at $8 each a month, along with post-call analysis, a dashboard and one support contact. Vapi includes 4 concurrent lines without a support package, sells more at $10 a line a month, and charges $2,000 a month for HIPAA. On your own stack each of those becomes a ticket in your backlog, and the pager is yours.
LiveKit Cloud sits in the middle and is a sensible first step off a platform. You keep hosted agents, turn detection and Agent Insights, which shows transcripts, traces, logs and recordings on one timeline for each session, and you pay $0.01 a minute for agent sessions after the included minutes. The Scale plan adds a SOC 2 Type II report and a signed BAA for HIPAA.
The custom column is the one this post is about, and its catch is fixed cost. You pay for servers whether or not calls arrive, and someone has to own upgrades, vendor changes and the 2am page. I model that as $600 a month of infrastructure plus two engineer-days of maintenance at a $150,000 loaded salary, about $1,754 a month, the same inputs as the voice AI build vs buy break-even.
- 8 to 14 cents a minute on published rates
- Platform fee of $0.05 to $0.055 before any model
- Concurrency, dashboard and support included or sold as add-ons
- The right answer below about 20,000 minutes a month
- $0.0479 a minute on the calculator default
- Agent session and observability at $0.01 a minute each
- Transcription, model and voice picked from a rate card
- A step off a platform without running servers
- About 2.5 cents a minute, measured
- No platform fee; about $1,754 a month fixed in my model
- Every vendor replaceable the week a better one ships
- Needs someone who owns incidents and upgrades
What does a voice AI agent cost at 20,000, 100,000 and 1 million minutes a month?
At 20,000 minutes a managed platform at 10 cents costs $2,000 and a custom stack $2,254 with fixed costs, so stay managed. At 100,000 it is $10,000 against $4,254, a saving of $5,746 a month. At 1 million it is $100,000 against $26,754, and the build repays itself in about a week.
The model behind those numbers: a managed rate of 10 cents, which is what my old platform cost; a custom rate of 2.5 cents; $1,754 a month of fixed infrastructure and maintenance; and a six-week build, about $17,300 at a $150,000 loaded salary. The break-even is $1,754 divided by the 7.5-cent gap, about 23,400 minutes a month. Against Retell's list price with its telephony, 12.5 cents, it falls to about 17,500. That is where the rule of thumb of roughly 20,000 minutes comes from.
Payback is the second number to check, because break-even ignores the build. At 100,000 minutes the $5,746 monthly saving repays $17,300 in about three months, or about two against Retell's list price. At 20,000 minutes the saving against Retell's list price is $246 a month, a 70-month payback, which is another way of saying never.
Run your own inputs in the voice AI cost calculator: minutes, vendor rate and engineering cost. If the case only works at a salary well below what you pay, or only with credits applied, it does not work.
| Monthly minutes | Managed at 10 cents | Retell list with telephony (12.5 cents) | Custom: 2.5 cents plus $1,754 fixed | Saving against 10 cents | Build payback |
|---|---|---|---|---|---|
| 20,000 | $2,000 | $2,500 | $2,254 | loses $254 | never |
| 50,000 | $5,000 | $6,250 | $3,004 | $1,996 | 8.7 months |
| 100,000 | $10,000 | $12,500 | $4,254 | $5,746 | 3.0 months |
| 250,000 | $25,000 | $31,250 | $8,004 | $16,996 | 1.0 month |
| 1,000,000 | $100,000 | $125,000 | $26,754 | $73,246 | about a week |
What changes as you grow from 20,000 to 1 million voice minutes a month?
The question moves from whether to build to which line to squeeze. At 20,000 minutes the fixed costs decide and a platform wins. At 100,000 the plan and the voice start to matter. At 1 million every tenth of a cent is $1,000 a month, vendor concurrency caps bind, and running your own SIP bridge is worth pricing.
At 20,000 minutes a month, about 670 a day, the busy hour carries about 1.3 simultaneous calls if 12% of a day's calls land in it, which is my assumption, and five channels keep blocking under 1%. Nothing in the stack is near a ceiling. Stay on a platform and use the time to log cost and latency per call, so a later move has a baseline.
At 100,000 minutes the plan matters. LiveKit's base fee plus third-party SIP minutes comes to $430 a month on Ship against $650 on Scale; Scale only wins on those two lines above about 320,000 minutes, so move for its SOC 2 report or its limits, not for price. The voice is now worth $2,100 a month between xAI and Cartesia on Ship. The busy hour carries about 6.7 simultaneous calls, or 14 channels, well inside Deepgram's pay-as-you-go limits and LiveKit Ship's 20 inference connections.
At 1 million minutes, the smallest line is worth money. Scale's SIP fee is $3,350 a month, enough to price running the Apache-licensed LiveKit server and SIP service yourself. The busy hour needs about 82 channels: Telnyx sells inbound channels at $12, $11 and $9 a month by tier, so 82 cost $848 against $3,200 of per-minute inbound. And 82 calls exceed Deepgram's published 45 text-to-speech streams on pay-as-you-go and, if you buy models through LiveKit, Scale's 50 inference connections, so plan for an enterprise tier or a second vendor. Bland, Vapi and LiveKit all list custom enterprise tiers, so treat list rates as a ceiling.
Fixed costs exceed the saving. Log cost and latency per call now so a later move has a baseline.
The monthly saving rises from nothing to about $5,746 against a 10-cent platform, and payback falls to about three months.
The voice alone is worth thousands a month, and LiveKit Scale beats Ship on SIP price above about 320,000 minutes.
Inbound channels instead of per-minute billing, enterprise concurrency, and your own LiveKit SIP service once its fee passes an engineer's time.
Every vendor in a custom chain needs its own agreement, and that is procurement time, not engineering time.
When should you stay on a managed voice AI platform?
When you run under about 20,000 minutes a month, when nobody on the team can own a voice pipeline and its pager, when a compliance agreement is needed this quarter, or when time to market matters more than margin. Those are common situations, and paying a platform through them is a sound decision.
I say that as someone who sells builds. Retell went into production in my interview pipeline in 48 hours, and the custom stack came later, when the volume made 7.5 cents a minute worth an engineer's attention. Both decisions were right when I made them. The mistake is not starting on a platform; it is staying on one long after the invoice has passed the cost of owning the pipeline.
If you do move, move gradually: run both stacks behind a routing layer, compare cost per completed call as well as cost per minute, and ramp traffic in stages. The full sequence is in migrating from Retell or Vapi to LiveKit, the architecture you are moving to is in how to build your own voice AI platform on LiveKit, and for a support line at thousands of calls a day the product version is an AI call center.
If you want the stack built rather than modelled, that is voice AI development, with voice agent builds from $12,000. It is built at $0: the work is split into checkpoints with acceptance criteria agreed before work starts, and each is invoiced only after you have seen it and accepted it.
AI voice agent cost per minute: common questions
→How much does an AI voice agent cost per minute?
Managed platforms land at about 8 to 14 cents a minute once telephony is included: Retell's default is $0.11 plus $0.015 for its telephony, and Bland lists $0.12 to $0.14 before telephony. A custom LiveKit stack costs about 2.5 cents a minute in my production system, down from about 10 cents on a managed platform.
→What is the most expensive part of a voice AI minute?
On a managed platform it is the platform fee: $0.055 of Retell's $0.11 default, or Vapi's $0.05 hosting fee. On a custom stack it is usually the voice, $0.009 a minute for xAI and $0.030 for Cartesia Sonic 3 at 600 characters a minute, unless the model resends a long prompt every turn.
→At what volume should I build my own voice AI stack?
Around 20,000 minutes a month. With $1,754 a month of fixed infrastructure and maintenance, a custom stack at 2.5 cents breaks even at about 23,400 minutes against a 10-cent platform and about 17,500 against Retell's 12.5 cents with telephony. At 100,000 minutes it saves about $5,746 a month and repays a six-week build in three months.
→Do per-minute voice AI prices include telephony?
Often not. Retell's $0.11 default assumes your own carrier, and its telephony adds $0.015 a minute. Bland bills telephony separately. On a custom stack the carrier is its own line: Telnyx lists inbound local at $0.0032 a minute and toll-free from $0.015, plus LiveKit Cloud's $0.004 SIP fee on the Ship plan when you bring your own trunk.
→Is a speech-to-speech model cheaper than a cascaded pipeline?
Not at list price. ElevenLabs lists its Speech Engine at $0.08 a minute and xAI lists speech-to-speech at $0.08 a minute of audio, about three times a tuned cascade at 2.5 cents. Google's Gemini 3.8 Live lists audio at $0.005 a minute in and $0.018 out, which is closer, before text tokens and telephony. Compare latency and control as well.
→Why do per-minute voice AI quotes differ so much?
Because three assumptions hide in every quote: how many characters a minute the agent speaks, how much of the call it talks, and how many tokens each turn resends. LiveKit assumes 600 characters and 3,000 input tokens a minute. Halving the talk share halves the voice line, and 6,800-token turns on Claude Haiku 4.5 cost eighteen times LiveKit's Flash-Lite default.
Open the article in your assistant with one click and ask it how this applies to your product.