Request a callbackBook a call
← All posts

Best TTS for Voice Agents in 2026: Real Cost Per Minute, Not Per Million Characters

TL;DR
  • TTS is the largest single line item in a voice agent's per-minute cost: $0.0300 of LiveKit's own $0.0672 default, more than STT, LLM, agent hosting and telephony combined.
  • Per-character list prices are unusable for budgeting. At LiveKit's assumed 600 characters per minute of agent speech, the market spans $0.009/min to $0.18/min, a 20x range for the same job.
  • You only pay for the agent's speech, roughly 40% of call time, so a 4-minute call bills about 1.6 TTS-minutes. Most published estimates silently assume 100% and are about 2.5x too high.
The whole category, converted to ¢/min
$ per minute of agent speech @ 600 chars/min (list prices, 24 Aug 2026)lower is better
xAI Text to Speech$15/1M chars$0.0090
Fish Audio S2.1 Pro$15/1M chars$0.0090
Deepgram Aura-1$0.015/1K chars$0.0090
Inworld Realtime TTS 1.5 Mini$15/1M chars$0.0090
Inworld Realtime TTS 2.0$25/1M chars$0.0150
Deepgram Aura-2$0.030/1K chars$0.0180
Rime Mist v3$0.03/1K chars$0.0180
Cartesia Sonic 3$50/1M chars$0.0300
Rime Coda$0.05/1K chars$0.0300
ElevenLabs Flash v2.5$150/1M chars$0.0900
ElevenLabs Multilingual v2$300/1M chars$0.1800
Same list prices you can read on the vendors' own pages and on LiveKit's published inference rate card, converted to a unit you can budget in. The spread is 20x between the cheapest credible option and ElevenLabs Multilingual v2, and if you're picking on a demo rather than this chart, that 20x is the number you're unknowingly agreeing to. Assumption: 600 characters per minute of agent speech, LiveKit's own calculator default. Substitute your own; the method matters more than my constant.

How much does TTS actually cost per minute of conversation?

Between about $0.009 and $0.18 per minute of agent speech, depending entirely on which voice you pick. Nobody in this category publishes that number, because they all publish dollars per million characters, which is a unit no one budgets a phone call in.

The conversion is trivial and the most useful thing in this post. Take the price per million characters, multiply by your characters per minute of synthesised speech, divide by a million. LiveKit's pricing calculator assumes 600 characters per minute of agent speech, and I use that throughout so my arithmetic reconciles with theirs. Rime's own page describes $0.03 per 1,000 characters as "~$0.03 per minute of audio," which implies roughly 1,000 characters per minute. Vendors don't even agree with each other on the constant.

That disagreement matters, because it's a 1.7x difference in every cost estimate depending on whose assumption you inherit. Measure your own: count the characters your agent actually synthesises per call from your logs, divide by agent speech minutes, and use that. It takes twenty minutes and makes every number below yours instead of mine.

Then apply the second correction, the one that saves people from over-budgeting. You pay for the agent's speech, not the call. On a four-minute call where the agent talks 40% of the time you're billing about 1.6 TTS-minutes, not 4. Estimates that assume the agent talks the whole call run roughly 2.5x too high, and I've seen that error survive all the way into a board deck.

Provider / modelList price$/min @ 600 chars$/min @ 1,000 charsPer 4-min call @ 40% talkSource
xAI Text to Speech$15 / 1M chars$0.0090$0.0150$0.0144docs.x.ai/developers/pricing
Fish Audio S2.1 Pro$15 / 1M chars$0.0090$0.0150$0.0144LiveKit inference rate card
Deepgram Aura-1$0.0150 / 1K chars$0.0090$0.0150$0.0144deepgram.com/pricing
Inworld Realtime 1.5 Mini$15 / 1M chars$0.0090$0.0150$0.0144LiveKit inference rate card
Deepgram Aura-2$0.030 / 1K chars$0.0180$0.0300$0.0288deepgram.com/pricing
Rime Mist v3$0.03 / 1K chars$0.0180$0.0300$0.0288rime.ai/pricing
Cartesia Sonic 3$50 / 1M chars$0.0300$0.0500$0.0480LiveKit inference rate card
Rime Coda$0.05 / 1K chars$0.0300$0.0500$0.0480rime.ai/pricing
ElevenLabs Flash v2.5$150 / 1M chars$0.0900$0.1500$0.1440LiveKit inference rate card
ElevenLabs Multilingual v2$300 / 1M chars$0.1800$0.3000$0.2880LiveKit inference rate card

Why is TTS the line item that decides your per-minute cost?

Because in LiveKit's own default calculator configuration, TTS at $0.0300 per minute is larger than STT, LLM, agent hosting and telephony combined. That's not my framing; it falls straight out of their published decomposition of a $0.0672 per-minute phone agent: agent session $0.0100, telephony $0.0100, LLM $0.0014, STT $0.0058, TTS $0.0300, observability $0.0100.

So the single highest-leverage cost decision in a voice stack is which voice you use, and it's usually made by whoever liked a demo. Moving from a $50-per-million-character voice to a $15-per-million-character one saves $0.021 per minute of agent speech. At 100,000 call-minutes a month with a 40% talk ratio, that's 40,000 TTS-minutes and about $840 a month, or roughly $10,000 a year, from one configuration line.

You can verify the arithmetic against LiveKit's own worked example on their pricing page. They price GPT-4o plus Nova-2 plus Eleven Multilingual v2 at about $0.2151 per minute, and $0.1800 of that, 84%, is the TTS, because Multilingual v2 lists at $300 per million characters. That example exists to show how the calculator works. It doubles as the clearest possible argument for reading this table before you pick a voice.

It's also the mechanical explanation for why a self-orchestrated stack lands near 2.5¢ a minute while LiveKit Cloud's own default lands at $0.0672. Three decisions do it: self-hosted agent workers remove the $0.0100 session and $0.0100 observability lines, third-party SIP replaces bundled telephony, and TTS selection replaces $0.0300 with $0.0090. Full working in the full LiveKit voice AI build guide.

LiveKit's own default phone-agent minute
$0.0672per minute
  • TTS$0.0300
  • Agent session$0.0100
  • Telephony$0.0100
  • Observability$0.0100
  • STT$0.0058
  • LLM$0.0014
LiveKit's published default calculator configuration for a phone agent. TTS is 45% of the minute on its own, three times STT and more than twenty times the LLM. Every team I've watched optimise a voice pipeline started with the model, which on this chart is the thinnest slice on the wheel.

Which TTS is fastest for voice agents?

Rime Mist v3, on published numbers, at 37ms time-to-first-audio p50 and 56ms p90. Rime is also the only vendor in this comparison publishing TTFA as a proper distribution with a stated measurement condition, and that condition matters enormously.

The condition is "at 1 concurrency." Rime states it plainly on their pricing page, an unusually honest disclosure that most aggregator posts strip out when they repeat the number. Your production concurrency isn't one. Published TTFA is a floor, not a forecast, and the gap widens exactly when you need the headroom.

Rime's other model, Coda, publishes 96ms p50 and 98ms p90: slower on the median but with a remarkably tight distribution, only 2ms between p50 and p90 against Mist v3's 19ms. For a voice agent, a tight tail is often worth more than a fast median, for the same reason p95 beats p50 everywhere else in the voice AI latency budget.

For self-hosting, Rime publishes sub-100ms model latency for Coda and roughly 70ms for Mist v3 on your own hardware, deployable via Docker Compose or Kubernetes. That's the only path in this category to removing network variance from your TTFA entirely, and at high volume it's where the next step-change in cost lives, though it converts a per-minute bill into a capacity-planning problem, a trade rather than a saving.

What actually differentiates a TTS for voice agents
 xAI TTSRime Mist v3Rime CodaCartesia Sonic 3ElevenLabs Flash v2.5
$ per 1M characters$15$30$50$50$150
$/min @ 600 chars$0.0090$0.0180$0.0300$0.0300$0.0900
Published TTFA p50not published37ms96msnot publishednot published
Published TTFA p90not published56ms98msnot publishednot published
Word-level timestampsnot publishednot publishednot published
Self-host optionDocker / KubernetesDocker / Kubernetes
HIPAA BAA availablenot publishedEnterpriseEnterprisenot publishednot published
Entry-tier concurrencysee xAI rate limits2020not publishednot published
Voicesnot published94184not publishednot published
Coda takes the win here, and not on price or median latency. It wins because it's the only column with word-level timestamps, which is what lets you truncate conversation context correctly after an interruption. Note how many cells read "not published": Rime is the only vendor in this table publishing a latency distribution, a timestamp capability and a concurrency limit on a page you can read without a sales call. That transparency is itself a selection criterion.

What actually breaks barge-in, and which providers handle it?

Mid-stream cancellation and word-level timestamps. A provider with better raw latency and no clean cancellation produces a worse-feeling agent than a slower one that stops when told, and this criterion only appears in comparisons written by people who've shipped an agent rather than benchmarked one.

Cancellation is the first requirement. When a caller interrupts, you must stop the in-flight synthesis, not merely stop playing the audio you already have. If you can't, you keep receiving audio for an utterance nobody wants, you keep paying for the characters, and depending on your buffering the caller may keep hearing it.

Word-level timestamps are the second, less obvious requirement. After an interruption you have to record what the caller actually heard, not what the model generated, or the agent refers back to information that was never delivered. Rime publishes this directly: Coda has word-level timestamps, Mist v3 doesn't. That makes it a barge-in correctness feature rather than a formatting nicety, and the reason I wouldn't choose a TTS on TTFA alone.

Both models stream over HTTP and WebSockets, which is table stakes. The thing to test before you commit is trivial and almost nobody does it: interrupt the agent at word five of a forty-word utterance and listen. The full mechanics, including the context-truncation code, are in turn detection and barge-in for voice agents.

swap_tts.py
from livekit.agents import AgentSession
from livekit.plugins import rime, cartesia, deepgram, elevenlabs

# Pick ONE. The rest of the session is identical in every case.
TTS = {
    # $0.0180/min @ 600 chars. 37ms TTFA p50. No word-level timestamps.
    "fast_cheap":  lambda: rime.TTS(model="mistv3"),

    # $0.0300/min. 96ms TTFA p50, 98ms p90 -- tight tail, timestamps: yes.
    "barge_in_safe": lambda: rime.TTS(model="coda"),

    # $0.0180/min @ 600 chars.
    "deepgram":    lambda: deepgram.TTS(model="aura-2"),

    # $0.0300/min.
    "cartesia":    lambda: cartesia.TTS(model="sonic-3"),

    # $0.0900/min -- 10x the cheapest option in this dict. Know why you chose it.
    "premium":     lambda: elevenlabs.TTS(model="eleven_flash_v2_5"),
}

session = AgentSession(
    tts=TTS["barge_in_safe"](),
    # ... stt, llm, vad, turn_handling unchanged
)

# Before committing: interrupt at word 5 of a 40-word utterance and listen.
# If audio continues, cancellation is broken and no benchmark number matters.
The architectural benefit of owning the orchestration layer, in one file: the TTS provider is a single argument. This is what makes the 20x price spread in the hero chart an opportunity rather than a lock-in risk. When a cheaper or better voice ships next quarter, adopting it is an afternoon.

Which TTS should you actually pick?

Ranked by binding constraint, because "it depends" is not an answer and a single winner would be dishonest. One: if cost is binding, xAI Text to Speech at $15 per million characters is the cheapest credible option on a published rate card, at $0.0090 per minute of agent speech. Two: if latency is binding, Rime Mist v3 at 37ms TTFA p50. Three: if barge-in correctness is binding, and for a real conversational agent it should be, Rime Coda, for the word-level timestamps.

Four: if you need HIPAA or on-premise, Rime is the only provider in this comparison publishing both a self-host path (Docker Compose or Kubernetes) and an Enterprise BAA plus SOC 2 Type II report. Five: if voice quality is genuinely the product, a consumer brand, a character, something people choose to listen to, ElevenLabs is the quality benchmark and the 10x premium over the cheapest option is a real product decision, not a mistake. Say so in the budget rather than discovering it in the invoice.

Six: if you're multilingual, check the language list before the price. Rime lists eight production-ready languages on Coda and four on Mist v3, and a cheap voice that doesn't speak your customers' language isn't cheap. Seven: if you don't know yet, pick anything with a LiveKit plugin and make the swap a one-line change, the configuration in the snippet above.

The honest caveat: I haven't independently measured TTFA for these providers, and I won't publish numbers I didn't measure. Everything in the latency rows above is what the vendor publishes, labelled as such. Before you commit, run twenty synthesis requests of a fixed sentence against your two finalists from your own region at your own concurrency and compare. That measurement beats every published table, including this one.

To size the decision in money rather than cents, put your own volume and talk ratio into the voice AI cost calculator; TTS is the input that moves the output most. And if you'd rather have the whole pipeline built with the provider abstraction and cost instrumentation already in place, that's what voice AI development is. Under about 20,000 minutes a month, though, stay on a managed platform and spend the attention elsewhere; the reasoning is in the voice AI build vs buy break-even.

Pick by constraint
Which TTS should you use for your voice agent?
Cost is the binding constraint
xAI Text to Speech

$15/1M chars: $0.0090/min at 600 chars/min, the cheapest credible published rate. Runner-up: Deepgram Aura-1 or Inworld 1.5 Mini at the same list price.

Latency is the binding constraint
Rime Mist v3

37ms TTFA p50, 56ms p90, published at one concurrency. Give up: word-level timestamps, which means give up correct context truncation on barge-in.

Barge-in correctness matters most
Rime Coda

96ms p50 with a 2ms gap to p90, plus word-level timestamps. Costs $0.0300/min against Mist v3's $0.0180. Worth it for genuinely conversational agents.

HIPAA, on-prem or data residency
Rime Enterprise

Self-host via Docker Compose or Kubernetes, BAA and SOC 2 Type II on Enterprise, sub-100ms self-hosted model latency for Coda and ~70ms for Mist.

Voice quality is the product
ElevenLabs

$0.0900/min for Flash v2.5 at 600 chars/min, 10x the cheapest option here. A legitimate choice for consumer brands, and a budgeting disaster if chosen by accident.

You genuinely do not know yet
Anything with a LiveKit plugin

Make the provider a one-line change, ship, then measure TTFA and cost on your own traffic and decide with data instead of with a demo.

Six constraints, five different answers, and the cheapest option wins only one. If a comparison table has the same provider winning every row, close the tab; that's a sales page with a border around it.

Best TTS for voice agents: common questions

What is the best TTS for a voice AI agent?

It depends on which constraint binds. For lowest cost, xAI Text to Speech at $15 per million characters, about $0.0090 per minute of agent speech at 600 characters per minute. For lowest latency, Rime Mist v3 at a published 37ms time-to-first-audio p50. For correct barge-in behaviour, Rime Coda, because it publishes word-level timestamps and Mist v3 doesn't. For voice quality as a product feature, ElevenLabs at roughly 10x the cheapest option.

How much does text-to-speech cost per minute?

Between roughly $0.009 and $0.18 per minute of agent speech in 2026, depending on the model. Convert any per-character price by multiplying by your characters per minute of synthesised speech and dividing by a million. LiveKit's pricing calculator assumes 600 characters per minute; Rime's own page implies closer to 1,000. Measure your own from logs, because the difference between those two constants is 1.7x on every estimate.

Which TTS has the lowest latency?

On published figures, Rime Mist v3 at 37ms time-to-first-audio p50 and 56ms p90, with Rime Coda at 96ms p50 and 98ms p90. Rime states these are measured at one concurrent stream, an unusually honest disclosure and an important one: production concurrency is never one, so treat published TTFA as a floor rather than a forecast.

Does TTS choice affect barge-in?

Yes, decisively. Barge-in requires cancelling an in-flight synthesis mid-stream, and correctly recovering afterwards requires word-level timestamps so you can truncate conversation context to the words the caller actually heard. Without timestamps the agent records the full generated response as spoken and refers back to information the caller never received. Rime publishes that Coda supports word-level timestamps and Mist v3 doesn't.

Can you self-host TTS for a voice agent?

Rime is the provider in this comparison publishing a self-host path, via Docker Compose or Kubernetes, with stated model latency of sub-100ms for Coda and roughly 70ms for Mist. Self-hosting removes network variance from time-to-first-audio and converts a per-minute bill into a capacity-planning problem, a real trade rather than an automatic saving: worth it at high, steady volume and rarely worth it below that.

Why is TTS more expensive than STT for voice agents?

Because synthesis is priced per character generated while streaming transcription is priced per minute of audio, and the per-minute equivalents are far apart. In LiveKit's own default calculator configuration TTS is $0.0300 per minute against $0.0058 for STT, roughly five times more, and streaming STT has fallen further since, with xAI listing speech-to-text at $0.20 per hour, which is $0.0033 per minute.

Ready to talk numbers?

Twenty minutes, straight to the engineer. No sales rep, no deck.