Request a callbackBook a call
← All posts

How to Build Your Own Voice AI Platform on LiveKit: Architecture, Costs and a Build Plan

TL;DR
  • Build your own voice AI platform when you sell voice agents to many customers, run past about 20,000 minutes a month, or must keep recordings and prompts in your own cloud. Fifty customers at 6,000 minutes each is 300,000 minutes: $30,000 a month at a managed platform's 10 cents, about $7,500 at 2.5 cents.
  • On LiveKit the platform is a SIP trunk, a dispatch rule, one room and one agent job per call, and a worker that loads the tenant's versioned settings from job metadata. Turn detection, tools, knowledge, recordings and metering sit around that loop.
  • My estimate is about ten engineer-weeks: six for the voice pipeline and four for tenants, the dashboard and metering. Load-test at 1.5 times your forecast peak before any tenant goes live.
Reference architecture
A multi-tenant voice AI platform on LiveKitcallSIP INVITEcreate roomdispatch with metadataload by tenantaudio, textretrieve, actrecordevery turnper calledit, version
Callersphone or browser
SIP trunkTelnyx or Twilio
LiveKit SIPinbound trunk, dispatch rule
Room per callLiveKit media server
Platform dashboardtenants, agents, calls
Tenant configprompt, voice, tools, versioned
Agent workerone job per call
STT, LLM, TTSplugins you can swap
Usage meteringminutes and cost per tenant
Call recordsPostgres: turns, tools, cost
Knowledge and toolspgvector, MCP, tenant APIs
RecordingsEgress to your bucket
Violet boxes are the platform you own; teal boxes are infrastructure and vendors you can swap. The tenant is settled before the first word: the dispatch rule's metadata or the dialled number tells the worker whose agent to load.

Who needs their own voice AI platform instead of Retell or Vapi?

Three kinds of company: software businesses and agencies that sell voice agents to many customers, operators whose minutes run past about 20,000 a month, and teams that must keep prompts, recordings and vendor choices in their own cloud. Everyone else should rent a managed platform and spend the engineering on their product.

The first group is where a platform pays fastest, because volume arrives with every customer. A company selling a booking agent to 50 clinics at 6,000 minutes each runs 300,000 minutes a month: $30,000 at a managed platform's 10 cents, about $7,500 at 2.5 cents on its own stack. That $22,500 gap is the gross margin of the business; on a rented platform you pay it out every month.

The other two groups build for control as much as cost. Owning the pipeline means choosing turn-taking behaviour per use case, adopting a cheaper or better voice the week it ships, and keeping recordings in a bucket you control. I built this stack in production for an AI interview platform serving 20+ enterprise clients after first shipping on Retell, and moved from about 10 cents a minute to about 2.5.

Who should not build: a single business with one use case under 20,000 minutes a month, a team with nobody who can own a voice pager, or anyone who needs a signed compliance agreement this quarter. A platform ships in days. The product version of what follows, with a white-label option for resellers, is the AI voice agent build.

Rent or build?
Should you build your own voice AI platform?
One business, one use case, under 20,000 minutes a month
Rent a managed platform

The fixed costs of a custom stack exceed the per-minute saving. Revisit when the invoice grows.

Selling voice agents to many customers
Build early

Per-call margin is the business model, and every new customer adds minutes.

Over about 20,000 minutes a month, two engineers to own it
Build

At 100,000 minutes a month the saving is about $5,700 against a 10-cent platform.

Recordings and prompts must stay in your cloud
Build, or rent from a vendor that signs

Self-hosted LiveKit keeps media in your own account; a platform needs the right agreement.

A compliance agreement needed inside 90 days
Rent from a vendor that signs

A custom chain needs an agreement from every vendor in it, which takes longer than the engineering.

Volume, resale and data control favour building. Speed and paperwork favour renting.

What does a voice AI platform actually do for its customers?

It lets each customer run their own agents on shared infrastructure: numbers with inbound and outbound calling, versioned agent settings (prompt, voice, language, turn-taking), knowledge and tools, transfers to people, recordings and transcripts, analytics, and usage billing per customer. The voice pipeline is one component; the platform is everything that makes it safe to share.

Read a managed platform's price list as a requirements document. Retell sells, as separate lines, telephony, a knowledge base, denoising, safety guardrails, PII removal, AI quality assurance, batch calling, branded caller ID and verified numbers (Retell pricing). Each is a feature your customers will ask for, and each is a line you either build, buy from another vendor, or decide to skip.

Three users define the product. The caller wants to interrupt, change their mind and finish the task in one call. The customer's operations lead wants to change a prompt, test it with a call and see why each call ended as it did. Your own engineers want every vendor replaceable and every call replayable. A platform that serves only the first user is a demo.

The platform, from the call up
Telephony

Numbers per tenant, SIP trunks, inbound and outbound calls, transfers and answering machine detection.

carrier
Media

One LiveKit room per call, bridged from SIP, or joined over WebRTC for browser callers.

LiveKit server
Agent runtime

Workers that load the tenant's settings and run voice activity detection, turn detection, transcription, the model and the voice.

one job per call
Knowledge and tools

Retrieval over each tenant's documents and typed tools into each tenant's systems.

tenant-scoped
Call records

Recording, transcript, tool calls, latency and cost for every call, searchable by outcome.

replayable
Control plane

Tenants, agent versions, numbers, keys and roles, edited in a dashboard and tested with a call.

versioned
Metering and billing

Minutes and cost per tenant, per agent and per call, exported to invoices.

your margin
The first five layers exist in any production voice agent. The last two turn one agent into a platform, and they are most of the product work.

What does the LiveKit architecture look like, component by component?

A call arrives on a SIP trunk, LiveKit's SIP service matches it to a dispatch rule and creates a room, and an agent worker is dispatched into that room with metadata naming the tenant. The worker runs the pipeline with plugins for transcription, the model and the voice, and writes every turn to your database.

Both the LiveKit server and the Agents framework are open source under Apache 2.0 (LiveKit server, LiveKit Agents), so you can run the whole stack in your own cloud or use LiveKit Cloud for media and keep only the workers. LiveKit handles transport, rooms and telephony ingress, and deliberately does not choose your transcription, model or voice. That separation is the opportunity: you pay each vendor directly and can replace any of them.

The worker is where your code lives. Each call becomes a job in its own process, so one failed call leaves the others alone (LiveKit deployments). LiveKit sizes a 4-core, 8 GB server at 10 to 25 concurrent calls, and its own test ran 30 agents on one at about 3.8 cores and 2.8 GB. Workers stop accepting jobs at a load threshold of 0.7 by default, so autoscale at 0.5 and give draining workers ten minutes or more to finish their calls.

ComponentWhat it doesMy pickCost signal
SIP trunkNumbers and the phone minuteTelnyx; Twilio for toll-free-heavy lines$0.0032 a minute inbound local on Telnyx
LiveKit SIP and roomsBridges each call into a room and dispatches an agentLiveKit Cloud first, self-hosted at high volume$0.004 a minute SIP fee on Ship; none self-hosted
Agent workersRun the pipeline, one process per callLiveKit Agents on compute-optimised instances10 to 25 calls per 4-core, 8 GB server
Speech to textStreaming transcriptionxAI or Deepgram$0.0033 to $0.0065 a minute
Language modelReplies and tool callsA Flash-Lite model on live turns, a larger one after the call$0.0004 to $0.0072 a minute
Text to speechThe agent's voicexAI for cost, a premium voice where the brand needs it$0.009 to $0.030 a minute at 600 characters
Turn detectionDecides when the caller has finishedLiveKit's audio turn detectorFree: v1 on LiveKit Cloud, v1-mini on your own CPU
Knowledge baseTenant documents for grounded answersPostgres with pgvectorYour database; no per-minute fee
RecordingsCall audio in your own storageLiveKit Egress to your bucket, or the carrier's recording$0.002 a minute on Telnyx
Control planeTenants, agent versions, numbers, usageA web dashboard on the same PostgresEngineering time, not minutes

How does each call reach the right tenant's agent?

Through the number that was dialled. Each tenant owns numbers; the SIP dispatch rule creates a room for the call and dispatches your agent with metadata; and the worker resolves the tenant and agent version from that metadata or the dialled number before the first word. Every tool, document and recording is then scoped to that tenant.

LiveKit gives you the pieces. Dispatch rules come in three types: individual, which creates a room per caller; direct, which sends callers to one room; and callee, which names the room after the number dialled (dispatch rules). A rule can be limited to specific trunks, can dispatch named agents through its room configuration, and passes its metadata and attributes to the SIP participant. Explicit dispatch by agent name is the recommended mode, and job metadata can be up to 512 KiB (agent dispatch). Carrier setup is in LiveKit SIP trunking with Twilio vs Telnyx.

Then enforce the boundary in code, not in the prompt. Tools read the tenant and the verified customer from the job's context, never from the model's arguments, so a caller cannot talk the agent into another tenant's data. Knowledge queries filter by tenant inside the database. Recordings land under a per-tenant prefix, ideally with a per-tenant encryption key, so deleting a customer is one operation.

Noisy neighbours are the multi-tenant failure specific to voice. One tenant's outbound campaign can take every speech stream your vendor tier allows and leave other tenants' callers in silence. Give each tenant a concurrency quota at admission, meter minutes and cost per call from the first day, and store agent settings as versions so a bad prompt edit rolls back in one click.

Tenant routing
One inbound call, tenant resolved before the first wordCallerSIP trunkLiveKit SIPAgent workerConfig storeCall records
dials a number owned by tenant A
SIP INVITE
match dispatch rule, create a room
individual rule: one room per caller
dispatch agent_name with metadata {tenant, agent_version}
load settings for tenant A, version 14
prompt, voice, tools, knowledge
settings
greeting in tenant A's voice
turn log, tool calls, cost, tenant id
every turn
The tenant is fixed from dispatch metadata before the greeting plays. Nothing the caller says afterwards can change it, and every record carries it.

How do turn detection, tools and the knowledge base fit into one turn?

In one streaming loop. Voice activity detection hears speech, the turn detector decides the caller has finished, retrieval runs on that finished turn, the model answers or calls a tool, and the voice starts on the first clause. Every stage overlaps the next, which is how a reply fits inside an 800-millisecond budget.

Turn detection decides whether the agent feels human. LiveKit's audio turn detector listens to intonation and rhythm rather than a transcript. Its v1 model runs on LiveKit Inference at no cost for agents deployed on LiveKit Cloud, and v1-mini runs on your own CPU, free in any context. It covers 14 languages, commits turns with a default minimum delay of 0.3 seconds and a maximum of 2.5, and needs the voice activity detector's silence setting at 0.25 seconds or more (turn detector). An interview needs longer pauses than a booking line; the tuning is in turn detection and barge-in.

Tools are typed functions with timeouts. LiveKit's framework can expose tools from MCP servers (Python only) and run long tools in the background while the agent keeps talking (tools). Writes need a read-back to the caller and an idempotency key built from the call and the arguments. If a tenant's systems already sit behind MCP servers, the gateway pattern in MCP in production applies unchanged.

Knowledge lookups run in the on_user_turn_completed hook, after the caller finishes and before the model answers, so retrieval adds no tool round trip (external data). Give it about 150 milliseconds at the 95th percentile, filter by tenant, take the top four to six passages, and have the agent say it is checking when a lookup runs long.

One conversational turn
  1. 1
    Caller speaks0 ms

    Audio arrives through SIP or WebRTC in the call's room, where the tenant's agent is already a participant.

  2. 2
    Streaming transcriptionpartial results

    Audio frames stream to speech-to-text while the caller talks. Nothing waits for silence.

  3. 3
    Turn detection commits0.3 s minimum by default

    The audio turn detector judges whether the thought is finished, not only whether the caller paused.

  4. 4
    Retrieval, then the modelfirst token matters

    Tenant-filtered passages are added in on_user_turn_completed, then the model streams a reply or calls a tool.

  5. 5
    Voice streams backfirst clause

    Synthesis starts on the first clause and plays into the room while the model is still writing.

  6. 6
    Barge-incancel everything

    If the caller speaks, stop playback, cancel synthesis, and keep only the words the caller heard in the history.

A pipeline that waits for each stage to finish before starting the next is slower by the sum of every wait. Streaming every stage is the cheapest latency win available.

What should recordings and observability capture on every call?

Enough to replay any call: the audio, a transcript with turn boundaries, every tool call and result, five latency timestamps per turn, the agent version, the tenant and the cost. Alert on the 95th percentile per stage and on rates, never on single calls. Without this, a customer complaint becomes an argument instead of a lookup.

Recording is cheap; consent is the work. Telnyx lists call recording at $0.002 a minute with free storage (Telnyx). LiveKit Egress writes session audio to your own storage, and LiveKit Cloud prices agent session recordings at $0.005 a minute after the included minutes. Put the recording notice in each tenant's greeting, because consent rules for recording differ by state and by country.

For observability, LiveKit Cloud's Agent Insights shows transcripts, traces, logs and recordings on one timeline for each session (Agent Insights). If you self-host, emit the same thing yourself: one structured line per turn with the call ID, tenant, agent version, five timestamps (speech end, turn committed, final transcript, first token, first audio) and token and character counts, so cost per call is a query, not an estimate.

Then give each tenant a view of their own calls: resolution and transfer rates, cost per call, and the calls that failed, filtered by intent. It is the screen your customers will open every day, so build it early rather than last.

The call record
What every call record needs
  • Tenant, agent version and call ID on every rowThe keys every other query hangs from
  • Recording under the tenant's prefix, with the consent line logged
  • Transcript with turn boundaries and interruptions markedKeep the words the caller heard, not the words generated
  • Every tool call with its arguments, result and latency
  • Five timestamps per turnSpeech end, turn committed, final transcript, first token, first audio
  • Characters synthesised and tokens used, priced at the day's rate cardCost per call becomes a query
  • An outcome set by code: resolved, transferred, abandoned or failedNot inferred from the transcript later
Seven fields. The last is the one teams skip, and it is the first number a customer's operations lead looks for.

What does a voice AI platform on LiveKit cost to run and to build?

About 2.5 cents a minute to run in my production stack, plus fixed infrastructure and maintenance I model at about $1,754 a month. To build, my estimate is about ten engineer-weeks: six for the voice pipeline and four for tenants, the dashboard and metering. Past about 20,000 minutes a month the saving on minutes pays for it.

The per-minute line is broken down in AI voice agent cost per minute: about $0.0199 of published list prices for the carrier, LiveKit's SIP fee, transcription, a Flash-Lite model and the voice, plus about $0.0051 of workers, Redis, logs and observability. Plans matter at the edges: LiveKit Cloud's Ship plan starts at $50 a month, and Scale at $500 with a SOC 2 Type II report and a signed BAA (LiveKit pricing).

The multi-tenant layer changes the fixed side more than the variable side. You need a warm floor of workers so a tenant's first call after a quiet hour is not a cold start, a staging copy for testing agent versions, and a database of call records that grows with every minute. None of that is large next to the minutes, but none of it goes away when traffic does.

Sell minutes above your cost and the difference is the business. Charge customers 8 cents a minute, at the low end of the managed platforms' published rates, run at 2.5, and a 300,000-minute month grosses $16,500 before support and sales. Model your own pricing, volume and build in the AI product cost estimator before you commit.

300,000 minutes a month, three ways
$ per month at 300,000 minuteslower is better
Retell default with its telephony, 12.5 cents$37,500
Managed platform at 10 cents$30,000
Own LiveKit stack at 2.5 cents plus $1,754 fixed$9,254
At this volume the fixed line barely registers. Every new customer adds minutes at 2.5 cents, not at a platform's rate, which is the whole case for owning the platform.

What breaks in production that never breaks in a voice AI demo?

Turn-taking with real callers, barge-in that does not cancel synthesis, vendor concurrency caps at peak, tail latency on the model, missing call records, cold workers, and one tenant starving the rest. Demos are single calls on good networks with cooperative speakers; production is none of those things.

Concurrency is the ceiling people meet first, and it is rarely their servers. Deepgram publishes 150 streaming transcription connections and 45 text-to-speech connections on pay-as-you-go (Deepgram pricing), and LiveKit's inference service allows 5 concurrent connections on Build, 20 on Ship and 50 on Scale. Hitting a cap does not throw a clean error: the call connects and the caller hears silence. For sizing at thousands of calls a day, see the AI call center build.

Most of the rest is fixed by habits rather than clever code: keep a warm worker per region above peak, write call state after every turn so a crashed worker loses one call, deploy outside the busy hour with drain periods longer than your longest calls, and load-test at 1.5 times forecast peak before a tenant launches.

Failure modeWhat the caller experiencesWhere to fix it
Silence-based endpointingAgent interrupts mid-sentenceAudio turn detector, with delays tuned per use case
Barge-in that does not cancel synthesisAgent talks over the callerStop playback and cancel in-flight synthesis on speech start
Vendor concurrency cap reachedCall connects to silence at peakKnow every cap; per-tenant quotas at admission
Tail latency on the modelLong pauses; the caller says hello?Alert on p95 per stage; fail over to a faster model on timeout
No per-call recordA complaint nobody can reproduceOne structured row per turn with call ID and tenant from day one
Cold workers after a quiet hourFirst call is slow to answerA warm floor of workers; prewarm prompts and indexes
One tenant's campaign starves the othersOther tenants' callers hear silencePer-tenant concurrency quotas and a separate outbound pool
Deploy restarts a busy workerCall drops mid-conversationDrain on SIGTERM with a grace period above your longest call

What would I build first, and in what order?

One tenant, one number, one call path, instrumented end to end on the cheapest components that work. Then turn detection and barge-in, then tools and knowledge, then recordings and evals, and only then the tenant model, dashboard and metering. Every step ships something a real caller can use, and nothing is tuned before it is measured.

Make the pipeline swappable from the first week. The most valuable property of your own stack is not today's cost; it is that when a better or cheaper voice ships next quarter, you adopt it with a configuration change and a test on a share of live calls. LiveKit's Agents framework includes a test framework with model judges, which makes that change reviewable rather than a leap of faith.

Leave multi-tenancy until the single-tenant path is boring, but design for it from the start: tenant and agent version on every record, tools that take the tenant from context, and settings stored as versions. Retrofitting those three is the expensive part of turning a voice agent into a platform.

If you want it built rather than planned, that is voice AI development, with voice agent builds from $12,000. It is built at $0: the work is split into checkpoints with acceptance criteria agreed before work starts, and each is invoiced only after you have seen it and accepted it.

About ten engineer-weeks, in order (my estimate)
  1. Weeks 1 to 2
    One call path, instrumented

    One tenant, one number, a SIP trunk, a dispatch rule and a worker with transcription, model and voice, logging five timestamps per turn from the first call.

  2. Weeks 3 to 4
    Turn-taking and barge-in

    Audio turn detector tuned for the use case, cancellation of in-flight synthesis, and history truncated to the words the caller heard.

  3. Week 5
    Tools and knowledge

    Typed tools with timeouts and read-back, retrieval in on_user_turn_completed, and a grounding check before answers are spoken.

  4. Week 6
    Recordings, evals and cutover

    Egress to your bucket, a regression set from real calls, and a staged ramp with a one-value rollback.

  5. Weeks 7 to 8
    Tenants and versions

    Tenant model, numbers and dispatch per tenant, versioned agent settings, and per-tenant concurrency quotas.

  6. Weeks 9 to 10
    Dashboard and metering

    Agent editor, call logs, analytics, minutes and cost per tenant, and a load test at 1.5 times forecast peak.

The first six weeks cover the same ground as the six-week pipeline estimate in the build vs buy analysis; the last four are the platform. A demo on a real number fits in the first two.

Building a voice AI platform on LiveKit: common questions

→What do you need to build your own voice AI platform?

A SIP trunk and numbers, LiveKit for media, SIP and agent dispatch, workers running streaming transcription, a language model and a voice, a turn detector, tools and a knowledge base per tenant, recordings and per-turn logs, and a control plane for tenants, agent versions and usage. My estimate is about six engineer-weeks for the pipeline and four for the platform layer.

→Is LiveKit free to use?

The LiveKit server and the Agents framework are open source under Apache 2.0, so you can self-host both and pay only for servers. LiveKit Cloud has a free Build plan, Ship from $50 a month and Scale from $500, with a third-party SIP fee of $0.004 or $0.003 a minute and agent sessions at $0.01 a minute for agents it hosts.

→How do you make a voice AI platform multi-tenant?

Resolve the tenant before the first word, from the dialled number or the dispatch rule's metadata, and load that tenant's versioned agent settings. Tools take the tenant from the job context, never from the model; knowledge queries filter by tenant in the database; recordings go under a per-tenant prefix; and each tenant gets a concurrency quota so one campaign cannot starve the rest.

→How fast does a voice agent need to respond?

Budget 800 milliseconds from the end of the caller's speech to the first audio of the reply, and judge it on the 95th percentile rather than the median. With LiveKit's audio turn detector the default minimum endpointing delay is 0.3 seconds, which leaves the rest of the budget for transcription, the model's first token and the voice's first audio.

→Should I build my own voice AI platform or use Retell or Vapi?

Use a managed platform for one business under about 20,000 minutes a month; Retell's default is $0.11 a minute and Vapi charges $0.05 before models and voices. Build when you resell agents to many customers, pass that volume, or must keep recordings in your own cloud. Fifty customers at 6,000 minutes each is $30,000 a month at 10 cents and about $7,500 at 2.5.

→What breaks first when a voice AI platform goes live?

Usually a vendor concurrency cap. Deepgram publishes 45 text-to-speech connections on pay-as-you-go and LiveKit's inference service allows 50 on Scale, so a busy hour can pass them long before your servers are full. The call connects and the caller hears silence, so list every cap, raise the smallest first, and load-test at 1.5 times peak.

Take this into your own chat

Open the article in your assistant with one click and ask it how this applies to your product.

Keep reading

See it in production: voice AI at 2.5¢ a minute, the case study

Ready to talk numbers?

Twenty minutes, straight to the engineer. No sales rep, no deck.