How to Build Your Own Voice AI Platform on LiveKit: Architecture, Costs and a Build Plan
- Build your own voice AI platform when you sell voice agents to many customers, run past about 20,000 minutes a month, or must keep recordings and prompts in your own cloud. Fifty customers at 6,000 minutes each is 300,000 minutes: $30,000 a month at a managed platform's 10 cents, about $7,500 at 2.5 cents.
- On LiveKit the platform is a SIP trunk, a dispatch rule, one room and one agent job per call, and a worker that loads the tenant's versioned settings from job metadata. Turn detection, tools, knowledge, recordings and metering sit around that loop.
- My estimate is about ten engineer-weeks: six for the voice pipeline and four for tenants, the dashboard and metering. Load-test at 1.5 times your forecast peak before any tenant goes live.
Who needs their own voice AI platform instead of Retell or Vapi?
Three kinds of company: software businesses and agencies that sell voice agents to many customers, operators whose minutes run past about 20,000 a month, and teams that must keep prompts, recordings and vendor choices in their own cloud. Everyone else should rent a managed platform and spend the engineering on their product.
The first group is where a platform pays fastest, because volume arrives with every customer. A company selling a booking agent to 50 clinics at 6,000 minutes each runs 300,000 minutes a month: $30,000 at a managed platform's 10 cents, about $7,500 at 2.5 cents on its own stack. That $22,500 gap is the gross margin of the business; on a rented platform you pay it out every month.
The other two groups build for control as much as cost. Owning the pipeline means choosing turn-taking behaviour per use case, adopting a cheaper or better voice the week it ships, and keeping recordings in a bucket you control. I built this stack in production for an AI interview platform serving 20+ enterprise clients after first shipping on Retell, and moved from about 10 cents a minute to about 2.5.
Who should not build: a single business with one use case under 20,000 minutes a month, a team with nobody who can own a voice pager, or anyone who needs a signed compliance agreement this quarter. A platform ships in days. The product version of what follows, with a white-label option for resellers, is the AI voice agent build.
The fixed costs of a custom stack exceed the per-minute saving. Revisit when the invoice grows.
Per-call margin is the business model, and every new customer adds minutes.
At 100,000 minutes a month the saving is about $5,700 against a 10-cent platform.
Self-hosted LiveKit keeps media in your own account; a platform needs the right agreement.
A custom chain needs an agreement from every vendor in it, which takes longer than the engineering.
What does a voice AI platform actually do for its customers?
It lets each customer run their own agents on shared infrastructure: numbers with inbound and outbound calling, versioned agent settings (prompt, voice, language, turn-taking), knowledge and tools, transfers to people, recordings and transcripts, analytics, and usage billing per customer. The voice pipeline is one component; the platform is everything that makes it safe to share.
Read a managed platform's price list as a requirements document. Retell sells, as separate lines, telephony, a knowledge base, denoising, safety guardrails, PII removal, AI quality assurance, batch calling, branded caller ID and verified numbers (Retell pricing). Each is a feature your customers will ask for, and each is a line you either build, buy from another vendor, or decide to skip.
Three users define the product. The caller wants to interrupt, change their mind and finish the task in one call. The customer's operations lead wants to change a prompt, test it with a call and see why each call ended as it did. Your own engineers want every vendor replaceable and every call replayable. A platform that serves only the first user is a demo.
Numbers per tenant, SIP trunks, inbound and outbound calls, transfers and answering machine detection.
carrierOne LiveKit room per call, bridged from SIP, or joined over WebRTC for browser callers.
LiveKit serverWorkers that load the tenant's settings and run voice activity detection, turn detection, transcription, the model and the voice.
one job per callRetrieval over each tenant's documents and typed tools into each tenant's systems.
tenant-scopedRecording, transcript, tool calls, latency and cost for every call, searchable by outcome.
replayableTenants, agent versions, numbers, keys and roles, edited in a dashboard and tested with a call.
versionedMinutes and cost per tenant, per agent and per call, exported to invoices.
your marginWhat does the LiveKit architecture look like, component by component?
A call arrives on a SIP trunk, LiveKit's SIP service matches it to a dispatch rule and creates a room, and an agent worker is dispatched into that room with metadata naming the tenant. The worker runs the pipeline with plugins for transcription, the model and the voice, and writes every turn to your database.
Both the LiveKit server and the Agents framework are open source under Apache 2.0 (LiveKit server, LiveKit Agents), so you can run the whole stack in your own cloud or use LiveKit Cloud for media and keep only the workers. LiveKit handles transport, rooms and telephony ingress, and deliberately does not choose your transcription, model or voice. That separation is the opportunity: you pay each vendor directly and can replace any of them.
The worker is where your code lives. Each call becomes a job in its own process, so one failed call leaves the others alone (LiveKit deployments). LiveKit sizes a 4-core, 8 GB server at 10 to 25 concurrent calls, and its own test ran 30 agents on one at about 3.8 cores and 2.8 GB. Workers stop accepting jobs at a load threshold of 0.7 by default, so autoscale at 0.5 and give draining workers ten minutes or more to finish their calls.
| Component | What it does | My pick | Cost signal |
|---|---|---|---|
| SIP trunk | Numbers and the phone minute | Telnyx; Twilio for toll-free-heavy lines | $0.0032 a minute inbound local on Telnyx |
| LiveKit SIP and rooms | Bridges each call into a room and dispatches an agent | LiveKit Cloud first, self-hosted at high volume | $0.004 a minute SIP fee on Ship; none self-hosted |
| Agent workers | Run the pipeline, one process per call | LiveKit Agents on compute-optimised instances | 10 to 25 calls per 4-core, 8 GB server |
| Speech to text | Streaming transcription | xAI or Deepgram | $0.0033 to $0.0065 a minute |
| Language model | Replies and tool calls | A Flash-Lite model on live turns, a larger one after the call | $0.0004 to $0.0072 a minute |
| Text to speech | The agent's voice | xAI for cost, a premium voice where the brand needs it | $0.009 to $0.030 a minute at 600 characters |
| Turn detection | Decides when the caller has finished | LiveKit's audio turn detector | Free: v1 on LiveKit Cloud, v1-mini on your own CPU |
| Knowledge base | Tenant documents for grounded answers | Postgres with pgvector | Your database; no per-minute fee |
| Recordings | Call audio in your own storage | LiveKit Egress to your bucket, or the carrier's recording | $0.002 a minute on Telnyx |
| Control plane | Tenants, agent versions, numbers, usage | A web dashboard on the same Postgres | Engineering time, not minutes |
How does each call reach the right tenant's agent?
Through the number that was dialled. Each tenant owns numbers; the SIP dispatch rule creates a room for the call and dispatches your agent with metadata; and the worker resolves the tenant and agent version from that metadata or the dialled number before the first word. Every tool, document and recording is then scoped to that tenant.
LiveKit gives you the pieces. Dispatch rules come in three types: individual, which creates a room per caller; direct, which sends callers to one room; and callee, which names the room after the number dialled (dispatch rules). A rule can be limited to specific trunks, can dispatch named agents through its room configuration, and passes its metadata and attributes to the SIP participant. Explicit dispatch by agent name is the recommended mode, and job metadata can be up to 512 KiB (agent dispatch). Carrier setup is in LiveKit SIP trunking with Twilio vs Telnyx.
Then enforce the boundary in code, not in the prompt. Tools read the tenant and the verified customer from the job's context, never from the model's arguments, so a caller cannot talk the agent into another tenant's data. Knowledge queries filter by tenant inside the database. Recordings land under a per-tenant prefix, ideally with a per-tenant encryption key, so deleting a customer is one operation.
Noisy neighbours are the multi-tenant failure specific to voice. One tenant's outbound campaign can take every speech stream your vendor tier allows and leave other tenants' callers in silence. Give each tenant a concurrency quota at admission, meter minutes and cost per call from the first day, and store agent settings as versions so a bad prompt edit rolls back in one click.
How do turn detection, tools and the knowledge base fit into one turn?
In one streaming loop. Voice activity detection hears speech, the turn detector decides the caller has finished, retrieval runs on that finished turn, the model answers or calls a tool, and the voice starts on the first clause. Every stage overlaps the next, which is how a reply fits inside an 800-millisecond budget.
Turn detection decides whether the agent feels human. LiveKit's audio turn detector listens to intonation and rhythm rather than a transcript. Its v1 model runs on LiveKit Inference at no cost for agents deployed on LiveKit Cloud, and v1-mini runs on your own CPU, free in any context. It covers 14 languages, commits turns with a default minimum delay of 0.3 seconds and a maximum of 2.5, and needs the voice activity detector's silence setting at 0.25 seconds or more (turn detector). An interview needs longer pauses than a booking line; the tuning is in turn detection and barge-in.
Tools are typed functions with timeouts. LiveKit's framework can expose tools from MCP servers (Python only) and run long tools in the background while the agent keeps talking (tools). Writes need a read-back to the caller and an idempotency key built from the call and the arguments. If a tenant's systems already sit behind MCP servers, the gateway pattern in MCP in production applies unchanged.
Knowledge lookups run in the on_user_turn_completed hook, after the caller finishes and before the model answers, so retrieval adds no tool round trip (external data). Give it about 150 milliseconds at the 95th percentile, filter by tenant, take the top four to six passages, and have the agent say it is checking when a lookup runs long.
- 1Caller speaks0 ms
Audio arrives through SIP or WebRTC in the call's room, where the tenant's agent is already a participant.
- 2Streaming transcriptionpartial results
Audio frames stream to speech-to-text while the caller talks. Nothing waits for silence.
- 3Turn detection commits0.3 s minimum by default
The audio turn detector judges whether the thought is finished, not only whether the caller paused.
- 4Retrieval, then the modelfirst token matters
Tenant-filtered passages are added in on_user_turn_completed, then the model streams a reply or calls a tool.
- 5Voice streams backfirst clause
Synthesis starts on the first clause and plays into the room while the model is still writing.
- 6Barge-incancel everything
If the caller speaks, stop playback, cancel synthesis, and keep only the words the caller heard in the history.
What should recordings and observability capture on every call?
Enough to replay any call: the audio, a transcript with turn boundaries, every tool call and result, five latency timestamps per turn, the agent version, the tenant and the cost. Alert on the 95th percentile per stage and on rates, never on single calls. Without this, a customer complaint becomes an argument instead of a lookup.
Recording is cheap; consent is the work. Telnyx lists call recording at $0.002 a minute with free storage (Telnyx). LiveKit Egress writes session audio to your own storage, and LiveKit Cloud prices agent session recordings at $0.005 a minute after the included minutes. Put the recording notice in each tenant's greeting, because consent rules for recording differ by state and by country.
For observability, LiveKit Cloud's Agent Insights shows transcripts, traces, logs and recordings on one timeline for each session (Agent Insights). If you self-host, emit the same thing yourself: one structured line per turn with the call ID, tenant, agent version, five timestamps (speech end, turn committed, final transcript, first token, first audio) and token and character counts, so cost per call is a query, not an estimate.
Then give each tenant a view of their own calls: resolution and transfer rates, cost per call, and the calls that failed, filtered by intent. It is the screen your customers will open every day, so build it early rather than last.
- Tenant, agent version and call ID on every rowThe keys every other query hangs from
- Recording under the tenant's prefix, with the consent line logged
- Transcript with turn boundaries and interruptions markedKeep the words the caller heard, not the words generated
- Every tool call with its arguments, result and latency
- Five timestamps per turnSpeech end, turn committed, final transcript, first token, first audio
- Characters synthesised and tokens used, priced at the day's rate cardCost per call becomes a query
- An outcome set by code: resolved, transferred, abandoned or failedNot inferred from the transcript later
What does a voice AI platform on LiveKit cost to run and to build?
About 2.5 cents a minute to run in my production stack, plus fixed infrastructure and maintenance I model at about $1,754 a month. To build, my estimate is about ten engineer-weeks: six for the voice pipeline and four for tenants, the dashboard and metering. Past about 20,000 minutes a month the saving on minutes pays for it.
The per-minute line is broken down in AI voice agent cost per minute: about $0.0199 of published list prices for the carrier, LiveKit's SIP fee, transcription, a Flash-Lite model and the voice, plus about $0.0051 of workers, Redis, logs and observability. Plans matter at the edges: LiveKit Cloud's Ship plan starts at $50 a month, and Scale at $500 with a SOC 2 Type II report and a signed BAA (LiveKit pricing).
The multi-tenant layer changes the fixed side more than the variable side. You need a warm floor of workers so a tenant's first call after a quiet hour is not a cold start, a staging copy for testing agent versions, and a database of call records that grows with every minute. None of that is large next to the minutes, but none of it goes away when traffic does.
Sell minutes above your cost and the difference is the business. Charge customers 8 cents a minute, at the low end of the managed platforms' published rates, run at 2.5, and a 300,000-minute month grosses $16,500 before support and sales. Model your own pricing, volume and build in the AI product cost estimator before you commit.
What breaks in production that never breaks in a voice AI demo?
Turn-taking with real callers, barge-in that does not cancel synthesis, vendor concurrency caps at peak, tail latency on the model, missing call records, cold workers, and one tenant starving the rest. Demos are single calls on good networks with cooperative speakers; production is none of those things.
Concurrency is the ceiling people meet first, and it is rarely their servers. Deepgram publishes 150 streaming transcription connections and 45 text-to-speech connections on pay-as-you-go (Deepgram pricing), and LiveKit's inference service allows 5 concurrent connections on Build, 20 on Ship and 50 on Scale. Hitting a cap does not throw a clean error: the call connects and the caller hears silence. For sizing at thousands of calls a day, see the AI call center build.
Most of the rest is fixed by habits rather than clever code: keep a warm worker per region above peak, write call state after every turn so a crashed worker loses one call, deploy outside the busy hour with drain periods longer than your longest calls, and load-test at 1.5 times forecast peak before a tenant launches.
| Failure mode | What the caller experiences | Where to fix it |
|---|---|---|
| Silence-based endpointing | Agent interrupts mid-sentence | Audio turn detector, with delays tuned per use case |
| Barge-in that does not cancel synthesis | Agent talks over the caller | Stop playback and cancel in-flight synthesis on speech start |
| Vendor concurrency cap reached | Call connects to silence at peak | Know every cap; per-tenant quotas at admission |
| Tail latency on the model | Long pauses; the caller says hello? | Alert on p95 per stage; fail over to a faster model on timeout |
| No per-call record | A complaint nobody can reproduce | One structured row per turn with call ID and tenant from day one |
| Cold workers after a quiet hour | First call is slow to answer | A warm floor of workers; prewarm prompts and indexes |
| One tenant's campaign starves the others | Other tenants' callers hear silence | Per-tenant concurrency quotas and a separate outbound pool |
| Deploy restarts a busy worker | Call drops mid-conversation | Drain on SIGTERM with a grace period above your longest call |
What would I build first, and in what order?
One tenant, one number, one call path, instrumented end to end on the cheapest components that work. Then turn detection and barge-in, then tools and knowledge, then recordings and evals, and only then the tenant model, dashboard and metering. Every step ships something a real caller can use, and nothing is tuned before it is measured.
Make the pipeline swappable from the first week. The most valuable property of your own stack is not today's cost; it is that when a better or cheaper voice ships next quarter, you adopt it with a configuration change and a test on a share of live calls. LiveKit's Agents framework includes a test framework with model judges, which makes that change reviewable rather than a leap of faith.
Leave multi-tenancy until the single-tenant path is boring, but design for it from the start: tenant and agent version on every record, tools that take the tenant from context, and settings stored as versions. Retrofitting those three is the expensive part of turning a voice agent into a platform.
If you want it built rather than planned, that is voice AI development, with voice agent builds from $12,000. It is built at $0: the work is split into checkpoints with acceptance criteria agreed before work starts, and each is invoiced only after you have seen it and accepted it.
- Weeks 1 to 2One call path, instrumented
One tenant, one number, a SIP trunk, a dispatch rule and a worker with transcription, model and voice, logging five timestamps per turn from the first call.
- Weeks 3 to 4Turn-taking and barge-in
Audio turn detector tuned for the use case, cancellation of in-flight synthesis, and history truncated to the words the caller heard.
- Week 5Tools and knowledge
Typed tools with timeouts and read-back, retrieval in on_user_turn_completed, and a grounding check before answers are spoken.
- Week 6Recordings, evals and cutover
Egress to your bucket, a regression set from real calls, and a staged ramp with a one-value rollback.
- Weeks 7 to 8Tenants and versions
Tenant model, numbers and dispatch per tenant, versioned agent settings, and per-tenant concurrency quotas.
- Weeks 9 to 10Dashboard and metering
Agent editor, call logs, analytics, minutes and cost per tenant, and a load test at 1.5 times forecast peak.
Building a voice AI platform on LiveKit: common questions
→What do you need to build your own voice AI platform?
A SIP trunk and numbers, LiveKit for media, SIP and agent dispatch, workers running streaming transcription, a language model and a voice, a turn detector, tools and a knowledge base per tenant, recordings and per-turn logs, and a control plane for tenants, agent versions and usage. My estimate is about six engineer-weeks for the pipeline and four for the platform layer.
→Is LiveKit free to use?
The LiveKit server and the Agents framework are open source under Apache 2.0, so you can self-host both and pay only for servers. LiveKit Cloud has a free Build plan, Ship from $50 a month and Scale from $500, with a third-party SIP fee of $0.004 or $0.003 a minute and agent sessions at $0.01 a minute for agents it hosts.
→How do you make a voice AI platform multi-tenant?
Resolve the tenant before the first word, from the dialled number or the dispatch rule's metadata, and load that tenant's versioned agent settings. Tools take the tenant from the job context, never from the model; knowledge queries filter by tenant in the database; recordings go under a per-tenant prefix; and each tenant gets a concurrency quota so one campaign cannot starve the rest.
→How fast does a voice agent need to respond?
Budget 800 milliseconds from the end of the caller's speech to the first audio of the reply, and judge it on the 95th percentile rather than the median. With LiveKit's audio turn detector the default minimum endpointing delay is 0.3 seconds, which leaves the rest of the budget for transcription, the model's first token and the voice's first audio.
→Should I build my own voice AI platform or use Retell or Vapi?
Use a managed platform for one business under about 20,000 minutes a month; Retell's default is $0.11 a minute and Vapi charges $0.05 before models and voices. Build when you resell agents to many customers, pass that volume, or must keep recordings in your own cloud. Fifty customers at 6,000 minutes each is $30,000 a month at 10 cents and about $7,500 at 2.5.
→What breaks first when a voice AI platform goes live?
Usually a vendor concurrency cap. Deepgram publishes 45 text-to-speech connections on pay-as-you-go and LiveKit's inference service allows 50 on Scale, so a busy hour can pass them long before your servers are full. The call connects and the caller hears silence, so list every cap, raise the smallest first, and load-test at 1.5 times peak.
Open the article in your assistant with one click and ask it how this applies to your product.