- TL;DR
- ElevenLabs pricing is a studio bill, not an API bill
- How ElevenLabs credits actually work
- Pricing, tier by tier
- Cost calculator: what your monthly minutes really cost
- The traps in the pricing page
- Alternatives, by use case
- ElevenLabs vs the cheap clouds: the math
- Latency and quality: what the premium buys
- Failure modes we have hit
- When ElevenLabs is worth the premium
- When a cheaper TTS API wins
- FAQ
- Related Cluster Intelligence
TL;DR
- 2026 self-serve pricing: Free ($0, 10k credits), Starter ($6, 30k), Creator ($22, 121k), Pro ($99, 600k), Scale ($299, 1.8M), Business ($990, 6M). Annual billing is ten months for twelve.
- A credit is not one fixed unit. Multilingual v2 spends 1 credit per character; Flash and Turbo spend about half. Roughly 1,000 characters is one minute of speech, so Pro through Business land near $0.17 per minute.
- The "$11 Creator" headline is a first-month discount. From the second month the plan is $22. Budget on $22.
- Above Starter you are not buying a lower rate, you are buying volume, seats, and cloning. The effective per-minute cost is nearly flat from Pro upward.
- Cheaper by use case: OpenAI TTS near $0.015/min for preset voices, Google WaveNet near $0.004/min, Azure Neural near $0.016/min. None of them clone a voice the way ElevenLabs does.
ElevenLabs pricing is a studio bill, not an API bill
Most ElevenLabs pricing posts quote the sticker and stop. That misses the design. ElevenLabs does not sell characters, it sells a workspace: text-to-speech, speech-to-text, voice cloning, dubbing, sound effects, and a conversational agent platform, all drawing from one monthly credit pool. The subscription is the price of admission. The credits are the real currency.
That structure explains the pricing page. It also explains why teams get surprised. A creator buying narration and an engineer shipping a voice agent look at the same six tiers and get very different value out of them. If all you need is synthesis, you are paying for a studio you will never open. If you need the studio, the alternatives stop being comparable, because most of them are synthesis and nothing else.
The rest of this piece separates those two buyers. First the credits and tiers, then a calculator for your own volume, then the places where a plain TTS API beats the subscription outright. If your real decision is dialog design versus telephony execution rather than raw synthesis, that half of the stack is covered in our Voiceflow vs Bland.ai comparison.
How ElevenLabs credits actually work
The credit decides your bill, and it is not a fixed unit.
For the Multilingual v2 text-to-speech model, one credit buys one character. Flash and Turbo are cheaper per character, which is why real-time products meter differently from narration. Speech-to-text, dubbing, sound effects, and music each set their own rate inside the same pool. Dubbing is the costly one, because a single minute runs transcription, translation, and synthesis behind the scenes.
The conversion that makes this manageable: about 1,000 characters is one minute of finished speech. A 300-word script is roughly 1,800 characters. That is enough to size a plan by hand. The Free plan's 10,000 credits is about ten minutes of narration, Creator's 121,000 is a couple of hours, and Business's 6,000,000 is a hundred.
Two rules matter for forecasting. Credits reset monthly and do not roll over, so spare budget is lost. And above Starter, paid plans can switch on usage-based billing, which keeps production running past the included credits instead of halting it; you pay an overage close to the plan's effective per-credit rate. Our production voice stack, which pairs synthesis with dialog logic, is documented in the agentic architecture blueprint if you want to see how the metering behaves under load.
Pricing, tier by tier
These are the published 2026 self-serve figures from elevenlabs.io/pricing — US pricing, before tax, checked in September 2026.
| Plan | Monthly | Annual (per mo) | Credits / mo | Seats | Commercial | Concurrency |
|---|---|---|---|---|---|---|
| Free | $0 | $0 | 10,000 | 1 | No | 2 |
| Starter | $6 | $5 | 30,000 | 1 | Yes | 3 |
| Creator | $22 | $18.33 | 121,000 | 1 | Yes | 5 |
| Pro | $99 | $82.50 | 600,000 | 1 | Yes | 10 |
| Scale | $299 | $249.17 | 1,800,000 | 3 | Yes | 15 |
| Business | $990 | $825 | 6,000,000 | 10 | Yes | 25 |
| Enterprise | Custom | Custom | Custom | Custom | Yes | Elevated |
Three jumps carry the whole ladder. Free to Starter ($6) buys the commercial license — you cannot publish monetized work on Free-tier audio. Starter to Creator ($22) buys professional voice cloning and the first real credit pool. Creator to Pro ($99) buys output quality and throughput: 44.1 kHz PCM over the API, higher concurrency, and usage-based billing for spikes.
Everything above Pro is volume and seats. Scale and Business add workspace seats and more concurrent requests. The per-minute rate on those tiers is essentially the same as Pro's.
Cost calculator: what your monthly minutes really cost
The tier list does not tell you what you will pay. Your monthly minutes do. Put in your real volume and compare an ElevenLabs plan against two cloud APIs that bill per character.
The traps in the pricing page
Four things about this page cost teams real money, and none of them are the headline number.
- The $11 Creator price is a first-month discount. From month two the plan is $22. Budget on $22 and treat the discount as a trial, not a rate.
- Credits do not roll over. A quiet month does not bank credits for a loud one. If your volume swings, either leave headroom in the tier or turn on usage-based billing so a spike bills overage instead of stopping production.
- Credits are shared across every product. Speech-to-text, sound effects, music, and dubbing all draw from the same pool. A team that runs dubbing alongside TTS finds its "minutes of speech" estimate overshoots badly, because dubbing burns credits far faster per minute.
- Conversational agent minutes bill separately. If you ship a voice agent, agent conversation minutes are metered on their own rate, outside the core credit pool. That is the line most teams miss when they size a plan off the TTS math alone.
Alternatives, by use case
ElevenLabs is the quality benchmark. It is not the budget option, and it is not always the fastest. Published per-character rates below are reconciled from vendor pricing pages and 2026 third-party comparisons; the per-minute column assumes the same ~1,000-characters-per-minute pacing, so it is an estimate, not a vendor figure.
| Alternative | Published rate | ~$/min | Voice cloning | Best for |
|---|---|---|---|---|
| OpenAI TTS | $15 / 1M chars | ~$0.015 | No (preset) | Simplest API; already inside OpenAI |
| Google Cloud WaveNet | $4 / 1M chars | ~$0.004 | No | Cheapest neural at volume; GCP stacks |
| Amazon Polly Neural | $16 / 1M chars | ~$0.016 | No | AWS-native pipelines |
| Microsoft Azure Neural | $16 / 1M chars | ~$0.016 | Custom (gated) | Compliance, on-prem containers, 140+ locales |
| Cartesia Sonic | ~$19–39 / 1M chars | ~$0.02–0.04 | Instant (Pro) | Sub-100ms real-time agents |
| Deepgram Aura-2 | ~$30 / 1M chars | ~$0.03 | No (preset) | On-prem/VPC; STT and TTS from one vendor |
| PlayHT (PlayAI) | from $39/mo | ~$0.07 | Instant | Now under Meta and winding down — migration risk |
Read the table as a shortlist, not a ranking. OpenAI TTS is the sensible default when you are already calling OpenAI and preset voices are fine. Google WaveNet and Amazon Polly are the cheap neural engines for high-volume narration. Azure is the one to reach for when infosec demands on-prem containers. Cartesia owns the sub-100ms slot that real-time agents care about. Deepgram is the enterprise pick when you want speech-to-text and speech synthesis from one vendor. PlayHT, once the obvious alternative, now carries real migration risk.
ElevenLabs vs the cheap clouds: the math
The gap is larger than the tier table suggests. Take a pipeline that produces 100 million characters a month, about 100,000 minutes of speech.
- ElevenLabs, at a Pro-to-Business effective rate near $0.17/min: roughly $17,000 a month.
- OpenAI TTS, at $15 per million characters: about $1,500 a month.
- Google WaveNet, at $4 per million characters: about $400 a month.
That is an order of magnitude, and on pure synthesis there is no arguing with it. The cloud APIs win the volume case outright. What they do not sell is a trained clone of a specific voice, the expressive v3 model, dubbing in one pass, or a managed agent platform. Those are the features you are actually buying on the ElevenLabs side of the ledger. If your workload is preset voices reading predictable text at scale, the clouds are correct and ElevenLabs is expensive. If voice is the product, the comparison inverts, because the alternatives would need three vendors stitched together to match one subscription.
Latency and quality: what the premium buys
Voice quality is the easiest thing to judge and the last thing that should decide your shortlist. Latency and deployment come first, because they kill or clear a shortlist before naturalness matters.
ElevenLabs splits its models by job. Eleven v3 is the expressive one — audio tags, multi-speaker dialogue, 70+ languages — but it is offline and caps a request at 5,000 characters. Flash v2.5 is the real-time variant that voice agents actually run, with a reported time-to-first-byte around 75ms and a 40,000-character cap. Choosing ElevenLabs for a live agent means running Flash, not v3, and understanding that trade.
Against that, the cloud engines sit at 300–500ms to first byte. Cartesia undercuts everyone at 40–90ms. OpenAI's realtime stack lands under 300ms when speech-to-speech. So ElevenLabs does not win on raw latency; it wins on expressive range and on the fact that cloning, dubbing, and agents live in the same account. A natural-sounding preset voice from any cloud can pass a blind listen test on easy copy. The ElevenLabs gap shows up on hard material — names, numbers, emotion, and languages it has already shipped for.
Failure modes we have hit
A few patterns show up again and again in production voice work:
- The credit cliff. No rollover, and on Starter-and-above a spike can exhaust the pool mid-month. Without usage-based billing, generation stops. Monitor burn in week one and set an upgrade trigger before you hit zero.
- The first-month trap. Teams budget the $11 Creator price and get a $22 bill in month two, then re-forecast everything. Model on the standard price from the start.
- Concurrency as the real ceiling. Credits can be fine while concurrency caps you at 2 to 25 simultaneous requests. A traffic spike queues even when the pool is healthy. Size concurrency, not just credits, against your peak.
- Agent minutes hidden outside the pool. A voice agent plan sized on TTS minutes underestimates, because agent conversation minutes bill on their own meter. Model the agent line separately.
- Betting on a shrinking vendor. PlayHT was the obvious alternative until it landed under Meta and began winding down. Pin your pipeline to a vendor with a clear roadmap before you rebuild on it.
When ElevenLabs is worth the premium
Buy it when voice is a recurring workflow or a shipped feature, not a one-off. The strongest cases: a creator producing narration weekly, a team localizing content through dubbing, a product that needs professional voice cloning from a short sample, or a voice agent where expressive range and multi-language coverage change the user experience. In each of those, the alternative is either a stack of three services or a human voice budget several times larger. Against that, $22 to $990 a month is cheap. The premium is real, and so is the work it replaces.
When a cheaper TTS API wins
Reach for a cloud API when the job is synthesis and only synthesis. Single language, preset voices, high volume, predictable copy — that is OpenAI TTS at $0.015/min, Google WaveNet at $0.004/min, or Azure and Polly at $0.016/min. If you already run inside one of those clouds, the integration cost is close to zero and the per-minute saving at scale is an order of magnitude. Skip ElevenLabs when no one will ever clone a voice, dub a video, or ship an agent. You would be paying studio prices for an API you already have.
FAQ
Ship voice agents that sound human
When your agents talk, latency and tone decide trust. Our ElevenLabs playbook covers voice-agent delivery patterns.
Get the ElevenLabs Voice Playbook →Download this guide’s assets
Get the configuration and data files referenced in this guide. Subscribe and we’ll send the bundle to your inbox.
Get the bundle →Is ElevenLabs free to use?
There is a real free tier: $0 a month for 10,000 credits, about ten minutes of text-to-speech, plus speech-to-text, sound effects, and music. It carries no commercial license, so anything you monetize needs a paid plan. Commercial use starts at Starter, $6/mo.
Why does my ElevenLabs bill say $22 when the page says $11?
The $11 Creator price is a first-month discount, currently 50% off. From the second month the Creator plan is $22. Model your budget on $22 and treat the first month as a trial.
How many characters is one ElevenLabs credit?
For the Multilingual v2 model, one credit buys one character. Flash and Turbo spend about half a credit per character. ElevenLabs' own estimate puts 10,000 credits at about ten minutes of speech, which is roughly 1,000 characters per finished minute.
What is the cheapest ElevenLabs alternative?
For plain synthesis, Google Cloud WaveNet at $4 per million characters is the cheapest neural option, near $0.004 per minute; Amazon Polly Standard is in the same range. OpenAI TTS at $15 per million is the simplest cheap API. None of them clone a voice or match the v3 expressive model, so they are cheaper for a different job.
Do ElevenLabs credits roll over month to month?
No. Credits reset every billing cycle and do not carry over, so unused budget is lost. If your volume swings, leave headroom in the tier or enable usage-based billing so a spike bills overage rather than stopping generation.
Related Cluster Intelligence
- Voiceflow vs Bland.ai: Voice Agent Comparison →
- Make.com 429 Rate Limit Fix: Circuit Breaker Protocol →
- Async AI Agent Architecture with Redis Queues →
- Multi-Agent Outbound Pipeline 2026: Enterprise Production Architecture & 3-Tier Agent Mesh →
- Smartlead Deliverability Repair Protocol →
- LLM Fallback Chain: Groq, Gemini, DeepSeek on OpenRouter →