Disclosure: This post contains affiliate links; we may earn a commission at no extra cost to you.
AI voice pricing is confusing because vendors sell different products with the same word “voice.” A YouTuber buys a monthly creator plan with credits. A training team buys seats and a pool of generated minutes. An app developer pays per character plus data transfer and language-model costs. A call center pays for connected agent minutes that bundle—or exclude—speech recognition, text-to-speech, an LLM, and telephony.
ElevenLabs is the best value for creators who prioritize natural, expressive narration and want one credit pool for text-to-speech, dubbing, sound effects, music, and other audio tools. Amazon Polly or Microsoft Azure Speech is far cheaper for high-volume, predictable API narration. Murf costs more than commodity APIs but includes a no-code studio and team workflow. The correct comparison is cost per approved finished minute after retakes, not the plan’s nominal character allowance.
ElevenLabs’ current lineup includes Free, Starter, Creator, Pro, Scale, Business, and custom Enterprise. Recent US pricing has listed Free with 10,000 monthly credits, Starter around $6 with 30,000, Creator around $22 with 121,000 (often a discounted first month), Pro around $99 with 600,000, Scale around $299 with 1.8 million, and Business around $990 with 6 million. Verify live rates, taxes, and annual discounts.
Our pick: ElevenLabs Creator
Creator is the practical tier for an individual who needs a Professional Voice Clone slot and materially more narration than Starter. Standard text-to-speech commonly costs about one credit per character, but models and products differ. Speech-to-text, dubbing, voice changing, music, agents, and sound effects have different burn rates. A headline “121,000 credits” is not 121,000 spoken words.
| Pricing model | Representative services | Easy to forecast? | Hidden cost |
|---|---|---|---|
| Subscription credit pool | ElevenLabs, Fliki, Descript | Moderate once each action’s rate is known | Different models/actions consume credits differently; expiry and overage |
| Characters per month | Many creator TTS plans and APIs | Good for scripted narration | Spaces/punctuation can count; Chinese/Japanese billing rules differ |
| Generated minutes | Murf and video/voice studios | Intuitive for finished voiceover | Regeneration can count again; downloaded vs generated time differs |
| Pure pay-as-you-go API | Amazon Polly, Azure Speech, Google Cloud TTS | Excellent at scale | Cloud account, storage/egress, support, and developer labor |
| Per connected conversation minute | ElevenAgents, voice-agent platforms | Only after call-duration testing | LLM, STT, TTS, telephony, silence, and minimum increments may be separate |
| Seat plus shared usage | Murf Enterprise, WellSaid, Descript, team products | Moderate for stable teams | Editor/reviewer seats, workspace minimums, SSO and support tiers |
| Custom voice training and hosting | Azure Custom Neural Voice, enterprise vendors | Poor without a quote | Talent, consent, data preparation, training, endpoint hosting, governance |
English narration typically runs around 130–160 spoken words per minute. An English word averages roughly five letters, plus a space and punctuation. A practical planning range is 800–1,000 billable characters per finished minute. Delivery, language, numbers, abbreviations, and markup change that ratio.
At 900 characters per minute, 100,000 characters yields about 111 raw minutes. That is not 111 approved minutes. If 20% of passages require a second take and 5% require a third, the same allowance may produce closer to 85–95 finished minutes. Expressive models with several auditions can use more.
Count the actual scripts before choosing a tier. In Microsoft Word, Google Docs, a code editor, or a simple script, total characters including spaces. Add titles, alternate intros, pronunciation variants, and legal lines. Then apply a revision factor:
Chinese characters and other scripts may be counted differently. Azure documentation, for example, notes that Chinese characters—including kanji, hanja, or hanzi in relevant languages—can count as two characters for billing. Use the vendor’s locale-specific rules rather than applying an English estimate globally.
ElevenLabs’ shared credit system is convenient if a creator uses several tools. Current approximate rates include one credit per TTS character on relevant models, hundreds of credits per transcription minute, roughly 900 per Eleven Music minute, a fixed amount per sound-effect generation, and higher per-minute costs for voice changing and dubbing. The current pricing page is authoritative because model discounts and product rates change.
Creator’s 121,000 credits could represent around two hours of raw English TTS under a one-credit-per-character model, but far less if spent on dubbing or repeated expressive takes. Pro has a lower effective unit price and much more capacity, but moving up solely to avoid careful script review can waste $77 monthly.
The Free plan is useful for auditions. It does not provide the same commercial license as paid output, and has no Professional Voice Clone slot. Starter is appropriate for short commercial experiments and Instant Voice Cloning under current terms. Creator adds a professional clone and higher-quality production capacity. Scale and Business become relevant for teams, high volume, multiple clone slots, seats, and enterprise operations.
ElevenLabs has introduced or expanded pay-as-you-go for developer products, reducing the need to prebuy a large app subscription in some API scenarios. App and API billing may not be identical. Confirm whether a monthly subscription, pay-as-you-go balance, or enterprise contract applies to the endpoint you will call.
Credits can expire or roll according to plan and current rules, and overage may require a top-up or usage-based setting. Do not assume cancellation preserves unused credits. Export masters and keep source text before ending service.
Amazon Polly charges by processed characters and offers Standard, Neural, Long-Form, and Generative engines with different voice availability, quality, price, and quotas. It has no creator timeline comparable to Murf. Developers send text or SSML through AWS and receive an audio stream or file.
Recent US pricing references have placed Standard around $4 per million characters, Neural around $16, Long-Form around $100, and Generative around $30, though region and engine availability matter. Use the official AWS pricing page and calculator for the deployment region. At 900 characters per minute, one million characters is roughly 1,100 raw minutes—illustrating why cloud APIs can undercut premium creator tools.
The bill is not the whole implementation. A developer must write integration, handle queues and throttling, store audio, manage pronunciation and SSML, monitor errors, secure AWS credentials, and support the application. Polly’s standard voices may sound less emotionally rich than ElevenLabs, and selected languages or engines have different voices.
Polly fits navigation, announcements, accessibility reads, high-volume articles, and app prompts where consistency and cost matter. It is less attractive for a flagship audiobook or emotionally performed commercial.
Azure Text to Speech also bills by character, with Free and Standard resources and model categories including neural, HD or custom options as the product evolves. Current pricing pages indicate free monthly allowances for relevant tiers—such as hundreds of thousands of neural characters or millions of standard characters—with paid pricing shown through Azure’s regional calculator.
Azure stands out for supported speaking styles, style degree, pronunciation, roles, multilingual voices, and enterprise integration. An organization already using Microsoft Azure can keep speech generation within established identity, billing, security, and procurement controls.
Custom Neural Voice is not simply another per-character option. Access is limited and governed. Costs can include a professional speaker, consent, recording studio, dataset preparation, training, deployment endpoint, and ongoing synthesis. A custom brand voice may save at scale, but a stock neural voice is far cheaper to launch.
Azure invoices can include connected cloud resources, support, networking, monitoring, and currency/tax effects. Build a budget alert. A leaked API key can generate a much larger problem than exceeding a creator plan by one video.
Murf sells a production interface rather than raw speech alone. The Studio combines script, voices, timing, pronunciation, emphasis, media, music, and export. Paid plans commonly start around the high-$20s per month and rise toward roughly $99 or enterprise quotes depending on generation capacity, seats, translation, cloning, brand tools, and support. Murf’s Free trial has offered around 10 minutes without downloads or commercial rights.
Minute-based limits are easy to understand until revisions begin. If a 10-minute module is regenerated twice, the account may consume more than 10 minutes. Check whether editing text, changing voice, previewing, or downloading triggers generation usage and whether unused time rolls forward.
Murf can still be cheaper than an API for a nontechnical team because it removes development. An instructional designer can synchronize voice to scenes without asking engineering to build a tool. Team plans also attach a cost to collaboration, shared projects, and approvals that a low API price excludes.
Murf Dub is a separate localization product with pay-as-you-go or plan pricing. Do not mix its minutes with ordinary Studio TTS when estimating.
Descript subscriptions include media minutes and AI credits. Custom AI Speaker generation, filler removal, Studio Sound, video regeneration, Underlord actions, transcription, and other tools can draw from different pools. This makes “cost of voice” hard to isolate, but the plan may replace a separate editor and transcription service.
The right calculation is cost per finished podcast or video. If Descript saves two editing hours, a higher TTS unit rate may be irrelevant. Top-up pricing can become expensive; official bundles have listed 3,000 AI credits around $180. Monitor the Usage page after a real episode.
Fliki also uses credits across voice and video. Current documentation has illustrated a standard voice costing roughly half a credit per generated minute, while avatars and premium media use more. Editing text can reprocess voice and consume again. The Standard plan has recently been listed around $28 monthly before annual promotions, with Premium around $88 before discounts, but verify live pricing and export limits.
Both tools are economical when their surrounding workflow is useful. Buying either solely as an API voice provider makes less sense than a dedicated service.
A live voice agent may incur:
Some vendors bundle several components into one per-minute rate; others pass through LLM and telephony costs. Silence, hold time, voicemail, minimum billable increments, and abandoned calls can count. ElevenLabs has published starting agent speech rates around cents per minute for relevant modes, but total conversation cost is higher after the rest of the stack.
Test 100 representative calls and calculate cost per resolved outcome, not per minute. A cheap agent that takes nine minutes and transfers half its calls can cost more than a better one that resolves in three. Include human escalation and quality review.
Four ten-minute videos are about 40 finished minutes monthly. At 900 characters per minute, the final text is around 36,000 characters. Apply a 1.4× expressive revision factor and budget about 50,000 characters.
ElevenLabs Starter’s 30,000 credits is likely too small under a one-credit-per-character model; Creator fits with room for other audio and professional cloning. Pro is unnecessary unless other shows or localization push the account toward hundreds of thousands of characters.
Polly Neural or Azure neural API may cost only a small fraction of a creator subscription for 50,000 characters, but will require a production interface and may not deliver the preferred expressive voice. Murf can fit if its included generation minutes cover the retries and the editor saves time. Record a human if 40 minutes can be performed and cleaned faster than prompting the model.
One hundred finished hours is 6,000 minutes, perhaps 5.4 million English characters before revisions. At this volume, vendor negotiation, batch APIs, and custom pricing matter. A cloud API can be dramatically cheaper in raw synthesis, while a premium clone may justify higher spend for learner engagement and consistent brand delivery.
Add translation, native review, captions, LMS packaging, content updates, and pronunciation QA. If 5% of lessons change annually, maintainable scene-based generation may save more than the initial voice rate. Ask for enterprise data protection, SSO, audit logs, clone governance, service levels, and indemnity terms.
Confirm whether usage is measured on input text, generated audio, downloads, or final duration. Ask what counts as a character, whether SSML counts, which model uses how many credits, whether failed generations are refunded, and whether previews consume quota. Check rollover, overage, top-ups, annual allotment timing, cancellation, commercial rights, clone slots, seats, API inclusion, concurrency, and rate limits.
Then inspect operational terms: can output remain published after cancellation, can the voice model be exported, how is consent verified, may data train models, what is retained, and how quickly can a compromised clone be disabled?
ElevenLabs Creator is the best premium self-service balance for expressive narration. Polly and Azure win raw high-volume API economics. Murf earns its premium through an accessible studio, while Descript and Fliki are valuable when the whole editing or video workflow is used. Count scripts, add retries, and compare approved minutes—otherwise every pricing page will make its own unit look cheaper than it really is.
Related reading: AI Voice API Pricing Compared · AI Voice Emotion Controls Compared · Voice Cloning + Avatar Workflow