Open ElevenLabs for the first time, generate a voiceover, and it sounds fine. Adjust the Stability slider because a tutorial told you to, generate again, and now it sounds worse — and you can't say why. That's the normal experience with these controls. They're labeled abstractly because what they're actually doing is statistical.
Three sliders sit in ElevenLabs' Voice Settings panel: Stability, Similarity, and Style Exaggeration. The useful way to think about them is as one system with three dials, where moving one changes what the other two are doing.
What Each Slider Actually Controls
- Stability governs how much the voice varies between generations and within a single take. Stability determines how stable the voice is and the randomness between generations. Lower values give the model more room to be performative — inflections shift, pacing breathes. Push it too low, and the generations become unstable or speak too fast. Push it too high, and you get a flat, monotone delivery.
- Similarity (labeled "Clarity + Similarity Enhancement" in the API) controls how closely the output sticks to the source voice. Pushing similarity higher makes the clone sound closer to the original speaker. The catch is source quality: if the original audio is noisy, the AI will faithfully reproduce that noise. For tips on cleaning up audio at the source, read our report on why 30 minutes of clean audio beats 3 hours of noisy recordings.
- Style Exaggeration amplifies the stylistic mannerisms of the source voice — vocal quirks and characteristic emphasis. It is computationally heavier and might increase latency. It is also the setting most likely to compound badly with the other two settings.
The Interaction Problem
These sliders don't operate independently:
- A high-stability, high-style combination fights itself — stability locks the delivery down while style pushes for exaggerated expression, resulting in a stiff, over-performed voice.
- A low-stability, high-style combination can tip into a caricature of the voice rather than a natural performance.
ElevenLabs' default guidance is stability around 50 and similarity near 75 as a baseline, with style left at 0 and adjusted only after you've heard how the base voice behaves.
Settings by Use Case
| Use Case | Stability | Similarity | Style | Why |
|---|---|---|---|---|
| Narrated explainer / tutorial | Mid-to-high | Mid-to-high | Low | Consistency and clarity outrank expressiveness |
| UGC-style ad / testimonial | Low-to-mid | Mid-to-high | Low-to-mid | Needs to sound spontaneous and unscripted |
| Branded narrator / host | Mid | High | Low | Needs voice identity match across videos |
| Character or dramatic VO | Low | Mid | Mid-to-high | Emotional range is the priority |
| IVR / accessibility audio | High | Mid | Very low | Predictability and neutrality are key |
Treat this table as a starting orientation. Generate a short test clip before committing a full script to render.
What Changed With Eleven v3
Eleven v3, ElevenLabs' newer expressive model, changes the Stability control. In v3, Stability isn't a continuous slider — it's three named modes:
- Creative is more emotional and expressive but prone to hallucinations.
- Natural is closest to the original voice recording and balanced.
- Robust is highly stable but less responsive to prompts, similar to v2.
This reflects an architectural difference: v3 is built around inline Audio Tags (e.g., `[whispers]` or `[excited]`) placed in the script.
The trade-off: Professional Voice Clones aren't fully optimized for Eleven v3 yet, which can result in lower clone quality compared to earlier models. If a brand has invested in a Professional Voice Clone, that's a real reason to stay on Multilingual v2 for now. For a complete walkthrough of setting up your microphone environment, read our guide on building your brand's voice clone: our recording-day checklist. Read more on this in our breakdown of instant vs professional voice cloning.
Where AI Voice Fits in a Video Workflow
The settings get you 80% of the way to a usable read. The remaining 20% — matching pacing to the cut, catching lines that need a retake, and polishing delivery — is human judgment.
If you are scripting these voiceover texts, read our report on switching our scripting workflow from ChatGPT to Claude.
A Quick Workflow Checklist
- Test on a 20-30 second clip first to avoid wasting credits.
- Check source audio quality before pushing Similarity high.
- Decide on consistency vs. expressiveness before adjusting sliders.
- Use Natural mode as a default starting point in v3.
- Keep a settings log per voice, as settings don't transfer between clones.
Summary
Stability, Similarity, and Style aren't independent dials — they're a system. ElevenLabs' default guidance is a starting point, not a formula.
Getting a voiceover that sounds right is still a production skill. Ready to get a voice that actually sounds like your brand? Get a free consultation and we'll walk through what your project needs.
FAQs
What is the best stability setting in ElevenLabs?
There isn't one. Default guidance is to start around 50 and adjust based on how the specific voice responds.
Why does my ElevenLabs voice sound robotic?
Robotic output usually comes from setting stability too high relative to the voice's natural range, combined with style exaggeration left at 0.
What's the difference between Stability and Similarity?
Stability controls how consistent the delivery is between generations. Similarity controls how closely the output matches the timbre of the source voice.
Should I use Eleven v3 or Multilingual v2?
Use Multilingual v2 if you use a Professional Voice Clone. Use v3 if you need emotional range or inline Audio Tags and can use an Instant Voice Clone.
Why does my cloned voice pick up background noise?
This occurs when the Similarity slider is set too high relative to the quality of the source recording, causing the AI to reproduce background artifacts.
What does Style Exaggeration do?
It amplifies the speaker's vocal quirks and characteristics. High values add computational cost and can make the output unstable.
Can I use the same settings across different cloned voices?
No. The correct settings combination is specific to the characteristics of the source audio.
Is Eleven v3 worth switching to right now?
Yes, if your project needs expressive, tag-directed performance and doesn't rely on a Professional Voice Clone.
