Lesson 4 of 7 · 5 min
Direct the delivery
You can shape how a voice performs with the stability, similarity, style, tier, and speaker-boost controls, and pick the right tier for drafts versus finals.
Video walkthrough coming soon
Audio Studio
A chosen voice is a starting point, not a finished performance. A few controls turn it from a flat read into the right read, and all of them ship with sensible defaults, so you can run without touching anything and tune only when a take feels off.
The rulers
Stability runs 0 to 1. Lower is more expressive and varied; higher is steadier and more even. Nudge it up for a calm brand read, down for an energetic ad. Similarity, also 0 to 1, sets how closely the output tracks the chosen voice's character. Style exaggeration pushes the voice's natural style harder, and the advice is to leave it at 0 for marketing: past that, it reads as a performance rather than a person talking.
Voice tier is about quality, not character
Tier is a separate choice, and it trades quality against speed and cost, not personality. There are four. Expressive gives the richest emotion and range. Versatile is a steady, natural read across many languages. Turbo is faster and cheaper. Draft is the fastest and cheapest, built for quick previews. The habit worth forming is the one a photographer uses with contact sheets: audition the script on Draft until the words and pacing are right, then rerun on Expressive or Versatile for the final, since the higher tiers cost more per run.
Speaker boost, and getting the same result twice
Speaker boost is on by default and sharpens presence; leave it on unless a read sounds over-processed. None of these settings are guessed for you, and none drift between runs. They stay exactly where you set them, so a voiceover you loved is reproducible: same voice, same tier, same rulers, same shape of result.