Fresh out of the oven
Spot a bug? These tools are new, so the odd thing may still wobble. If something looks wrong, it probably is. Tell us and we will sort it..
Fresh out of the oven
Spot a bug? These tools are new, so the odd thing may still wobble. If something looks wrong, it probably is. Tell us and we will sort it..
Fresh out of the oven
Spot a bug? These tools are new, so the odd thing may still wobble. If something looks wrong, it probably is. Tell us and we will sort it..
Pinnora's AI voice generator turns a script into one MP3 read. Paste finished words and they are spoken close to verbatim; paste a rough brief and it writes the spoken lines first from your project's research, leaving out any claim or URL that is not already in your grounding. Billing is per character spoken.
The tool is open above. This is the short version of what to do with it.
8 to 5,000 characters. Finished words are read close to verbatim; a rough brief is rewritten for the ear first, short sentences and one CTA.
Voice tier picks Expressive, Versatile, Turbo or Draft. Voice under it lists ten stock reads plus the brand voice connected to your org.
Stability opens at 0.5 and Similarity at 0.75. Leave Style exaggeration at 0 for marketing reads. Speaker boost is on by default.
One MP3 lands with the spoken script beside it. Billing is per character spoken, so the same text, voice and tier in one project bills once.
Short answers to what people ask before they start.
The Script box takes 8 characters minimum and 5,000 maximum, roughly 800 words. The script pass paces at about 145 words a minute, so a full box is around five and a half minutes of audio. Go past 5,000 and the render is refused with a note to split it into separate renders. There is no duration control, so the audio runs as long as the words do.
The script pass works out which one it is holding. A finished read comes back essentially verbatim. A brief gets written into spoken lines against the project's research, strategy, past winners and brand memory, under a rule that every claim, price and URL must already appear in that grounding. A domain that is not in the data is left unsaid rather than guessed. The spoken script is saved with the MP3.
Yes. Project / brand default is the first entry in Voice, and it resolves to the default voice on your org's connected audio account, a cloned voice included. With nothing connected it falls through to the stock library instead. A fresh connect takes effect straight away rather than waiting out the five minute cache. The ten named voices stay there for a specific read.
Per character of the script that gets spoken, not per attempt. Stop is checked before the script pass and again right before the audio call, so a stop that lands in time spends nothing. The charge key is built from the voice, the tier and the exact text, so re-running the same read in the same project bills once rather than twice. Turbo and Draft sit at half the per-character rate of Expressive and Versatile.
Expressive reads inline square-bracket tags such as [warmly] or [pause] as performance directions. Versatile, Turbo and Draft would say the brackets out loud, so they are stripped before the read along with the spacing they leave behind. Round brackets are spoken literally on every tier, which is why the script pass is told to use ellipses, dashes and CAPS for emphasis instead.
The rest of the family. Every one is free to try and runs on the same engine.