Voices
Audition candidate voices against your own script, then save the one you want and point a workflow at it.
Choosing a voice is three decisions, not one: a provider, a model within that provider, and a voice from that model's catalogue. You make all three in one place.
- Voice Bench — the audition surface. Type a script, hear it read, tune the provider's settings, and compare two candidates side by side.
- Voice Profile — a saved, named voice, owned by your organization. Profiles are reusable: one profile can serve every workflow that should sound the same.
- Fallback voice — a second configuration stored on the profile, used if the primary one fails to start when a call begins.
Providers
The text-to-speech providers wired into the platform today are Deepgram, OpenAI, Google, ElevenLabs, Cartesia, Camb, Rime, Sarvam, Speaches and Azure.
Each exposes its own voice list and its own tunable settings. Speech rate is available broadly;
expressive controls are provider-specific — Azure's styles and Rime's sampling controls on its
arcana model appear only when you pick that provider.
If a provider's parameter is not in the voice settings panel, it is not available from the builder today, whatever the provider itself supports. The panel is the source of truth for what you can control.
Auditioning a voice
Voice Bench has two modes.
- Single renders your script through one configuration, so you can tune settings and hear each change.
- A/B Compare renders the same script through two configurations, so the difference is the only variable.
Phone quality (8kHz) re-renders the preview at the sample rate a phone line actually carries. A voice that shines at full bandwidth can lose its edges there, so judge on the phone-quality preview before committing.
Saving a profile
Once a configuration sounds right, save it as a named Voice Profile. Tick Include Config B as a fallback voice first to store the second A/B candidate as that profile's fallback.
Name profiles after the agent they belong to rather than the voice — "Support line", not "Cartesia Sonic". The provider behind a profile can change later; what the profile is for does not.
Using a profile in a workflow
A profile does nothing until a workflow points at it. Open the workflow's model configuration and pick a Voice Profile; its settings are expanded into the workflow's speech configuration at call time.
Two consequences worth knowing:
- The reference is live. Edit a profile and every workflow pointing at it changes on the next call, with nothing to republish.
- The choice is per workflow, and it applies to every node in that workflow that speaks. There is no per-node voice override.
Expressive Mode sits beside the profile picker. It adds the selected provider's own emotion markup to the model's instructions, so the agent can direct its delivery from what the caller said. The markup language differs per provider and one provider's tags are read aloud as words by another, so the toggle is disabled unless the profile's model supports expressive delivery.
Next steps
| Goal | Guide |
|---|---|
| Decide what the voice actually says | Prompts |
| Set the opening line as a recording instead | Nodes |
| Hear a voice on a real call | First call |
Common questions
What is Woise?
Woise lets you build AI agents that answer and make phone calls and chat on your website. You lay out the conversation on a visual canvas, try it in your browser, then connect it to a phone number.
Do I need a phone number to start?
No. You can build and test an agent in the browser without one. When you are ready, connect a phone number and choose which agent answers it.
How does billing work?
You pay for the minutes your agents use, from your credit balance. There are no seat fees and no per-agent fees, so an agent that is not taking calls costs nothing.
Which languages can an agent speak?
Agents can listen and speak in more than 40 languages, including English, Spanish, Hindi, Tamil and Arabic. You pick the language and the voice for each agent.
Can an agent connect to my CRM or calendar?
Yes. An agent can look things up and make changes in 26 apps, such as Salesforce, HubSpot, Google Calendar, Slack and WhatsApp, while the caller is still on the line. You connect each app once.
Do I need engineers to build an agent?
No. Everything is built in a visual editor, so you can create and change an agent without writing code. If your team prefers code, there is also a REST API and an MCP server.
What happens to call recordings and transcripts?
Every call keeps a transcript and its outcome, and audio is only recorded if you turn it on. Only people in your account can see your calls, and your account lives in the region you chose when you signed up.
What do I have to set up or run?
None. We run the speech, the language models, the recordings and the phone connections, and scale them with your call volume. There is nothing for you to set up or keep running.
How do I reach the team?
Email contact@woise.ai, call or WhatsApp +91 83412 34080, or use the form on our contact page. A person reads every message.
Still stuck? Talk to our team.