---
name: voiceover-tts
description: Turn a script into a finished voice-over — narration, ad reads, podcast lines, audiobook passages, Russian or any of 30+ languages — with Qwen3-TTS and Gemini TTS through the Twelver MCP server. Use when the user asks to voice, narrate or read a text aloud, wants an MP3/WAV of a script, needs the same voice again for a follow-up line, wants a voice designed from a description, or asks for a two-speaker dialogue (озвучка текста, закадровый голос, начитка, синтез речи).
license: MIT
metadata:
  author: Twelver
  homepage: https://twelver.ru/skills/ru/voiceover-tts/
  version: "1"
---

# Voice-over from a script

Two engines, one job: get the user's exact words recorded in the right voice, and
pay for each take once.

## Connect the tools

Twelver runs one remote MCP server per toolset (Streamable HTTP, no local install):

- Qwen3-TTS: `https://twelver.ru/api/mcp/v1/qwen-audio-speech`
- Gemini TTS: `https://twelver.ru/api/mcp/v1/google-audio-speech`

Use `twelver.app` instead of `twelver.ru` for accounts registered on
twelver.app; they are separate account systems. Auth is OAuth sign-in (needs a
registered account) or `Authorization: Bearer sk-…` with a key from
`https://twelver.ru/api-keys`. Connect Qwen first, since it handles most jobs.

Each tool returns a text summary plus a link to the audio file. Give the user the
link and do not try to inline the audio.

## Tools

| Tool | Use it for |
|---|---|
| `qwen-list-voices` / `google-list-voices` | The user's existing voices, free. Call before creating a new one |
| `qwen-text-to-speech` | Default read. zh/en/fr/de/ru/it/es/pt/ja/ko only; up to 10,000 chars, split and stitched automatically |
| `qwen-voice-clone` | A voice from the user's own 3–60 s recording |
| `google-text-to-speech` | Other languages, directed delivery (emotion, whisper, character) via `deliveryNotes`, tags like `<laugh>`, `<sigh>`, `<short pause>` in the text; up to 7,000 chars, not split |
| `google-voice-design` | A new voice from a text description, with a ~10 s preview |
| `google-text-to-speech-dialogue` | Two speakers in one track (Google only, exactly two) |

## Procedure

1. **Confirm the language of the text before the first take.** The engine
   choice depends on it.
2. **List existing voices first** if the user has voiced anything before (both
   list tools are free), and offer those by description.
3. **Pick one engine and bill once.** Never synthesise the same text on both to
   compare. Qwen is the default: about 4× cheaper per character than Google.
   Qwen already has a voice for each common house style: Neil for a news desk,
   Ethan for an energetic ad, Cherry for warm and conversational, Bellona for
   authoritative.
4. **Switch to Google** when the language is not one Qwen speaks, when it is a
   two-speaker dialogue, or when the delivery has to be *directed*: an emotion,
   a whisper, a named character. "Narrator", "friendly" or "professional" are a
   choice of voice, not direction.
5. **Send the script exactly as the user wrote it, in one call.** Do not
   re-punctuate it, remove stress marks or fix its spelling. A silent edit
   changes the audio and the user cannot see why. If something must change to be
   pronounceable, say so before synthesising.
6. **Keep the returned voice id** (`customVoiceFileId` / `voiceFileId`) and pass
   it on every later line of the same job.
7. **Name the voice you used** in one line of the reply, so the user can ask for
   it again.

## Pitfalls

- Never re-synthesise text that has already been voiced word for word. The
  user pays again for an identical file.
- Qwen's `instructions` works only for English and Chinese text on
  instruct-capable voices. In Russian it is silently dropped. Put the direction
  into the choice of voice and into the punctuation, or use Google.
- A voice *from a description* goes to `google-voice-design`, never to
  `qwen-voice-design`, which costs about 25× more for the same result.
- There are two kinds of "I don't like it". A wrong *voice* (age, gender,
  character) needs a different voice. A wrong *take* (a stumble, a misplaced
  stress, one phrase too fast) needs the same voice read again; the engines are
  not deterministic. Ask which one it is.
- To fix one phrase in an approved take, synthesise that phrase alone and splice
  it in. Do not re-read the whole script.
- Google refuses child voices with an empty success. Say so and offer an adult
  voice instead of retrying.
- If the audio is for a video where someone is *seen* speaking, do not
  synthesise it first. Lip-synced video engines want a short recording of the
  voice plus the line in the video prompt; audio added over finished footage
  does not match the mouth.
- Voice only the text the user asked for. Your own commentary costs the same
  per character.

## Done when

The user's text exists as audio in one voice, word for word as written, and the
reply names that voice. Extra variants of a take nobody complained about are not
"done"; they spend the user's balance.
