Навык для агента · voiceover-tts
Навык озвучки для ИИ-агента: текст в голос через MCP
SKILL.md для Claude, Codex и других агентов: закадровый голос, начитка и диалоги на Qwen3-TTS и Gemini TTS через MCP Twelver — один голос на всю работу, каждый дубль оплачивается один раз.
Что делает агент
- Выбирает движок по языку и задаче: Qwen по умолчанию, Gemini для режиссуры и диалогов.
- Отправляет сценарий дословно одним вызовом и сохраняет голос для следующих реплик.
- Отличает «не тот голос» от «не тот дубль» и не озвучивает один текст дважды.
- Проектирует голос по описанию дешёвым способом — через Gemini.
Запросы, на которые он срабатывает
- «Озвучь этот рекламный текст энергичным мужским голосом»
- «Сделай начитку главы аудиокниги и сохрани голос для следующих»
- «Нужен диалог двух ведущих для подкаста»
Установка
- Скачайте SKILL.md и положите его в папку с именем
voiceover-ttsв каталоге навыков вашего агента — для Claude Code это~/.claude/skills/voiceover-tts/SKILL.md. - Подключите эти MCP-серверы Twelver (вход по OAuth или API-ключ):
https://twelver.ru/api/mcp/v1/qwen-audio-speechhttps://twelver.ru/api/mcp/v1/google-audio-speech
- Попросите агента о задаче своими словами — навык он загрузит сам.
Файл навыка
---
name: voiceover-tts
description: Turn a script into a finished voice-over — narration, ad reads, podcast lines, audiobook passages, Russian or any of 30+ languages — with Qwen3-TTS and Gemini TTS through the Twelver MCP server. Use when the user asks to voice, narrate or read a text aloud, wants an MP3/WAV of a script, needs the same voice again for a follow-up line, wants a voice designed from a description, or asks for a two-speaker dialogue (озвучка текста, закадровый голос, начитка, синтез речи).
license: MIT
metadata:
author: Twelver
homepage: https://twelver.ru/skills/ru/voiceover-tts/
version: "1"
---
# Voice-over from a script
Two engines, one job: get the user's exact words recorded in the right voice, and
pay for each take once.
## Connect the tools
Twelver runs one remote MCP server per toolset (Streamable HTTP, no local install):
- Qwen3-TTS: `https://twelver.ru/api/mcp/v1/qwen-audio-speech`
- Gemini TTS: `https://twelver.ru/api/mcp/v1/google-audio-speech`
Use `twelver.app` instead of `twelver.ru` for accounts registered on
twelver.app; they are separate account systems. Auth is OAuth sign-in (needs a
registered account) or `Authorization: Bearer sk-…` with a key from
`https://twelver.ru/api-keys`. Connect Qwen first, since it handles most jobs.
Each tool returns a text summary plus a link to the audio file. Give the user the
link and do not try to inline the audio.
## Tools
| Tool | Use it for |
|---|---|
| `qwen-list-voices` / `google-list-voices` | The user's existing voices, free. Call before creating a new one |
| `qwen-text-to-speech` | Default read. zh/en/fr/de/ru/it/es/pt/ja/ko only; up to 10,000 chars, split and stitched automatically |
| `qwen-voice-clone` | A voice from the user's own 3–60 s recording |
| `google-text-to-speech` | Other languages, directed delivery (emotion, whisper, character) via `deliveryNotes`, tags like `<laugh>`, `<sigh>`, `<short pause>` in the text; up to 7,000 chars, not split |
| `google-voice-design` | A new voice from a text description, with a ~10 s preview |
| `google-text-to-speech-dialogue` | Two speakers in one track (Google only, exactly two) |
## Procedure
1. **Confirm the language of the text before the first take.** The engine
choice depends on it.
2. **List existing voices first** if the user has voiced anything before (both
list tools are free), and offer those by description.
3. **Pick one engine and bill once.** Never synthesise the same text on both to
compare. Qwen is the default: about 4× cheaper per character than Google.
Qwen already has a voice for each common house style: Neil for a news desk,
Ethan for an energetic ad, Cherry for warm and conversational, Bellona for
authoritative.
4. **Switch to Google** when the language is not one Qwen speaks, when it is a
two-speaker dialogue, or when the delivery has to be *directed*: an emotion,
a whisper, a named character. "Narrator", "friendly" or "professional" are a
choice of voice, not direction.
5. **Send the script exactly as the user wrote it, in one call.** Do not
re-punctuate it, remove stress marks or fix its spelling. A silent edit
changes the audio and the user cannot see why. If something must change to be
pronounceable, say so before synthesising.
6. **Keep the returned voice id** (`customVoiceFileId` / `voiceFileId`) and pass
it on every later line of the same job.
7. **Name the voice you used** in one line of the reply, so the user can ask for
it again.
## Pitfalls
- Never re-synthesise text that has already been voiced word for word. The
user pays again for an identical file.
- Qwen's `instructions` works only for English and Chinese text on
instruct-capable voices. In Russian it is silently dropped. Put the direction
into the choice of voice and into the punctuation, or use Google.
- A voice *from a description* goes to `google-voice-design`, never to
`qwen-voice-design`, which costs about 25× more for the same result.
- There are two kinds of "I don't like it". A wrong *voice* (age, gender,
character) needs a different voice. A wrong *take* (a stumble, a misplaced
stress, one phrase too fast) needs the same voice read again; the engines are
not deterministic. Ask which one it is.
- To fix one phrase in an approved take, synthesise that phrase alone and splice
it in. Do not re-read the whole script.
- Google refuses child voices with an empty success. Say so and offer an adult
voice instead of retrying.
- If the audio is for a video where someone is *seen* speaking, do not
synthesise it first. Lip-synced video engines want a short recording of the
voice plus the line in the video prompt; audio added over finished footage
does not match the mouth.
- Voice only the text the user asked for. Your own commentary costs the same
per character.
## Done when
The user's text exists as audio in one voice, word for word as written, and the
reply names that voice. Extra variants of a take nobody complained about are not
"done"; they spend the user's balance.