Design any voice without writing the prompt yourself
Claude × Fish Audio
Hey, it's Adam. The soccer commentator, the children's storyteller, the tiny chirpy bird — I got all three back in one conversation, without leaving Claude, without browsing a voice library, and without writing a single detailed voice description myself. That last part is the actual trick. Voice design lives or dies on how well the voice is described, and describing voices is a skill most people don't have. So you don't do it — you make Claude do it. This guide is that whole setup, plus the prompt library. Fish Audio partners with me on this content; the workflow is the one I run daily. You commented FISH, so here's everything.
Delegate the prompt, not just the generation 🎙
Here's the thing nobody says about designing voices: the model is rarely the bottleneck. Your description is.
"A deep male voice" gets you a shrug. What actually works reads like a casting note — persona, age, register, the physical texture of the voice, the pacing, the situation it's speaking in. Four or five sentences of specific, audible detail. Most people don't write that, so they conclude the tool is mediocre and go back to browsing a library.
Connect Fish Audio to Claude and that problem disappears, because you stop writing the description. You say "a tired detective at 3am" — Claude expands it into the full spec, writes a preview line that proves it, sends it to Fish Audio, and hands you back candidates to pick from. One line in, finished voices out, no app-switching.
That's the whole workflow: you bring the idea, Claude brings the vocabulary, Fish Audio brings the voice.
Connect it — one command 🔌
This is genuinely the setup step in full.
claude mcp add --transport http fish-audio https://api.fish.audio/mcpThen run /mcp inside Claude Code to sign in.
In the Claude app: Settings → Connectors → Add custom connector → paste https://api.fish.audio/mcp, and it walks you through sign-in.
Worth knowing: the MCP connection signs you in with your Fish Audio account — you don't paste an API key. (An API key is only needed if you're calling the REST API directly from your own code, which the last section covers.) Once connected, Claude can search the voice library, generate speech, and transcribe audio using your account.
Test: ask Claude "Generate a short greeting with a warm English voice." If audio comes back, you're done setting up and everything below just works.
Try Fish Audio — free development tier available
Give Claude the job description, once 📋
Paste this into Claude once — into a project's instructions, or a saved skill, so it persists. Everything after this is one-liners.
You are my voice director for Fish Audio. When I give you a short voice idea — even three words — do all of this without asking me questions:1. Expand it into a full voice-design instruction, 3-5 sentences of natural language, covering: - Persona and context: who is speaking, where, and why - Profile: age, gender, register, accent - Timbre and texture: physical audible qualities (breathy, gravelly, nasal, velvety, thin, resonant) — never personality words like "nice" or "professional" - Pacing and delivery: speed, pauses, articulation - Emotional colour and what kind of script it's built for2. Write a preview line (1-2 sentences) that ONLY this character would plausibly say. Never generic filler like "Have a wonderful day." Add inline emotion and delivery cues to the preview line so I hear the performance, not just the timbre.3. Generate it through Fish Audio, returning 3 candidates in one request so I can compare.4. Show me the instruction text you used, so I can edit and re-run.Then wait. If I say "colder", "older", "less polished", "more tired" — adjust the instruction and regenerate, don't start over.
That last rule is the one that saves you time. Voice work is iterative, and you want to nudge a description, not rewrite it from scratch every round.
Test: say "a tired detective at 3am" and check what Claude wrote before it generated. If the instruction it produced doesn't contain physical texture words, tell it so — it'll correct and stay corrected.
The prompt library 🎭
These are the descriptions Claude should be producing for you. Use them as-is, or as the quality bar for what you accept back.
Instruction: A live football commentator in the final seconds of a decisive match. Male, mid-forties, bright forward-placed tenor that thins and strains at the top when he pushes volume. Pacing accelerates as play builds, words crowding together, then breaks open into long held vowels. Euphoric and slightly hoarse — a man who has been shouting for ninety minutes and has one shout left.Preview line: He's through — nobody's tracking him — and OH, he's done it, right at the death!
Instruction: Someone reading a bedtime story to a child who is nearly asleep. Female, thirties, warm mid-range voice with a soft rounded timbre and a gentle smile audible in the vowels. Pacing is slow and lilting with generous pauses at the end of each phrase, and she leans into the interesting words rather than rushing past them. Tender and unhurried, built for narration a child will fall asleep to.Preview line: And the little fox stopped right at the edge of the water... and looked at the moon sitting there in it.
Instruction: A very small bird character with an outsized personality. Extremely high, bright, thin soprano with a light nasal chirp and quick flutters at the ends of words. Rapid, bouncy pacing with almost no pauses, words tumbling out in excitable bursts. Relentlessly cheerful and slightly nosy, built for animated character dialogue.Preview line: Oh! Oh oh oh — you're new here, aren't you? I've seen everyone, and I have definitely not seen you!
Instruction: A detective at the end of a very long night shift. Male, late forties, low grainy baritone with a dry rasp and audible fatigue softening the consonants. Pacing is slow and uneven, with pauses that feel like thinking rather than timing, and sentences that trail off before they finish. Weary and quietly certain, built for noir narration and game dialogue.Preview line: Nothing about this room is wrong. That's exactly what's wrong with it.
Instruction: A support agent who has heard this exact problem before and isn't annoyed by it. Female, thirties, warm mid alto with a clean rounded timbre and no broadcast polish. Even, calm pacing that slows deliberately on instructions, with crisp articulation on numbers and names. Patient and competent, built for phone support and real-time voice agents where clarity matters more than character.Preview line: I can see exactly what happened, and it's a quick fix. Give me about two minutes.
Instruction: A creator explaining a tool they actually use. Late twenties, neutral accent, clear mid-range voice with a slightly dry, matter-of-fact texture and no announcer polish. Brisk conversational pacing with short pauses before the point rather than after it. Direct and unimpressed by hype, built for vertical short-form where the viewer decides in two seconds whether to stay.Preview line: Most people are using this completely wrong. It takes about a minute to fix.
Instruction: A villain who never needs to raise his voice. Male, fifties, low resonant baritone, smooth on the surface with a cold edge underneath. Slow unhurried pacing with long confident pauses, every word fully finished, articulation precise to the point of clipped. Amused and entirely unthreatened, built for game dialogue where menace comes from calm.Preview line: You made it further than the others. That isn't a compliment — it just takes slightly longer this way.
Instruction: An elderly documentary narrator with decades of broadcasting behind him. Male, seventies, low gravelled bass with a dry papery rasp at the edges of long vowels. Unhurried, weighted pacing with deliberate pauses before key facts, landing consonants softly rather than crisply. Authoritative but warm — a man telling you something he genuinely finds remarkable.Preview line: For eleven months of the year, this valley holds nothing at all. And then, almost overnight, it holds everything.
🔑 The pattern in all eight: persona, then physical texture, then pacing, then what it's built for — and a preview line that would sound absurd in anyone else's mouth. That last part is the real test. If the line works for a generic voice, the description isn't specific enough yet.
The dials worth knowing 🎚
Ask Claude for these by name and it'll pass them through:
- Candidates per request — you can ask for several voices from one instruction and compare. This is why "three voices" comes back at once, and it costs no more than asking for one.
- Language — set it explicitly for the voice you want, rather than letting it infer.
- Speed — adjust pacing without rewriting the description.
- Seed — the underrated one. Note the seed of a keeper and you can return to that neighbourhood later instead of hunting blind.
- Guidance — how tightly it sticks to your instruction versus interpreting freely. Nudge it when results feel either too literal or too loose.
Test: generate three candidates, keep one, then regenerate that same instruction with the seed noted. Different results tell you how much variance to expect before you commit a voice to a project.
Keep the one you like 💾
Candidates are exploratory — they're for finding the voice, not for building on. Once one is right, save it as a proper voice model so it becomes a stable, reusable voice with an ID you can call in every future generation.
This matters more than it sounds. Regenerating from the same description later gets you a cousin, not the same character. If a voice is going into a series, a game, or an agent, save it the moment it's right.
🔑 What makes voice design work
- 1Don't write the description — make Claude write it, then edit.
- 2Physical, audible adjectives only. "Gravelly," never "professional."
- 3The preview line must be one only that character would say.
- 4Three candidates per idea. Compare, don't accept the first.
- 5Nudge the instruction, don't restart it. "Colder. Older. Less polished."
- 6Save the keeper as a voice model. Immediately.
- 7Note your seeds on anything you might need to return to.
The honest bits ⚠️
- What's actually free, precisely. Fish Audio publishes a free development model — the same quality as their production model at $0, under fair-use limits, with no latency or uptime guarantees, aimed at testing, prototyping and smaller businesses. Production workloads sit on the paid model. So "free for developers" is real, and it's the tier you'll be on while you experiment.
- It's exploratory by design. Candidates are short and variable. The same instruction won't give you the same voice twice, which is a feature while you're hunting and a problem once you've decided — hence saving keepers.
- Expect a few rounds. Two or three nudges to land a character is normal, not failure. Budget the iterations rather than judging the first result.
- Designed voices have no owner. Cloned ones do. Designing from a description creates something new. Cloning needs a voice you own or have written permission to use — that's not a grey area, and it's the one rule here with real consequences.
- Disclose where it matters. Narration nobody assumes is human is fine. Anything that could be taken for a specific real person speaking needs to say what it is.
- Going beyond Claude. If you want this in your own product rather than in a conversation, the same thing runs over the REST API — that's where you'd create an actual API key, and where the parameters above become request fields.
Links 🔗
Fish Audio