Skip to main content

Try it live in the API playground

Drop text with the markers below into the text field and send a real request to hear the emotion.

Overview

Fish Audio models support 64+ emotional expressions and voice styles that can be controlled through text markers in your input. Add natural pauses, laughter, and other human-like elements to make speech more engaging and realistic.
This page shows S2 usage with [bracket] cues. If you use the legacy S1 model, wrap markers in parentheses instead — see S1 (legacy) syntax below for the full list, or the Models Overview.

How It Works

Add emotional or stylistic cues in square brackets within your text:
The S2 TTS models will interpret these markers and adjust the voice accordingly.

Complete Emotion Reference

Basic Emotions (24 expressions)

Advanced Emotions (25 expressions)

Tone Markers (5 expressions)

Control volume and intensity:

Audio Effects (10 expressions)

Add natural human sounds:

Special Effects

Additional markers for atmosphere and context: You can also use natural expressions like “Ha,ha,ha” for laughter without tags.

Usage Guidelines

Placement Rules

For S2:
  • Sentence-level emotion cues usually work best at the beginning of sentences
  • Tone controls can go anywhere in the text
  • Sound effects can go anywhere in the text
  • Bracket cues can use natural language descriptions and are not limited to a fixed set of tags
Correct:

Advanced Techniques

Combining Effects

You can layer multiple emotions for complex expressions:

Emotion Transitions

Create natural emotional progressions:

Background Effects

Add atmospheric sounds:

Intensity Modifiers

Fine-tune emotional intensity with descriptive modifiers:

Language Support

All 13 supported languages can use emotion markers. For sentence-level control, cues usually work best at the sentence start in these languages:
  • English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish, Portuguese

Best Practices

Do’s

  • Use one primary emotion per sentence
  • Test different emotion combinations
  • Match emotions to context logically
  • Add appropriate text after sound effects (e.g., “Ha ha” after laughing)
  • Use natural expressions when possible
  • Space out emotional changes for realism

Don’ts

  • Don’t overuse emotion tags in short text
  • Don’t mix conflicting emotions
  • Don’t make bracket descriptions so long that they interrupt readability
  • Don’t forget brackets
  • Don’t place sentence-level emotion cues far from the sentence they control

Common Use Cases

Customer Service

Storytelling

Educational Content

Marketing & Sales

Troubleshooting

Emotion Not Working?

  1. Check placement - Put the cue where the emotion or effect should begin
  2. Keep wording clear - Use concise natural language descriptions
  3. Use the right syntax - S2 cues use square brackets; S1 cues must use parentheses

Unnatural Sound?

  • Space out emotional changes
  • Use appropriate intensity
  • Test with different voices
  • Add context text after sound effects

Performance Notes

  • Emotion markers don’t count toward token limits
  • No additional latency for emotion processing
  • All emotions available on all pricing tiers
  • Maximum of 3 combined emotions per sentence recommended

Quick Reference Tables

Emotion Intensity Scale

Common Combinations

S1 (legacy) syntax

The default S2-Pro model uses [bracket] cues with free-form natural language. The previous-generation S1 model uses the same emotion names but requires (parentheses) and a fixed tag set:

See Also