All models

s2.1-pro

Fish AudioAudio
Get your API key
s2.1-pro

Enhance natural voiceover workflows with next-generation voice capabilities

Fish Audio S2.1 Pro is the production voice model recommended by Fish Audio for text-to-speech and voice cloning. Official materials state that it improves quality, latency, and throughput over S2 Pro, making it suitable for applications that require natural voices and ongoing voiceover production. This platform explicitly selects this model through model: s2.1-pro, while retaining the interfaces for text, voice references, output formats, and task callbacks.

Fish AudioModel brand
Text-to-speechModel type
Voice cloningTask capability
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API hostapi.acedata.cloud
models2.1-pro

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Clarify the model, inputs and outputs, and invocation method before selecting.

Model selection
HTTP request header model: s2.1-pro
Platform API
POST /fish/tts; text is a required input
Audio delivery
Returns audio_url synchronously; supports asynchronous callback_url callbacks
Output formats
mp3, wav, pcm; both wav and pcm return a WAV container
Voice reference
reference_id or a single references sample; the two are mutually exclusive
Prosody control
prosody.speed controls speech rate, and prosody.volume controls volume gain

The S2 series can use [bracket] natural-language expression prompts in text, such as [whisper]; actual results should be confirmed by listening. Capability descriptions combine public model materials and this platform's documentation; actual parameters, outputs, and billing rules are subject to the corresponding API and pricing sections.

Core capabilities

Learn about the voice tasks that s2.1-pro is suited to handle.

Production-oriented voice upgrade

Officially, S2.1 Pro is positioned as the recommended production model, with improvements in quality, latency, and throughput over S2 Pro. When selecting a model, evaluate these changes using your own scripts and voice samples; do not turn official directional descriptions into fixed latency or concurrency commitments.

Reuse existing voice assets

You can use an existing voice through reference_id, or provide a reference recording and an accurate verbatim transcript in a single request. Reusing voice materials helps compare old and new models, but generated results still need to be checked for tone, pronunciation, and character consistency.

Organize production tasks in one API

Submit synthesis using text and format settings, then retrieve audio_url when complete. For long scripts, you can configure callback_url, save the task ID first, and then receive the completed result; applications do not need to present a pending task as delivered audio.

Use Cases

Choose based on specific content and delivery methods.

Scalable Content Voiceovers

For frequently updated knowledge content, product explanations, or product tutorials, first select representative scripts for evaluation, then gradually migrate the production workflow. Save model and voice information to make it easier to identify settings when differences in expression occur.

Application Interactions with Natural Voices

Convert navigation, explanations, and help text into playable audio, focusing on the clarity and waiting experience that real users hear. This entry point delivers audio links; real-time conversational applications require separate design for playback and buffering.

Upgrade Evaluation for Existing Voiceover Workflows

Compare S2 Pro and S2.1 Pro using the same voice and text, checking proper nouns, long sentences, pauses, and post-processing effort. Decide the migration scope through small-batch listening tests, avoiding direct bulk replacement of existing finished content.

How to Choose This Model

Compare based on scripts, voices, and production costs.

Explicitly Use the New-Generation Model

This entry point still defaults to s2-pro; to use S2.1 Pro, set model: s2.1-pro. Do not assume that a request with no specified model has already used it merely based on its official positioning as the “recommended production model.”

Measure Upgrade Benefits with Your Own Samples

Official materials do not provide a fixed improvement percentage that can be guaranteed for this page. Compare audio quality, task wait time, review, and revision costs before deciding whether to replace an existing workflow; S1 can also remain a candidate for stable long-form reading.

Getting Started: Evaluate the Actual Benefits of the New Model for Series Voiceovers

Arrange the inputs first, then connect them to the corresponding application workflow.

Prepare Inputs

Keep the same voice, script, format, and speaking rate fixed, and select short samples containing numbers, long sentences, and emotional changes.

Organize Calls and Follow-Up Workflows

Submit text to /fish/tts, select s2.1-pro through the model request header, and provide either reference_id or a single recording reference. After completion, listen using audio_url; for series content, save the model, voice, speaking rate, and format.

Practical Task Example: Evaluate the Actual Benefits of the New Model for Series Voiceovers

Design tasks directly from the following inputs and acceptance priorities.

Suggested Task

Explicitly specify model: s2.1-pro, conduct a side-by-side listening test with s2-pro using the same script, and record the actual wait time.

Key Checks

Verify naturalness, pronunciation, sentence endings, and stability, and compare rework volume; official quality and efficiency improvements do not constitute a fixed latency commitment for this platform.

Usage Limits

Learn about synthesis methods and delivery scope.

  • This endpoint is for text-to-speech and is not equivalent to speech recognition, music generation, or native real-time audio streaming. Long-form content should be produced in segments and reviewed for numbers, abbreviations, and proper nouns.
  • One-time cloning accepts only one publicly accessible HTTPS MP3/WAV recording sample and its accurate transcript. Base64, data URIs, or URLs with credentials are not accepted; long-term voice profiles are not automatically saved.
  • format=pcm returns a WAV container and cannot be handled directly as raw PCM bytes; opus is not supported. For supported generation parameters and actual billing, see this platform's API documentation and pricing.

Frequently Asked Questions

Answers to common questions about using s2.1-pro.

What has changed in S2.1 Pro compared with S2 Pro?

The official description states that S2.1 Pro improves quality, latency, and throughput, and lists it as the recommended production model. No specific degree of improvement is promised here; migration should be evaluated using your own scripts, voices, and actual call results.

Where should the model be specified?

Specify s2.1-pro through the HTTP request header model. If not specified, the default is s2-pro; the voice reference_id belongs to voice selection and is configured separately from the synthesis model.

How should I choose between saving a voice and one-time cloning?

For recurring characters or series content, use reference_id to reuse a voice; for one-time projects, use references to provide a recording and transcript. The two are mutually exclusive, and one-time cloning accepts only one reference sample.

How do I control speaking speed and output format?

prosody.speed=1.0 indicates the original speed, volume uses dB, and 0 means no volume change. You can choose mp3, wav, or pcm; the latter two both use WAV containers, and MP3 bitrates can be 64, 128, or 192.

How do I track completion results for long scripts?

After setting callback_url, first save task_id and started_at, then wait for the completion callback; you can also query by task ID. Receiving audio_url means it is ready for playback or editing.