What has changed in S2.1 Pro compared with S2 Pro?
The official description states that S2.1 Pro improves quality, latency, and throughput, and lists it as the recommended production model. No specific degree of improvement is promised here; migration should be evaluated using your own scripts, voices, and actual call results.
Where should the model be specified?
Specify s2.1-pro through the HTTP request header model. If not specified, the default is s2-pro; the voice reference_id belongs to voice selection and is configured separately from the synthesis model.
How should I choose between saving a voice and one-time cloning?
For recurring characters or series content, use reference_id to reuse a voice; for one-time projects, use references to provide a recording and transcript. The two are mutually exclusive, and one-time cloning accepts only one reference sample.
How do I control speaking speed and output format?
prosody.speed=1.0 indicates the original speed, volume uses dB, and 0 means no volume change. You can choose mp3, wav, or pcm; the latter two both use WAV containers, and MP3 bitrates can be 64, 128, or 192.
How do I track completion results for long scripts?
After setting callback_url, first save task_id and started_at, then wait for the completion callback; you can also query by task ID. Receiving audio_url means it is ready for playback or editing.