Speech (Text-to-Speech)
Synthesize natural-sounding speech from text. The endpoint streams the audio bytes directly back in the chosen response_format — there is no JSON wrapper.
POST
/v1/audio/speechRequest
📋JSON body
1{
2 "model": "your-tts-model",
3 "input": "Hello! Welcome to Pawan.Krd.",
4 "voice": "alloy",
5 "response_format": "mp3",
6 "speed": 1.0
7}Parameters
modelstringrequired
TTS model id (type: text-to-speech).
inputstringrequired
Text to synthesize. Length limits depend on the model.
voicestringrequired
Voice id from the model's voice enum.
response_formatstring
"mp3" | "opus" | "aac" | "flac" | "wav" | "pcm".
speednumber
Playback speed multiplier (typically 0.25–4.0).
sample_rateinteger
Output sample rate in Hz (when supported).
voice_referencefile
Reference audio for voice cloning (multipart only).
voice_reference_textstring
Optional transcript of the reference audio for cloning models.
Saving the audio
💻curl
1curl https://api.pawan.krd/v1/audio/speech \
2 -H "Authorization: Bearer pk-your_key" \
3 -H "Content-Type: application/json" \
4 -o speech.mp3 \
5 -d '{
6 "model": "your-tts-model",
7 "input": "Hello world",
8 "voice": "alloy"
9 }'Voice cloning
Models with the voiceReference capability accept a reference audio clip via multipart upload. Optionally include the transcript via voice_reference_text.
💻curl (multipart)
1curl https://api.pawan.krd/v1/audio/speech \
2 -H "Authorization: Bearer pk-your_key" \
3 -F model="your-clone-model" \
4 -F input="Hello in my own voice" \
5 -F voice_reference=@reference.wav \
6 -F voice_reference_text="The exact transcript of reference.wav" \
7 -o speech.mp3✓
Discovering voices
Available voices for a model live in
parameters.voice.values on the model object — fetch it via GET /v1/models/{modelId}.