Transcriptions (Speech-to-Text)
Convert audio to text. STT models can transcribe in the source language; models with the translation capability can additionally translate non-English audio into English text via the dedicated translations endpoint.
Transcribe
POST
/v1/audio/transcriptions💻curl
1curl https://api.pawan.krd/v1/audio/transcriptions \
2 -H "Authorization: Bearer pk-your_key" \
3 -F model="your-stt-model" \
4 -F file=@audio.mp3 \
5 -F response_format="verbose_json" \
6 -F language="en" \
7 -F "timestamp_granularities[]=word"Parameters
modelstringrequired
STT model id (type: speech-to-text).
filefilerequired
Audio file (multipart). Up to 200 MB.
languagestring
ISO-639-1 language hint, e.g. "en". Improves accuracy when known.
promptstring
Optional context hint to bias the decoder (when supported).
response_formatstring
"json" | "text" | "srt" | "verbose_json" | "vtt".
temperaturenumber
Sampling temperature for the decoder.
timestamp_granularities[]string[]
"segment" and/or "word" — requires verbose_json.
Verbose response
With response_format=verbose_json and timestamp_granularities[]=word, you get rich segment- and word-level timestamps:
📋200 OK
1{
2 "task": "transcribe",
3 "language": "en",
4 "duration": 12.34,
5 "text": "Hello and welcome to the show.",
6 "segments": [
7 { "id": 0, "start": 0.0, "end": 2.1, "text": "Hello" },
8 { "id": 1, "start": 2.1, "end": 4.0, "text": "and welcome to the show." }
9 ],
10 "words": [
11 { "word": "Hello", "start": 0.0, "end": 0.4 }
12 ]
13}Translate
POST
/v1/audio/translationsSame shape as transcribe, but the output is always English. Only models with the translation capability accept this endpoint.
💻curl
1curl https://api.pawan.krd/v1/audio/translations \
2 -H "Authorization: Bearer pk-your_key" \
3 -F model="your-stt-model" \
4 -F file=@spanish.mp3 \
5 -F response_format="json"✓
Subtitle workflows
Use
response_format=srt or vtt to get a subtitle-ready string back — you can drop the body straight into a .srt / .vtt file.