curl --request POST \
--url https://geoff.ai/api/v1/audio/transcribe \
--header 'Authorization: Bearer <token>' \
--header 'X-Stack-Id: stk_jmqog66ha0mugmro' \
--header 'Content-Type: application/json' \
--data '{
"audio_url": "https://files.geoff.ai/audio/interview.mp3",
"granularity": "word"
}'
import requests
response = requests.post(
"https://geoff.ai/api/v1/audio/transcribe",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"audio_url": "https://files.geoff.ai/audio/interview.mp3",
"granularity": "word",
"language": "en",
},
)
data = response.json()["data"]
print(data["text"])
for w in data.get("words", []):
print(f' {w["start_s"]:.2f}–{w["end_s"]:.2f}: {w["word"]}')
{
"data": {
"text": "Welcome to the show. Today we're talking about...",
"has_speech": true,
"language": "en",
"duration_s": 124.3,
"granularity": "word",
"segments": [
{ "start_s": 0.0, "end_s": 30.0, "text": "Welcome to the show. Today we're..." }
],
"words": [
{ "word": "Welcome", "start_s": 0.12, "end_s": 0.58 },
{ "word": "to", "start_s": 0.61, "end_s": 0.73 }
]
},
"trace_id": "04ede0ab069fb1ba8be5156a24b1e081",
"extra_info": {
"audio_megabytes": 2.4
}
}
Speech to Text
Transcribe Audio
Convert speech audio to text with optional segment- or word-level timestamps. Auto-detects language.
POST
/
v1
/
audio
/
transcribe
curl --request POST \
--url https://geoff.ai/api/v1/audio/transcribe \
--header 'Authorization: Bearer <token>' \
--header 'X-Stack-Id: stk_jmqog66ha0mugmro' \
--header 'Content-Type: application/json' \
--data '{
"audio_url": "https://files.geoff.ai/audio/interview.mp3",
"granularity": "word"
}'
import requests
response = requests.post(
"https://geoff.ai/api/v1/audio/transcribe",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"audio_url": "https://files.geoff.ai/audio/interview.mp3",
"granularity": "word",
"language": "en",
},
)
data = response.json()["data"]
print(data["text"])
for w in data.get("words", []):
print(f' {w["start_s"]:.2f}–{w["end_s"]:.2f}: {w["word"]}')
{
"data": {
"text": "Welcome to the show. Today we're talking about...",
"has_speech": true,
"language": "en",
"duration_s": 124.3,
"granularity": "word",
"segments": [
{ "start_s": 0.0, "end_s": 30.0, "text": "Welcome to the show. Today we're..." }
],
"words": [
{ "word": "Welcome", "start_s": 0.12, "end_s": 0.58 },
{ "word": "to", "start_s": 0.61, "end_s": 0.73 }
]
},
"trace_id": "04ede0ab069fb1ba8be5156a24b1e081",
"extra_info": {
"audio_megabytes": 2.4
}
}
Speech-to-text with long-form support. Returns the transcript plus
optional segment timestamps (default) or word-level timestamps. Long
audio is segmented server-side and stitched in the response.
Authorization
string
required
Bearer token.
Bearer API_key.Request Body
string
required
URL of the audio to transcribe. Any common codec accepted
(mp3 / wav / m4a / ogg / mp4); auto-converted to 16 kHz mono.
string
ISO 639-1 language hint. Omit for auto-detect.
string
segment (default) returns ~30 s buckets; word returns per-word
start/end timestamps when the backing engine supports it.boolean
When
true, returns the English translation alongside the original
transcript. Default: false.curl --request POST \
--url https://geoff.ai/api/v1/audio/transcribe \
--header 'Authorization: Bearer <token>' \
--header 'X-Stack-Id: stk_jmqog66ha0mugmro' \
--header 'Content-Type: application/json' \
--data '{
"audio_url": "https://files.geoff.ai/audio/interview.mp3",
"granularity": "word"
}'
import requests
response = requests.post(
"https://geoff.ai/api/v1/audio/transcribe",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"audio_url": "https://files.geoff.ai/audio/interview.mp3",
"granularity": "word",
"language": "en",
},
)
data = response.json()["data"]
print(data["text"])
for w in data.get("words", []):
print(f' {w["start_s"]:.2f}–{w["end_s"]:.2f}: {w["word"]}')
{
"data": {
"text": "Welcome to the show. Today we're talking about...",
"has_speech": true,
"language": "en",
"duration_s": 124.3,
"granularity": "word",
"segments": [
{ "start_s": 0.0, "end_s": 30.0, "text": "Welcome to the show. Today we're..." }
],
"words": [
{ "word": "Welcome", "start_s": 0.12, "end_s": 0.58 },
{ "word": "to", "start_s": 0.61, "end_s": 0.73 }
]
},
"trace_id": "04ede0ab069fb1ba8be5156a24b1e081",
"extra_info": {
"audio_megabytes": 2.4
}
}
Notes
has_speech: falseindicates the engine detected no speech in the audio (e.g. instrumental music, silence). Thetextfield will be empty in that case.granularity: 'word'may fall back to segment-level timing when the backing engine doesn’t expose word timestamps; the response includesgranularity_fallbackwith the reason.translate: truekeeps the original-languagetextand adds atranslationfield with the English version.
string
default:"stk_jmqog66ha0mugmro"
required
Supplied automatically by the documentation playground.