Realtime API
WebSocket
A transcription session is JSON over one WebSocket. Send session.update before your first audio, then input_audio_buffer.append with base64 audio in chunks of any size.
Connect
- WSS
wss://api.dotwave.ai/v1/realtimeThe server greets you withsession.created.
- From your server
Authorization: Bearer <API key>. The OpenAI SDK sends it for you.- From a browser
?token=<client secret>, with a secret fromPOST /v1/realtime/client_secrets.?model=- Optional, as the OpenAI SDK sends it; it must name the session’s model.
Without an SDK
This example also asks for shorter segments, completed after 800 ms of silence.
import asyncio, base64, json, os, time, wave
import websockets
URL = "wss://api.dotwave.ai/v1/realtime"
HEADERS = {"Authorization": f"Bearer {os.environ['DOTWAVE_API_KEY']}"}
async def send_wav(ws, path):
# 24 kHz mono PCM16, paced against a clock at the speed of speech.
with wave.open(path, "rb") as wav:
started, sent = time.monotonic(), 0
while chunk := wav.readframes(1920): # 1,920 samples at a time
await ws.send(json.dumps({
"type": "input_audio_buffer.append",
"audio": base64.b64encode(chunk).decode(),
}))
sent += len(chunk) // 2
await asyncio.sleep(max(0, started + sent / 24000 - time.monotonic()))
async def print_transcripts(ws):
async for raw in ws:
event = json.loads(raw)
kind = event.get("type")
if kind == "conversation.item.input_audio_transcription.delta":
print(event["delta"], end="", flush=True)
elif kind == "conversation.item.input_audio_transcription.completed":
print(f" [{event.get('audio_start_ms')}-{event.get('audio_end_ms')} ms]")
elif kind == "error":
print("\nerror:", event["error"]["message"])
async def main():
async with websockets.connect(URL, additional_headers=HEADERS) as ws:
await ws.send(json.dumps({
"type": "session.update",
"session": {
"type": "transcription",
"audio": {"input": {
"format": {"type": "audio/pcm", "rate": 24000},
"transcription": {
"model": "nemotron-asr-streaming",
"language": "pt-BR",
},
"turn_detection": {
"type": "server_vad",
"silence_duration_ms": 800,
},
}},
},
}))
receiver = asyncio.create_task(print_transcripts(ws))
await send_wav(ws, "fala-24k.wav")
await asyncio.sleep(2) # the last segment completes after 800 ms of silence
receiver.cancel()
asyncio.run(main())Audio formats
| Audio | audio.input.format |
|---|---|
| PCM16, 24 kHz | {"type": "audio/pcm", "rate": 24000}, the default |
| PCM16, 16 kHz | {"type": "audio/pcm", "rate": 16000} |
| G.711 μ-law, 8 kHz | {"type": "audio/pcmu"} |
| G.711 A-law, 8 kHz | {"type": "audio/pcma"} |
Audio is mono, and PCM16 is little-endian. Chunks can be any size. A format the model does not take is refused with unsupported_audio_format, and the session stays open.
Send at the speed of speech
Audio is transcribed against a clock that starts with your first audio. The server holds up to 640 ms of audio ahead of that clock; audio sent faster ends the session with close code 1013, and audio that arrives after its moment has passed is transcribed as silence. A live microphone meets both limits by itself. For a recording, pace the chunks against a clock, as the examples do, rather than sleeping a fixed time between sends.
Segments and timing
- Deltas
conversation.item.input_audio_transcription.deltacarries new text as soon as it is recognized. Text is final when it arrives: it is never revised.- Segments
- A segment completes after a pause, with
conversation.item.input_audio_transcription.completedand its full transcript.turn_detection.silence_duration_mssets the pause, from 400 ms to 3200 ms (the default). Withturn_detectionset tonull, a segment completes only when you sendinput_audio_buffer.commit. - Timestamps
- Each completed segment carries
audio_start_msandaudio_end_ms, its place in your audio. - Quiet sessions
- A session closes after 30 seconds without audio. To keep a quiet session open, send silence. A client secret can set
turn_detection.idle_timeout_msfrom 10000 to 300000. - Session length
- A session runs for up to 24 hours.
Close codes
On the socket, a problem with one event arrives as an error event, and the session continues. When a session ends, session.end gives the reason, and the socket closes with one of these codes:
| Close code | Meaning | What to do |
|---|---|---|
| 1000 | Closed normally: by you, after 30 seconds without audio, or after 24 hours. | Nothing, or open a new session. |
| 1011 | A server error. | Open a new session. |
| 1012 | The server is restarting. | Open a new session. |
| 1013 | Audio arrived faster than real time. | Open a new session, and pace the audio. |
| 4400 | A protocol error, such as a malformed message or an invalid parameter. | Fix the client. |
| 4401 | The credential is invalid, expired or already used. | Use your API key, or a new client secret. |
| 4413 | The service is not ready. | Retry shortly. |
| 4429 | No capacity when you connected. | Retry after a short wait. |