Docs / Realtime API / WebSocket

Realtime API

WebSocket

A transcription session is JSON over one WebSocket. Send session.update before your first audio, then input_audio_buffer.append with base64 audio in chunks of any size.

Connect

WSS
wss://api.dotwave.ai/v1/realtimeThe server greets you with session.created.
From your server
Authorization: Bearer <API key>. The OpenAI SDK sends it for you.
From a browser
?token=<client secret>, with a secret from POST /v1/realtime/client_secrets.
?model=
Optional, as the OpenAI SDK sends it; it must name the session’s model.

Without an SDK

This example also asks for shorter segments, completed after 800 ms of silence.

import asyncio, base64, json, os, time, wave
import websockets

URL = "wss://api.dotwave.ai/v1/realtime"
HEADERS = {"Authorization": f"Bearer {os.environ['DOTWAVE_API_KEY']}"}

async def send_wav(ws, path):
    # 24 kHz mono PCM16, paced against a clock at the speed of speech.
    with wave.open(path, "rb") as wav:
        started, sent = time.monotonic(), 0
        while chunk := wav.readframes(1920):  # 1,920 samples at a time
            await ws.send(json.dumps({
                "type": "input_audio_buffer.append",
                "audio": base64.b64encode(chunk).decode(),
            }))
            sent += len(chunk) // 2
            await asyncio.sleep(max(0, started + sent / 24000 - time.monotonic()))

async def print_transcripts(ws):
    async for raw in ws:
        event = json.loads(raw)
        kind = event.get("type")
        if kind == "conversation.item.input_audio_transcription.delta":
            print(event["delta"], end="", flush=True)
        elif kind == "conversation.item.input_audio_transcription.completed":
            print(f"  [{event.get('audio_start_ms')}-{event.get('audio_end_ms')} ms]")
        elif kind == "error":
            print("\nerror:", event["error"]["message"])

async def main():
    async with websockets.connect(URL, additional_headers=HEADERS) as ws:
        await ws.send(json.dumps({
            "type": "session.update",
            "session": {
                "type": "transcription",
                "audio": {"input": {
                    "format": {"type": "audio/pcm", "rate": 24000},
                    "transcription": {
                        "model": "nemotron-asr-streaming",
                        "language": "pt-BR",
                    },
                    "turn_detection": {
                        "type": "server_vad",
                        "silence_duration_ms": 800,
                    },
                }},
            },
        }))
        receiver = asyncio.create_task(print_transcripts(ws))
        await send_wav(ws, "fala-24k.wav")
        await asyncio.sleep(2)  # the last segment completes after 800 ms of silence
        receiver.cancel()

asyncio.run(main())

Audio formats

Audioaudio.input.format
PCM16, 24 kHz{"type": "audio/pcm", "rate": 24000}, the default
PCM16, 16 kHz{"type": "audio/pcm", "rate": 16000}
G.711 μ-law, 8 kHz{"type": "audio/pcmu"}
G.711 A-law, 8 kHz{"type": "audio/pcma"}

Audio is mono, and PCM16 is little-endian. Chunks can be any size. A format the model does not take is refused with unsupported_audio_format, and the session stays open.

Send at the speed of speech

Audio is transcribed against a clock that starts with your first audio. The server holds up to 640 ms of audio ahead of that clock; audio sent faster ends the session with close code 1013, and audio that arrives after its moment has passed is transcribed as silence. A live microphone meets both limits by itself. For a recording, pace the chunks against a clock, as the examples do, rather than sleeping a fixed time between sends.

Segments and timing

Deltas
conversation.item.input_audio_transcription.delta carries new text as soon as it is recognized. Text is final when it arrives: it is never revised.
Segments
A segment completes after a pause, with conversation.item.input_audio_transcription.completed and its full transcript. turn_detection.silence_duration_ms sets the pause, from 400 ms to 3200 ms (the default). With turn_detection set to null, a segment completes only when you send input_audio_buffer.commit.
Timestamps
Each completed segment carries audio_start_ms and audio_end_ms, its place in your audio.
Quiet sessions
A session closes after 30 seconds without audio. To keep a quiet session open, send silence. A client secret can set turn_detection.idle_timeout_ms from 10000 to 300000.
Session length
A session runs for up to 24 hours.

Close codes

On the socket, a problem with one event arrives as an error event, and the session continues. When a session ends, session.end gives the reason, and the socket closes with one of these codes:

Close codeMeaningWhat to do
1000Closed normally: by you, after 30 seconds without audio, or after 24 hours.Nothing, or open a new session.
1011A server error.Open a new session.
1012The server is restarting.Open a new session.
1013Audio arrived faster than real time.Open a new session, and pace the audio.
4400A protocol error, such as a malformed message or an invalid parameter.Fix the client.
4401The credential is invalid, expired or already used.Use your API key, or a new client secret.
4413The service is not ready.Retry shortly.
4429No capacity when you connected.Retry after a short wait.