# WebSocket

A transcription session is JSON over one WebSocket. Send `session.update` before your first audio, then `input_audio_buffer.append` with base64 audio in chunks of any size.

## Connect

- `WSS wss://api.dotwave.ai/v1/realtime`: The server greets you with `session.created`.

- **From your server**: `Authorization: Bearer <API key>`. The OpenAI SDK sends it for you.
- **From a browser**: `?token=<client secret>`, with a secret from [`POST /v1/realtime/client_secrets`](https://dotwave.ai/docs/realtime/client-secrets/).
- **`?model=`**: Optional, as the OpenAI SDK sends it; it must name the session’s model.

## Without an SDK

This example also asks for shorter segments, completed after 800 ms of silence.

```python
import asyncio, base64, json, os, time, wave
import websockets

URL = "wss://api.dotwave.ai/v1/realtime"
HEADERS = {"Authorization": f"Bearer {os.environ['DOTWAVE_API_KEY']}"}

async def send_wav(ws, path):
    # 24 kHz mono PCM16, paced against a clock at the speed of speech.
    with wave.open(path, "rb") as wav:
        started, sent = time.monotonic(), 0
        while chunk := wav.readframes(1920):  # 1,920 samples at a time
            await ws.send(json.dumps({
                "type": "input_audio_buffer.append",
                "audio": base64.b64encode(chunk).decode(),
            }))
            sent += len(chunk) // 2
            await asyncio.sleep(max(0, started + sent / 24000 - time.monotonic()))

async def print_transcripts(ws):
    async for raw in ws:
        event = json.loads(raw)
        kind = event.get("type")
        if kind == "conversation.item.input_audio_transcription.delta":
            print(event["delta"], end="", flush=True)
        elif kind == "conversation.item.input_audio_transcription.completed":
            print(f"  [{event.get('audio_start_ms')}-{event.get('audio_end_ms')} ms]")
        elif kind == "error":
            print("\nerror:", event["error"]["message"])

async def main():
    async with websockets.connect(URL, additional_headers=HEADERS) as ws:
        await ws.send(json.dumps({
            "type": "session.update",
            "session": {
                "type": "transcription",
                "audio": {"input": {
                    "format": {"type": "audio/pcm", "rate": 24000},
                    "transcription": {
                        "model": "nemotron-asr-streaming",
                        "language": "pt-BR",
                    },
                    "turn_detection": {
                        "type": "server_vad",
                        "silence_duration_ms": 800,
                    },
                }},
            },
        }))
        receiver = asyncio.create_task(print_transcripts(ws))
        await send_wav(ws, "fala-24k.wav")
        await asyncio.sleep(2)  # the last segment completes after 800 ms of silence
        receiver.cancel()

asyncio.run(main())
```

## Audio formats

| Audio | `audio.input.format` |
| --- | --- |
| PCM16, 24 kHz | `{"type": "audio/pcm", "rate": 24000}`, the default |
| PCM16, 16 kHz | `{"type": "audio/pcm", "rate": 16000}` |
| G.711 μ-law, 8 kHz | `{"type": "audio/pcmu"}` |
| G.711 A-law, 8 kHz | `{"type": "audio/pcma"}` |

Audio is mono, and PCM16 is little-endian. Chunks can be any size. A format the model does not take is refused with `unsupported_audio_format`, and the session stays open.

### Send at the speed of speech

Audio is transcribed against a clock that starts with your first audio. The server holds up to 640 ms of audio ahead of that clock; audio sent faster ends the session with close code 1013, and audio that arrives after its moment has passed is transcribed as silence. A live microphone meets both limits by itself. For a recording, pace the chunks against a clock, as the examples do, rather than sleeping a fixed time between sends.

## Segments and timing

- **Deltas**: `conversation.item.input_audio_transcription.delta` carries new text as soon as it is recognized. Text is final when it arrives: it is never revised.
- **Segments**: A segment completes after a pause, with `conversation.item.input_audio_transcription.completed` and its full transcript. `turn_detection.silence_duration_ms` sets the pause, from 400 ms to 3200 ms (the default). With `turn_detection` set to `null`, a segment completes only when you send `input_audio_buffer.commit`.
- **Timestamps**: Each completed segment carries `audio_start_ms` and `audio_end_ms`, its place in your audio.
- **Quiet sessions**: A session closes after 30 seconds without audio. To keep a quiet session open, send silence. A client secret can set `turn_detection.idle_timeout_ms` from 10000 to 300000.
- **Session length**: A session runs for up to 24 hours.

## Close codes

On the socket, a problem with one event arrives as an `error` event, and the session continues. When a session ends, `session.end` gives the reason, and the socket closes with one of these codes:

| Close code | Meaning | What to do |
| --- | --- | --- |
| 1000 | Closed normally: by you, after 30 seconds without audio, or after 24 hours. | Nothing, or open a new session. |
| 1011 | A server error. | Open a new session. |
| 1012 | The server is restarting. | Open a new session. |
| 1013 | Audio arrived faster than real time. | Open a new session, and pace the audio. |
| 4400 | A protocol error, such as a malformed message or an invalid parameter. | Fix the client. |
| 4401 | The credential is invalid, expired or already used. | Use your API key, or a new client secret. |
| 4413 | The service is not ready. | Retry shortly. |
| 4429 | No capacity when you connected. | Retry after a short wait. |

> After an interruption, open a new session and continue with new audio. Do not resend audio from the previous session.
