# Live API

Full-duplex speech-to-speech sessions over one WebSocket. Stream audio in and receive the model’s speech, with transcripts of both sides, while the model listens and speaks at the same time.

[Create an account](https://api.dotwave.ai/auth/signup) · [Models](https://dotwave.ai/docs/models/)

> **Public beta.** Events and limits may change before general availability.

> **Coming from the OpenAI Realtime API?** The Live API follows the Live convention, which is not Realtime-shaped: there is no `input_audio_buffer`, speech does not arrive inside a `response`, and the two transcripts are their own event streams. A Realtime client will not run a Live session unchanged. For transcription in the Realtime format, use the [Realtime API](https://dotwave.ai/docs/realtime/).

## Endpoints

- `POST https://api.dotwave.ai/v1/live/client_secrets`: Creates a client secret, so a browser can open a session. [Client secrets](https://dotwave.ai/docs/live/client-secrets/)
- `WSS wss://api.dotwave.ai/v1/live/sessions`: The session socket. [WebSocket](https://dotwave.ai/docs/live/websocket/)

## When to use it

Use the Live API when the model should hold the conversation itself: it hears the user while it speaks, takes and gives the turn on its own, and answers with speech. To turn speech into text for your own pipeline, use the [Realtime API](https://dotwave.ai/docs/realtime/) or the [Deepgram-compatible API](https://dotwave.ai/docs/deepgram/).

## The session in eight lines

- **Connect**: `wss://api.dotwave.ai/v1/live/sessions`, with `Authorization: Bearer <API key>` from a server, or a [client secret](https://dotwave.ai/docs/live/client-secrets/)
- **Send**: `session.start`: must be the first event, or the socket refuses it
- **Receive**: `session.started`: carries `session.id`, the resolved model and audio config
- **Send**: `session.input_audio.append`: base64 24 kHz PCM16, or the same chunks as binary messages
- **Receive**: `session.output_audio.delta`: base64 by default, binary with `output.encoding`
- **Receive**: `session.input_transcript.delta` / `session.output_transcript.delta`
- **Receive**: `session.turn.event`: when the user or the model takes or gives the turn
- **Send**: `session.close` → `session.closed`

## A first session

This example streams a WAV file into a session, prints both transcripts, and saves the model’s reply. Install `websockets` and set `DOTWAVE_API_KEY` first.

```python
import asyncio, base64, json, os, time, wave
import websockets

URL = "wss://api.dotwave.ai/v1/live/sessions"
HEADERS = {"Authorization": f"Bearer {os.environ['DOTWAVE_API_KEY']}"}

async def send_audio(ws, path):
    # 24 kHz mono PCM16, then 5 s of silence while the model answers.
    with wave.open(path, "rb") as f:
        assert f.getframerate() == 24000 and f.getsampwidth() == 2
        assert f.getnchannels() == 1
        audio = f.readframes(f.getnframes()) + bytes(48000 * 5)
    # 3,840 bytes at a time, paced against a clock at the speed of speech.
    started = time.monotonic()
    for offset in range(0, len(audio), 3840):
        await ws.send(json.dumps({
            "type": "session.input_audio.append",
            "audio": base64.b64encode(audio[offset:offset + 3840]).decode(),
        }))
        await asyncio.sleep(max(0, started + (offset + 3840) / 48000 - time.monotonic()))
    await ws.send(json.dumps({"type": "session.close"}))

async def main():
    async with websockets.connect(URL, additional_headers=HEADERS) as ws:
        # session.start must be the first event, or the socket refuses it.
        await ws.send(json.dumps({
            "type": "session.start",
            "session": {"model": "nemotron-voicechat"},
        }))
        sender = asyncio.create_task(send_audio(ws, "speech-24000-mono.wav"))

        reply = bytearray()
        async for raw in ws:
            event = json.loads(raw)
            kind = event.get("type", "")
            if kind == "session.output_audio.delta":
                reply += base64.b64decode(event["delta"])
            elif kind.endswith("_transcript.delta"):  # both sides' transcripts
                print(event.get("delta", ""), end="", flush=True)
            elif kind in ("session.closed", "error"):
                break
        sender.cancel()

    with wave.open("reply.wav", "wb") as f:
        f.setnchannels(1); f.setsampwidth(2); f.setframerate(24000)
        f.writeframes(bytes(reply))

asyncio.run(main())
```

## Models and limits

Session length, voices and supported fields depend on the model: see its page in [Models](https://dotwave.ai/docs/models/), and [Pricing and limits](https://dotwave.ai/docs/pricing-and-limits/). A field the model does not support is refused, never silently ignored.

## Reference

- [Client secrets](https://dotwave.ai/docs/live/client-secrets/) (Live API): Let a browser open one session without your API key.
- [WebSocket](https://dotwave.ai/docs/live/websocket/) (Live API): Connect, start the session, stream audio and close.
- [Client events](https://dotwave.ai/docs/live/client-events/) (Live API): The events you send, and the ones the socket refuses.
- [Server events](https://dotwave.ai/docs/live/server-events/) (Live API): Speech, transcripts, turns, usage and errors.
- [Python](https://dotwave.ai/docs/guides/python/) (Guide): Live microphone capture and playback.
- [TypeScript](https://dotwave.ai/docs/guides/typescript/) (Guide): A server-side sender and receiver.
