Docs / Live API

Live API

Live API

Full-duplex speech-to-speech sessions over one WebSocket. Stream audio in and receive the model’s speech, with transcripts of both sides, while the model listens and speaks at the same time.

Endpoints

POST
https://api.dotwave.ai/v1/live/client_secretsCreates a client secret, so a browser can open a session. Client secrets
WSS
wss://api.dotwave.ai/v1/live/sessionsThe session socket. WebSocket

When to use it

Use the Live API when the model should hold the conversation itself: it hears the user while it speaks, takes and gives the turn on its own, and answers with speech. To turn speech into text for your own pipeline, use the Realtime API or the Deepgram-compatible API.

The session in eight lines

Connect
wss://api.dotwave.ai/v1/live/sessions, with Authorization: Bearer <API key> from a server, or a client secret
Send
session.start: must be the first event, or the socket refuses it
Receive
session.started: carries session.id, the resolved model and audio config
Send
session.input_audio.append: base64 24 kHz PCM16, or the same chunks as binary messages
Receive
session.output_audio.delta: base64 by default, binary with output.encoding
Receive
session.input_transcript.delta / session.output_transcript.delta
Receive
session.turn.event: when the user or the model takes or gives the turn
Send
session.close → session.closed

A first session

This example streams a WAV file into a session, prints both transcripts, and saves the model’s reply. Install websockets and set DOTWAVE_API_KEY first.

import asyncio, base64, json, os, time, wave
import websockets

URL = "wss://api.dotwave.ai/v1/live/sessions"
HEADERS = {"Authorization": f"Bearer {os.environ['DOTWAVE_API_KEY']}"}

async def send_audio(ws, path):
    # 24 kHz mono PCM16, then 5 s of silence while the model answers.
    with wave.open(path, "rb") as f:
        assert f.getframerate() == 24000 and f.getsampwidth() == 2
        assert f.getnchannels() == 1
        audio = f.readframes(f.getnframes()) + bytes(48000 * 5)
    # 3,840 bytes at a time, paced against a clock at the speed of speech.
    started = time.monotonic()
    for offset in range(0, len(audio), 3840):
        await ws.send(json.dumps({
            "type": "session.input_audio.append",
            "audio": base64.b64encode(audio[offset:offset + 3840]).decode(),
        }))
        await asyncio.sleep(max(0, started + (offset + 3840) / 48000 - time.monotonic()))
    await ws.send(json.dumps({"type": "session.close"}))

async def main():
    async with websockets.connect(URL, additional_headers=HEADERS) as ws:
        # session.start must be the first event, or the socket refuses it.
        await ws.send(json.dumps({
            "type": "session.start",
            "session": {"model": "nemotron-voicechat"},
        }))
        sender = asyncio.create_task(send_audio(ws, "speech-24000-mono.wav"))

        reply = bytearray()
        async for raw in ws:
            event = json.loads(raw)
            kind = event.get("type", "")
            if kind == "session.output_audio.delta":
                reply += base64.b64decode(event["delta"])
            elif kind.endswith("_transcript.delta"):  # both sides' transcripts
                print(event.get("delta", ""), end="", flush=True)
            elif kind in ("session.closed", "error"):
                break
        sender.cancel()

    with wave.open("reply.wav", "wb") as f:
        f.setnchannels(1); f.setsampwidth(2); f.setframerate(24000)
        f.writeframes(bytes(reply))

asyncio.run(main())

Models and limits

Session length, voices and supported fields depend on the model: see its page in Models, and Pricing and limits. A field the model does not support is refused, never silently ignored.

Reference