Docs / Quickstart

Get started

Quickstart

Create an API key, check it, then make a first call. Each example runs from a terminal with Python and a 24 kHz mono WAV file.

1. Create an API key

Create an account. Your API key is shown once after signup; new accounts start with free credits, no card required. Keep the key on your server, in an environment variable:

export DOTWAVE_API_KEY="wk_…"

2. Check the key

List the models your key can use:

curl https://api.dotwave.ai/v1/models \
  -H "Authorization: Bearer $DOTWAVE_API_KEY"

3. Make a first call

Choose the tab for your API. The Live API and Deepgram-compatible API examples use the websockets package; the Realtime API example uses the OpenAI SDK, pip install "openai[realtime]".

import asyncio, base64, json, os, time, wave
import websockets

URL = "wss://api.dotwave.ai/v1/live/sessions"
HEADERS = {"Authorization": f"Bearer {os.environ['DOTWAVE_API_KEY']}"}

async def send_audio(ws, path):
    # 24 kHz mono PCM16, then 5 s of silence while the model answers.
    with wave.open(path, "rb") as f:
        assert f.getframerate() == 24000 and f.getsampwidth() == 2
        assert f.getnchannels() == 1
        audio = f.readframes(f.getnframes()) + bytes(48000 * 5)
    # 3,840 bytes at a time, paced against a clock at the speed of speech.
    started = time.monotonic()
    for offset in range(0, len(audio), 3840):
        await ws.send(json.dumps({
            "type": "session.input_audio.append",
            "audio": base64.b64encode(audio[offset:offset + 3840]).decode(),
        }))
        await asyncio.sleep(max(0, started + (offset + 3840) / 48000 - time.monotonic()))
    await ws.send(json.dumps({"type": "session.close"}))

async def main():
    async with websockets.connect(URL, additional_headers=HEADERS) as ws:
        # session.start must be the first event, or the socket refuses it.
        await ws.send(json.dumps({
            "type": "session.start",
            "session": {"model": "nemotron-voicechat"},
        }))
        sender = asyncio.create_task(send_audio(ws, "speech-24000-mono.wav"))

        reply = bytearray()
        async for raw in ws:
            event = json.loads(raw)
            kind = event.get("type", "")
            if kind == "session.output_audio.delta":
                reply += base64.b64decode(event["delta"])
            elif kind.endswith("_transcript.delta"):  # both sides' transcripts
                print(event.get("delta", ""), end="", flush=True)
            elif kind in ("session.closed", "error"):
                break
        sender.cancel()

    with wave.open("reply.wav", "wb") as f:
        f.setnchannels(1); f.setsampwidth(2); f.setframerate(24000)
        f.writeframes(bytes(reply))

asyncio.run(main())

Next

Live API
Events, audio and close codes of a full-duplex session: Live API.
Realtime API
Formats, segments and languages of a transcription session: Realtime API.
Deepgram-compatible API
Query parameters, messages and plugins: Deepgram-compatible API.
Browsers
Create a client secret on your server: Authentication.