Docs / Realtime API

Realtime API

Realtime API

Streaming transcription in the OpenAI Realtime transcription format. Stream audio over a WebSocket and receive text while the speaker is still talking.

Endpoints

POST
https://api.dotwave.ai/v1/realtime/client_secretsCreates a client secret, so a browser can open a session. Client secrets
WSS
wss://api.dotwave.ai/v1/realtimeThe transcription socket. WebSocket

When to use it

Use the Realtime API when your client speaks OpenAI’s Realtime transcription events, or when you start with the OpenAI SDK: it connects by changing its base URL. If your code or framework was built for Deepgram, use the Deepgram-compatible API. For a model that talks back, use the Live API.

The session in six lines

Connect
wss://api.dotwave.ai/v1/realtime, with Authorization: Bearer <API key>, or a client secret from POST /v1/realtime/client_secrets
Send
session.update: type: "transcription", and the language, before the first audio
Receive
session.updated: the configuration in effect
Send
input_audio_buffer.append: base64 24 kHz PCM16 or G.711, at the speed of speech
Receive
conversation.item.input_audio_transcription.delta: new text as it is recognized
Receive
conversation.item.input_audio_transcription.completed: each segment’s full transcript

A first session

Point the OpenAI SDK at https://api.dotwave.ai/v1: nothing else changes. Install it with its WebSocket support and set your key:

python -m pip install "openai[realtime]"
export DOTWAVE_API_KEY="wk_…"

This example transcribes a 24 kHz mono WAV file, sent at the speed of speech. Name the language with its tag, here pt-BR, or leave it out to let the model detect the language.

import asyncio, base64, os, time, wave
from openai import AsyncOpenAI

client = AsyncOpenAI(
    api_key=os.environ["DOTWAVE_API_KEY"],
    base_url="https://api.dotwave.ai/v1",
)

async def send_wav(conn, path):
    # 24 kHz mono PCM16, paced against a clock at the speed of speech.
    with wave.open(path, "rb") as wav:
        started, sent = time.monotonic(), 0
        while chunk := wav.readframes(1920):  # 1,920 samples at a time
            audio = base64.b64encode(chunk).decode()
            await conn.input_audio_buffer.append(audio=audio)
            sent += len(chunk) // 2
            await asyncio.sleep(max(0, started + sent / 24000 - time.monotonic()))

async def print_transcripts(conn):
    async for event in conn:
        if event.type == "conversation.item.input_audio_transcription.delta":
            print(event.delta, end="", flush=True)
        elif event.type == "conversation.item.input_audio_transcription.completed":
            print()

async def main():
    async with client.realtime.connect(model="nemotron-asr-streaming") as conn:
        await conn.session.update(session={
            "type": "transcription",
            "audio": {"input": {
                "format": {"type": "audio/pcm", "rate": 24000},
                "transcription": {
                    "model": "nemotron-asr-streaming",
                    "language": "pt-BR",
                },
            }},
        })
        printer = asyncio.create_task(print_transcripts(conn))
        await send_wav(conn, "fala-24k.wav")
        await asyncio.sleep(4)  # the last segment completes after 3.2 s of silence
        printer.cancel()

asyncio.run(main())

At a glance

Audio
Mono PCM16 at 24 kHz, the default, or 16 kHz; or G.711 μ-law or A-law. See audio formats.
Languages
Automatic detection by default; naming the language is more accurate. See Languages.
Text
Deltas are final when they arrive: text is never revised. A segment completes after a pause. See server events.
Pricing
Billed per minute of session. See Pricing and limits.

Reference