# Realtime API

Streaming transcription in the OpenAI Realtime transcription format. Stream audio over a WebSocket and receive text while the speaker is still talking.

[Create an account](https://api.dotwave.ai/auth/signup) · [OpenAI SDK guide](https://dotwave.ai/docs/guides/openai-sdk/)

> **Public beta.** We publish no accuracy figures per language yet, so test with audio like yours before you rely on it.

## Endpoints

- `POST https://api.dotwave.ai/v1/realtime/client_secrets`: Creates a client secret, so a browser can open a session. [Client secrets](https://dotwave.ai/docs/realtime/client-secrets/)
- `WSS wss://api.dotwave.ai/v1/realtime`: The transcription socket. [WebSocket](https://dotwave.ai/docs/realtime/websocket/)

## When to use it

Use the Realtime API when your client speaks OpenAI’s Realtime transcription events, or when you start with the OpenAI SDK: it connects by changing its base URL. If your code or framework was built for Deepgram, use the [Deepgram-compatible API](https://dotwave.ai/docs/deepgram/). For a model that talks back, use the [Live API](https://dotwave.ai/docs/live/).

## The session in six lines

- **Connect**: `wss://api.dotwave.ai/v1/realtime`, with `Authorization: Bearer <API key>`, or a client secret from `POST /v1/realtime/client_secrets`
- **Send**: `session.update`: `type: "transcription"`, and the language, before the first audio
- **Receive**: `session.updated`: the configuration in effect
- **Send**: `input_audio_buffer.append`: base64 24 kHz PCM16 or G.711, at the speed of speech
- **Receive**: `conversation.item.input_audio_transcription.delta`: new text as it is recognized
- **Receive**: `conversation.item.input_audio_transcription.completed`: each segment’s full transcript

## A first session

Point the OpenAI SDK at `https://api.dotwave.ai/v1`: nothing else changes. Install it with its WebSocket support and set your key:

```bash
python -m pip install "openai[realtime]"
export DOTWAVE_API_KEY="wk_…"
```

This example transcribes a 24 kHz mono WAV file, sent at the speed of speech. Name the language with its tag, here `pt-BR`, or leave it out to let the model detect the language.

```python
import asyncio, base64, os, time, wave
from openai import AsyncOpenAI

client = AsyncOpenAI(
    api_key=os.environ["DOTWAVE_API_KEY"],
    base_url="https://api.dotwave.ai/v1",
)

async def send_wav(conn, path):
    # 24 kHz mono PCM16, paced against a clock at the speed of speech.
    with wave.open(path, "rb") as wav:
        started, sent = time.monotonic(), 0
        while chunk := wav.readframes(1920):  # 1,920 samples at a time
            audio = base64.b64encode(chunk).decode()
            await conn.input_audio_buffer.append(audio=audio)
            sent += len(chunk) // 2
            await asyncio.sleep(max(0, started + sent / 24000 - time.monotonic()))

async def print_transcripts(conn):
    async for event in conn:
        if event.type == "conversation.item.input_audio_transcription.delta":
            print(event.delta, end="", flush=True)
        elif event.type == "conversation.item.input_audio_transcription.completed":
            print()

async def main():
    async with client.realtime.connect(model="nemotron-asr-streaming") as conn:
        await conn.session.update(session={
            "type": "transcription",
            "audio": {"input": {
                "format": {"type": "audio/pcm", "rate": 24000},
                "transcription": {
                    "model": "nemotron-asr-streaming",
                    "language": "pt-BR",
                },
            }},
        })
        printer = asyncio.create_task(print_transcripts(conn))
        await send_wav(conn, "fala-24k.wav")
        await asyncio.sleep(4)  # the last segment completes after 3.2 s of silence
        printer.cancel()

asyncio.run(main())
```

## At a glance

- **Audio**: Mono PCM16 at 24 kHz, the default, or 16 kHz; or G.711 μ-law or A-law. See [audio formats](https://dotwave.ai/docs/realtime/websocket/#audio-title).
- **Languages**: Automatic detection by default; naming the language is more accurate. See [Languages](https://dotwave.ai/docs/realtime/languages/).
- **Text**: Deltas are final when they arrive: text is never revised. A segment completes after a pause. See [server events](https://dotwave.ai/docs/realtime/server-events/).
- **Pricing**: Billed per minute of session. See [Pricing and limits](https://dotwave.ai/docs/pricing-and-limits/).

## Reference

- [Client secrets](https://dotwave.ai/docs/realtime/client-secrets/) (Realtime API): Let a browser open one session without your API key.
- [WebSocket](https://dotwave.ai/docs/realtime/websocket/) (Realtime API): Connect, configure, stream audio at the speed of speech, and close.
- [Client events](https://dotwave.ai/docs/realtime/client-events/) (Realtime API): Configure the session and send audio.
- [Server events](https://dotwave.ai/docs/realtime/server-events/) (Realtime API): Transcript deltas, completed segments and errors.
- [Languages](https://dotwave.ai/docs/realtime/languages/) (Realtime API): Name the language, or let the model detect it.
- [OpenAI SDK](https://dotwave.ai/docs/guides/openai-sdk/) (Guide): Python and TypeScript, with only the base URL changed.
