Realtime API
Realtime API
Streaming transcription in the OpenAI Realtime transcription format. Stream audio over a WebSocket and receive text while the speaker is still talking.
Endpoints
- POST
https://api.dotwave.ai/v1/realtime/client_secretsCreates a client secret, so a browser can open a session. Client secrets- WSS
wss://api.dotwave.ai/v1/realtimeThe transcription socket. WebSocket
When to use it
Use the Realtime API when your client speaks OpenAI’s Realtime transcription events, or when you start with the OpenAI SDK: it connects by changing its base URL. If your code or framework was built for Deepgram, use the Deepgram-compatible API. For a model that talks back, use the Live API.
The session in six lines
- Connect
wss://api.dotwave.ai/v1/realtime, withAuthorization: Bearer <API key>, or a client secret fromPOST /v1/realtime/client_secrets- Send
session.update:type: "transcription", and the language, before the first audio- Receive
session.updated: the configuration in effect- Send
input_audio_buffer.append: base64 24 kHz PCM16 or G.711, at the speed of speech- Receive
conversation.item.input_audio_transcription.delta: new text as it is recognized- Receive
conversation.item.input_audio_transcription.completed: each segment’s full transcript
A first session
Point the OpenAI SDK at https://api.dotwave.ai/v1: nothing else changes. Install it with its WebSocket support and set your key:
python -m pip install "openai[realtime]"
export DOTWAVE_API_KEY="wk_…"This example transcribes a 24 kHz mono WAV file, sent at the speed of speech. Name the language with its tag, here pt-BR, or leave it out to let the model detect the language.
import asyncio, base64, os, time, wave
from openai import AsyncOpenAI
client = AsyncOpenAI(
api_key=os.environ["DOTWAVE_API_KEY"],
base_url="https://api.dotwave.ai/v1",
)
async def send_wav(conn, path):
# 24 kHz mono PCM16, paced against a clock at the speed of speech.
with wave.open(path, "rb") as wav:
started, sent = time.monotonic(), 0
while chunk := wav.readframes(1920): # 1,920 samples at a time
audio = base64.b64encode(chunk).decode()
await conn.input_audio_buffer.append(audio=audio)
sent += len(chunk) // 2
await asyncio.sleep(max(0, started + sent / 24000 - time.monotonic()))
async def print_transcripts(conn):
async for event in conn:
if event.type == "conversation.item.input_audio_transcription.delta":
print(event.delta, end="", flush=True)
elif event.type == "conversation.item.input_audio_transcription.completed":
print()
async def main():
async with client.realtime.connect(model="nemotron-asr-streaming") as conn:
await conn.session.update(session={
"type": "transcription",
"audio": {"input": {
"format": {"type": "audio/pcm", "rate": 24000},
"transcription": {
"model": "nemotron-asr-streaming",
"language": "pt-BR",
},
}},
})
printer = asyncio.create_task(print_transcripts(conn))
await send_wav(conn, "fala-24k.wav")
await asyncio.sleep(4) # the last segment completes after 3.2 s of silence
printer.cancel()
asyncio.run(main())At a glance
- Audio
- Mono PCM16 at 24 kHz, the default, or 16 kHz; or G.711 μ-law or A-law. See audio formats.
- Languages
- Automatic detection by default; naming the language is more accurate. See Languages.
- Text
- Deltas are final when they arrive: text is never revised. A segment completes after a pause. See server events.
- Pricing
- Billed per minute of session. See Pricing and limits.
Reference
Let a browser open one session without your API key.
Open → Realtime APIWebSocketConnect, configure, stream audio at the speed of speech, and close.
Open → Realtime APIClient eventsConfigure the session and send audio.
Open → Realtime APIServer eventsTranscript deltas, completed segments and errors.
Open → Realtime APILanguagesName the language, or let the model detect it.
Open → GuideOpenAI SDKPython and TypeScript, with only the base URL changed.
Open →