Docs / Deepgram-compatible API

Deepgram-compatible API

Deepgram-compatible API

Streaming transcription for clients built for Deepgram’s live transcription API: configuration in the query string, Authorization: Token, and Results messages of whole words. The Deepgram plugins of LiveKit Agents and Pipecat connect by changing their base URL.

Endpoint

WSS
wss://api.dotwave.ai/v1/listenConfigured by query parameters; authenticated with Authorization: Token <API key>.

When to use it

Use the Deepgram-compatible API when your code or framework already talks to Deepgram, such as the Deepgram plugins of LiveKit Agents and Pipecat. With the OpenAI SDK, use the Realtime API.

The session in six lines

Connect
wss://api.dotwave.ai/v1/listen?encoding=linear16&sample_rate=24000, with Authorization: Token <API key>
Receive
Metadata, on connect
Send
Binary audio, in chunks of any size, at the speed of speech
Receive
Results: whole words, with is_final, speech_final and from_finalize
Receive
SpeechStarted and UtteranceEnd, around each segment
Send
KeepAlive, Finalize, and CloseStream to end

A first session

Any WebSocket client works. This example sends a 24 kHz mono WAV file as binary audio at the speed of speech and prints each final transcript. Install websockets and set DOTWAVE_API_KEY first.

import asyncio, json, os, time, wave
import websockets

URL = ("wss://api.dotwave.ai/v1/listen"
       "?encoding=linear16&sample_rate=24000&language=pt-BR")
HEADERS = {"Authorization": f"Token {os.environ['DOTWAVE_API_KEY']}"}

async def send_wav(ws, path):
    # 24 kHz mono PCM16 as binary messages, paced at the speed of speech.
    with wave.open(path, "rb") as wav:
        started, sent = time.monotonic(), 0
        while chunk := wav.readframes(1920):  # 1,920 samples at a time
            await ws.send(chunk)
            sent += len(chunk) // 2
            await asyncio.sleep(max(0, started + sent / 24000 - time.monotonic()))
    # Send what is still open as a final, then close.
    await ws.send(json.dumps({"type": "CloseStream"}))

async def main():
    async with websockets.connect(URL, additional_headers=HEADERS) as ws:
        sender = asyncio.create_task(send_wav(ws, "fala-24k.wav"))
        async for raw in ws:
            message = json.loads(raw)
            if message.get("type") == "Results" and message["is_final"]:
                print(message["channel"]["alternatives"][0]["transcript"])
        await sender

asyncio.run(main())

At a glance

Audio
Mono PCM16 at 24 or 16 kHz, or G.711 μ-law or A-law at 8 kHz. See parameters.
Languages
Name the language, or leave it out or set language=multi for automatic detection. See Languages.
Not supported
diarize, multichannel and keywords are refused with an error that names the parameter.
Pricing
Billed per minute of session. See Pricing and limits.

Reference