Deepgram-compatible API
Deepgram-compatible API
Streaming transcription for clients built for Deepgram’s live transcription API: configuration in the query string, Authorization: Token, and Results messages of whole words. The Deepgram plugins of LiveKit Agents and Pipecat connect by changing their base URL.
Endpoint
- WSS
wss://api.dotwave.ai/v1/listenConfigured by query parameters; authenticated withAuthorization: Token <API key>.
When to use it
Use the Deepgram-compatible API when your code or framework already talks to Deepgram, such as the Deepgram plugins of LiveKit Agents and Pipecat. With the OpenAI SDK, use the Realtime API.
The session in six lines
- Connect
wss://api.dotwave.ai/v1/listen?encoding=linear16&sample_rate=24000, withAuthorization: Token <API key>- Receive
Metadata, on connect- Send
- Binary audio, in chunks of any size, at the speed of speech
- Receive
Results: whole words, withis_final,speech_finalandfrom_finalize- Receive
SpeechStartedandUtteranceEnd, around each segment- Send
KeepAlive,Finalize, andCloseStreamto end
A first session
Any WebSocket client works. This example sends a 24 kHz mono WAV file as binary audio at the speed of speech and prints each final transcript. Install websockets and set DOTWAVE_API_KEY first.
import asyncio, json, os, time, wave
import websockets
URL = ("wss://api.dotwave.ai/v1/listen"
"?encoding=linear16&sample_rate=24000&language=pt-BR")
HEADERS = {"Authorization": f"Token {os.environ['DOTWAVE_API_KEY']}"}
async def send_wav(ws, path):
# 24 kHz mono PCM16 as binary messages, paced at the speed of speech.
with wave.open(path, "rb") as wav:
started, sent = time.monotonic(), 0
while chunk := wav.readframes(1920): # 1,920 samples at a time
await ws.send(chunk)
sent += len(chunk) // 2
await asyncio.sleep(max(0, started + sent / 24000 - time.monotonic()))
# Send what is still open as a final, then close.
await ws.send(json.dumps({"type": "CloseStream"}))
async def main():
async with websockets.connect(URL, additional_headers=HEADERS) as ws:
sender = asyncio.create_task(send_wav(ws, "fala-24k.wav"))
async for raw in ws:
message = json.loads(raw)
if message.get("type") == "Results" and message["is_final"]:
print(message["channel"]["alternatives"][0]["transcript"])
await sender
asyncio.run(main())At a glance
- Audio
- Mono PCM16 at 24 or 16 kHz, or G.711 μ-law or A-law at 8 kHz. See parameters.
- Languages
- Name the language, or leave it out or set
language=multifor automatic detection. See Languages. - Not supported
diarize,multichannelandkeywordsare refused with an error that names the parameter.- Pricing
- Billed per minute of session. See Pricing and limits.
Reference
Audio format, language and segmentation, in the query string.
Open → Deepgram-compatible APIMessagesResults, UtteranceEnd, SpeechStarted, and the control messages you send.
LiveKit Agents and Pipecat, with only the base URL changed.
Open → TranscriptionLanguagesName the language, or let the model detect it.
Open →