# Deepgram-compatible API

Streaming transcription for clients built for Deepgram’s live transcription API: configuration in the query string, `Authorization: Token`, and `Results` messages of whole words. The Deepgram plugins of LiveKit Agents and Pipecat connect by changing their base URL.

[Create an account](https://api.dotwave.ai/auth/signup) · [LiveKit Agents and Pipecat](https://dotwave.ai/docs/deepgram/plugins/)

> **Public beta.** We publish no accuracy figures per language yet, so test with audio like yours before you rely on it.

## Endpoint

- `WSS wss://api.dotwave.ai/v1/listen`: Configured by [query parameters](https://dotwave.ai/docs/deepgram/parameters/); authenticated with `Authorization: Token <API key>`.

## When to use it

Use the Deepgram-compatible API when your code or framework already talks to Deepgram, such as the Deepgram plugins of LiveKit Agents and Pipecat. With the OpenAI SDK, use the [Realtime API](https://dotwave.ai/docs/realtime/).

## The session in six lines

- **Connect**: `wss://api.dotwave.ai/v1/listen?encoding=linear16&sample_rate=24000`, with `Authorization: Token <API key>`
- **Receive**: `Metadata`, on connect
- **Send**: Binary audio, in chunks of any size, at the speed of speech
- **Receive**: `Results`: whole words, with `is_final`, `speech_final` and `from_finalize`
- **Receive**: `SpeechStarted` and `UtteranceEnd`, around each segment
- **Send**: `KeepAlive`, `Finalize`, and `CloseStream` to end

## A first session

Any WebSocket client works. This example sends a 24 kHz mono WAV file as binary audio at the speed of speech and prints each final transcript. Install `websockets` and set `DOTWAVE_API_KEY` first.

```python
import asyncio, json, os, time, wave
import websockets

URL = ("wss://api.dotwave.ai/v1/listen"
       "?encoding=linear16&sample_rate=24000&language=pt-BR")
HEADERS = {"Authorization": f"Token {os.environ['DOTWAVE_API_KEY']}"}

async def send_wav(ws, path):
    # 24 kHz mono PCM16 as binary messages, paced at the speed of speech.
    with wave.open(path, "rb") as wav:
        started, sent = time.monotonic(), 0
        while chunk := wav.readframes(1920):  # 1,920 samples at a time
            await ws.send(chunk)
            sent += len(chunk) // 2
            await asyncio.sleep(max(0, started + sent / 24000 - time.monotonic()))
    # Send what is still open as a final, then close.
    await ws.send(json.dumps({"type": "CloseStream"}))

async def main():
    async with websockets.connect(URL, additional_headers=HEADERS) as ws:
        sender = asyncio.create_task(send_wav(ws, "fala-24k.wav"))
        async for raw in ws:
            message = json.loads(raw)
            if message.get("type") == "Results" and message["is_final"]:
                print(message["channel"]["alternatives"][0]["transcript"])
        await sender

asyncio.run(main())
```

## At a glance

- **Audio**: Mono PCM16 at 24 or 16 kHz, or G.711 μ-law or A-law at 8 kHz. See [parameters](https://dotwave.ai/docs/deepgram/parameters/).
- **Languages**: Name the language, or leave it out or set `language=multi` for automatic detection. See [Languages](https://dotwave.ai/docs/realtime/languages/).
- **Not supported**: `diarize`, `multichannel` and `keywords` are refused with an error that names the parameter.
- **Pricing**: Billed per minute of session. See [Pricing and limits](https://dotwave.ai/docs/pricing-and-limits/).

## Reference

- [Parameters](https://dotwave.ai/docs/deepgram/parameters/) (Deepgram-compatible API): Audio format, language and segmentation, in the query string.
- [Messages](https://dotwave.ai/docs/deepgram/messages/) (Deepgram-compatible API): `Results`, `UtteranceEnd`, `SpeechStarted`, and the control messages you send.
- [Plugins](https://dotwave.ai/docs/deepgram/plugins/) (Deepgram-compatible API): LiveKit Agents and Pipecat, with only the base URL changed.
- [Languages](https://dotwave.ai/docs/realtime/languages/) (Transcription): Name the language, or let the model detect it.
