NVIDIA
Nemotron VoiceChat 11B
A full-duplex voice model that listens while it speaks. Stream audio in and receive speech, plus transcripts of both sides of the conversation. In public beta.
Model ID nemotron-voicechat
- Price
- $0.005 / minuteIntroductory price, billed while the session is open
- Modality
- Speech to speechFull duplex
- Language
- Englishen-US
- Session limit
- 120 seconds
- Status
- Public betaMay change before general availability
Use this model
Set DOTWAVE_API_KEY on your server, then open and configure a session. See the guides below for audio capture, playback, and transcripts.
import asyncio, json, os
import websockets
LIVE_URL = "wss://api.dotwave.ai/v1/live/sessions"
async def main():
async with websockets.connect(
LIVE_URL,
additional_headers={"Authorization": f"Bearer {os.environ['DOTWAVE_API_KEY']}"},
) as ws:
# session.start must be the first message, or the socket refuses it.
await ws.send(json.dumps({
"type": "session.start",
"session": {
"model": "nemotron-voicechat",
"audio": {
"input": {
"format": {"type": "audio/pcm", "rate": 24000}
},
"output": {
"format": {"type": "audio/pcm", "rate": 24000},
"encoding": "base64",
},
},
},
}))
# Send session.input_audio.append and handle
# session.output_audio.delta events here.
asyncio.run(main())import WebSocket from 'ws';
const ws = new WebSocket('wss://api.dotwave.ai/v1/live/sessions', {
headers: { Authorization: `Bearer ${process.env.DOTWAVE_API_KEY}` },
});
// session.start must be the first message, or the socket refuses it.
ws.on('open', () => {
ws.send(JSON.stringify({
type: 'session.start',
session: {
model: 'nemotron-voicechat',
audio: {
input: { format: { type: 'audio/pcm', rate: 24000 } },
output: {
format: { type: 'audio/pcm', rate: 24000 },
encoding: 'base64',
},
},
},
}));
});
ws.on('message', raw => {
const event = JSON.parse(raw.toString());
if (event.type === 'session.output_audio.delta') {
playPcm24k(Buffer.from(event.delta, 'base64'));
}
});curl https://api.dotwave.ai/v1/live/client_secrets \
-H "Authorization: Bearer $DOTWAVE_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"expires_after": {
"anchor": "created_at",
"seconds": 60
},
"session": {
"type": "live",
"model": "nemotron-voicechat",
"audio": {
"input": {
"format": {"type": "audio/pcm", "rate": 24000}
},
"output": {
"format": {"type": "audio/pcm", "rate": 24000},
"encoding": "base64"
}
}
}
}'Specifications
- Author
- NVIDIA
- Parameters
- 11B
- Input
- Audio
- Output
- Audio and transcripts
- Interaction
- Full-duplex conversation
- Access
- Live API
Limits
Sessions are capped at 120 seconds. VoiceChat uses its built-in persona and speaks in its one voice, aria; a session.start naming any other voice is refused. Custom instructions, text input, and tools are not supported.