Docs / Models / NVIDIA

NVIDIA

Nemotron 3.5 ASR Streaming 0.6B

A streaming speech recognition model. Stream audio in and receive punctuated text while the speaker is still talking, in 32 languages or with automatic language detection. In public beta.

Model ID nemotron-asr-streaming

Price
$0.00045 / minuteBilled from the first audio until the session closes
Modality
Speech to textStreaming
Languages
32Or automatic detection
Status
Public betaMay change before general availability

Use this model

Set DOTWAVE_API_KEY on your server. Point the OpenAI SDK at https://api.dotwave.ai/v1 and open a transcription session, or point the Deepgram plugin of LiveKit Agents or Pipecat at .wave. For a browser, create a short-lived client secret on your server, as in the curl example.

import asyncio, os
from openai import AsyncOpenAI

client = AsyncOpenAI(
    api_key=os.environ["DOTWAVE_API_KEY"],
    base_url="https://api.dotwave.ai/v1",
)

async def main():
    async with client.realtime.connect(model="nemotron-asr-streaming") as conn:
        await conn.session.update(session={
            "type": "transcription",
            "audio": {"input": {
                "format": {"type": "audio/pcm", "rate": 24000},
                "transcription": {
                    "model": "nemotron-asr-streaming",
                    "language": "pt-BR",  # leave it out to detect the language
                },
            }},
        })
        # Stream 24 kHz mono PCM16 as it is captured, with
        # conn.input_audio_buffer.append(audio=base64_chunk).
        async for event in conn:
            if event.type == "conversation.item.input_audio_transcription.delta":
                print(event.delta, end="", flush=True)
            elif event.type == "conversation.item.input_audio_transcription.completed":
                print()

asyncio.run(main())

Languages

Set audio.input.transcription.language, or language on the Deepgram route, to one of these tags. Leave it out and the model detects the language automatically; when you know the language, naming it is more accurate. We publish no accuracy figures per language yet, so test with audio like yours before you rely on it.

  • Arabic ar-AR
  • Bulgarian bg-BG
  • Chinese, Mandarin Simplified zh-CN
  • Croatian hr-HR
  • Czech cs-CZ
  • Danish da-DK
  • Dutch nl-NL
  • English, United Kingdom en-GB
  • English, United States en-US
  • Estonian et-EE
  • Finnish fi-FI
  • French, Canada fr-CA
  • French, France fr-FR
  • German de-DE
  • Hindi hi-IN
  • Hungarian hu-HU
  • Italian it-IT
  • Japanese ja-JP
  • Korean ko-KR
  • Norwegian Bokmål nb-NO
  • Polish pl-PL
  • Portuguese, Brazil pt-BR
  • Portuguese, Portugal pt-PT
  • Romanian ro-RO
  • Russian ru-RU
  • Slovak sk-SK
  • Spanish, Spain es-ES
  • Spanish, United States es-US
  • Swedish sv-SE
  • Turkish tr-TR
  • Ukrainian uk-UA
  • Vietnamese vi-VN

Specifications

Author
NVIDIA
Parameters
0.6B
Input
Mono audio: PCM16 at 24 or 16 kHz, or G.711
Output
Punctuated, capitalized text
Interaction
Streaming

Limits

Stream audio at the speed of speech: the server accepts up to 640 ms of audio ahead of real time, and a client that sends faster is disconnected. A session closes after 30 seconds without audio and runs for up to 24 hours. Transcription prompts, noise reduction, speaker diarization, multichannel audio and keyword boosting are not supported.

Errors and close codes