Models NVIDIA

Nemotron 3.5 ASR Streaming 0.6B

Multilingual speech recognition for live audio. With .wave, one H100 supports 4,800 concurrent streams at the 80 ms chunk setting—20× NVIDIA’s published concurrency.

A Renaissance scribe in a blue robe transcribing with a quill in a dark stone workshop.
Measured

4,800

Concurrent live streams

On one H100 SXM 80 GB.

Price

$0.00045

Per minute of audio

Same model on Together AI: $0.0045/min.

Model

40

Language locales

32 available directly; 8 require fine-tuning.

About the model

Nemotron 3.5 ASR is NVIDIA’s 600M-parameter model for streaming speech recognition. It turns incoming speech into punctuated text while the speaker is still talking, using one multilingual checkpoint.

Its cache-aware FastConformer encoder reuses earlier context instead of repeatedly processing overlapping audio. An RNN-T decoder generates the transcript. Audio chunks can be set to 80, 160, 320, 560 or 1,120 ms at inference time, without retraining.

On .wave, the model runs unchanged. Our benchmark supports 4,800 concurrent streams on one H100 SXM 80 GB at the 80 ms chunk setting.

NVIDIA model card

Key capabilities

Native streaming

Transcribe new audio as it arrives, retaining the context from earlier chunks.

Configurable chunks

Adjust chunk duration to balance responsiveness and recognition accuracy. Chunk size is distinct from end-to-end latency.

Multilingual recognition

Provide the target locale, or let the model detect the spoken language automatically.

Formatted transcripts

Produce punctuation and capitalization as part of the model’s output.

Price comparison

Rates in USD per minute of audio.
ServiceModelPer minute
.waveNemotron 3.5 ASR$0.00045
Together AINemotron 3.5 ASR$0.0045
Together AIWhisper Large v3 (Streaming)$0.0035

Together AI rates checked 16 September 2026.

Applications & use cases

Multilingual voice agents

Provide live transcripts for customer support, in-car assistants and retail kiosks serving speakers of different languages.

Healthcare & clinical documentation

Turn medical dictation into text for clinical notes and documentation. Private deployment can support workflows that need to keep audio within a controlled environment.

Live captioning & transcription

Add captions to meetings, events and broadcasts as people speak. Make live conversations more accessible and easier to follow.

Post-call & offline analytics

Transcribe recorded calls for contact-center quality review, sales analysis and searchable conversation archives.

Languages

One checkpoint covers 40 locales. NVIDIA lists 32 for transcription out of the box and 8 that need fine-tuning. Regional variants include French (France and Canada), Portuguese (Brazil and Portugal), and Spanish (Spain and the US).

  • English en-US
  • English (UK) en-GB
  • Spanish es-ES
  • Spanish (US) es-US
  • French fr-FR
  • French (Canada) fr-CA
  • German de-DE
  • Italian it-IT
  • Portuguese (Brazil) pt-BR
  • Portuguese pt-PT
  • Dutch nl-NL
  • Polish pl-PL
  • Russian ru-RU
  • Ukrainian uk-UA
  • Czech cs-CZ
  • Slovak sk-SK
  • Hungarian hu-HU
  • Romanian ro-RO
  • Bulgarian bg-BG
  • Croatian hr-HR
  • Slovenian sl-SIFine-tuning
  • Greek el-GRFine-tuning
  • Swedish sv-SE
  • Danish da-DK
  • Finnish fi-FI
  • Norwegian (Bokmål) nb-NO
  • Norwegian (Nynorsk) nn-NOFine-tuning
  • Estonian et-EE
  • Latvian lv-LVFine-tuning
  • Lithuanian lt-LTFine-tuning
  • Maltese mt-MTFine-tuning
  • Turkish tr-TR
  • Arabic ar-AR
  • Hebrew he-ILFine-tuning
  • Hindi hi-IN
  • Chinese zh-CN
  • Japanese ja-JP
  • Korean ko-KR
  • Vietnamese vi-VN
  • Thai th-THFine-tuning

Language coverage from NVIDIA

Working on a streaming model?

Tell us about your model, expected traffic and timing requirements.

Talk to an engineer