4,800
Concurrent live streams
On one H100 SXM 80 GB.
Models NVIDIA
Multilingual speech recognition for live audio. With .wave, one H100 supports 4,800 concurrent streams at the 80 ms chunk setting—20× NVIDIA’s published concurrency.
4,800
On one H100 SXM 80 GB.
$0.00045
Same model on Together AI: $0.0045/min.
40
32 available directly; 8 require fine-tuning.
Nemotron 3.5 ASR is NVIDIA’s 600M-parameter model for streaming speech recognition. It turns incoming speech into punctuated text while the speaker is still talking, using one multilingual checkpoint.
Its cache-aware FastConformer encoder reuses earlier context instead of repeatedly processing overlapping audio. An RNN-T decoder generates the transcript. Audio chunks can be set to 80, 160, 320, 560 or 1,120 ms at inference time, without retraining.
On .wave, the model runs unchanged. Our benchmark supports 4,800 concurrent streams on one H100 SXM 80 GB at the 80 ms chunk setting.
Transcribe new audio as it arrives, retaining the context from earlier chunks.
Adjust chunk duration to balance responsiveness and recognition accuracy. Chunk size is distinct from end-to-end latency.
Provide the target locale, or let the model detect the spoken language automatically.
Produce punctuation and capitalization as part of the model’s output.
| Service | Model | Per minute |
|---|---|---|
| .wave | Nemotron 3.5 ASR | $0.00045 |
| Together AI | Nemotron 3.5 ASR | $0.0045 |
| Together AI | Whisper Large v3 (Streaming) | $0.0035 |
Together AI rates checked 16 September 2026.
Provide live transcripts for customer support, in-car assistants and retail kiosks serving speakers of different languages.
Turn medical dictation into text for clinical notes and documentation. Private deployment can support workflows that need to keep audio within a controlled environment.
Add captions to meetings, events and broadcasts as people speak. Make live conversations more accessible and easier to follow.
Transcribe recorded calls for contact-center quality review, sales analysis and searchable conversation archives.
One checkpoint covers 40 locales. NVIDIA lists 32 for transcription out of the box and 8 that need fine-tuning. Regional variants include French (France and Canada), Portuguese (Brazil and Portugal), and Spanish (Spain and the US).
en-USen-GBes-ESes-USfr-FRfr-CAde-DEit-ITpt-BRpt-PTnl-NLpl-PLru-RUuk-UAcs-CZsk-SKhu-HUro-RObg-BGhr-HRsl-SIFine-tuningel-GRFine-tuningsv-SEda-DKfi-FInb-NOnn-NOFine-tuninget-EElv-LVFine-tuninglt-LTFine-tuningmt-MTFine-tuningtr-TRar-ARhe-ILFine-tuninghi-INzh-CNja-JPko-KRvi-VNth-THFine-tuningTell us about your model, expected traffic and timing requirements.