Realtime API
Client events
The events you send on the Realtime socket. They are OpenAI’s Realtime transcription events; session.close is an addition of .wave’s.
Events
| Event | What it does |
|---|---|
session.update | Sets the session type, audio format, language and segmentation, before the first audio. Answered by session.updated. |
input_audio_buffer.append | Base64 audio in audio, in chunks of any size. |
input_audio_buffer.commit | Ends the current segment now. |
input_audio_buffer.clear | Drops audio sent but not yet transcribed. |
session.close | Ends the session. An addition of .wave’s; closing the socket ends it too. |
Configure with session.update
{"type": "session.update",
"session": {"type": "transcription",
"audio": {"input": {"format": {"type": "audio/pcm", "rate": 24000},
"transcription": {"model": "nemotron-asr-streaming",
"language": "pt-BR"},
"turn_detection": {"type": "server_vad",
"silence_duration_ms": 3200}}}}}session.typetranscription.audio.input.format- One of the audio formats; 24 kHz PCM16 by default.
audio.input.transcription.model- The model, such as
nemotron-asr-streaming. audio.input.transcription.language- A tag from the language list. Leave it out to let the model detect the language. Set it before the first audio.
audio.input.turn_detection{"type": "server_vad", "silence_duration_ms": …}, from 400 to 3200; ornullto end segments only withinput_audio_buffer.commit.- Fields without effect
transcription.prompt,noise_reduction, and thethresholdandprefix_padding_msofturn_detectionare accepted and returned asnull.semantic_vadis refused withinvalid_parameter.