Skip to main content

Streaming transcription and translation

Two experimental functions send raw audio to a model as it arrives and stream results back.

func ExperimentalStreamTranscribe(ctx context.Context, opts StreamTranscribeOptions) (*StreamTranscriptionResult, error)
func ExperimentalStreamTranslate(ctx context.Context, opts StreamTranslateOptions) (*StreamTranslationResult, error)

FullStream() on each result is single-consumer. Read it before any other result accessor. An accessor such as ProviderMetadata() consumes the stream, and FullStream() is unavailable afterwards.

Audio sources​

FunctionDescription
ai.NewChannelAudioStream(chunks <-chan []byte, onCancel func(reason error)) provider.AudioStreamWraps a channel of raw audio chunks, for live input. The producer closes the channel at the end of the audio. onCancel runs at most once if the consumer cancels.
ai.NewReaderAudioStream(r io.Reader, chunkSize int) provider.AudioStreamReads audio from r in chunkSize-byte chunks. chunkSize defaults to 4096. Cancel closes r when it is an io.Closer.

StreamTranscribeOptions​

FieldTypeDescription
Modelprovider.TranscriptionModelModel must implement provider.TranscriptionStreamer.
Audioprovider.AudioStreamAudio is the source of raw audio chunks to transcribe.
InputAudioFormatprovider.AudioFormatInputAudioFormat is the input audio format for the raw audio chunks.
ProviderOptionsmap[string]interface{}ProviderOptions contains provider-specific request options.
Headersmap[string]stringAdditional HTTP/WebSocket headers to send when supported by the provider.
IncludeRawChunksboolIncludeRawChunks requests provider raw chunks in the stream.
Telemetry*TelemetrySettingsTelemetry configures observability for this operation. When both Telemetry and ExperimentalTelemetry are set, Telemetry wins.
ExperimentalTelemetry*TelemetrySettingsExperimentalTelemetry configures observability for this operation. Deprecated: use Telemetry.

StreamTranslateOptions​

FieldTypeDescription
Modelprovider.SpeechTranslationModelModel to use for speech-to-speech translation.
Audioprovider.AudioStreamAudio is the source of raw audio chunks to translate.
InputAudioFormatprovider.AudioFormatInputAudioFormat is the input audio format for the raw audio chunks.
TargetLanguagestringTargetLanguage is the language to translate the audio into, as a BCP-47-style language tag (e.g. "en", "es", "fr-CA").
SourceLanguagestringSourceLanguage is the language of the source audio. Auto-detected when empty.
OutputAudioFormat*provider.AudioFormatOutputAudioFormat is the desired audio format for translated audio chunks. When nil, the provider default output format is used.
ProviderOptionsmap[string]interface{}ProviderOptions contains provider-specific request options.
Headersmap[string]stringAdditional HTTP/WebSocket headers to send when supported by the provider.
IncludeRawChunksboolIncludeRawChunks requests provider raw chunks in the stream.

Stream types​

ai.TranscriptionStream and ai.TranslationStream have the same shape as provider.TextStream: Next(), Err() and Close().

TranscriptionStreamPart​

Type is one of transcript-delta, transcript-partial, transcript-final, raw or error.

FieldTypeDescription
TypestringType is one of "transcript-delta", "transcript-partial", "transcript-final", "raw", or "error".
IDstring
Deltastring
Textstring
StartSecond*float64
EndSecond*float64
DurationInSeconds*float64
ChannelIndex*int
ProviderMetadatamap[string]interface{}
RawValueinterface{}RawValue carries the provider's raw chunk when Type is "raw".
Errinterface{}Err carries the streamed error payload when Type is "error".

TranslationStreamPart​

Type is one of audio, output-text-delta, output-text-final, source-transcript-delta, source-transcript-partial, source-transcript-final, raw or error.

FieldTypeDescription
TypestringType is one of "audio", "output-text-delta", "output-text-final", "source-transcript-delta", "source-transcript-partial", "source-transcript-final", "raw", or "error".
IDstring
AudioData[]byte
Deltastring
Textstring
StartSecond*float64
EndSecond*float64
ChannelIndex*int
ProviderMetadatamap[string]interface{}
RawValueinterface{}RawValue carries the provider's raw chunk when Type is "raw".
Errinterface{}Err carries the streamed error payload when Type is "error".

SpeechTranslationModelResponseMetadata​

FieldTypeDescription
Timestamptime.Time
ModelIDstring
Headersmap[string]string

Errors​

ai.IsNoTranslationGeneratedError(err) reports whether err is a *ai.NoTranslationGeneratedError.