GenerateSpeech
Generates speech audio from text.
Signature
func GenerateSpeech(ctx context.Context, opts GenerateSpeechOptions) (*GenerateSpeechResult, error)
GenerateSpeechOptions
| Field | Type | Required | Description |
|---|---|---|---|
Model | provider.SpeechModel | Yes | Speech model |
Text | string | Yes | Input text |
Voice | string | No | Voice identifier |
Speed | *float64 | No | Speech speed |
OutputFormat | string | No | Desired audio format, such as mp3 or wav |
Instructions | string | No | Provider-dependent delivery instructions |
Language | string | No | Speech language code or provider-supported automatic language mode |
ProviderOptions | map[string]interface{} | No | Provider-specific options keyed by provider name |
Headers | map[string]string | No | Extra request headers |
MaxRetries | *int | No | Maximum retries per speech call. Defaults to 2; set to 0 to disable |
Telemetry | *ai.TelemetrySettings | No | Observability configuration for this call. ExperimentalTelemetry is a deprecated alias; Telemetry wins when both are set |
Telemetry
When a telemetry integration is registered (global via telemetry.RegisterTelemetryIntegration, or per-call via Telemetry.Integrations), GenerateSpeech emits an ai.generateSpeech root span covering the whole call (including retries), with gen_ai.output.type: "speech", the input text (ai.request.text, input-gated), the generated audio's size/media type/format (ai.response.audio.*, output-gated), and the provider's reported usage flattened into numeric ai.usage.* / gen_ai.usage.* attributes (e.g. ai.usage.characters for a provider that reports {"characters": N}; see the Telemetry guide for the key-normalization rules). A failed call — including a provider response with no audio — ends the span with codes.Error status.
GenerateSpeechResult
| Field | Type | Description |
|---|---|---|
Audio | ai.GeneratedAudioFile | Generated audio file payload |
Warnings | []types.Warning | Provider warnings (if exposed) |
Responses | []ai.SpeechModelResponseMetadata | Response metadata entries (timestamp, modelId, optional headers, optional body) |
ProviderMetadata | map[string]interface{} | Provider-specific metadata |
GeneratedAudioFile
type GeneratedAudioFile struct {
Data []byte `json:"data,omitempty"`
URL string `json:"url,omitempty"`
MediaType string `json:"mediaType"`
Format string `json:"format"`
}
func (f GeneratedAudioFile) Base64() string
func (f GeneratedAudioFile) Uint8Array() []byte
Deprecated Compatibility Aliases
The TypeScript SDK still exports deprecated experimental_generateSpeech and
Experimental_SpeechResult names. Go exposes equivalent migration aliases:
func ExperimentalGenerateSpeech(ctx context.Context, opts GenerateSpeechOptions) (*GenerateSpeechResult, error)
type Experimental_SpeechResult = GenerateSpeechResult
Example
result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{
Model: speechModel,
Text: "Hello from Go AI SDK",
Voice: "alloy",
})
if err != nil {
// handle
}
_ = os.WriteFile("output.mp3", result.Audio.Data, 0644)