Skip to main content

ElevenLabs Provider

ElevenLabs provides text-to-speech models. The current Go provider implements the shared provider.SpeechModel interface and is used through ai.GenerateSpeech.

Setup​

import (
"context"
"log"
"os"

"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/elevenlabs"
)
provider := elevenlabs.New(elevenlabs.Config{
APIKey: os.Getenv("ELEVENLABS_API_KEY"),
})

model, err := provider.SpeechModel("eleven_multilingual_v2")
if err != nil {
log.Fatal(err)
}

Generate Speech​

result, err := ai.GenerateSpeech(context.Background(), ai.GenerateSpeechOptions{
Model: model,
Text: "Welcome to the Go-AI SDK.",
Voice: "Rachel",
OutputFormat: "mp3",
})
if err != nil {
log.Fatal(err)
}

if err := os.WriteFile("welcome.mp3", result.Audio.Data, 0644); err != nil {
log.Fatal(err)
}

Provider Options​

Pass ElevenLabs-specific request options under the "elevenlabs" key when the provider supports them.

result, err := ai.GenerateSpeech(context.Background(), ai.GenerateSpeechOptions{
Model: model,
Text: "A more expressive delivery.",
Voice: "Rachel",
ProviderOptions: map[string]interface{}{
"elevenlabs": map[string]interface{}{
"voiceSettings": map[string]interface{}{
"stability": 0.3,
"similarityBoost": 0.8,
"style": 0.5,
"useSpeakerBoost": true,
},
},
},
})

Models​

Common model IDs include:

Model IDBest For
eleven_multilingual_v2Multilingual speech
eleven_flash_v2_5Low-latency speech
eleven_turbo_v2_5Fast speech synthesis

Transcription​

provider.TranscriptionModel(modelID) supports both batch (scribe_v1, scribe_v2) and realtime (scribe_v2_realtime) models:

transcriptionModel, err := provider.TranscriptionModel("scribe_v2")
if err != nil {
log.Fatal(err)
}

result, err := ai.Transcribe(ctx, ai.TranscribeOptions{
Model: transcriptionModel,
Audio: audioBytes,
})

Streaming Transcription (Scribe v2 Realtime)​

scribe_v2_realtime is a streaming-only model — DoTranscribe (the batch path) rejects it. Use ai.ExperimentalStreamTranscribe instead:

liveModel, err := provider.TranscriptionModel("scribe_v2_realtime")
if err != nil {
log.Fatal(err)
}

stream, err := ai.ExperimentalStreamTranscribe(ctx, ai.StreamTranscribeOptions{
Model: liveModel,
Audio: audioStream,
})

ELEVENLABS_API_KEY is read automatically when no API key is passed to Config.

Workflow Serialization​

ElevenLabs speech and transcription models can cross a workflow boundary with provider.SerializeSpeechModel / DeserializeSpeechModel and provider.SerializeTranscriptionModel / DeserializeTranscriptionModel. See Provider Serialization for the mechanism.

See Also​