Skip to main content

Google Provider

Google provides the Gemini family of models through Google AI Studio. Gemini models offer exceptional multimodal capabilities, massive context windows (up to 1M+ tokens), and competitive pricing.

Setup​

Installation​

The Google provider is included in the Go-AI SDK:

import (
"github.com/digitallysavvy/go-ai/pkg/ai"
providerapi "github.com/digitallysavvy/go-ai/pkg/provider"
"github.com/digitallysavvy/go-ai/pkg/providers/google"
)

Configuration​

googleProvider := google.New(google.Config{
APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"),
})

model, err := googleProvider.LanguageModel("gemini-2.0-flash")
if err != nil {
log.Fatal(err)
}

Get API Key​

  1. Visit makersuite.google.com/app/apikey
  2. Create new API key
  3. Set environment variable:
export GOOGLE_GENERATIVE_AI_API_KEY=AI...

Available Models​

Gemini 3 / 2.5 Series​

Model IDBest For
gemini-3.1-pro-previewAdvanced text generation
gemini-3.1-pro-preview-customtoolsAdvanced text generation with custom tools
gemini-3-flash-previewFast preview generation
gemini-3-pro-previewAdvanced preview generation
gemini-2.5-proComplex reasoning and multimodal tasks
gemini-2.5-flashFast multimodal tasks
gemini-2.5-flash-liteCost-effective high-volume tasks

Gemini 2.0 Series​

Model IDContextInput PriceOutput PriceBest For
gemini-2.0-flash1M$0.075/1M (≤128K)
$0.15/1M (>128K)
$0.30/1M (≤128K)
$0.60/1M (>128K)
Fast, multimodal
gemini-2.0-flash-thinking1M$0.075/1M (≤128K)
$0.15/1M (>128K)
$0.30/1M (≤128K)
$0.60/1M (>128K)
Reasoning tasks

Gemini 1.5 Series​

Model IDContextInput PriceOutput PriceBest For
gemini-1.5-pro2M$1.25/1M (≤128K)
$2.50/1M (>128K)
$5.00/1M (≤128K)
$10.00/1M (>128K)
Complex analysis
gemini-1.5-flash1M$0.075/1M (≤128K)
$0.15/1M (>128K)
$0.30/1M (≤128K)
$0.60/1M (>128K)
Fast responses
gemini-1.5-flash-8b1M$0.0375/1M (≤128K)
$0.075/1M (>128K)
$0.15/1M (≤128K)
$0.30/1M (>128K)
Cost-effective

Gemini 1.0 Series (Legacy)​

Model IDContextInput PriceOutput PriceBest For
gemini-1.0-pro32K$0.50/1M$1.50/1MGeneral purpose

Image Generation Models​

Breaking change: non-Gemini Imagen models (imagen-4.0-*, etc.) were removed from the Google provider, matching the TypeScript SDK. Calling provider.ImageModel(id) with a model ID that doesn't start with gemini- now fails with: "Google image models other than Gemini are no longer supported. Use a model ID that starts with gemini-." Use Google Vertex AI if you still need Imagen.

Model IDAspect RatiosPriceBest For
gemini-2.5-flash-image1:1, 3:4, 4:3, 9:16, 16:9Per imageFast Gemini generation
gemini-3-pro-image-preview1:1, 3:4, 4:3, 9:16, 16:9Per imageAdvanced generation

Gemini image models ignore N (they generate one image per call and are auto-batched by ai.GenerateImage to reach the requested count) instead of erroring on N > 1.

Gemini image models can use Google Search grounding during image generation through provider options:

result, err := imageModel.DoGenerate(ctx, &provider.ImageGenerateOptions{
Prompt: "create a current-events infographic",
ProviderOptions: map[string]interface{}{
"google": map[string]interface{}{
"googleSearch": map[string]interface{}{},
},
},
})

googleSearch is only supported for Gemini image models — the only image models this provider exposes now that Imagen has been removed.

Embedding Models​

Model IDDimensionsPriceBest For
gemini-embedding-0013072See Google pricingMultimodal semantic search
gemini-embedding-23072See Google pricingStable Gemini embedding workloads
gemini-embedding-2-previewProvider-definedSee Google pricingPreview embedding workloads

Gemini TTS Speech Models​

Google Gemini TTS is exposed through Provider.SpeechModel and Provider.Speech.

Model IDGo constant
gemini-2.5-flash-preview-ttsgoogle.ModelGemini25FlashTTS
gemini-2.5-pro-preview-ttsgoogle.ModelGemini25ProTTS
gemini-3.1-flash-tts-previewgoogle.ModelGemini31FlashTTSPreview
speechModel, err := googleProvider.SpeechModel(google.ModelGemini25FlashTTS)
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{
Model: speechModel,
Text: "Welcome to Go-AI.",
Voice: "Kore",
})
if err != nil {
log.Fatal(err)
}

fmt.Println(result.Audio.MediaType)

Gemini TTS defaults to the Kore voice when no voice is supplied. Use ProviderOptions["google"] with google.GoogleSpeechModelOptions to pass Gemini-specific speech options such as multi-speaker voice configuration.

Realtime Models​

Gemini Live is exposed through Provider.RealtimeModel, ExperimentalRealtimeModel, and GetRealtimeToken. Model IDs are strings, matching the TypeScript GoogleRealtimeModelId surface.

model, err := googleProvider.RealtimeModel("gemini-3.1-flash-live-preview")
if err != nil {
log.Fatal(err)
}

secretCreator, ok := model.(providerapi.RealtimeClientSecretCreator)
if !ok {
log.Fatal("model does not support client secret creation")
}
token, err := secretCreator.DoCreateClientSecret(ctx, providerapi.ClientSecretOptions{})
if err != nil {
log.Fatal(err)
}

wsConfigProvider, ok := model.(providerapi.RealtimeWebSocketConfigProvider)
if !ok {
log.Fatal("model does not support client-side WebSocket config")
}
ws := wsConfigProvider.GetWebSocketConfig(token.Token, token.URL)
fmt.Println(ws.URL)

Use ai.ConnectRealtime for a provider-neutral Go session helper. Browser-only TypeScript pieces such as BrowserRealtimeTransport, BrowserRealtimeAudio, and useRealtime are not part of the Go runtime.

Gemini Live Translate configuration can be passed through ProviderOptions["google"]["translationConfig"]. The provider merges it into the realtime session generationConfig.translationConfig, matching the TypeScript Live Translate request shape.

Background-reasoning Live models (model IDs matching gemini-<major>.<minor>-live...thinking, e.g. gemini-3.8-live-extended-thinking) require a thinkingLevel or thinkingBudget in the session setup. When neither is set — via the typed ProviderOptions["google"]["thinkingConfig"] or a raw ProviderOptions["generationConfig"]["thinkingConfig"] — the provider adds thinkingLevel: "low" while keeping any other fields (e.g. includeThoughts) and without disturbing an explicit thinkingBudget: 0:

setup := map[string]interface{}{
"providerOptions": map[string]interface{}{
"google": map[string]interface{}{
"thinkingConfig": map[string]interface{}{"includeThoughts": true},
},
},
}
// -> generationConfig.thinkingConfig == {"includeThoughts": true, "thinkingLevel": "low"}

Input transcription events (input-transcription-completed) accumulate consecutive fragments for one user utterance under a stable synthetic item ID, and allocate a new ID once Google signals the utterance is finished or the turn completes; this keeps multi-turn and barge-in (interrupted) conversations from overwriting or misgrouping earlier user messages.

Speech Translation (experimental)​

Provider.SpeechTranslationModel(modelID) (alias Provider.Translation) returns a speech-to-speech translation model over the same Live API BidiGenerateContent WebSocket, for use with ai.ExperimentalStreamTranslate:

translationModel, err := googleProvider.SpeechTranslationModel("gemini-3.5-live-translate-preview")
if err != nil {
log.Fatal(err)
}

result, err := ai.ExperimentalStreamTranslate(ctx, ai.StreamTranslateOptions{
Model: translationModel,
Audio: audioStream,
TargetLanguage: "es",
})
if err != nil {
log.Fatal(err)
}

stream, err := result.FullStream()
if err != nil {
log.Fatal(err)
}
defer stream.Close()

for {
part, err := stream.Next()
if err == io.EOF {
break
}
if err != nil {
log.Fatal(err)
}
if part.Type == "audio" {
// part.AudioData carries translated PCM audio.
}
}

Provider-Specific Features​

Service Tier​

For the Gemini API, pass serviceTier in Google provider options:

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "hello",
ProviderOptions: map[string]interface{}{
"google": map[string]interface{}{
"serviceTier": "priority",
},
},
})

Massive Context Windows​

Gemini supports up to 2M tokens of context:

// Process entire codebases or books
hugeDocument := loadDocument("book.txt") // up to 2M tokens

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: fmt.Sprintf("Summarize this book:\n\n%s", hugeDocument),
})

Multimodal Capabilities​

Gemini excels at processing text, images, video, and audio together:

// Image understanding
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Describe this image in detail"},
types.FileContent{
URL: "https://example.com/photo.jpg",
MediaType: "image/jpeg",
},
},
},
},
})

// Multiple images
result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Compare these images"},
types.FileContent{
URL: "https://example.com/before.jpg",
MediaType: "image/jpeg",
},
types.FileContent{
URL: "https://example.com/after.jpg",
MediaType: "image/jpeg",
},
},
},
},
})

// Video understanding
result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "What happens in this video?"},
types.FileContent{
URL: "https://example.com/video.mp4",
MediaType: "video/mp4",
},
},
},
},
})

Function Calling​

Gemini supports sophisticated function calling:

weatherTool := types.Tool{
Name: "get_weather",
Description: "Get current weather for a location",
Parameters: map[string]interface{}{
"type": "object",
"properties": map[string]interface{}{
"location": map[string]interface{}{
"type": "string",
"description": "City and country",
},
"unit": map[string]interface{}{
"type": "string",
"enum": []string{"celsius", "fahrenheit"},
},
},
"required": []string{"location"},
},
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "What's the weather in Tokyo?",
Tools: []types.Tool{weatherTool},
StopWhen: []ai.StopCondition{ai.IsStepCount(5)},
})

for _, call := range result.ToolCalls {
fmt.Printf("Function: %s\nArgs: %v\n", call.ToolName, call.Arguments)
}

Interactions API​

The Google provider exposes the Gemini Interactions API through Provider.Interactions(modelID) and agent presets through Provider.InteractionsAgent(agent). Interactions models implement the standard provider.LanguageModel interface, so they can be used with DoGenerate, DoStream, and higher-level AI SDK helpers.

store := false
interactionsModel, err := googleProvider.Interactions(google.InteractionsModelGemini25Flash)
if err != nil {
log.Fatal(err)
}

result, err := interactionsModel.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: types.Prompt{Text: "Give a concise answer."},
ProviderOptions: map[string]interface{}{
"google": google.GoogleInteractionsProviderOptions{
Store: &store,
ResponseModalities: []string{"text"},
ThinkingLevel: "low",
},
},
})
if err != nil {
log.Fatal(err)
}

fmt.Println(result.Text)

Use PreviousInteractionID with Store for stateful follow-up calls. Agent models run as background interactions and are polled until a terminal status; PollingTimeoutMs controls that wait. Agent model IDs include deep-research-pro-preview-12-2025, deep-research-preview-04-2026, deep-research-max-preview-04-2026, and antigravity-preview-05-2026.

agentModel, err := googleProvider.InteractionsAgent(google.InteractionsAgentDeepResearchProPreview)
if err != nil {
log.Fatal(err)
}

result, err := agentModel.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: types.Prompt{Text: "Research the main tradeoffs in retrieval augmented generation."},
ProviderOptions: map[string]interface{}{
"google": google.GoogleInteractionsProviderOptions{
PollingTimeoutMs: 120000,
},
},
})

Interactions responses preserve ordered text, reasoning, files, sources, and tool-call content. Google metadata is available under ProviderMetadata["google"], including the interactionId when returned by the API.

Connect Gemini to real-time web information:

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "What are the latest developments in quantum computing?",
Tools: []types.Tool{google.GoogleSearchTool()},
})

// Response includes citations
for _, source := range result.Sources {
fmt.Printf("Source: %s\n", source.URL)
}

JSON Mode​

Structured output generation:

import "github.com/digitallysavvy/go-ai/pkg/schema"

bookSchema := schema.NewSimpleJSONSchema(map[string]interface{}{
"type": "object",
"properties": map[string]interface{}{
"title": map[string]string{"type": "string"},
"author": map[string]string{"type": "string"},
"year": map[string]string{"type": "integer"},
"genre": map[string]interface{}{
"type": "array",
"items": map[string]string{"type": "string"},
},
},
"required": []string{"title", "author", "year"},
})

result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{
Model: model,
Schema: bookSchema,
Prompt: "Extract book metadata: '1984 by George Orwell, published 1949'",
})

Streaming​

Real-time response streaming:

stream, err := ai.StreamText(ctx, ai.StreamTextOptions{
Model: model,
Prompt: "Write a long story",
})
if err != nil {
log.Fatal(err)
}
defer stream.Close()

for chunk := range stream.Chunks() {
fmt.Print(chunk.Text)
}

Thinking Mode​

Gemini 2.0 Flash Thinking uses extended reasoning:

model, err := googleProvider.LanguageModel(google.ModelGemini20FlashThinkingExp)

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Solve this logic puzzle: Three friends...",
})

// Reasoning content is preserved in ordered content parts when returned.
for _, part := range result.Content {
if reasoning, ok := part.(types.ReasoningContent); ok {
fmt.Println("Reasoning:", reasoning.Text)
}
}
fmt.Println("Answer:", result.Text)

Gemini 3 thinking and tool-call streams preserve thoughtSignature in provider metadata so multi-step calls can safely round-trip model-generated reasoning and function calls. Streaming no-argument function calls are emitted with {} input instead of being dropped.

Usage includes per-modality token details when the API returns them. Text, image, audio, and video token counts are preserved in the result usage details rather than collapsed into only total input/output tokens.

Use the "google" provider key for AI Studio options. Vertex options use the "googleVertex" key (see the Google Vertex AI provider for its lookup order) — the hyphenated "google-vertex" string is only the Vertex model's Provider() identity, not a ProviderOptions key.

Examples​

Basic Text Generation​

package main

import (
"context"
"fmt"
"log"
"os"

"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/google"
)

func main() {
ctx := context.Background()
googleProvider := google.New(google.Config{
APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"),
})

model, err := googleProvider.LanguageModel("gemini-2.0-flash")
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Explain machine learning in simple terms",
})
if err != nil {
log.Fatal(err)
}

fmt.Println(result.Text)
fmt.Printf("Tokens: %d\n", result.Usage.GetTotalTokens())
}

Image Analysis​

imageData, err := os.ReadFile("chart.png")
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Analyze this chart and extract key insights"},
types.FileContent{
Data: imageData,
MediaType: "image/png",
},
},
},
},
})

fmt.Println(result.Text)

Video Analysis​

// Upload video file
videoData, err := os.ReadFile("demo.mp4")
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Describe what happens in this video"},
types.FileContent{
Data: videoData,
MediaType: "video/mp4",
},
},
},
},
})

fmt.Println(result.Text)

Document Processing with Large Context​

// Load multiple documents
docs := []string{
loadFile("doc1.txt"),
loadFile("doc2.txt"),
loadFile("doc3.txt"),
}
combined := strings.Join(docs, "\n\n---\n\n")

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: fmt.Sprintf("Analyze these documents and identify common themes:\n\n%s", combined),
})

Grounded Search Responses​

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "What are the current COVID-19 vaccination rates in major countries?",
Tools: []types.Tool{google.GoogleSearchTool()},
})

fmt.Println("Answer:", result.Text)

fmt.Println("\nSources:")
for _, source := range result.Sources {
fmt.Printf("- %s: %s\n", source.Title, source.URL)
}

Text Embeddings​

embeddingModel, err := googleProvider.EmbeddingModel(google.EmbeddingModelGeminiEmbedding001)
if err != nil {
log.Fatal(err)
}

dimensions := 768
result, err := embeddingModel.DoEmbed(ctx, "Go is a programming language", &provider.EmbedModelOptions{
ProviderOptions: map[string]interface{}{
"google": google.GoogleEmbeddingProviderOptions{
TaskType: "RETRIEVAL_DOCUMENT",
OutputDimensionality: &dimensions,
},
},
})
if err != nil {
log.Fatal(err)
}

fmt.Printf("Dimensions: %d\n", len(result.Embedding))

Google embeddings also support multimodal content and provider-hosted files through fileData:

result, err := embeddingModel.DoEmbed(ctx, "Summarize this document for retrieval.", &provider.EmbedModelOptions{
ProviderOptions: map[string]interface{}{
"google": google.GoogleEmbeddingProviderOptions{
TaskType: "RETRIEVAL_DOCUMENT",
Content: [][]google.EmbeddingPart{
{
google.FileDataEmbeddingPart{
MimeType: "application/pdf",
FileURI: "files/abc123",
},
},
},
},
},
})
if err != nil {
log.Fatal(err)
}

Multi-Turn Conversation​

messages := []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "What is the capital of France?"},
},
},
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: messages,
})
messages = append(messages,
types.Message{
Role: types.RoleAssistant,
Content: []types.ContentPart{
types.TextContent{Text: result.Text},
},
},
types.Message{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "What's the population?"},
},
},
)

result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: messages,
})

Image Generation​

Google's text-to-image generation is now Gemini-only — Imagen model IDs (imagen-4.0-*) were removed from this provider; use Google Vertex AI for Imagen.

Gemini Image Models offer fast generation:

// Gemini 2.5 Flash - Fast image generation
geminiImageModel, err := googleProvider.ImageModel("gemini-2.5-flash-image")
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{
Model: geminiImageModel,
Prompt: "Abstract colorful art with geometric shapes and vibrant patterns",
Size: "1920x1080", // 16:9 aspect ratio
})
if err != nil {
log.Fatal(err)
}

fmt.Printf("Image generated: %d bytes\n", len(result.Images[0].Data))
os.WriteFile("gemini_output.png", result.Images[0].Data, 0644)

Supported Aspect Ratios:

  • 1:1 - Square (1024x1024, 512x512)
  • 4:3 - Standard (1024x768)
  • 3:4 - Portrait (768x1024)
  • 16:9 - Widescreen (1920x1080, 1792x1024)
  • 9:16 - Vertical (1080x1920, 1024x1792)

Model Comparison:

ModelSpeedQualityUse Case
gemini-2.5-flash-imageVery FastGoodQuick iterations
gemini-3-pro-image-previewFastHighAdvanced generation

Note: For Imagen models and image editing, use Google Vertex AI.

Video Generation​

provider.VideoModel(modelID) submits to Veo's :predictLongRunning endpoint and polls the long-running operation until it completes (MaxVideosPerCall() returns 4 — Veo accepts up to 4 videos per call, and ai.GenerateVideo fans out across multiple calls automatically when you ask for more):

videoModel, err := googleProvider.VideoModel("veo-3.1-generate-preview")
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{Text: "A drone shot flying over a coastline at sunrise"},
})
if err != nil {
log.Fatal(err)
}

for _, video := range result.Videos {
fmt.Println(video.URL)
}

VideoModel also implements provider.VideoModelStarter / VideoModelStatusChecker, so it works with the async ai.ExperimentalStartVideo / ai.ExperimentalGetVideoStatus calls.

Advanced Configuration​

Custom HTTP Client​

googleProvider := google.New(google.Config{
APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"),
HTTPClient: &http.Client{
Timeout: time.Minute * 5,
Transport: &http.Transport{
MaxIdleConns: 100,
IdleConnTimeout: 90 * time.Second,
},
},
})

Safety Settings​

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: prompt,
ProviderOptions: map[string]interface{}{
"google": map[string]interface{}{
"safetySettings": []map[string]string{
{"category": "HARM_CATEGORY_HATE_SPEECH", "threshold": "BLOCK_LOW_AND_ABOVE"},
{"category": "HARM_CATEGORY_DANGEROUS_CONTENT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"},
{"category": "HARM_CATEGORY_SEXUALLY_EXPLICIT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"},
{"category": "HARM_CATEGORY_HARASSMENT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"},
},
},
},
})

Generation Config​

temperature := 0.9
topP := 0.95
topK := 40
maxTokens := 2048
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: prompt,
Temperature: &temperature,
TopP: &topP,
TopK: &topK,
MaxTokens: &maxTokens,
})

Error Handling​

Common Errors​

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: prompt,
})
if err != nil {
var providerErr *providererrors.ProviderError
if errors.As(err, &providerErr) {
switch providerErr.StatusCode {
case 400:
log.Printf("Invalid request: %s", providerErr.Message)
case 401:
log.Fatal("Invalid API key")
case 403:
log.Fatal("Permission denied or quota exceeded")
case 404:
log.Fatal("Model not found")
case 429:
log.Println("Rate limited, retrying...")
time.Sleep(time.Second * 5)
case 500, 503:
log.Println("Service error, retrying...")
time.Sleep(time.Second * 10)
default:
log.Printf("Google AI error: %d - %s", providerErr.StatusCode, providerErr.Message)
}
}
return nil, err
}

Content Filtering​

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: prompt,
})
if err != nil {
var providerErr *providererrors.ProviderError
if errors.As(err, &providerErr) {
if providerErr.StatusCode == 400 && strings.Contains(providerErr.Message, "SAFETY") {
log.Println("Content blocked by safety filters")
// Try with adjusted safety settings
}
}
}

Best Practices​

  1. Model Selection

    • Use gemini-2.0-flash for fast, multimodal tasks
    • Use gemini-1.5-pro for complex reasoning with huge context
    • Use gemini-1.5-flash-8b for cost-effective high-volume tasks
  2. Context Management

    • Take advantage of massive context windows
    • Include full documents rather than chunking
    • Use structured prompts for better results
  3. Multimodal Input

    • Compress images appropriately (max 20MB)
    • Use clear, specific instructions for vision tasks
    • Combine multiple modalities for richer context
  4. Cost Optimization

    • Monitor context length (pricing changes after 128K)
    • Use flash-8b for simple tasks
    • Cache repeated queries on your side
  5. Safety

    • Adjust safety settings based on use case
    • Handle content filtering gracefully
    • Log filtered content for review

Rate Limits & Pricing​

Rate Limits (Free Tier)​

ModelRPMTPMRPD
Gemini 2.0 Flash151M1,500
Gemini 1.5 Pro232K50
Gemini 1.5 Flash151M1,500

RPM = Requests per minute, TPM = Tokens per minute, RPD = Requests per day

Paid tier offers significantly higher limits.

Context-Based Pricing​

func calculateGeminiCost(inputTokens, outputTokens int, model string) float64 {
// Pricing changes based on context length
var inputRate, outputRate float64

if inputTokens <= 128000 {
// Low context pricing
inputRate = 0.075 / 1_000_000 // for flash models
outputRate = 0.30 / 1_000_000
} else {
// High context pricing (2x)
inputRate = 0.15 / 1_000_000
outputRate = 0.60 / 1_000_000
}

return float64(inputTokens)*inputRate + float64(outputTokens)*outputRate
}

Workflow Serialization​

Google embedding, image, and (unary) transcription models can cross a workflow boundary with providerutils.SerializeModel / DeserializeModel (language models could already be serialized). See Provider Serialization for the mechanism; speech, speech-translation, video, and realtime models are not yet serializable.

See Also​

May 2026 parity updates​

Model refresh and token usage​

The Google provider tracks the refreshed Gemini model IDs in pkg/providers/google/model_ids.go. Usage details include per-modality token fields through types.Usage.InputDetails.TextTokens and types.Usage.InputDetails.ImageTokens when Google returns modality-specific counts.

Thought signatures and no-arg tool calls​

The provider preserves Google thoughtSignature metadata on text and tool-call parts so multi-turn requests can forward model reasoning state. Streaming tool calls with no arguments are emitted with an empty argument object, matching the TypeScript SDK.

Provider options​

Use the standard ProviderOptions map for Google-specific options such as thinking configuration. The current Go provider accepts the same nested thinkingConfig request shape that the TypeScript provider sends to Gemini.

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Use Google search if useful.",
ProviderOptions: map[string]interface{}{
"google": map[string]interface{}{
"thinkingConfig": map[string]interface{}{
"thinkingBudget": 512,
"includeThoughts": true,
},
},
},
})