Google Provider
Google provides the Gemini family of models through Google AI Studio. Gemini models offer exceptional multimodal capabilities, massive context windows (up to 1M+ tokens), and competitive pricing.
Setup
Installation
The Google provider is included in the Go-AI SDK:
import (
"github.com/digitallysavvy/go-ai/pkg/ai"
providerapi "github.com/digitallysavvy/go-ai/pkg/provider"
"github.com/digitallysavvy/go-ai/pkg/providers/google"
)
Configuration
googleProvider := google.New(google.Config{
APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"),
})
model, err := googleProvider.LanguageModel("gemini-2.0-flash")
if err != nil {
log.Fatal(err)
}
Get API Key
- Visit makersuite.google.com/app/apikey
- Create new API key
- Set environment variable:
export GOOGLE_GENERATIVE_AI_API_KEY=AI...
Available Models
Gemini 3 / 2.5 Series
| Model ID | Best For |
|---|---|
| gemini-3.1-pro-preview | Advanced text generation |
| gemini-3.1-pro-preview-customtools | Advanced text generation with custom tools |
| gemini-3-flash-preview | Fast preview generation |
| gemini-3-pro-preview | Advanced preview generation |
| gemini-2.5-pro | Complex reasoning and multimodal tasks |
| gemini-2.5-flash | Fast multimodal tasks |
| gemini-2.5-flash-lite | Cost-effective high-volume tasks |
Gemini 2.0 Series
| Model ID | Context | Input Price | Output Price | Best For |
|---|---|---|---|---|
| gemini-2.0-flash | 1M | $0.075/1M (≤128K) $0.15/1M (>128K) | $0.30/1M (≤128K) $0.60/1M (>128K) | Fast, multimodal |
| gemini-2.0-flash-thinking | 1M | $0.075/1M (≤128K) $0.15/1M (>128K) | $0.30/1M (≤128K) $0.60/1M (>128K) | Reasoning tasks |
Gemini 1.5 Series
| Model ID | Context | Input Price | Output Price | Best For |
|---|---|---|---|---|
| gemini-1.5-pro | 2M | $1.25/1M (≤128K) $2.50/1M (>128K) | $5.00/1M (≤128K) $10.00/1M (>128K) | Complex analysis |
| gemini-1.5-flash | 1M | $0.075/1M (≤128K) $0.15/1M (>128K) | $0.30/1M (≤128K) $0.60/1M (>128K) | Fast responses |
| gemini-1.5-flash-8b | 1M | $0.0375/1M (≤128K) $0.075/1M (>128K) | $0.15/1M (≤128K) $0.30/1M (>128K) | Cost-effective |
Gemini 1.0 Series (Legacy)
| Model ID | Context | Input Price | Output Price | Best For |
|---|---|---|---|---|
| gemini-1.0-pro | 32K | $0.50/1M | $1.50/1M | General purpose |
Image Generation Models
Breaking change: non-Gemini Imagen models (
imagen-4.0-*, etc.) were removed from the Google provider, matching the TypeScript SDK. Callingprovider.ImageModel(id)with a model ID that doesn't start withgemini-now fails with: "Google image models other than Gemini are no longer supported. Use a model ID that starts withgemini-." Use Google Vertex AI if you still need Imagen.
| Model ID | Aspect Ratios | Price | Best For |
|---|---|---|---|
| gemini-2.5-flash-image | 1:1, 3:4, 4:3, 9:16, 16:9 | Per image | Fast Gemini generation |
| gemini-3-pro-image-preview | 1:1, 3:4, 4:3, 9:16, 16:9 | Per image | Advanced generation |
Gemini image models ignore N (they generate one image per call and are
auto-batched by ai.GenerateImage to reach the requested count) instead of
erroring on N > 1.
Gemini image models can use Google Search grounding during image generation through provider options:
result, err := imageModel.DoGenerate(ctx, &provider.ImageGenerateOptions{
Prompt: "create a current-events infographic",
ProviderOptions: map[string]interface{}{
"google": map[string]interface{}{
"googleSearch": map[string]interface{}{},
},
},
})
googleSearch is only supported for Gemini image models — the only image models this provider exposes now that Imagen has been removed.
Embedding Models
| Model ID | Dimensions | Price | Best For |
|---|---|---|---|
| gemini-embedding-001 | 3072 | See Google pricing | Multimodal semantic search |
| gemini-embedding-2 | 3072 | See Google pricing | Stable Gemini embedding workloads |
| gemini-embedding-2-preview | Provider-defined | See Google pricing | Preview embedding workloads |
Gemini TTS Speech Models
Google Gemini TTS is exposed through Provider.SpeechModel and Provider.Speech.
| Model ID | Go constant |
|---|---|
| gemini-2.5-flash-preview-tts | google.ModelGemini25FlashTTS |
| gemini-2.5-pro-preview-tts | google.ModelGemini25ProTTS |
| gemini-3.1-flash-tts-preview | google.ModelGemini31FlashTTSPreview |
speechModel, err := googleProvider.SpeechModel(google.ModelGemini25FlashTTS)
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{
Model: speechModel,
Text: "Welcome to Go-AI.",
Voice: "Kore",
})
if err != nil {
log.Fatal(err)
}
fmt.Println(result.Audio.MediaType)
Gemini TTS defaults to the Kore voice when no voice is supplied. Use ProviderOptions["google"] with google.GoogleSpeechModelOptions to pass Gemini-specific speech options such as multi-speaker voice configuration.
Realtime Models
Gemini Live is exposed through Provider.RealtimeModel,
ExperimentalRealtimeModel, and GetRealtimeToken. Model IDs are strings, matching
the TypeScript GoogleRealtimeModelId surface.
model, err := googleProvider.RealtimeModel("gemini-3.1-flash-live-preview")
if err != nil {
log.Fatal(err)
}
secretCreator, ok := model.(providerapi.RealtimeClientSecretCreator)
if !ok {
log.Fatal("model does not support client secret creation")
}
token, err := secretCreator.DoCreateClientSecret(ctx, providerapi.ClientSecretOptions{})
if err != nil {
log.Fatal(err)
}
wsConfigProvider, ok := model.(providerapi.RealtimeWebSocketConfigProvider)
if !ok {
log.Fatal("model does not support client-side WebSocket config")
}
ws := wsConfigProvider.GetWebSocketConfig(token.Token, token.URL)
fmt.Println(ws.URL)
Use ai.ConnectRealtime for a provider-neutral Go session helper. Browser-only
TypeScript pieces such as BrowserRealtimeTransport, BrowserRealtimeAudio,
and useRealtime are not part of the Go runtime.
Gemini Live Translate configuration can be passed through ProviderOptions["google"]["translationConfig"]. The provider merges it into the realtime session generationConfig.translationConfig, matching the TypeScript Live Translate request shape.
Background-reasoning Live models (model IDs matching gemini-<major>.<minor>-live...thinking, e.g. gemini-3.8-live-extended-thinking) require a thinkingLevel or thinkingBudget in the session setup. When neither is set — via the typed ProviderOptions["google"]["thinkingConfig"] or a raw ProviderOptions["generationConfig"]["thinkingConfig"] — the provider adds thinkingLevel: "low" while keeping any other fields (e.g. includeThoughts) and without disturbing an explicit thinkingBudget: 0:
setup := map[string]interface{}{
"providerOptions": map[string]interface{}{
"google": map[string]interface{}{
"thinkingConfig": map[string]interface{}{"includeThoughts": true},
},
},
}
// -> generationConfig.thinkingConfig == {"includeThoughts": true, "thinkingLevel": "low"}
Input transcription events (input-transcription-completed) accumulate consecutive fragments for one user utterance under a stable synthetic item ID, and allocate a new ID once Google signals the utterance is finished or the turn completes; this keeps multi-turn and barge-in (interrupted) conversations from overwriting or misgrouping earlier user messages.
Speech Translation (experimental)
Provider.SpeechTranslationModel(modelID) (alias Provider.Translation)
returns a speech-to-speech translation model over the same Live API
BidiGenerateContent WebSocket, for use with ai.ExperimentalStreamTranslate:
translationModel, err := googleProvider.SpeechTranslationModel("gemini-3.5-live-translate-preview")
if err != nil {
log.Fatal(err)
}
result, err := ai.ExperimentalStreamTranslate(ctx, ai.StreamTranslateOptions{
Model: translationModel,
Audio: audioStream,
TargetLanguage: "es",
})
if err != nil {
log.Fatal(err)
}
stream, err := result.FullStream()
if err != nil {
log.Fatal(err)
}
defer stream.Close()
for {
part, err := stream.Next()
if err == io.EOF {
break
}
if err != nil {
log.Fatal(err)
}
if part.Type == "audio" {
// part.AudioData carries translated PCM audio.
}
}
Provider-Specific Features
Service Tier
For the Gemini API, pass serviceTier in Google provider options:
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "hello",
ProviderOptions: map[string]interface{}{
"google": map[string]interface{}{
"serviceTier": "priority",
},
},
})
Massive Context Windows
Gemini supports up to 2M tokens of context:
// Process entire codebases or books
hugeDocument := loadDocument("book.txt") // up to 2M tokens
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: fmt.Sprintf("Summarize this book:\n\n%s", hugeDocument),
})
Multimodal Capabilities
Gemini excels at processing text, images, video, and audio together:
// Image understanding
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Describe this image in detail"},
types.FileContent{
URL: "https://example.com/photo.jpg",
MediaType: "image/jpeg",
},
},
},
},
})
// Multiple images
result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Compare these images"},
types.FileContent{
URL: "https://example.com/before.jpg",
MediaType: "image/jpeg",
},
types.FileContent{
URL: "https://example.com/after.jpg",
MediaType: "image/jpeg",
},
},
},
},
})
// Video understanding
result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "What happens in this video?"},
types.FileContent{
URL: "https://example.com/video.mp4",
MediaType: "video/mp4",
},
},
},
},
})
Function Calling
Gemini supports sophisticated function calling:
weatherTool := types.Tool{
Name: "get_weather",
Description: "Get current weather for a location",
Parameters: map[string]interface{}{
"type": "object",
"properties": map[string]interface{}{
"location": map[string]interface{}{
"type": "string",
"description": "City and country",
},
"unit": map[string]interface{}{
"type": "string",
"enum": []string{"celsius", "fahrenheit"},
},
},
"required": []string{"location"},
},
}
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "What's the weather in Tokyo?",
Tools: []types.Tool{weatherTool},
StopWhen: []ai.StopCondition{ai.IsStepCount(5)},
})
for _, call := range result.ToolCalls {
fmt.Printf("Function: %s\nArgs: %v\n", call.ToolName, call.Arguments)
}
Interactions API
The Google provider exposes the Gemini Interactions API through Provider.Interactions(modelID) and agent presets through Provider.InteractionsAgent(agent). Interactions models implement the standard provider.LanguageModel interface, so they can be used with DoGenerate, DoStream, and higher-level AI SDK helpers.
store := false
interactionsModel, err := googleProvider.Interactions(google.InteractionsModelGemini25Flash)
if err != nil {
log.Fatal(err)
}
result, err := interactionsModel.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: types.Prompt{Text: "Give a concise answer."},
ProviderOptions: map[string]interface{}{
"google": google.GoogleInteractionsProviderOptions{
Store: &store,
ResponseModalities: []string{"text"},
ThinkingLevel: "low",
},
},
})
if err != nil {
log.Fatal(err)
}
fmt.Println(result.Text)
Use PreviousInteractionID with Store for stateful follow-up calls. Agent models run as background interactions and are polled until a terminal status; PollingTimeoutMs controls that wait. Agent model IDs include deep-research-pro-preview-12-2025, deep-research-preview-04-2026, deep-research-max-preview-04-2026, and antigravity-preview-05-2026.
agentModel, err := googleProvider.InteractionsAgent(google.InteractionsAgentDeepResearchProPreview)
if err != nil {
log.Fatal(err)
}
result, err := agentModel.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: types.Prompt{Text: "Research the main tradeoffs in retrieval augmented generation."},
ProviderOptions: map[string]interface{}{
"google": google.GoogleInteractionsProviderOptions{
PollingTimeoutMs: 120000,
},
},
})
Interactions responses preserve ordered text, reasoning, files, sources, and tool-call content. Google metadata is available under ProviderMetadata["google"], including the interactionId when returned by the API.
Grounding with Google Search
Connect Gemini to real-time web information:
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "What are the latest developments in quantum computing?",
Tools: []types.Tool{google.GoogleSearchTool()},
})
// Response includes citations
for _, source := range result.Sources {
fmt.Printf("Source: %s\n", source.URL)
}
JSON Mode
Structured output generation:
import "github.com/digitallysavvy/go-ai/pkg/schema"
bookSchema := schema.NewSimpleJSONSchema(map[string]interface{}{
"type": "object",
"properties": map[string]interface{}{
"title": map[string]string{"type": "string"},
"author": map[string]string{"type": "string"},
"year": map[string]string{"type": "integer"},
"genre": map[string]interface{}{
"type": "array",
"items": map[string]string{"type": "string"},
},
},
"required": []string{"title", "author", "year"},
})
result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{
Model: model,
Schema: bookSchema,
Prompt: "Extract book metadata: '1984 by George Orwell, published 1949'",
})
Streaming
Real-time response streaming:
stream, err := ai.StreamText(ctx, ai.StreamTextOptions{
Model: model,
Prompt: "Write a long story",
})
if err != nil {
log.Fatal(err)
}
defer stream.Close()
for chunk := range stream.Chunks() {
fmt.Print(chunk.Text)
}
Thinking Mode
Gemini 2.0 Flash Thinking uses extended reasoning:
model, err := googleProvider.LanguageModel(google.ModelGemini20FlashThinkingExp)
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Solve this logic puzzle: Three friends...",
})
// Reasoning content is preserved in ordered content parts when returned.
for _, part := range result.Content {
if reasoning, ok := part.(types.ReasoningContent); ok {
fmt.Println("Reasoning:", reasoning.Text)
}
}
fmt.Println("Answer:", result.Text)
Gemini 3 thinking and tool-call streams preserve thoughtSignature in provider metadata so multi-step calls can safely round-trip model-generated reasoning and function calls. Streaming no-argument function calls are emitted with {} input instead of being dropped.
Usage includes per-modality token details when the API returns them. Text, image, audio, and video token counts are preserved in the result usage details rather than collapsed into only total input/output tokens.
Use the "google" provider key for AI Studio options. Vertex options use the "googleVertex" key (see the Google Vertex AI provider for its lookup order) — the hyphenated "google-vertex" string is only the Vertex model's Provider() identity, not a ProviderOptions key.
Examples
Basic Text Generation
package main
import (
"context"
"fmt"
"log"
"os"
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/google"
)
func main() {
ctx := context.Background()
googleProvider := google.New(google.Config{
APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"),
})
model, err := googleProvider.LanguageModel("gemini-2.0-flash")
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Explain machine learning in simple terms",
})
if err != nil {
log.Fatal(err)
}
fmt.Println(result.Text)
fmt.Printf("Tokens: %d\n", result.Usage.GetTotalTokens())
}
Image Analysis
imageData, err := os.ReadFile("chart.png")
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Analyze this chart and extract key insights"},
types.FileContent{
Data: imageData,
MediaType: "image/png",
},
},
},
},
})
fmt.Println(result.Text)
Video Analysis
// Upload video file
videoData, err := os.ReadFile("demo.mp4")
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Describe what happens in this video"},
types.FileContent{
Data: videoData,
MediaType: "video/mp4",
},
},
},
},
})
fmt.Println(result.Text)
Document Processing with Large Context
// Load multiple documents
docs := []string{
loadFile("doc1.txt"),
loadFile("doc2.txt"),
loadFile("doc3.txt"),
}
combined := strings.Join(docs, "\n\n---\n\n")
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: fmt.Sprintf("Analyze these documents and identify common themes:\n\n%s", combined),
})
Grounded Search Responses
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "What are the current COVID-19 vaccination rates in major countries?",
Tools: []types.Tool{google.GoogleSearchTool()},
})
fmt.Println("Answer:", result.Text)
fmt.Println("\nSources:")
for _, source := range result.Sources {
fmt.Printf("- %s: %s\n", source.Title, source.URL)
}
Text Embeddings
embeddingModel, err := googleProvider.EmbeddingModel(google.EmbeddingModelGeminiEmbedding001)
if err != nil {
log.Fatal(err)
}
dimensions := 768
result, err := embeddingModel.DoEmbed(ctx, "Go is a programming language", &provider.EmbedModelOptions{
ProviderOptions: map[string]interface{}{
"google": google.GoogleEmbeddingProviderOptions{
TaskType: "RETRIEVAL_DOCUMENT",
OutputDimensionality: &dimensions,
},
},
})
if err != nil {
log.Fatal(err)
}
fmt.Printf("Dimensions: %d\n", len(result.Embedding))
Google embeddings also support multimodal content and provider-hosted files through fileData:
result, err := embeddingModel.DoEmbed(ctx, "Summarize this document for retrieval.", &provider.EmbedModelOptions{
ProviderOptions: map[string]interface{}{
"google": google.GoogleEmbeddingProviderOptions{
TaskType: "RETRIEVAL_DOCUMENT",
Content: [][]google.EmbeddingPart{
{
google.FileDataEmbeddingPart{
MimeType: "application/pdf",
FileURI: "files/abc123",
},
},
},
},
},
})
if err != nil {
log.Fatal(err)
}
Multi-Turn Conversation
messages := []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "What is the capital of France?"},
},
},
}
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: messages,
})
messages = append(messages,
types.Message{
Role: types.RoleAssistant,
Content: []types.ContentPart{
types.TextContent{Text: result.Text},
},
},
types.Message{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "What's the population?"},
},
},
)
result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: messages,
})
Image Generation
Google's text-to-image generation is now Gemini-only — Imagen model IDs
(imagen-4.0-*) were removed from this provider; use Google Vertex AI for
Imagen.
Gemini Image Models offer fast generation:
// Gemini 2.5 Flash - Fast image generation
geminiImageModel, err := googleProvider.ImageModel("gemini-2.5-flash-image")
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{
Model: geminiImageModel,
Prompt: "Abstract colorful art with geometric shapes and vibrant patterns",
Size: "1920x1080", // 16:9 aspect ratio
})
if err != nil {
log.Fatal(err)
}
fmt.Printf("Image generated: %d bytes\n", len(result.Images[0].Data))
os.WriteFile("gemini_output.png", result.Images[0].Data, 0644)
Supported Aspect Ratios:
1:1- Square (1024x1024, 512x512)4:3- Standard (1024x768)3:4- Portrait (768x1024)16:9- Widescreen (1920x1080, 1792x1024)9:16- Vertical (1080x1920, 1024x1792)
Model Comparison:
| Model | Speed | Quality | Use Case |
|---|---|---|---|
| gemini-2.5-flash-image | Very Fast | Good | Quick iterations |
| gemini-3-pro-image-preview | Fast | High | Advanced generation |
Note: For Imagen models and image editing, use Google Vertex AI.
Video Generation
provider.VideoModel(modelID) submits to Veo's :predictLongRunning
endpoint and polls the long-running operation until it completes
(MaxVideosPerCall() returns 4 — Veo accepts up to 4 videos per call, and
ai.GenerateVideo fans out across multiple calls automatically when you ask
for more):
videoModel, err := googleProvider.VideoModel("veo-3.1-generate-preview")
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{Text: "A drone shot flying over a coastline at sunrise"},
})
if err != nil {
log.Fatal(err)
}
for _, video := range result.Videos {
fmt.Println(video.URL)
}
VideoModel also implements provider.VideoModelStarter /
VideoModelStatusChecker, so it works with the async
ai.ExperimentalStartVideo / ai.ExperimentalGetVideoStatus calls.
Advanced Configuration
Custom HTTP Client
googleProvider := google.New(google.Config{
APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"),
HTTPClient: &http.Client{
Timeout: time.Minute * 5,
Transport: &http.Transport{
MaxIdleConns: 100,
IdleConnTimeout: 90 * time.Second,
},
},
})
Safety Settings
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: prompt,
ProviderOptions: map[string]interface{}{
"google": map[string]interface{}{
"safetySettings": []map[string]string{
{"category": "HARM_CATEGORY_HATE_SPEECH", "threshold": "BLOCK_LOW_AND_ABOVE"},
{"category": "HARM_CATEGORY_DANGEROUS_CONTENT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"},
{"category": "HARM_CATEGORY_SEXUALLY_EXPLICIT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"},
{"category": "HARM_CATEGORY_HARASSMENT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"},
},
},
},
})
Generation Config
temperature := 0.9
topP := 0.95
topK := 40
maxTokens := 2048
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: prompt,
Temperature: &temperature,
TopP: &topP,
TopK: &topK,
MaxTokens: &maxTokens,
})
Error Handling
Common Errors
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: prompt,
})
if err != nil {
var providerErr *providererrors.ProviderError
if errors.As(err, &providerErr) {
switch providerErr.StatusCode {
case 400:
log.Printf("Invalid request: %s", providerErr.Message)
case 401:
log.Fatal("Invalid API key")
case 403:
log.Fatal("Permission denied or quota exceeded")
case 404:
log.Fatal("Model not found")
case 429:
log.Println("Rate limited, retrying...")
time.Sleep(time.Second * 5)
case 500, 503:
log.Println("Service error, retrying...")
time.Sleep(time.Second * 10)
default:
log.Printf("Google AI error: %d - %s", providerErr.StatusCode, providerErr.Message)
}
}
return nil, err
}
Content Filtering
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: prompt,
})
if err != nil {
var providerErr *providererrors.ProviderError
if errors.As(err, &providerErr) {
if providerErr.StatusCode == 400 && strings.Contains(providerErr.Message, "SAFETY") {
log.Println("Content blocked by safety filters")
// Try with adjusted safety settings
}
}
}
Best Practices
-
Model Selection
- Use
gemini-2.0-flashfor fast, multimodal tasks - Use
gemini-1.5-profor complex reasoning with huge context - Use
gemini-1.5-flash-8bfor cost-effective high-volume tasks
- Use
-
Context Management
- Take advantage of massive context windows
- Include full documents rather than chunking
- Use structured prompts for better results
-
Multimodal Input
- Compress images appropriately (max 20MB)
- Use clear, specific instructions for vision tasks
- Combine multiple modalities for richer context
-
Cost Optimization
- Monitor context length (pricing changes after 128K)
- Use flash-8b for simple tasks
- Cache repeated queries on your side
-
Safety
- Adjust safety settings based on use case
- Handle content filtering gracefully
- Log filtered content for review
Rate Limits & Pricing
Rate Limits (Free Tier)
| Model | RPM | TPM | RPD |
|---|---|---|---|
| Gemini 2.0 Flash | 15 | 1M | 1,500 |
| Gemini 1.5 Pro | 2 | 32K | 50 |
| Gemini 1.5 Flash | 15 | 1M | 1,500 |
RPM = Requests per minute, TPM = Tokens per minute, RPD = Requests per day
Paid tier offers significantly higher limits.
Context-Based Pricing
func calculateGeminiCost(inputTokens, outputTokens int, model string) float64 {
// Pricing changes based on context length
var inputRate, outputRate float64
if inputTokens <= 128000 {
// Low context pricing
inputRate = 0.075 / 1_000_000 // for flash models
outputRate = 0.30 / 1_000_000
} else {
// High context pricing (2x)
inputRate = 0.15 / 1_000_000
outputRate = 0.60 / 1_000_000
}
return float64(inputTokens)*inputRate + float64(outputTokens)*outputRate
}
Workflow Serialization
Google embedding, image, and (unary) transcription models can cross a
workflow boundary with providerutils.SerializeModel / DeserializeModel
(language models could already be serialized). See
Provider Serialization
for the mechanism; speech, speech-translation, video, and realtime models are
not yet serializable.
See Also
- API Reference: GenerateText
- Google Vertex AI Provider - Enterprise deployment
- Google AI Studio Documentation
- Gemini API Reference
May 2026 parity updates
Model refresh and token usage
The Google provider tracks the refreshed Gemini model IDs in pkg/providers/google/model_ids.go. Usage details include per-modality token fields through types.Usage.InputDetails.TextTokens and types.Usage.InputDetails.ImageTokens when Google returns modality-specific counts.
Thought signatures and no-arg tool calls
The provider preserves Google thoughtSignature metadata on text and tool-call parts so multi-turn requests can forward model reasoning state. Streaming tool calls with no arguments are emitted with an empty argument object, matching the TypeScript SDK.
Provider options
Use the standard ProviderOptions map for Google-specific options such as thinking configuration. The current Go provider accepts the same nested thinkingConfig request shape that the TypeScript provider sends to Gemini.
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Use Google search if useful.",
ProviderOptions: map[string]interface{}{
"google": map[string]interface{}{
"thinkingConfig": map[string]interface{}{
"thinkingBudget": 512,
"includeThoughts": true,
},
},
},
})