# Google Provider

Google provides the Gemini family of models through Google AI Studio. Gemini models offer exceptional multimodal capabilities, massive context windows (up to 1M+ tokens), and competitive pricing.

## Setup

### Installation

The Google provider is included in the Go-AI SDK:

```go
import (
    "github.com/digitallysavvy/go-ai/pkg/ai"
    providerapi "github.com/digitallysavvy/go-ai/pkg/provider"
    "github.com/digitallysavvy/go-ai/pkg/providers/google"
)
```

### Configuration

```go
googleProvider := google.New(google.Config{
    APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"),
})

model, err := googleProvider.LanguageModel("gemini-2.0-flash")
if err != nil {
    log.Fatal(err)
}
```

### Get API Key

1. Visit [makersuite.google.com/app/apikey](https://makersuite.google.com/app/apikey)
2. Create new API key
3. Set environment variable:

```bash
export GOOGLE_GENERATIVE_AI_API_KEY=AI...
```

## Available Models

### Gemini 3 / 2.5 Series

| Model ID | Best For |
|----------|----------|
| gemini-3.1-pro-preview | Advanced text generation |
| gemini-3.1-pro-preview-customtools | Advanced text generation with custom tools |
| gemini-3-flash-preview | Fast preview generation |
| gemini-3-pro-preview | Advanced preview generation |
| gemini-2.5-pro | Complex reasoning and multimodal tasks |
| gemini-2.5-flash | Fast multimodal tasks |
| gemini-2.5-flash-lite | Cost-effective high-volume tasks |

### Gemini 2.0 Series

| Model ID | Context | Input Price | Output Price | Best For |
|----------|---------|-------------|--------------|----------|
| gemini-2.0-flash | 1M | $0.075/1M (≤128K)<br/>$0.15/1M (>128K) | $0.30/1M (≤128K)<br/>$0.60/1M (>128K) | Fast, multimodal |
| gemini-2.0-flash-thinking | 1M | $0.075/1M (≤128K)<br/>$0.15/1M (>128K) | $0.30/1M (≤128K)<br/>$0.60/1M (>128K) | Reasoning tasks |

### Gemini 1.5 Series

| Model ID | Context | Input Price | Output Price | Best For |
|----------|---------|-------------|--------------|----------|
| gemini-1.5-pro | 2M | $1.25/1M (≤128K)<br/>$2.50/1M (>128K) | $5.00/1M (≤128K)<br/>$10.00/1M (>128K) | Complex analysis |
| gemini-1.5-flash | 1M | $0.075/1M (≤128K)<br/>$0.15/1M (>128K) | $0.30/1M (≤128K)<br/>$0.60/1M (>128K) | Fast responses |
| gemini-1.5-flash-8b | 1M | $0.0375/1M (≤128K)<br/>$0.075/1M (>128K) | $0.15/1M (≤128K)<br/>$0.30/1M (>128K) | Cost-effective |

### Gemini 1.0 Series (Legacy)

| Model ID | Context | Input Price | Output Price | Best For |
|----------|---------|-------------|--------------|----------|
| gemini-1.0-pro | 32K | $0.50/1M | $1.50/1M | General purpose |

### Image Generation Models

> **Breaking change:** non-Gemini Imagen models (`imagen-4.0-*`, etc.) were
> removed from the Google provider, matching the TypeScript SDK. Calling
> `provider.ImageModel(id)` with a model ID that doesn't start with
> `gemini-` now fails with: "Google image models other than Gemini are no
> longer supported. Use a model ID that starts with `gemini-`." Use Google
> Vertex AI if you still need Imagen.

| Model ID | Aspect Ratios | Price | Best For |
|----------|--------------|-------|----------|
| gemini-2.5-flash-image | 1:1, 3:4, 4:3, 9:16, 16:9 | Per image | Fast Gemini generation |
| gemini-3-pro-image-preview | 1:1, 3:4, 4:3, 9:16, 16:9 | Per image | Advanced generation |

Gemini image models ignore `N` (they generate one image per call and are
auto-batched by `ai.GenerateImage` to reach the requested count) instead of
erroring on `N > 1`.

Gemini image models can use Google Search grounding during image generation through provider options:

```go
result, err := imageModel.DoGenerate(ctx, &provider.ImageGenerateOptions{
    Prompt: "create a current-events infographic",
    ProviderOptions: map[string]interface{}{
        "google": map[string]interface{}{
            "googleSearch": map[string]interface{}{},
        },
    },
})
```

`googleSearch` is only supported for Gemini image models — the only image models this provider exposes now that Imagen has been removed.

### Embedding Models

| Model ID | Dimensions | Price | Best For |
|----------|-----------|-------|----------|
| gemini-embedding-001 | 3072 | See Google pricing | Multimodal semantic search |
| gemini-embedding-2 | 3072 | See Google pricing | Stable Gemini embedding workloads |
| gemini-embedding-2-preview | Provider-defined | See Google pricing | Preview embedding workloads |

### Gemini TTS Speech Models

Google Gemini TTS is exposed through `Provider.SpeechModel` and `Provider.Speech`.

| Model ID | Go constant |
|----------|-------------|
| gemini-2.5-flash-preview-tts | `google.ModelGemini25FlashTTS` |
| gemini-2.5-pro-preview-tts | `google.ModelGemini25ProTTS` |
| gemini-3.1-flash-tts-preview | `google.ModelGemini31FlashTTSPreview` |

```go
speechModel, err := googleProvider.SpeechModel(google.ModelGemini25FlashTTS)
if err != nil {
    log.Fatal(err)
}

result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{
    Model: speechModel,
    Text:  "Welcome to Go-AI.",
    Voice: "Kore",
})
if err != nil {
    log.Fatal(err)
}

fmt.Println(result.Audio.MediaType)
```

Gemini TTS defaults to the `Kore` voice when no voice is supplied. Use `ProviderOptions["google"]` with `google.GoogleSpeechModelOptions` to pass Gemini-specific speech options such as multi-speaker voice configuration.

### Realtime Models

Gemini Live is exposed through `Provider.RealtimeModel`,
`ExperimentalRealtimeModel`, and `GetRealtimeToken`. Model IDs are strings, matching
the TypeScript `GoogleRealtimeModelId` surface.

```go
model, err := googleProvider.RealtimeModel("gemini-3.1-flash-live-preview")
if err != nil {
    log.Fatal(err)
}

secretCreator, ok := model.(providerapi.RealtimeClientSecretCreator)
if !ok {
    log.Fatal("model does not support client secret creation")
}
token, err := secretCreator.DoCreateClientSecret(ctx, providerapi.ClientSecretOptions{})
if err != nil {
    log.Fatal(err)
}

wsConfigProvider, ok := model.(providerapi.RealtimeWebSocketConfigProvider)
if !ok {
    log.Fatal("model does not support client-side WebSocket config")
}
ws := wsConfigProvider.GetWebSocketConfig(token.Token, token.URL)
fmt.Println(ws.URL)
```

Use `ai.ConnectRealtime` for a provider-neutral Go session helper. Browser-only
TypeScript pieces such as `BrowserRealtimeTransport`, `BrowserRealtimeAudio`,
and `useRealtime` are not part of the Go runtime.

Gemini Live Translate configuration can be passed through `ProviderOptions["google"]["translationConfig"]`. The provider merges it into the realtime session `generationConfig.translationConfig`, matching the TypeScript Live Translate request shape.

Background-reasoning Live models (model IDs matching `gemini-<major>.<minor>-live...thinking`, e.g. `gemini-3.8-live-extended-thinking`) require a `thinkingLevel` or `thinkingBudget` in the session setup. When neither is set — via the typed `ProviderOptions["google"]["thinkingConfig"]` or a raw `ProviderOptions["generationConfig"]["thinkingConfig"]` — the provider adds `thinkingLevel: "low"` while keeping any other fields (e.g. `includeThoughts`) and without disturbing an explicit `thinkingBudget: 0`:

```go
setup := map[string]interface{}{
    "providerOptions": map[string]interface{}{
        "google": map[string]interface{}{
            "thinkingConfig": map[string]interface{}{"includeThoughts": true},
        },
    },
}
// -> generationConfig.thinkingConfig == {"includeThoughts": true, "thinkingLevel": "low"}
```

Input transcription events (`input-transcription-completed`) accumulate consecutive fragments for one user utterance under a stable synthetic item ID, and allocate a new ID once Google signals the utterance is finished or the turn completes; this keeps multi-turn and barge-in (interrupted) conversations from overwriting or misgrouping earlier user messages.

### Speech Translation (experimental)

`Provider.SpeechTranslationModel(modelID)` (alias `Provider.Translation`)
returns a speech-to-speech translation model over the same Live API
`BidiGenerateContent` WebSocket, for use with `ai.ExperimentalStreamTranslate`:

```go
translationModel, err := googleProvider.SpeechTranslationModel("gemini-3.5-live-translate-preview")
if err != nil {
    log.Fatal(err)
}

result, err := ai.ExperimentalStreamTranslate(ctx, ai.StreamTranslateOptions{
    Model:          translationModel,
    Audio:          audioStream,
    TargetLanguage: "es",
})
if err != nil {
    log.Fatal(err)
}

stream, err := result.FullStream()
if err != nil {
    log.Fatal(err)
}
defer stream.Close()

for {
    part, err := stream.Next()
    if err == io.EOF {
        break
    }
    if err != nil {
        log.Fatal(err)
    }
    if part.Type == "audio" {
        // part.AudioData carries translated PCM audio.
    }
}
```

## Provider-Specific Features

### Service Tier

For the Gemini API, pass `serviceTier` in Google provider options:

```go
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: "hello",
    ProviderOptions: map[string]interface{}{
        "google": map[string]interface{}{
            "serviceTier": "priority",
        },
    },
})
```

### Massive Context Windows

Gemini supports up to 2M tokens of context:

```go
// Process entire codebases or books
hugeDocument := loadDocument("book.txt") // up to 2M tokens

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: fmt.Sprintf("Summarize this book:\n\n%s", hugeDocument),
})
```

### Multimodal Capabilities

Gemini excels at processing text, images, video, and audio together:

```go
// Image understanding
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model: model,
    Messages: []types.Message{
        {
            Role: types.RoleUser,
            Content: []types.ContentPart{
                types.TextContent{Text: "Describe this image in detail"},
                types.FileContent{
                    URL:       "https://example.com/photo.jpg",
                    MediaType: "image/jpeg",
                },
            },
        },
    },
})

// Multiple images
result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model: model,
    Messages: []types.Message{
        {
            Role: types.RoleUser,
            Content: []types.ContentPart{
                types.TextContent{Text: "Compare these images"},
                types.FileContent{
                    URL:       "https://example.com/before.jpg",
                    MediaType: "image/jpeg",
                },
                types.FileContent{
                    URL:       "https://example.com/after.jpg",
                    MediaType: "image/jpeg",
                },
            },
        },
    },
})

// Video understanding
result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model: model,
    Messages: []types.Message{
        {
            Role: types.RoleUser,
            Content: []types.ContentPart{
                types.TextContent{Text: "What happens in this video?"},
                types.FileContent{
                    URL:       "https://example.com/video.mp4",
                    MediaType: "video/mp4",
                },
            },
        },
    },
})
```

### Function Calling

Gemini supports sophisticated function calling:

```go
weatherTool := types.Tool{
    Name:        "get_weather",
    Description: "Get current weather for a location",
    Parameters: map[string]interface{}{
        "type": "object",
        "properties": map[string]interface{}{
            "location": map[string]interface{}{
                "type":        "string",
                "description": "City and country",
            },
            "unit": map[string]interface{}{
                "type": "string",
                "enum": []string{"celsius", "fahrenheit"},
            },
        },
        "required": []string{"location"},
    },
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: "What's the weather in Tokyo?",
    Tools:  []types.Tool{weatherTool},
    StopWhen: []ai.StopCondition{ai.IsStepCount(5)},
})

for _, call := range result.ToolCalls {
    fmt.Printf("Function: %s\nArgs: %v\n", call.ToolName, call.Arguments)
}
```

### Interactions API

The Google provider exposes the Gemini Interactions API through `Provider.Interactions(modelID)` and agent presets through `Provider.InteractionsAgent(agent)`. Interactions models implement the standard `provider.LanguageModel` interface, so they can be used with `DoGenerate`, `DoStream`, and higher-level AI SDK helpers.

```go
store := false
interactionsModel, err := googleProvider.Interactions(google.InteractionsModelGemini25Flash)
if err != nil {
    log.Fatal(err)
}

result, err := interactionsModel.DoGenerate(ctx, &provider.GenerateOptions{
    Prompt: types.Prompt{Text: "Give a concise answer."},
    ProviderOptions: map[string]interface{}{
        "google": google.GoogleInteractionsProviderOptions{
            Store:              &store,
            ResponseModalities: []string{"text"},
            ThinkingLevel:      "low",
        },
    },
})
if err != nil {
    log.Fatal(err)
}

fmt.Println(result.Text)
```

Use `PreviousInteractionID` with `Store` for stateful follow-up calls. Agent models run as background interactions and are polled until a terminal status; `PollingTimeoutMs` controls that wait. Agent model IDs include `deep-research-pro-preview-12-2025`, `deep-research-preview-04-2026`, `deep-research-max-preview-04-2026`, and `antigravity-preview-05-2026`.

```go
agentModel, err := googleProvider.InteractionsAgent(google.InteractionsAgentDeepResearchProPreview)
if err != nil {
    log.Fatal(err)
}

result, err := agentModel.DoGenerate(ctx, &provider.GenerateOptions{
    Prompt: types.Prompt{Text: "Research the main tradeoffs in retrieval augmented generation."},
    ProviderOptions: map[string]interface{}{
        "google": google.GoogleInteractionsProviderOptions{
            PollingTimeoutMs: 120000,
        },
    },
})
```

Interactions responses preserve ordered text, reasoning, files, sources, and tool-call content. Google metadata is available under `ProviderMetadata["google"]`, including the `interactionId` when returned by the API.

### Grounding with Google Search

Connect Gemini to real-time web information:

```go
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: "What are the latest developments in quantum computing?",
    Tools:  []types.Tool{google.GoogleSearchTool()},
})

// Response includes citations
for _, source := range result.Sources {
    fmt.Printf("Source: %s\n", source.URL)
}
```

### JSON Mode

Structured output generation:

```go
import "github.com/digitallysavvy/go-ai/pkg/schema"

bookSchema := schema.NewSimpleJSONSchema(map[string]interface{}{
    "type": "object",
    "properties": map[string]interface{}{
        "title":  map[string]string{"type": "string"},
        "author": map[string]string{"type": "string"},
        "year":   map[string]string{"type": "integer"},
        "genre":  map[string]interface{}{
            "type": "array",
            "items": map[string]string{"type": "string"},
        },
    },
    "required": []string{"title", "author", "year"},
})

result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{
    Model:  model,
    Schema: bookSchema,
    Prompt: "Extract book metadata: '1984 by George Orwell, published 1949'",
})
```

### Streaming

Real-time response streaming:

```go
stream, err := ai.StreamText(ctx, ai.StreamTextOptions{
    Model:  model,
    Prompt: "Write a long story",
})
if err != nil {
    log.Fatal(err)
}
defer stream.Close()

for chunk := range stream.Chunks() {
    fmt.Print(chunk.Text)
}
```

### Thinking Mode

Gemini 2.0 Flash Thinking uses extended reasoning:

```go
model, err := googleProvider.LanguageModel(google.ModelGemini20FlashThinkingExp)

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: "Solve this logic puzzle: Three friends...",
})

// Reasoning content is preserved in ordered content parts when returned.
for _, part := range result.Content {
    if reasoning, ok := part.(types.ReasoningContent); ok {
        fmt.Println("Reasoning:", reasoning.Text)
    }
}
fmt.Println("Answer:", result.Text)
```

Gemini 3 thinking and tool-call streams preserve `thoughtSignature` in provider metadata so multi-step calls can safely round-trip model-generated reasoning and function calls. Streaming no-argument function calls are emitted with `{}` input instead of being dropped.

Usage includes per-modality token details when the API returns them. Text, image, audio, and video token counts are preserved in the result usage details rather than collapsed into only total input/output tokens.

Use the `"google"` provider key for AI Studio options. Vertex options use the `"googleVertex"` key (see the [Google Vertex AI provider](https://goaisdk.com/docs/providers/google-vertex.md#provider-specific-options) for its lookup order) — the hyphenated `"google-vertex"` string is only the Vertex model's `Provider()` identity, not a `ProviderOptions` key.

## Examples

### Basic Text Generation

```go
package main

import (
    "context"
    "fmt"
    "log"
    "os"

    "github.com/digitallysavvy/go-ai/pkg/ai"
    "github.com/digitallysavvy/go-ai/pkg/providers/google"
)

func main() {
    ctx := context.Background()
    googleProvider := google.New(google.Config{
        APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"),
    })

    model, err := googleProvider.LanguageModel("gemini-2.0-flash")
    if err != nil {
        log.Fatal(err)
    }

    result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
        Model:  model,
        Prompt: "Explain machine learning in simple terms",
    })
    if err != nil {
        log.Fatal(err)
    }

    fmt.Println(result.Text)
    fmt.Printf("Tokens: %d\n", result.Usage.GetTotalTokens())
}
```

### Image Analysis

```go
imageData, err := os.ReadFile("chart.png")
if err != nil {
    log.Fatal(err)
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model: model,
    Messages: []types.Message{
        {
            Role: types.RoleUser,
            Content: []types.ContentPart{
                types.TextContent{Text: "Analyze this chart and extract key insights"},
                types.FileContent{
                    Data:      imageData,
                    MediaType: "image/png",
                },
            },
        },
    },
})

fmt.Println(result.Text)
```

### Video Analysis

```go
// Upload video file
videoData, err := os.ReadFile("demo.mp4")
if err != nil {
    log.Fatal(err)
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model: model,
    Messages: []types.Message{
        {
            Role: types.RoleUser,
            Content: []types.ContentPart{
                types.TextContent{Text: "Describe what happens in this video"},
                types.FileContent{
                    Data:      videoData,
                    MediaType: "video/mp4",
                },
            },
        },
    },
})

fmt.Println(result.Text)
```

### Document Processing with Large Context

```go
// Load multiple documents
docs := []string{
    loadFile("doc1.txt"),
    loadFile("doc2.txt"),
    loadFile("doc3.txt"),
}
combined := strings.Join(docs, "\n\n---\n\n")

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: fmt.Sprintf("Analyze these documents and identify common themes:\n\n%s", combined),
})
```

### Grounded Search Responses

```go
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: "What are the current COVID-19 vaccination rates in major countries?",
    Tools:  []types.Tool{google.GoogleSearchTool()},
})

fmt.Println("Answer:", result.Text)

fmt.Println("\nSources:")
for _, source := range result.Sources {
    fmt.Printf("- %s: %s\n", source.Title, source.URL)
}
```

### Text Embeddings

```go
embeddingModel, err := googleProvider.EmbeddingModel(google.EmbeddingModelGeminiEmbedding001)
if err != nil {
    log.Fatal(err)
}

dimensions := 768
result, err := embeddingModel.DoEmbed(ctx, "Go is a programming language", &provider.EmbedModelOptions{
    ProviderOptions: map[string]interface{}{
        "google": google.GoogleEmbeddingProviderOptions{
            TaskType:             "RETRIEVAL_DOCUMENT",
            OutputDimensionality: &dimensions,
        },
    },
})
if err != nil {
    log.Fatal(err)
}

fmt.Printf("Dimensions: %d\n", len(result.Embedding))
```

Google embeddings also support multimodal content and provider-hosted files through `fileData`:

```go
result, err := embeddingModel.DoEmbed(ctx, "Summarize this document for retrieval.", &provider.EmbedModelOptions{
    ProviderOptions: map[string]interface{}{
        "google": google.GoogleEmbeddingProviderOptions{
            TaskType: "RETRIEVAL_DOCUMENT",
            Content: [][]google.EmbeddingPart{
                {
                    google.FileDataEmbeddingPart{
                        MimeType: "application/pdf",
                        FileURI:  "files/abc123",
                    },
                },
            },
        },
    },
})
if err != nil {
    log.Fatal(err)
}
```

### Multi-Turn Conversation

```go
messages := []types.Message{
    {
        Role: types.RoleUser,
        Content: []types.ContentPart{
            types.TextContent{Text: "What is the capital of France?"},
        },
    },
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:    model,
    Messages: messages,
})
messages = append(messages,
    types.Message{
        Role: types.RoleAssistant,
        Content: []types.ContentPart{
            types.TextContent{Text: result.Text},
        },
    },
    types.Message{
        Role: types.RoleUser,
        Content: []types.ContentPart{
            types.TextContent{Text: "What's the population?"},
        },
    },
)

result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:    model,
    Messages: messages,
})
```

### Image Generation

Google's text-to-image generation is now **Gemini-only** — Imagen model IDs
(`imagen-4.0-*`) were removed from this provider; use Google Vertex AI for
Imagen.

**Gemini Image Models** offer fast generation:

```go
// Gemini 2.5 Flash - Fast image generation
geminiImageModel, err := googleProvider.ImageModel("gemini-2.5-flash-image")
if err != nil {
    log.Fatal(err)
}

result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{
    Model:  geminiImageModel,
    Prompt: "Abstract colorful art with geometric shapes and vibrant patterns",
    Size:   "1920x1080", // 16:9 aspect ratio
})
if err != nil {
    log.Fatal(err)
}

fmt.Printf("Image generated: %d bytes\n", len(result.Images[0].Data))
os.WriteFile("gemini_output.png", result.Images[0].Data, 0644)
```

**Supported Aspect Ratios:**
- `1:1` - Square (1024x1024, 512x512)
- `4:3` - Standard (1024x768)
- `3:4` - Portrait (768x1024)
- `16:9` - Widescreen (1920x1080, 1792x1024)
- `9:16` - Vertical (1080x1920, 1024x1792)

**Model Comparison:**

| Model | Speed | Quality | Use Case |
|-------|-------|---------|----------|
| gemini-2.5-flash-image | Very Fast | Good | Quick iterations |
| gemini-3-pro-image-preview | Fast | High | Advanced generation |

**Note:** For Imagen models and image editing, use Google Vertex AI.

## Video Generation

`provider.VideoModel(modelID)` submits to Veo's `:predictLongRunning`
endpoint and polls the long-running operation until it completes
(`MaxVideosPerCall()` returns 4 — Veo accepts up to 4 videos per call, and
`ai.GenerateVideo` fans out across multiple calls automatically when you ask
for more):

```go
videoModel, err := googleProvider.VideoModel("veo-3.1-generate-preview")
if err != nil {
    log.Fatal(err)
}

result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{
    Model:  videoModel,
    Prompt: ai.VideoPrompt{Text: "A drone shot flying over a coastline at sunrise"},
})
if err != nil {
    log.Fatal(err)
}

for _, video := range result.Videos {
    fmt.Println(video.URL)
}
```

`VideoModel` also implements `provider.VideoModelStarter` /
`VideoModelStatusChecker`, so it works with the async
`ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus` calls.

## Advanced Configuration

### Custom HTTP Client

```go
googleProvider := google.New(google.Config{
    APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"),
    HTTPClient: &http.Client{
        Timeout: time.Minute * 5,
        Transport: &http.Transport{
            MaxIdleConns:    100,
            IdleConnTimeout: 90 * time.Second,
        },
    },
})
```

### Safety Settings

```go
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: prompt,
    ProviderOptions: map[string]interface{}{
        "google": map[string]interface{}{
            "safetySettings": []map[string]string{
                {"category": "HARM_CATEGORY_HATE_SPEECH", "threshold": "BLOCK_LOW_AND_ABOVE"},
                {"category": "HARM_CATEGORY_DANGEROUS_CONTENT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"},
                {"category": "HARM_CATEGORY_SEXUALLY_EXPLICIT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"},
                {"category": "HARM_CATEGORY_HARASSMENT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"},
            },
        },
    },
})
```

### Generation Config

```go
temperature := 0.9
topP := 0.95
topK := 40
maxTokens := 2048
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:       model,
    Prompt:      prompt,
    Temperature: &temperature,
    TopP:        &topP,
    TopK:        &topK,
    MaxTokens:   &maxTokens,
})
```

## Error Handling

### Common Errors

```go
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: prompt,
})
if err != nil {
    var providerErr *providererrors.ProviderError
    if errors.As(err, &providerErr) {
        switch providerErr.StatusCode {
        case 400:
            log.Printf("Invalid request: %s", providerErr.Message)
        case 401:
            log.Fatal("Invalid API key")
        case 403:
            log.Fatal("Permission denied or quota exceeded")
        case 404:
            log.Fatal("Model not found")
        case 429:
            log.Println("Rate limited, retrying...")
            time.Sleep(time.Second * 5)
        case 500, 503:
            log.Println("Service error, retrying...")
            time.Sleep(time.Second * 10)
        default:
            log.Printf("Google AI error: %d - %s", providerErr.StatusCode, providerErr.Message)
        }
    }
    return nil, err
}
```

### Content Filtering

```go
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: prompt,
})
if err != nil {
    var providerErr *providererrors.ProviderError
    if errors.As(err, &providerErr) {
        if providerErr.StatusCode == 400 && strings.Contains(providerErr.Message, "SAFETY") {
            log.Println("Content blocked by safety filters")
            // Try with adjusted safety settings
        }
    }
}
```

## Best Practices

1. **Model Selection**
   - Use `gemini-2.0-flash` for fast, multimodal tasks
   - Use `gemini-1.5-pro` for complex reasoning with huge context
   - Use `gemini-1.5-flash-8b` for cost-effective high-volume tasks

2. **Context Management**
   - Take advantage of massive context windows
   - Include full documents rather than chunking
   - Use structured prompts for better results

3. **Multimodal Input**
   - Compress images appropriately (max 20MB)
   - Use clear, specific instructions for vision tasks
   - Combine multiple modalities for richer context

4. **Cost Optimization**
   - Monitor context length (pricing changes after 128K)
   - Use flash-8b for simple tasks
   - Cache repeated queries on your side

5. **Safety**
   - Adjust safety settings based on use case
   - Handle content filtering gracefully
   - Log filtered content for review

## Rate Limits & Pricing

### Rate Limits (Free Tier)

| Model | RPM | TPM | RPD |
|-------|-----|-----|-----|
| Gemini 2.0 Flash | 15 | 1M | 1,500 |
| Gemini 1.5 Pro | 2 | 32K | 50 |
| Gemini 1.5 Flash | 15 | 1M | 1,500 |

RPM = Requests per minute, TPM = Tokens per minute, RPD = Requests per day

Paid tier offers significantly higher limits.

### Context-Based Pricing

```go
func calculateGeminiCost(inputTokens, outputTokens int, model string) float64 {
    // Pricing changes based on context length
    var inputRate, outputRate float64

    if inputTokens <= 128000 {
        // Low context pricing
        inputRate = 0.075 / 1_000_000  // for flash models
        outputRate = 0.30 / 1_000_000
    } else {
        // High context pricing (2x)
        inputRate = 0.15 / 1_000_000
        outputRate = 0.60 / 1_000_000
    }

    return float64(inputTokens)*inputRate + float64(outputTokens)*outputRate
}
```

## Workflow Serialization

Google embedding, image, and (unary) transcription models can cross a
workflow boundary with `providerutils.SerializeModel` / `DeserializeModel`
(language models could already be serialized). See
[Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization)
for the mechanism; speech, speech-translation, video, and realtime models are
not yet serializable.

## See Also

- [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md)
- [Google Vertex AI Provider](https://goaisdk.com/docs/providers/google-vertex.md) - Enterprise deployment
- [Google AI Studio Documentation](https://ai.google.dev/docs)
- [Gemini API Reference](https://ai.google.dev/api/rest)

## May 2026 parity updates

### Model refresh and token usage

The Google provider tracks the refreshed Gemini model IDs in `pkg/providers/google/model_ids.go`. Usage details include per-modality token fields through `types.Usage.InputDetails.TextTokens` and `types.Usage.InputDetails.ImageTokens` when Google returns modality-specific counts.

### Thought signatures and no-arg tool calls

The provider preserves Google `thoughtSignature` metadata on text and tool-call parts so multi-turn requests can forward model reasoning state. Streaming tool calls with no arguments are emitted with an empty argument object, matching the TypeScript SDK.

### Provider options

Use the standard `ProviderOptions` map for Google-specific options such as thinking configuration. The current Go provider accepts the same nested `thinkingConfig` request shape that the TypeScript provider sends to Gemini.

```go
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: "Use Google search if useful.",
    ProviderOptions: map[string]interface{}{
        "google": map[string]interface{}{
            "thinkingConfig": map[string]interface{}{
                "thinkingBudget": 512,
                "includeThoughts": true,
            },
        },
    },
})
```
