Google Vertex AI Provider
Google Vertex AI provides enterprise-grade AI platform with Gemini models, custom model training, and MLOps capabilities. Offers enhanced security, compliance, and integration with Google Cloud.
Setup
Installation
import (
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/googlevertex"
)
Configuration
provider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "us-central1",
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}
model, err := provider.LanguageModel("gemini-2.0-flash")
Authentication
- Install Google Cloud SDK
- Authenticate:
gcloud auth application-default login
- Set project:
export GOOGLE_CLOUD_PROJECT=your-project-id
The Vertex provider reuses auth per provider instance. You can provide an explicit access token or token generator to override default Google credential discovery, which is useful for tests, service-account impersonation, and custom credential brokers.
Available Models
Gemini Models
| Model ID | Context | Price | Best For |
|---|---|---|---|
| gemini-2.0-flash | 1M | $0.075/1M (≤128K) | Fast, multimodal |
| gemini-1.5-pro | 2M | $1.25/1M (≤128K) | Complex tasks |
| gemini-1.5-flash | 1M | $0.075/1M (≤128K) | General purpose |
Grok models are not reachable through Provider.LanguageModel — see
xAI (Grok) Sub-Provider and
Vertex MaaS and Grok below for the MaaS model IDs
(e.g. xai/grok-4.20-reasoning).
PaLM Models (Legacy)
| Model ID | Context | Price | Best For |
|---|---|---|---|
| text-bison | 8K | $0.50/1M | Text generation |
| chat-bison | 8K | $0.50/1M | Chat |
Embedding Models
| Model ID | Dimensions | Price | Best For |
|---|---|---|---|
| textembedding-gecko | 768 | Free | Semantic search |
Gemini TTS Speech Models
Vertex reuses the Gemini speech model implementation and exposes it through Provider.SpeechModel and Provider.Speech.
Vertex Chirp transcription is exposed through Provider.TranscriptionModel. Chirp requests use the Vertex Speech-to-Text v2 recognize endpoint and support language codes, word time offsets, punctuation, and billed-duration metadata.
| Model ID | Go constant |
|---|---|
gemini-2.5-flash-tts | googlevertex.SpeechModelGemini25FlashTTS |
gemini-2.5-pro-tts | googlevertex.SpeechModelGemini25ProTTS |
gemini-2.5-flash-lite-preview-tts | googlevertex.SpeechModelGemini25FlashLitePreviewTTS |
gemini-3.1-flash-tts-preview | googlevertex.SpeechModelGemini31FlashTTSPreview |
provider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "us-central1",
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}
speechModel, err := provider.SpeechModel(googlevertex.SpeechModelGemini25FlashTTS)
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{
Model: speechModel,
Text: "Vertex Gemini can synthesize speech.",
Voice: "Kore",
})
Provider metadata for Vertex speech is keyed as google, matching the TypeScript SDK's reuse of GoogleSpeechModel.
Image Generation Models
Breaking change: non-Gemini Imagen models (
imagen-3.0-*, etc.) were removed from the Vertex provider too, matching the TypeScript SDK. Callingprovider.ImageModel(id)with a model ID that doesn't start withgemini-now fails with: "Google image models other than Gemini are no longer supported. Use a model ID that starts withgemini-."
| Model ID | Aspect Ratios | Price | Best For |
|---|---|---|---|
| gemini-2.5-flash-image | 1:1, 3:4, 4:3, 9:16, 16:9 | Per image | Fast Gemini generation |
| gemini-3-pro-image-preview | 1:1, 3:4, 4:3, 9:16, 16:9 | Per image | Advanced generation |
Gemini Transcription
Provider.TranscriptionModel(modelID) dispatches on the model ID: any ID beginning with gemini uses a Gemini
transcription model instead of Chirp.
- Unary (
gemini-3.5-transcribe): callsgenerateContentonce and returns the full transcript, usable withai.Transcribe. - Live (any
-livesuffix, e.g.gemini-3.5-transcribe-live): streams over the Vertex Live WebSocket and only supportsai.ExperimentalStreamTranscribe(DoStream); it returns an error if you call the unary path. - Any other model ID still routes to Chirp (Speech-to-Text v2
recognize).
transcriptionModel, err := provider.TranscriptionModel("gemini-3.5-transcribe")
if err != nil {
log.Fatal(err)
}
result, err := ai.Transcribe(ctx, ai.TranscribeOptions{
Model: transcriptionModel,
Audio: audioBytes,
})
liveModel, err := provider.TranscriptionModel("gemini-3.5-transcribe-live")
if err != nil {
log.Fatal(err)
}
stream, err := ai.ExperimentalStreamTranscribe(ctx, ai.StreamTranscribeOptions{
Model: liveModel,
Audio: audioStream,
})
Vertex transcription (Chirp and Gemini) does not support Express Mode API keys — use standard Google Cloud credentials.
Provider-Specific Features
EU And US Multi-Region Routing
When Location is eu or us, standard Gemini, MaaS, and Anthropic-on-Vertex requests use the regional REP hosts:
aiplatform.eu.rep.googleapis.comaiplatform.us.rep.googleapis.com
Explicit BaseURL values still take precedence.
Location (and the Cloud Speech-to-Text region transcription option) is interpolated directly into the request host. A value that isn't a single DNS label (letters, digits, and hyphens) is rejected with an InvalidArgumentError before any request is sent; this only applies when no explicit BaseURL is configured, since that makes Location unused.
PayGo Request Type
Vertex AI uses request headers for PayGo tier selection. Set sharedRequestType and optionally requestType in Vertex provider options:
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "hello",
ProviderOptions: map[string]interface{}{
"vertex": map[string]interface{}{
"sharedRequestType": "priority",
"requestType": "shared",
},
},
})
serviceTier is a Gemini API body option and is ignored for Vertex requests with a warning.
Enterprise Security
Configure IAM, VPC Service Controls, CMEK, and audit logging in Google Cloud.
The SDK sends requests through the Vertex endpoint selected by Project,
Location, and optional BaseURL.
Model Garden
Access hundreds of models from Model Garden:
// Deploy from Model Garden
model, err := provider.LanguageModel("publishers/google/models/gemini-2.0-flash")
Batch Prediction And Custom Training
Use Google Cloud Vertex AI client libraries for batch prediction and custom training jobs. Go-AI focuses on the AI SDK provider surfaces for generation, embedding, image, speech, MaaS, and Anthropic-on-Vertex requests.
Image Generation
Vertex AI's text-to-image generation is now Gemini-only — Imagen model
IDs (imagen-3.0-*) were removed from this provider, matching the
TypeScript SDK.
// Create Vertex AI provider
provider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: os.Getenv("GOOGLE_VERTEX_LOCATION"),
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}
Gemini Image Generation:
// Gemini 2.5 Flash Image - Fast generation
geminiModel, err := provider.ImageModel("gemini-2.5-flash-image")
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{
Model: geminiModel,
Prompt: "A magical forest with glowing mushrooms and ethereal lighting",
Size: "1024x1024", // Square 1:1 aspect ratio
})
if err != nil {
log.Fatal(err)
}
os.WriteFile("vertex_gemini_image.png", result.Image.Data, 0644)
Supported Aspect Ratios:
1:1- Square (1024x1024)4:3- Standard (1024x768)3:4- Portrait (768x1024)16:9- Widescreen (1920x1080)9:16- Vertical (1080x1920)
Model Comparison:
| Model | Speed | Quality | Best For |
|---|---|---|---|
| gemini-2.5-flash-image | Very Fast | Good | Quick generation |
| gemini-3-pro-image-preview | Fast | High | Advanced generation |
Enterprise Features for Image Generation:
- VPC network support for secure generation
- Customer-managed encryption keys (CMEK)
- Audit logging for compliance
- Regional data residency
- Batch image generation for large-scale operations
Video Generation
provider.VideoModel(modelID) submits to Veo's :predictLongRunning
endpoint (MaxVideosPerCall() returns 4) and works with both the polling
ai.GenerateVideo and the async ai.ExperimentalStartVideo /
ai.ExperimentalGetVideoStatus calls:
videoModel, err := provider.VideoModel("veo-3.1-generate-preview")
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{Text: "An aerial shot of a mountain range at dawn"},
})
if err != nil {
log.Fatal(err)
}
for _, video := range result.Videos {
fmt.Println(video.URL)
}
xAI (Grok) Sub-Provider
pkg/providers/googlevertex/xai is a standalone package (mirroring TS's
separate @ai-sdk/google-vertex/xai) that serves Grok models through
Vertex's Model-as-a-Service OpenAI-compatible endpoint. It is not reachable
from googlevertex.Provider — construct it directly:
import vertexxai "github.com/digitallysavvy/go-ai/pkg/providers/googlevertex/xai"
xaiProvider, err := vertexxai.New(vertexxai.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "global", // MaaS Grok endpoint default
})
if err != nil {
log.Fatal(err)
}
model, err := xaiProvider.LanguageModel(vertexxai.ModelGrok420Reasoning)
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Explain the Vertex MaaS Grok endpoint.",
})
Auth resolves the same way as googlevertex.Config: AccessToken,
AuthToken, or TokenSource, falling back to Application Default
Credentials when none are set. This is a separate code path from the
generic googlevertex.NewMaaS wrapper below — either works for Grok, but
googlevertex/xai is the TS-parity implementation for new code.
Provider-Specific Options
Pass Vertex AI-specific options through ProviderOptions under the
"googleVertex" key (the model's Provider() method itself still returns
"google-vertex", which is only the model identity/error-wrapping name —
not the options map key).
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Analyze this document for compliance issues.",
ProviderOptions: map[string]interface{}{
"googleVertex": map[string]interface{}{
// Vertex-specific safety settings
"safetySettings": []map[string]interface{}{
{
"category": "HARM_CATEGORY_DANGEROUS_CONTENT",
"threshold": "BLOCK_ONLY_HIGH",
},
},
},
},
})
Vertex-specific options are read from the googleVertex key (falling back
to vertex, then google), matching the TypeScript SDK's own
providerOptions.googleVertex convention. The legacy "vertex" key is
still accepted, and a plain "google" key works as the last fallback so
options shared between the AI Studio and Vertex providers don't need
duplicating. Note that the hyphenated "google-vertex" string is only the
model's Provider() identity/error-wrapping name — it was never a
ProviderOptions key.
| Provider | ProviderOptions key(s), in lookup order |
|---|---|
Google AI Studio (providers/google) | "google" |
Google Vertex AI (providers/googlevertex) | "googleVertex" → "vertex" → "google" |
Vertex streaming can request function-call argument deltas with streamFunctionCallArguments; the SDK only sends this field for streaming Vertex requests. It is omitted for non-streaming DoGenerate calls, matching the TypeScript provider.
ProviderOptions: map[string]interface{}{
"googleVertex": map[string]interface{}{
"streamFunctionCallArguments": true,
},
}
Examples
Basic Text Generation
package main
import (
"context"
"fmt"
"log"
"os"
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/googlevertex"
)
func main() {
ctx := context.Background()
provider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "us-central1",
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}
model, err := provider.LanguageModel("gemini-2.0-flash")
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain Vertex AI"})
if err != nil {
log.Fatal(err)
}
fmt.Println(result.Text)
}
Multi-Region Deployment
// Deploy across regions for high availability
usProvider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "us",
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}
euProvider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "eu",
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}
// Try primary region first
result, err := generateWithFallback(usProvider, euProvider, prompt)
Best Practices
-
Enterprise Features
- Use VPC Service Controls
- Enable audit logging
- Configure IAM policies
- Use customer-managed encryption keys
-
Cost Management
- Monitor usage in Cloud Console
- Set up budget alerts
- Use committed use discounts
- Optimize batch operations
-
Performance
- Choose regions close to users
- Use batch prediction for large datasets
- Cache predictions when appropriate
-
Compliance
- Leverage Google Cloud compliance certifications
- Configure data residency requirements
- Enable audit logs for governance
Rate Limits
Varies by model and quota:
| Model | Default QPM | Max QPM |
|---|---|---|
| Gemini 2.0 Flash | 60 | 1000+ |
| Gemini 1.5 Pro | 60 | 1000+ |
Request quota increases through Cloud Console.
Workflow Serialization
Vertex embedding, image, and transcription models can cross a workflow
boundary with providerutils.SerializeModel / DeserializeModel (language
models could already be serialized). See
Provider Serialization
for the mechanism; speech, video, and the xai sub-provider are not yet
serializable.
See Also
- Google Provider - Google AI Studio
- Vertex AI Documentation
- Model Garden
May 2026 parity updates
Auth reuse
Vertex providers reuse Google authentication per provider instance. Configure googlevertex.New once and reuse it for language, embedding, and image models instead of creating a new provider for every call.
Vertex MaaS and Grok
Use googlevertex.NewMaaS for Vertex Model-as-a-Service models. Grok MaaS constants include MaaSModelGrok420Reasoning, MaaSModelGrok420NonReasoning, MaaSModelGrok41FastReasoning, and MaaSModelGrok41FastNonReasoning.
maas := googlevertex.NewMaaS(googlevertex.MaaSConfig{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "us-central1",
})
model, err := maas.LanguageModel(string(googlevertex.MaaSModelGrok420Reasoning))
Auth token overrides
Vertex Anthropic auth token generation can be overridden through MaaS configuration when a custom token source is required.