Skip to main content

Google Vertex AI Provider

Google Vertex AI provides enterprise-grade AI platform with Gemini models, custom model training, and MLOps capabilities. Offers enhanced security, compliance, and integration with Google Cloud.

Setup​

Installation​

import (
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/googlevertex"
)

Configuration​

provider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "us-central1",
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}

model, err := provider.LanguageModel("gemini-2.0-flash")

Authentication​

  1. Install Google Cloud SDK
  2. Authenticate:
gcloud auth application-default login
  1. Set project:
export GOOGLE_CLOUD_PROJECT=your-project-id

The Vertex provider reuses auth per provider instance. You can provide an explicit access token or token generator to override default Google credential discovery, which is useful for tests, service-account impersonation, and custom credential brokers.

Available Models​

Gemini Models​

Model IDContextPriceBest For
gemini-2.0-flash1M$0.075/1M (≤128K)Fast, multimodal
gemini-1.5-pro2M$1.25/1M (≤128K)Complex tasks
gemini-1.5-flash1M$0.075/1M (≤128K)General purpose

Grok models are not reachable through Provider.LanguageModel — see xAI (Grok) Sub-Provider and Vertex MaaS and Grok below for the MaaS model IDs (e.g. xai/grok-4.20-reasoning).

PaLM Models (Legacy)​

Model IDContextPriceBest For
text-bison8K$0.50/1MText generation
chat-bison8K$0.50/1MChat

Embedding Models​

Model IDDimensionsPriceBest For
textembedding-gecko768FreeSemantic search

Gemini TTS Speech Models​

Vertex reuses the Gemini speech model implementation and exposes it through Provider.SpeechModel and Provider.Speech.

Vertex Chirp transcription is exposed through Provider.TranscriptionModel. Chirp requests use the Vertex Speech-to-Text v2 recognize endpoint and support language codes, word time offsets, punctuation, and billed-duration metadata.

Model IDGo constant
gemini-2.5-flash-ttsgooglevertex.SpeechModelGemini25FlashTTS
gemini-2.5-pro-ttsgooglevertex.SpeechModelGemini25ProTTS
gemini-2.5-flash-lite-preview-ttsgooglevertex.SpeechModelGemini25FlashLitePreviewTTS
gemini-3.1-flash-tts-previewgooglevertex.SpeechModelGemini31FlashTTSPreview
provider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "us-central1",
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}

speechModel, err := provider.SpeechModel(googlevertex.SpeechModelGemini25FlashTTS)
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{
Model: speechModel,
Text: "Vertex Gemini can synthesize speech.",
Voice: "Kore",
})

Provider metadata for Vertex speech is keyed as google, matching the TypeScript SDK's reuse of GoogleSpeechModel.

Image Generation Models​

Breaking change: non-Gemini Imagen models (imagen-3.0-*, etc.) were removed from the Vertex provider too, matching the TypeScript SDK. Calling provider.ImageModel(id) with a model ID that doesn't start with gemini- now fails with: "Google image models other than Gemini are no longer supported. Use a model ID that starts with gemini-."

Model IDAspect RatiosPriceBest For
gemini-2.5-flash-image1:1, 3:4, 4:3, 9:16, 16:9Per imageFast Gemini generation
gemini-3-pro-image-preview1:1, 3:4, 4:3, 9:16, 16:9Per imageAdvanced generation

Gemini Transcription​

Provider.TranscriptionModel(modelID) dispatches on the model ID: any ID beginning with gemini uses a Gemini transcription model instead of Chirp.

  • Unary (gemini-3.5-transcribe): calls generateContent once and returns the full transcript, usable with ai.Transcribe.
  • Live (any -live suffix, e.g. gemini-3.5-transcribe-live): streams over the Vertex Live WebSocket and only supports ai.ExperimentalStreamTranscribe (DoStream); it returns an error if you call the unary path.
  • Any other model ID still routes to Chirp (Speech-to-Text v2 recognize).
transcriptionModel, err := provider.TranscriptionModel("gemini-3.5-transcribe")
if err != nil {
log.Fatal(err)
}

result, err := ai.Transcribe(ctx, ai.TranscribeOptions{
Model: transcriptionModel,
Audio: audioBytes,
})
liveModel, err := provider.TranscriptionModel("gemini-3.5-transcribe-live")
if err != nil {
log.Fatal(err)
}

stream, err := ai.ExperimentalStreamTranscribe(ctx, ai.StreamTranscribeOptions{
Model: liveModel,
Audio: audioStream,
})

Vertex transcription (Chirp and Gemini) does not support Express Mode API keys — use standard Google Cloud credentials.

Provider-Specific Features​

EU And US Multi-Region Routing​

When Location is eu or us, standard Gemini, MaaS, and Anthropic-on-Vertex requests use the regional REP hosts:

  • aiplatform.eu.rep.googleapis.com
  • aiplatform.us.rep.googleapis.com

Explicit BaseURL values still take precedence.

Location (and the Cloud Speech-to-Text region transcription option) is interpolated directly into the request host. A value that isn't a single DNS label (letters, digits, and hyphens) is rejected with an InvalidArgumentError before any request is sent; this only applies when no explicit BaseURL is configured, since that makes Location unused.

PayGo Request Type​

Vertex AI uses request headers for PayGo tier selection. Set sharedRequestType and optionally requestType in Vertex provider options:

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "hello",
ProviderOptions: map[string]interface{}{
"vertex": map[string]interface{}{
"sharedRequestType": "priority",
"requestType": "shared",
},
},
})

serviceTier is a Gemini API body option and is ignored for Vertex requests with a warning.

Enterprise Security​

Configure IAM, VPC Service Controls, CMEK, and audit logging in Google Cloud. The SDK sends requests through the Vertex endpoint selected by Project, Location, and optional BaseURL.

Model Garden​

Access hundreds of models from Model Garden:

// Deploy from Model Garden
model, err := provider.LanguageModel("publishers/google/models/gemini-2.0-flash")

Batch Prediction And Custom Training​

Use Google Cloud Vertex AI client libraries for batch prediction and custom training jobs. Go-AI focuses on the AI SDK provider surfaces for generation, embedding, image, speech, MaaS, and Anthropic-on-Vertex requests.

Image Generation​

Vertex AI's text-to-image generation is now Gemini-only — Imagen model IDs (imagen-3.0-*) were removed from this provider, matching the TypeScript SDK.

// Create Vertex AI provider
provider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: os.Getenv("GOOGLE_VERTEX_LOCATION"),
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}

Gemini Image Generation:

// Gemini 2.5 Flash Image - Fast generation
geminiModel, err := provider.ImageModel("gemini-2.5-flash-image")
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{
Model: geminiModel,
Prompt: "A magical forest with glowing mushrooms and ethereal lighting",
Size: "1024x1024", // Square 1:1 aspect ratio
})
if err != nil {
log.Fatal(err)
}

os.WriteFile("vertex_gemini_image.png", result.Image.Data, 0644)

Supported Aspect Ratios:

  • 1:1 - Square (1024x1024)
  • 4:3 - Standard (1024x768)
  • 3:4 - Portrait (768x1024)
  • 16:9 - Widescreen (1920x1080)
  • 9:16 - Vertical (1080x1920)

Model Comparison:

ModelSpeedQualityBest For
gemini-2.5-flash-imageVery FastGoodQuick generation
gemini-3-pro-image-previewFastHighAdvanced generation

Enterprise Features for Image Generation:

  • VPC network support for secure generation
  • Customer-managed encryption keys (CMEK)
  • Audit logging for compliance
  • Regional data residency
  • Batch image generation for large-scale operations

Video Generation​

provider.VideoModel(modelID) submits to Veo's :predictLongRunning endpoint (MaxVideosPerCall() returns 4) and works with both the polling ai.GenerateVideo and the async ai.ExperimentalStartVideo / ai.ExperimentalGetVideoStatus calls:

videoModel, err := provider.VideoModel("veo-3.1-generate-preview")
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{Text: "An aerial shot of a mountain range at dawn"},
})
if err != nil {
log.Fatal(err)
}

for _, video := range result.Videos {
fmt.Println(video.URL)
}

xAI (Grok) Sub-Provider​

pkg/providers/googlevertex/xai is a standalone package (mirroring TS's separate @ai-sdk/google-vertex/xai) that serves Grok models through Vertex's Model-as-a-Service OpenAI-compatible endpoint. It is not reachable from googlevertex.Provider — construct it directly:

import vertexxai "github.com/digitallysavvy/go-ai/pkg/providers/googlevertex/xai"

xaiProvider, err := vertexxai.New(vertexxai.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "global", // MaaS Grok endpoint default
})
if err != nil {
log.Fatal(err)
}

model, err := xaiProvider.LanguageModel(vertexxai.ModelGrok420Reasoning)
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Explain the Vertex MaaS Grok endpoint.",
})

Auth resolves the same way as googlevertex.Config: AccessToken, AuthToken, or TokenSource, falling back to Application Default Credentials when none are set. This is a separate code path from the generic googlevertex.NewMaaS wrapper below — either works for Grok, but googlevertex/xai is the TS-parity implementation for new code.

Provider-Specific Options​

Pass Vertex AI-specific options through ProviderOptions under the "googleVertex" key (the model's Provider() method itself still returns "google-vertex", which is only the model identity/error-wrapping name — not the options map key).

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Analyze this document for compliance issues.",
ProviderOptions: map[string]interface{}{
"googleVertex": map[string]interface{}{
// Vertex-specific safety settings
"safetySettings": []map[string]interface{}{
{
"category": "HARM_CATEGORY_DANGEROUS_CONTENT",
"threshold": "BLOCK_ONLY_HIGH",
},
},
},
},
})

Vertex-specific options are read from the googleVertex key (falling back to vertex, then google), matching the TypeScript SDK's own providerOptions.googleVertex convention. The legacy "vertex" key is still accepted, and a plain "google" key works as the last fallback so options shared between the AI Studio and Vertex providers don't need duplicating. Note that the hyphenated "google-vertex" string is only the model's Provider() identity/error-wrapping name — it was never a ProviderOptions key.

ProviderProviderOptions key(s), in lookup order
Google AI Studio (providers/google)"google"
Google Vertex AI (providers/googlevertex)"googleVertex" → "vertex" → "google"

Vertex streaming can request function-call argument deltas with streamFunctionCallArguments; the SDK only sends this field for streaming Vertex requests. It is omitted for non-streaming DoGenerate calls, matching the TypeScript provider.

ProviderOptions: map[string]interface{}{
"googleVertex": map[string]interface{}{
"streamFunctionCallArguments": true,
},
}

Examples​

Basic Text Generation​

package main

import (
"context"
"fmt"
"log"
"os"

"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/googlevertex"
)

func main() {
ctx := context.Background()
provider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "us-central1",
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}

model, err := provider.LanguageModel("gemini-2.0-flash")
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain Vertex AI"})
if err != nil {
log.Fatal(err)
}

fmt.Println(result.Text)
}

Multi-Region Deployment​

// Deploy across regions for high availability
usProvider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "us",
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}

euProvider, err := googlevertex.New(googlevertex.Config{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "eu",
AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"),
})
if err != nil {
log.Fatal(err)
}

// Try primary region first
result, err := generateWithFallback(usProvider, euProvider, prompt)

Best Practices​

  1. Enterprise Features

    • Use VPC Service Controls
    • Enable audit logging
    • Configure IAM policies
    • Use customer-managed encryption keys
  2. Cost Management

    • Monitor usage in Cloud Console
    • Set up budget alerts
    • Use committed use discounts
    • Optimize batch operations
  3. Performance

    • Choose regions close to users
    • Use batch prediction for large datasets
    • Cache predictions when appropriate
  4. Compliance

    • Leverage Google Cloud compliance certifications
    • Configure data residency requirements
    • Enable audit logs for governance

Rate Limits​

Varies by model and quota:

ModelDefault QPMMax QPM
Gemini 2.0 Flash601000+
Gemini 1.5 Pro601000+

Request quota increases through Cloud Console.

Workflow Serialization​

Vertex embedding, image, and transcription models can cross a workflow boundary with providerutils.SerializeModel / DeserializeModel (language models could already be serialized). See Provider Serialization for the mechanism; speech, video, and the xai sub-provider are not yet serializable.

See Also​

May 2026 parity updates​

Auth reuse​

Vertex providers reuse Google authentication per provider instance. Configure googlevertex.New once and reuse it for language, embedding, and image models instead of creating a new provider for every call.

Vertex MaaS and Grok​

Use googlevertex.NewMaaS for Vertex Model-as-a-Service models. Grok MaaS constants include MaaSModelGrok420Reasoning, MaaSModelGrok420NonReasoning, MaaSModelGrok41FastReasoning, and MaaSModelGrok41FastNonReasoning.

maas := googlevertex.NewMaaS(googlevertex.MaaSConfig{
Project: os.Getenv("GOOGLE_VERTEX_PROJECT"),
Location: "us-central1",
})

model, err := maas.LanguageModel(string(googlevertex.MaaSModelGrok420Reasoning))

Auth token overrides​

Vertex Anthropic auth token generation can be overridden through MaaS configuration when a custom token source is required.