Skip to main content

AI Gateway Provider

The AI Gateway provider enables you to access multiple AI providers through a unified interface with automatic provider selection, failover, and cost optimization.

Features​

  • Unified Interface: Single API for multiple AI providers
  • Automatic Provider Selection: Gateway selects the best available provider
  • Failover Handling: Automatic failover to alternative providers
  • Cost Optimization: Intelligentrouting based on cost and performance
  • Zero Data Retention: Option to not log or retain request data
  • Observability: Built-in metrics and monitoring
  • Rate Limiting: Automatic rate limit handling across providers

Installation​

The Gateway provider is included in the core Go-AI SDK:

go get github.com/digitallysavvy/go-ai

Setup​

Environment Variables​

export AI_GATEWAY_API_KEY="your-api-key-here"

AI_GATEWAY_API_KEY may contain either an AI Gateway API key or a Vercel access token. When a Vercel access token can access multiple teams, set TeamIDOrSlug; the provider sends it as x-vercel-ai-gateway-team.

Basic Configuration​

package main

import (
"github.com/digitallysavvy/go-ai/pkg/providers/gateway"
)

func main() {
// Create provider with API key from environment
provider, err := gateway.New(gateway.Config{
APIKey: os.Getenv("AI_GATEWAY_API_KEY"),
})
if err != nil {
log.Fatal(err)
}
}

Advanced Configuration​

provider, err := gateway.New(gateway.Config{
APIKey: "your-api-key-or-vercel-access-token",
TeamIDOrSlug: "team-slug-or-id",
BaseURL: "https://ai-gateway.vercel.sh/v4/ai", // Custom base URL
Headers: map[string]string{
"X-Custom-Header": "value",
},
ZeroDataRetention: true, // Enable zero data retention
DisallowPromptTraining: true,
QuotaEntityID: "tenant-123",
MetadataCacheRefreshMillis: 300000, // 5 minutes
})

Text Generation​

import (
"context"
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/gateway"
)

// Create provider
provider, _ := gateway.New(gateway.Config{
APIKey: os.Getenv("AI_GATEWAY_API_KEY"),
})

// Create language model
// Gateway routes to the best available provider
model, _ := provider.LanguageModel(string(gateway.GatewayLanguageModelAnthropicClaudeSonnet45))

// Generate text
maxTokens := 200
result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{
Model: model,
Prompt: "Explain quantum computing in simple terms.",
MaxTokens: &maxTokens,
})

Model ID catalog​

pkg/providers/gateway/model_ids.go is generated from the upstream Gateway catalog (go generate ./pkg/providers/gateway) and gives compile-time constants such as gateway.GatewayLanguageModelAnthropicClaudeSonnet45, or you can pass the raw Gateway model ID string directly.

Breaking change: all xai/* model IDs were renamed to spacexai/* upstream (GatewayLanguageModelXai* constants are now GatewayLanguageModelSpacexai*, e.g. gateway.GatewayLanguageModelSpacexaiGrok420MultiAgent / "spacexai/grok-4.20-multi-agent"), and 44 language, 4 image, and 4 video IDs that no longer exist upstream were removed from the generated file.

Video Generation​

The Gateway provider supports video generation with automatic provider selection:

// Create video model
videoModel, err := provider.VideoModel("google/veo-3.1-generate-001")
if err != nil {
log.Fatal(err)
}

// Generate video
duration := 4.0
result, err := ai.GenerateVideo(context.Background(), ai.GenerateVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{
Text: "A cat playing with a ball of yarn",
},
AspectRatio: "16:9",
Duration: &duration,
})

Video Generation Features​

  • Text-to-Video: Generate videos from text descriptions
  • Image-to-Video: Animate static images
  • Batch Generation: Generate multiple videos in one request
  • Provider Selection: Automatic routing to best video provider
  • Custom Parameters: Aspect ratio, duration, FPS, resolution

Supported Video Providers​

The gateway can route video requests to:

  • Google Vertex AI (Veo)
  • FAL AI
  • Replicate
  • Other configured providers

Video Generation Example​

videoModel, _ := provider.VideoModel("google/veo-3.1-generate-001")

duration := 4.0
fps := 30
result, err := ai.GenerateVideo(context.Background(), ai.GenerateVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{
Text: "A serene ocean sunset with waves",
},
AspectRatio: "16:9",
Duration: &duration,
FPS: &fps,
ProviderOptions: map[string]interface{}{
"enhancePrompt": true,
},
})

// Access generated videos
for _, video := range result.Videos {
fmt.Printf("Video URL: %s\n", video.URL)
fmt.Printf("Media Type: %s\n", video.MediaType)
}

Image Generation​

imageModel, _ := provider.ImageModel(string(gateway.GatewayImageModelOpenaiGptImage2))

n := 1
result, err := ai.GenerateImage(context.Background(), ai.GenerateImageOptions{
Model: imageModel,
Prompt: "A futuristic cityscape at sunset",
N: &n,
Size: "1024x1024",
})

// The Gateway image model sets its own retryability signal (*bool) on a
// failed/empty generation; ai.GenerateImage reads it internally through
// the retry loop, so this does not appear as a field on the returned
// GenerateImageResult.

Embeddings​

embeddingModel, _ := provider.EmbeddingModel(string(gateway.GatewayEmbeddingModelOpenaiTextEmbedding3Small))

result, err := ai.EmbedMany(context.Background(), ai.EmbedManyOptions{
Model: embeddingModel,
Inputs: []string{
"The cat sat on the mat",
"A dog ran through the park",
},
})

Provider Options​

Customize provider selection and behavior:

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Hello, world!",
ProviderOptions: map[string]interface{}{
"gateway": map[string]interface{}{
"only": []string{"anthropic", "openai", "google"},
"tags": []string{"production", "fast"},
"order": []string{"openai", "anthropic"},
"sort": "cost",
},
},
})

Provider Option Fields​

  • gateway.only: Restrict candidate providers.
  • gateway.order: Explicit provider order.
  • gateway.sort: Gateway sort strategy.
  • gateway.tags: Routing tags.
  • gateway.models: Ordered fallback model list. Build entries with gateway.GatewayModel(id) for a plain fallback, or gateway.GatewayConditionalModelFallback(id, when) for an evaluation-only conditional fallback (only the first entry may be conditional). JSON also accepts a plain model-ID string in place of either helper.
  • gateway.has: Restrict routing to models satisfying capability tags (gateway.GatewayHasImplicitCaching, GatewayHasReasoning, GatewayHasStructuredOutput, GatewayHasToolUse, GatewayHasVision, or a weight-format condition from GatewayHasQuantization / GatewayHasNotQuantization).
  • gateway.zeroDataRetention: Per-call zero retention.
  • gateway.disallowPromptTraining: Restrict to no-prompt-training providers.
  • gateway.quotaEntityId: Quota identity for tenant/account-level tracking.
  • gateway.serviceTier: Unified service tier intent. Supported values are "flex" and "priority"; Gateway translates this to the provider-specific tier and overrides lower-level provider tier options.

Breaking change: gateway.hipaaCompliant / Config.HIPAACompliant were removed; the Gateway API no longer accepts this field.

Reranking​

import goprovider "github.com/digitallysavvy/go-ai/pkg/provider"

reranker, _ := gatewayProvider.RerankingModel("cohere/rerank-v3.5")
topN := 3
res, err := reranker.DoRerank(ctx, &goprovider.RerankOptions{
Query: "gateway routing",
Documents: []string{
"Provider fallback order can reduce latency variance.",
"Embedding vectors power semantic search.",
"HIPAA compliance restricts provider choice.",
},
TopN: &topN,
})

Timeout Handling​

The Gateway SDK provides enhanced timeout error handling with troubleshooting guidance.

Typed Error Classification​

Gateway response errors are decoded into typed errors under pkg/providers/gateway/errors. The common GatewayError interface exposes status code, public type, generation ID, and retryability. Unknown Gateway error types are classified the same way as the TypeScript SDK while preserving the raw type through GatewayErrorDetails.

var gatewayErr gatewayerrors.GatewayError
if errors.As(err, &gatewayErr) {
fmt.Println(gatewayErr.GetType(), gatewayErr.GetStatusCode(), gatewayErr.IsRetryable())
}

var details gatewayerrors.GatewayErrorDetails
if errors.As(err, &details) {
fmt.Println(details.GetRawType(), details.GetCode(), details.GetParam())
}

Inline []byte file data in Gateway language-model requests is base64-encoded exactly once for file, reasoning-file, and tool-result file parts. URL, provider-reference, and text file data are preserved in their provider-native form.

Realtime Runtime Client Secrets​

Gateway realtime is exposed through provider.ExperimentalRealtime(modelID). Use GetRealtimeToken or MintRealtimeClientSecret on a server to create the short-lived client secret used by browser WebSocket clients.

gw, err := gateway.New(gateway.Config{
APIKey: os.Getenv("AI_GATEWAY_API_KEY"),
TeamIDOrSlug: os.Getenv("VERCEL_TEAM_ID"),
})
if err != nil {
return err
}

expiresAfterSeconds := 600
secret, err := gw.GetRealtimeToken(ctx, provider.RealtimeFactoryGetTokenOptions{
Model: "openai/gpt-realtime",
ClientSecretOptions: provider.ClientSecretOptions{
ExpiresAfterSeconds: &expiresAfterSeconds,
},
})
if err != nil {
return err
}

model := gw.ExperimentalRealtime("openai/gpt-realtime")
ws := model.GetWebSocketConfig(secret.Token, secret.URL)
fmt.Println(ws.URL, ws.Protocols)

MintRealtimeClientSecret sends POST /v1/realtime/client-secrets at the Gateway origin, not under /v4/ai. The realtime WebSocket protocols include ai-gateway-realtime.v1, ai-gateway-auth.<token>, and, when configured, a base64url encoded ai-gateway-team.<team> entry. Realtime event parsing, client event serialization, and session config building are identity mappings for provider-native realtime payloads.

Setting Timeouts​

import (
"context"
"time"
gatewayerrors "github.com/digitallysavvy/go-ai/pkg/providers/gateway/errors"
)

// Text generation: 30-60 seconds
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()

options.Model = model
result, err := ai.GenerateText(ctx, options)

// Video generation: 5-10 minutes
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Minute)
defer cancel()

options.Model = videoModel
result, err := ai.GenerateVideo(ctx, options)

Detecting Timeout Errors​

options.Model = videoModel
result, err := ai.GenerateVideo(ctx, options)
if err != nil {
if gatewayerrors.IsGatewayTimeoutError(err) {
// This is a client-side timeout
// Error message includes troubleshooting guidance
fmt.Printf("Timeout error: %v\n", err)
} else {
// Other error
fmt.Printf("Error: %v\n", err)
}
}

Timeout Best Practices​

Text Generation:

  • Simple prompts: 30-60 seconds
  • Complex generation: 2-5 minutes

Video Generation:

  • Short videos (2-4s): 5-10 minutes
  • Long videos: 15-30 minutes
  • Batch generation: 20-40 minutes

Image Generation:

  • Single image: 1-2 minutes
  • Batch images: 3-5 minutes

Zero Data Retention​

Enable zero data retention mode to ensure requests are not logged:

provider, err := gateway.New(gateway.Config{
APIKey: os.Getenv("AI_GATEWAY_API_KEY"),
ZeroDataRetention: true,
})

Observability​

The Gateway automatically includes observability headers when running on Vercel:

// Automatically includes:
// - ai-o11y-deployment-id
// - ai-o11y-environment
// - ai-o11y-region
// - ai-o11y-request-id

Available Models​

Query available models and providers:

metadata, err := provider.GetAvailableModels(context.Background())
if err != nil {
log.Fatal(err)
}

for _, model := range metadata.Models {
fmt.Printf("Provider: %s\n", model.Specification.Provider)
fmt.Printf(" - %s (%s)\n", model.Name, model.ID)
fmt.Printf(" Specification: %v\n", model.Specification)
}

Credits​

Check your Gateway credit usage:

credits, err := provider.GetCredits(context.Background())
if err != nil {
log.Fatal(err)
}

fmt.Printf("Balance: %s\n", credits.Balance)
fmt.Printf("Total Used: %s\n", credits.TotalUsed)

/v1/credits reads the team from a teamId/slug query parameter rather than the x-vercel-ai-gateway-team header every other endpoint uses. When Config.TeamIDOrSlug (or a per-call x-vercel-ai-gateway-team header) is set, GetCredits forwards it as that query parameter too, so credentials that can access multiple teams (such as Vercel access tokens) are scoped correctly instead of getting a 401.

Evaluation model fallbacks​

EvaluationFallbackCondition.Question is optional. Without it, ConfidenceBelow checks every Choice/Score question and ProbabilityBetween checks every Boolean question, instead of a single named question:

confidenceBelow := 0.6
when := gateway.EvaluationFallbackCondition{ConfidenceBelow: &confidenceBelow}
fallback := gateway.GatewayConditionalModelFallback("openai/gpt-5.6-sol", when)

A condition still needs a check: {} or a condition with only Question set is rejected.

Error Handling​

The Gateway provides detailed error types:

import gatewayerrors "github.com/digitallysavvy/go-ai/pkg/providers/gateway/errors"

options.Model = model
result, err := ai.GenerateText(ctx, options)
if err != nil {
switch {
case gatewayerrors.IsGatewayTimeoutError(err):
// Handle timeout with troubleshooting guidance
fmt.Printf("Request timed out: %v\n", err)

case providererrors.IsRateLimitError(err):
// Handle rate limiting
fmt.Printf("Rate limited: %v\n", err)

case providererrors.IsProviderError(err):
// Handle provider errors
fmt.Printf("Provider error: %v\n", err)

default:
// Handle other errors
fmt.Printf("Error: %v\n", err)
}
}

Configuration Reference​

Config​

type Config struct {
// APIKey is the AI Gateway API key or Vercel access token (required)
// Can also be set via AI_GATEWAY_API_KEY environment variable
APIKey string

// TeamIDOrSlug scopes Vercel access-token requests to a team.
TeamIDOrSlug string

// BaseURL is the base URL for the AI Gateway API
// Default: https://ai-gateway.vercel.sh/v4/ai
BaseURL string

// Headers are custom headers to include in requests
Headers map[string]string

// MetadataCacheRefreshMillis is how frequently to refresh metadata cache
// Default: 300000 (5 minutes)
MetadataCacheRefreshMillis int64

// HTTPClient is a custom HTTP client
HTTPClient *http.Client

// ZeroDataRetention enables zero data retention mode
// When true, requests are not logged or retained
ZeroDataRetention bool

// ProjectID is forwarded as the "ai-o11y-project-id" header. Can also be
// set via VERCEL_PROJECT_ID.
ProjectID *string

// DisallowPromptTraining filters routing to providers that do not train on prompts
DisallowPromptTraining bool

// QuotaEntityID identifies the entity against which quota is tracked
QuotaEntityID string
}

Workflow Serialization​

Gateway embedding and image models can cross a workflow boundary with provider.SerializeEmbeddingModel / DeserializeEmbeddingModel and provider.SerializeImageModel / DeserializeImageModel (language models could already be serialized with providerutils.SerializeModel / DeserializeModel). See Provider Serialization for the mechanism; speech, transcription, and video models are not yet serializable.

Limitations​

  • Speech synthesis and transcription not directly supported (use specific providers)
  • Some provider-specific features may not be available through Gateway

Examples​

Learn More​

May 2026 parity updates​

Reranking​

Gateway supports reranking through provider.RerankingModel and the /reranking-model endpoint.

ranker, err := gatewayProvider.RerankingModel("cohere/rerank-v3.5")
topN := 3
result, err := ranker.DoRerank(ctx, &provider.RerankOptions{
Query: "go concurrency",
Documents: []string{"Goroutines are lightweight.", "SQL joins tables."},
TopN: &topN,
})

Routing, quota, and compliance​

gateway.GatewayProviderOptions includes Sort, QuotaEntityID, DisallowPromptTraining, ZeroDataRetention, ServiceTier, Has, provider order, model filters, and BYOK options.

zdr := true
opts := gateway.GatewayProviderOptions{
Sort: "cost",
QuotaEntityID: "tenant_123",
ServiceTier: "priority",
ZeroDataRetention: &zdr,
Models: []gateway.GatewayModelFallback{
gateway.GatewayModel("openai/gpt-5.1"),
gateway.GatewayModel("anthropic/claude-sonnet-4-6"),
},
}

Unknown model types​

Gateway metadata parsing is resilient to unknown model types and recognizes reranking as a valid model type.