AI Gateway Provider
The AI Gateway provider enables you to access multiple AI providers through a unified interface with automatic provider selection, failover, and cost optimization.
Features
- Unified Interface: Single API for multiple AI providers
- Automatic Provider Selection: Gateway selects the best available provider
- Failover Handling: Automatic failover to alternative providers
- Cost Optimization: Intelligentrouting based on cost and performance
- Zero Data Retention: Option to not log or retain request data
- Observability: Built-in metrics and monitoring
- Rate Limiting: Automatic rate limit handling across providers
Installation
The Gateway provider is included in the core Go-AI SDK:
go get github.com/digitallysavvy/go-ai
Setup
Environment Variables
export AI_GATEWAY_API_KEY="your-api-key-here"
AI_GATEWAY_API_KEY may contain either an AI Gateway API key or a Vercel access token. When a Vercel access token can access multiple teams, set TeamIDOrSlug; the provider sends it as x-vercel-ai-gateway-team.
Basic Configuration
package main
import (
"github.com/digitallysavvy/go-ai/pkg/providers/gateway"
)
func main() {
// Create provider with API key from environment
provider, err := gateway.New(gateway.Config{
APIKey: os.Getenv("AI_GATEWAY_API_KEY"),
})
if err != nil {
log.Fatal(err)
}
}
Advanced Configuration
provider, err := gateway.New(gateway.Config{
APIKey: "your-api-key-or-vercel-access-token",
TeamIDOrSlug: "team-slug-or-id",
BaseURL: "https://ai-gateway.vercel.sh/v4/ai", // Custom base URL
Headers: map[string]string{
"X-Custom-Header": "value",
},
ZeroDataRetention: true, // Enable zero data retention
DisallowPromptTraining: true,
QuotaEntityID: "tenant-123",
MetadataCacheRefreshMillis: 300000, // 5 minutes
})
Text Generation
import (
"context"
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/gateway"
)
// Create provider
provider, _ := gateway.New(gateway.Config{
APIKey: os.Getenv("AI_GATEWAY_API_KEY"),
})
// Create language model
// Gateway routes to the best available provider
model, _ := provider.LanguageModel(string(gateway.GatewayLanguageModelAnthropicClaudeSonnet45))
// Generate text
maxTokens := 200
result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{
Model: model,
Prompt: "Explain quantum computing in simple terms.",
MaxTokens: &maxTokens,
})
Model ID catalog
pkg/providers/gateway/model_ids.go is generated from the upstream Gateway
catalog (go generate ./pkg/providers/gateway) and gives compile-time
constants such as gateway.GatewayLanguageModelAnthropicClaudeSonnet45, or
you can pass the raw Gateway model ID string directly.
Breaking change: all
xai/*model IDs were renamed tospacexai/*upstream (GatewayLanguageModelXai*constants are nowGatewayLanguageModelSpacexai*, e.g.gateway.GatewayLanguageModelSpacexaiGrok420MultiAgent/"spacexai/grok-4.20-multi-agent"), and 44 language, 4 image, and 4 video IDs that no longer exist upstream were removed from the generated file.
Video Generation
The Gateway provider supports video generation with automatic provider selection:
// Create video model
videoModel, err := provider.VideoModel("google/veo-3.1-generate-001")
if err != nil {
log.Fatal(err)
}
// Generate video
duration := 4.0
result, err := ai.GenerateVideo(context.Background(), ai.GenerateVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{
Text: "A cat playing with a ball of yarn",
},
AspectRatio: "16:9",
Duration: &duration,
})
Video Generation Features
- Text-to-Video: Generate videos from text descriptions
- Image-to-Video: Animate static images
- Batch Generation: Generate multiple videos in one request
- Provider Selection: Automatic routing to best video provider
- Custom Parameters: Aspect ratio, duration, FPS, resolution
Supported Video Providers
The gateway can route video requests to:
- Google Vertex AI (Veo)
- FAL AI
- Replicate
- Other configured providers
Video Generation Example
videoModel, _ := provider.VideoModel("google/veo-3.1-generate-001")
duration := 4.0
fps := 30
result, err := ai.GenerateVideo(context.Background(), ai.GenerateVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{
Text: "A serene ocean sunset with waves",
},
AspectRatio: "16:9",
Duration: &duration,
FPS: &fps,
ProviderOptions: map[string]interface{}{
"enhancePrompt": true,
},
})
// Access generated videos
for _, video := range result.Videos {
fmt.Printf("Video URL: %s\n", video.URL)
fmt.Printf("Media Type: %s\n", video.MediaType)
}
Image Generation
imageModel, _ := provider.ImageModel(string(gateway.GatewayImageModelOpenaiGptImage2))
n := 1
result, err := ai.GenerateImage(context.Background(), ai.GenerateImageOptions{
Model: imageModel,
Prompt: "A futuristic cityscape at sunset",
N: &n,
Size: "1024x1024",
})
// The Gateway image model sets its own retryability signal (*bool) on a
// failed/empty generation; ai.GenerateImage reads it internally through
// the retry loop, so this does not appear as a field on the returned
// GenerateImageResult.
Embeddings
embeddingModel, _ := provider.EmbeddingModel(string(gateway.GatewayEmbeddingModelOpenaiTextEmbedding3Small))
result, err := ai.EmbedMany(context.Background(), ai.EmbedManyOptions{
Model: embeddingModel,
Inputs: []string{
"The cat sat on the mat",
"A dog ran through the park",
},
})
Provider Options
Customize provider selection and behavior:
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Hello, world!",
ProviderOptions: map[string]interface{}{
"gateway": map[string]interface{}{
"only": []string{"anthropic", "openai", "google"},
"tags": []string{"production", "fast"},
"order": []string{"openai", "anthropic"},
"sort": "cost",
},
},
})
Provider Option Fields
- gateway.only: Restrict candidate providers.
- gateway.order: Explicit provider order.
- gateway.sort: Gateway sort strategy.
- gateway.tags: Routing tags.
- gateway.models: Ordered fallback model list. Build entries with
gateway.GatewayModel(id)for a plain fallback, orgateway.GatewayConditionalModelFallback(id, when)for an evaluation-only conditional fallback (only the first entry may be conditional). JSON also accepts a plain model-ID string in place of either helper. - gateway.has: Restrict routing to models satisfying capability tags
(
gateway.GatewayHasImplicitCaching,GatewayHasReasoning,GatewayHasStructuredOutput,GatewayHasToolUse,GatewayHasVision, or a weight-format condition fromGatewayHasQuantization/GatewayHasNotQuantization). - gateway.zeroDataRetention: Per-call zero retention.
- gateway.disallowPromptTraining: Restrict to no-prompt-training providers.
- gateway.quotaEntityId: Quota identity for tenant/account-level tracking.
- gateway.serviceTier: Unified service tier intent. Supported values are
"flex"and"priority"; Gateway translates this to the provider-specific tier and overrides lower-level provider tier options.
Breaking change:
gateway.hipaaCompliant/Config.HIPAACompliantwere removed; the Gateway API no longer accepts this field.
Reranking
import goprovider "github.com/digitallysavvy/go-ai/pkg/provider"
reranker, _ := gatewayProvider.RerankingModel("cohere/rerank-v3.5")
topN := 3
res, err := reranker.DoRerank(ctx, &goprovider.RerankOptions{
Query: "gateway routing",
Documents: []string{
"Provider fallback order can reduce latency variance.",
"Embedding vectors power semantic search.",
"HIPAA compliance restricts provider choice.",
},
TopN: &topN,
})
Timeout Handling
The Gateway SDK provides enhanced timeout error handling with troubleshooting guidance.
Typed Error Classification
Gateway response errors are decoded into typed errors under pkg/providers/gateway/errors. The common GatewayError interface exposes status code, public type, generation ID, and retryability. Unknown Gateway error types are classified the same way as the TypeScript SDK while preserving the raw type through GatewayErrorDetails.
var gatewayErr gatewayerrors.GatewayError
if errors.As(err, &gatewayErr) {
fmt.Println(gatewayErr.GetType(), gatewayErr.GetStatusCode(), gatewayErr.IsRetryable())
}
var details gatewayerrors.GatewayErrorDetails
if errors.As(err, &details) {
fmt.Println(details.GetRawType(), details.GetCode(), details.GetParam())
}
Inline []byte file data in Gateway language-model requests is base64-encoded exactly once for file, reasoning-file, and tool-result file parts. URL, provider-reference, and text file data are preserved in their provider-native form.
Realtime Runtime Client Secrets
Gateway realtime is exposed through provider.ExperimentalRealtime(modelID).
Use GetRealtimeToken or MintRealtimeClientSecret on a server to create the
short-lived client secret used by browser WebSocket clients.
gw, err := gateway.New(gateway.Config{
APIKey: os.Getenv("AI_GATEWAY_API_KEY"),
TeamIDOrSlug: os.Getenv("VERCEL_TEAM_ID"),
})
if err != nil {
return err
}
expiresAfterSeconds := 600
secret, err := gw.GetRealtimeToken(ctx, provider.RealtimeFactoryGetTokenOptions{
Model: "openai/gpt-realtime",
ClientSecretOptions: provider.ClientSecretOptions{
ExpiresAfterSeconds: &expiresAfterSeconds,
},
})
if err != nil {
return err
}
model := gw.ExperimentalRealtime("openai/gpt-realtime")
ws := model.GetWebSocketConfig(secret.Token, secret.URL)
fmt.Println(ws.URL, ws.Protocols)
MintRealtimeClientSecret sends POST /v1/realtime/client-secrets at the
Gateway origin, not under /v4/ai. The realtime WebSocket protocols include
ai-gateway-realtime.v1, ai-gateway-auth.<token>, and, when configured, a
base64url encoded ai-gateway-team.<team> entry. Realtime event parsing,
client event serialization, and session config building are identity mappings
for provider-native realtime payloads.
Setting Timeouts
import (
"context"
"time"
gatewayerrors "github.com/digitallysavvy/go-ai/pkg/providers/gateway/errors"
)
// Text generation: 30-60 seconds
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
options.Model = model
result, err := ai.GenerateText(ctx, options)
// Video generation: 5-10 minutes
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Minute)
defer cancel()
options.Model = videoModel
result, err := ai.GenerateVideo(ctx, options)
Detecting Timeout Errors
options.Model = videoModel
result, err := ai.GenerateVideo(ctx, options)
if err != nil {
if gatewayerrors.IsGatewayTimeoutError(err) {
// This is a client-side timeout
// Error message includes troubleshooting guidance
fmt.Printf("Timeout error: %v\n", err)
} else {
// Other error
fmt.Printf("Error: %v\n", err)
}
}
Timeout Best Practices
Text Generation:
- Simple prompts: 30-60 seconds
- Complex generation: 2-5 minutes
Video Generation:
- Short videos (2-4s): 5-10 minutes
- Long videos: 15-30 minutes
- Batch generation: 20-40 minutes
Image Generation:
- Single image: 1-2 minutes
- Batch images: 3-5 minutes
Zero Data Retention
Enable zero data retention mode to ensure requests are not logged:
provider, err := gateway.New(gateway.Config{
APIKey: os.Getenv("AI_GATEWAY_API_KEY"),
ZeroDataRetention: true,
})
Observability
The Gateway automatically includes observability headers when running on Vercel:
// Automatically includes:
// - ai-o11y-deployment-id
// - ai-o11y-environment
// - ai-o11y-region
// - ai-o11y-request-id
Available Models
Query available models and providers:
metadata, err := provider.GetAvailableModels(context.Background())
if err != nil {
log.Fatal(err)
}
for _, model := range metadata.Models {
fmt.Printf("Provider: %s\n", model.Specification.Provider)
fmt.Printf(" - %s (%s)\n", model.Name, model.ID)
fmt.Printf(" Specification: %v\n", model.Specification)
}
Credits
Check your Gateway credit usage:
credits, err := provider.GetCredits(context.Background())
if err != nil {
log.Fatal(err)
}
fmt.Printf("Balance: %s\n", credits.Balance)
fmt.Printf("Total Used: %s\n", credits.TotalUsed)
/v1/credits reads the team from a teamId/slug query parameter rather than the x-vercel-ai-gateway-team header every other endpoint uses. When Config.TeamIDOrSlug (or a per-call x-vercel-ai-gateway-team header) is set, GetCredits forwards it as that query parameter too, so credentials that can access multiple teams (such as Vercel access tokens) are scoped correctly instead of getting a 401.
Evaluation model fallbacks
EvaluationFallbackCondition.Question is optional. Without it, ConfidenceBelow checks every Choice/Score question and ProbabilityBetween checks every Boolean question, instead of a single named question:
confidenceBelow := 0.6
when := gateway.EvaluationFallbackCondition{ConfidenceBelow: &confidenceBelow}
fallback := gateway.GatewayConditionalModelFallback("openai/gpt-5.6-sol", when)
A condition still needs a check: {} or a condition with only Question set is rejected.
Error Handling
The Gateway provides detailed error types:
import gatewayerrors "github.com/digitallysavvy/go-ai/pkg/providers/gateway/errors"
options.Model = model
result, err := ai.GenerateText(ctx, options)
if err != nil {
switch {
case gatewayerrors.IsGatewayTimeoutError(err):
// Handle timeout with troubleshooting guidance
fmt.Printf("Request timed out: %v\n", err)
case providererrors.IsRateLimitError(err):
// Handle rate limiting
fmt.Printf("Rate limited: %v\n", err)
case providererrors.IsProviderError(err):
// Handle provider errors
fmt.Printf("Provider error: %v\n", err)
default:
// Handle other errors
fmt.Printf("Error: %v\n", err)
}
}
Configuration Reference
Config
type Config struct {
// APIKey is the AI Gateway API key or Vercel access token (required)
// Can also be set via AI_GATEWAY_API_KEY environment variable
APIKey string
// TeamIDOrSlug scopes Vercel access-token requests to a team.
TeamIDOrSlug string
// BaseURL is the base URL for the AI Gateway API
// Default: https://ai-gateway.vercel.sh/v4/ai
BaseURL string
// Headers are custom headers to include in requests
Headers map[string]string
// MetadataCacheRefreshMillis is how frequently to refresh metadata cache
// Default: 300000 (5 minutes)
MetadataCacheRefreshMillis int64
// HTTPClient is a custom HTTP client
HTTPClient *http.Client
// ZeroDataRetention enables zero data retention mode
// When true, requests are not logged or retained
ZeroDataRetention bool
// ProjectID is forwarded as the "ai-o11y-project-id" header. Can also be
// set via VERCEL_PROJECT_ID.
ProjectID *string
// DisallowPromptTraining filters routing to providers that do not train on prompts
DisallowPromptTraining bool
// QuotaEntityID identifies the entity against which quota is tracked
QuotaEntityID string
}
Workflow Serialization
Gateway embedding and image models can cross a workflow boundary with
provider.SerializeEmbeddingModel / DeserializeEmbeddingModel and
provider.SerializeImageModel / DeserializeImageModel (language models
could already be serialized with providerutils.SerializeModel /
DeserializeModel). See
Provider Serialization
for the mechanism; speech, transcription, and video models are not yet
serializable.
Limitations
- Speech synthesis and transcription not directly supported (use specific providers)
- Some provider-specific features may not be available through Gateway
Examples
Related Documentation
Learn More
May 2026 parity updates
Reranking
Gateway supports reranking through provider.RerankingModel and the /reranking-model endpoint.
ranker, err := gatewayProvider.RerankingModel("cohere/rerank-v3.5")
topN := 3
result, err := ranker.DoRerank(ctx, &provider.RerankOptions{
Query: "go concurrency",
Documents: []string{"Goroutines are lightweight.", "SQL joins tables."},
TopN: &topN,
})
Routing, quota, and compliance
gateway.GatewayProviderOptions includes Sort, QuotaEntityID, DisallowPromptTraining, ZeroDataRetention, ServiceTier, Has, provider order, model filters, and BYOK options.
zdr := true
opts := gateway.GatewayProviderOptions{
Sort: "cost",
QuotaEntityID: "tenant_123",
ServiceTier: "priority",
ZeroDataRetention: &zdr,
Models: []gateway.GatewayModelFallback{
gateway.GatewayModel("openai/gpt-5.1"),
gateway.GatewayModel("anthropic/claude-sonnet-4-6"),
},
}
Unknown model types
Gateway metadata parsing is resilient to unknown model types and recognizes reranking as a valid model type.