xAI Provider
xAI provides Grok models with real-time knowledge through X (Twitter) integration, image/video generation, and strong reasoning capabilities.
Setup
Installation
import (
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/xai"
)
Configuration
provider := xai.New(xai.Config{
APIKey: os.Getenv("XAI_API_KEY"),
})
// Responses API (the only language model API xAI supports)
model, err := provider.LanguageModel("grok-3")
Get API key
- Sign up at x.ai
- Request API access
- Set environment variable:
export XAI_API_KEY=xai-...
Available models
Language models
| Model ID | Description |
|---|---|
grok-3 | Latest Grok 3 model |
grok-3-latest | Alias for latest Grok 3 |
grok-3-mini | Smaller, faster Grok 3 variant |
grok-4 | Latest Grok 4 model |
grok-4-latest | Alias for latest Grok 4 |
grok-4.3 | Current Grok 4.3 model |
grok-4.20-multi-agent | Multi-agent orchestration |
grok-4.20-reasoning | Reasoning-optimized |
grok-4.20-non-reasoning | Fast non-reasoning |
Image models
| Model ID | Description |
|---|---|
grok-imagine-image | Image generation |
grok-imagine-image-pro | Higher quality image generation |
Video models
| Model ID | Description |
|---|---|
grok-imagine-video | Video generation |
Audio models
xAI speech and transcription use provider factory methods rather than named model IDs. The Go provider method keeps the shared provider interface shape, but the model ID parameter is ignored to match the TypeScript provider default.
speechModel, _ := provider.SpeechModel("")
transcriptionModel, _ := provider.TranscriptionModel("")
Responses API (only)
LanguageModel() returns a Responses API model. Starting in v0.4.0 this was the default (with the legacy Chat Completions API available as an opt-in escape hatch); as of this release the Chat Completions API (ChatCompletionsLanguageModel()) has been removed entirely, matching the upstream TypeScript SDK. See the migration guide if you were still using ChatCompletionsLanguageModel().
model, _ := provider.LanguageModel("grok-3")
result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "What's happening in AI today?",
})
Provider options
Provider-executed web search
Use xai.WebSearch with the Responses API. EnableImageSearch maps to xAI's enable_image_search request field.
enableImageSearch := true
result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Show me images of SpaceX Starship on the launch pad.",
Tools: []types.Tool{
xai.WebSearch(xai.WebSearchConfig{
EnableImageSearch: &enableImageSearch,
}),
},
})
The legacy chat-completions SearchParameters option was removed along with ChatCompletionsLanguageModel(). Use the provider-executed WebSearch or XSearch tools on the Responses API instead.
Reasoning summary
Control reasoning output in Responses API:
result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Explain quantum computing",
ProviderOptions: map[string]interface{}{
"xai": map[string]interface{}{
"reasoningSummary": "detailed", // "auto", "concise", or "detailed"
},
},
})
File URL inputs
The xAI Responses API path accepts public file URLs for non-image content, including application/pdf and text/* media types. The SDK serializes these as input_file with file_url, matching the TypeScript provider. Inline non-image bytes remain unsupported by xAI; use a public URL or provider file reference for PDFs and text files.
result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Summarize this report."},
types.FileContent{
MediaType: "application/pdf",
FileData: types.FileData{
Type: types.FileDataTypeURL,
URL: "https://example.com/report.pdf",
},
},
},
}},
})
Cost metadata
xAI Responses usage metadata is exposed under ProviderMetadata["xai"]. costInUsdTicks is forwarded directly, and token costs are available in the nested cost map using TypeScript-compatible camelCase keys.
meta := result.ProviderMetadata["xai"].(map[string]interface{})
fmt.Println(meta["costInUsdTicks"])
if cost, ok := meta["cost"].(map[string]interface{}); ok {
fmt.Println(cost["inputTokensCost"])
fmt.Println(cost["outputTokensCost"])
}
Logprobs
result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Hello",
ProviderOptions: map[string]interface{}{
"xai": map[string]interface{}{
"logprobs": true,
"topLogprobs": 5,
},
},
})
Image generation
imageModel, _ := provider.ImageModel("grok-imagine-image")
result, _ := ai.GenerateImage(ctx, ai.GenerateImageOptions{
Model: imageModel,
Prompt: "A futuristic cityscape at sunset",
ProviderOptions: map[string]interface{}{
"xai": map[string]interface{}{
"quality": "high", // "low", "medium", "high"
},
},
})
Supports multiple input images for editing, b64_json output format, and CostInUsdTicks in provider metadata.
Video generation
videoModel, _ := provider.VideoModel("grok-imagine-video")
result, _ := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{
Model: videoModel,
Prompt: "A cat playing piano",
})
Returns ModerationError if content is filtered. CostInUsdTicks available in provider metadata.
Async video
xai.VideoModel implements provider.VideoModelStarter / VideoModelStatusChecker, so it also works with the fire-and-forget ai.ExperimentalStartVideo / ai.ExperimentalGetVideoStatus calls instead of the polling ai.GenerateVideo:
started, _ := ai.ExperimentalStartVideo(ctx, ai.StartVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{Text: "A cat playing piano"},
})
// Persist started.Operation, then later (even from another process):
status, _ := ai.ExperimentalGetVideoStatus(ctx, videoModel, ai.GetVideoStatusOptions{
Operation: started.Operation,
})
if status.Status == provider.VideoOperationStatusCompleted {
fmt.Println(status.Videos)
}
Speech and transcription
SpeechModel("") maps to xAI's /v1/tts endpoint. It defaults to voice
"eve", language "auto", and MP3 output, matching the TypeScript provider.
Use provider options under the "xai" key for xAI-specific audio controls.
speechModel, _ := provider.SpeechModel("")
result, _ := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{
Model: speechModel,
Text: "Hello from Go-AI.",
Voice: "ara",
OutputFormat: "wav",
ProviderOptions: map[string]interface{}{
"xai": map[string]interface{}{
"sampleRate": 44100,
"optimizeStreamingLatency": 1,
"textNormalization": true,
},
},
})
TranscriptionModel("") maps to /v1/stt. Multipart request fields match the
TypeScript provider, including repeated keyterm fields and the audio file
part last.
transcriptionModel, _ := provider.TranscriptionModel("")
transcript, _ := ai.Transcribe(ctx, ai.TranscribeOptions{
Model: transcriptionModel,
Audio: audioBytes,
MimeType: "audio/wav",
ProviderOptions: map[string]interface{}{
"xai": map[string]interface{}{
"language": "en",
"diarize": true,
"keyterm": []string{"Go-AI", "Grok"},
"fillerWords": true,
},
},
})
fmt.Println(transcript.Text)
Best practices
- Use the Responses API — it is the default and recommended path
- Leverage ReasoningSummary — get insight into model reasoning without full chain-of-thought overhead
- Check CostInUsdTicks — available in image and video provider metadata for cost tracking
See also
May 2026 parity updates
Non-image file URLs
The xAI Responses provider accepts non-image file URLs by serializing URL file parts as input_file with file_url.
messages := []types.Message{{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Summarize this report."},
types.FileContent{FileData: types.FileData{
Type: types.FileDataTypeURL,
URL: "https://example.com/report.pdf",
MediaType: "application/pdf",
}},
},
}}
Cost and ZDR metadata
xAI response, image, and video calls surface costInUsdTicks under provider metadata when the API returns it. Encrypted reasoning is preserved for zero-data-retention follow-up turns.
File uploads
Use ai.UploadFile with xai.New(...). The returned ProviderReference is map[string]string{"xai": fileID}.
Workflow serialization
xAI's image, speech, and transcription models can cross a workflow boundary with providerutils.SerializeModel / DeserializeModel (language models could already be serialized). See Provider Serialization for the mechanism; video models are not yet serializable.