Skip to main content

xAI Provider

xAI provides Grok models with real-time knowledge through X (Twitter) integration, image/video generation, and strong reasoning capabilities.

Setup​

Installation​

import (
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/xai"
)

Configuration​

provider := xai.New(xai.Config{
APIKey: os.Getenv("XAI_API_KEY"),
})

// Responses API (the only language model API xAI supports)
model, err := provider.LanguageModel("grok-3")

Get API key​

  1. Sign up at x.ai
  2. Request API access
  3. Set environment variable:
export XAI_API_KEY=xai-...

Available models​

Language models​

Model IDDescription
grok-3Latest Grok 3 model
grok-3-latestAlias for latest Grok 3
grok-3-miniSmaller, faster Grok 3 variant
grok-4Latest Grok 4 model
grok-4-latestAlias for latest Grok 4
grok-4.3Current Grok 4.3 model
grok-4.20-multi-agentMulti-agent orchestration
grok-4.20-reasoningReasoning-optimized
grok-4.20-non-reasoningFast non-reasoning

Image models​

Model IDDescription
grok-imagine-imageImage generation
grok-imagine-image-proHigher quality image generation

Video models​

Model IDDescription
grok-imagine-videoVideo generation

Audio models​

xAI speech and transcription use provider factory methods rather than named model IDs. The Go provider method keeps the shared provider interface shape, but the model ID parameter is ignored to match the TypeScript provider default.

speechModel, _ := provider.SpeechModel("")
transcriptionModel, _ := provider.TranscriptionModel("")

Responses API (only)​

LanguageModel() returns a Responses API model. Starting in v0.4.0 this was the default (with the legacy Chat Completions API available as an opt-in escape hatch); as of this release the Chat Completions API (ChatCompletionsLanguageModel()) has been removed entirely, matching the upstream TypeScript SDK. See the migration guide if you were still using ChatCompletionsLanguageModel().

model, _ := provider.LanguageModel("grok-3")

result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "What's happening in AI today?",
})

Provider options​

Use xai.WebSearch with the Responses API. EnableImageSearch maps to xAI's enable_image_search request field.

enableImageSearch := true
result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Show me images of SpaceX Starship on the launch pad.",
Tools: []types.Tool{
xai.WebSearch(xai.WebSearchConfig{
EnableImageSearch: &enableImageSearch,
}),
},
})

The legacy chat-completions SearchParameters option was removed along with ChatCompletionsLanguageModel(). Use the provider-executed WebSearch or XSearch tools on the Responses API instead.

Reasoning summary​

Control reasoning output in Responses API:

result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Explain quantum computing",
ProviderOptions: map[string]interface{}{
"xai": map[string]interface{}{
"reasoningSummary": "detailed", // "auto", "concise", or "detailed"
},
},
})

File URL inputs​

The xAI Responses API path accepts public file URLs for non-image content, including application/pdf and text/* media types. The SDK serializes these as input_file with file_url, matching the TypeScript provider. Inline non-image bytes remain unsupported by xAI; use a public URL or provider file reference for PDFs and text files.

result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Summarize this report."},
types.FileContent{
MediaType: "application/pdf",
FileData: types.FileData{
Type: types.FileDataTypeURL,
URL: "https://example.com/report.pdf",
},
},
},
}},
})

Cost metadata​

xAI Responses usage metadata is exposed under ProviderMetadata["xai"]. costInUsdTicks is forwarded directly, and token costs are available in the nested cost map using TypeScript-compatible camelCase keys.

meta := result.ProviderMetadata["xai"].(map[string]interface{})
fmt.Println(meta["costInUsdTicks"])

if cost, ok := meta["cost"].(map[string]interface{}); ok {
fmt.Println(cost["inputTokensCost"])
fmt.Println(cost["outputTokensCost"])
}

Logprobs​

result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Hello",
ProviderOptions: map[string]interface{}{
"xai": map[string]interface{}{
"logprobs": true,
"topLogprobs": 5,
},
},
})

Image generation​

imageModel, _ := provider.ImageModel("grok-imagine-image")

result, _ := ai.GenerateImage(ctx, ai.GenerateImageOptions{
Model: imageModel,
Prompt: "A futuristic cityscape at sunset",
ProviderOptions: map[string]interface{}{
"xai": map[string]interface{}{
"quality": "high", // "low", "medium", "high"
},
},
})

Supports multiple input images for editing, b64_json output format, and CostInUsdTicks in provider metadata.

Video generation​

videoModel, _ := provider.VideoModel("grok-imagine-video")

result, _ := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{
Model: videoModel,
Prompt: "A cat playing piano",
})

Returns ModerationError if content is filtered. CostInUsdTicks available in provider metadata.

Async video​

xai.VideoModel implements provider.VideoModelStarter / VideoModelStatusChecker, so it also works with the fire-and-forget ai.ExperimentalStartVideo / ai.ExperimentalGetVideoStatus calls instead of the polling ai.GenerateVideo:

started, _ := ai.ExperimentalStartVideo(ctx, ai.StartVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{Text: "A cat playing piano"},
})

// Persist started.Operation, then later (even from another process):
status, _ := ai.ExperimentalGetVideoStatus(ctx, videoModel, ai.GetVideoStatusOptions{
Operation: started.Operation,
})
if status.Status == provider.VideoOperationStatusCompleted {
fmt.Println(status.Videos)
}

Speech and transcription​

SpeechModel("") maps to xAI's /v1/tts endpoint. It defaults to voice "eve", language "auto", and MP3 output, matching the TypeScript provider. Use provider options under the "xai" key for xAI-specific audio controls.

speechModel, _ := provider.SpeechModel("")

result, _ := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{
Model: speechModel,
Text: "Hello from Go-AI.",
Voice: "ara",
OutputFormat: "wav",
ProviderOptions: map[string]interface{}{
"xai": map[string]interface{}{
"sampleRate": 44100,
"optimizeStreamingLatency": 1,
"textNormalization": true,
},
},
})

TranscriptionModel("") maps to /v1/stt. Multipart request fields match the TypeScript provider, including repeated keyterm fields and the audio file part last.

transcriptionModel, _ := provider.TranscriptionModel("")

transcript, _ := ai.Transcribe(ctx, ai.TranscribeOptions{
Model: transcriptionModel,
Audio: audioBytes,
MimeType: "audio/wav",
ProviderOptions: map[string]interface{}{
"xai": map[string]interface{}{
"language": "en",
"diarize": true,
"keyterm": []string{"Go-AI", "Grok"},
"fillerWords": true,
},
},
})

fmt.Println(transcript.Text)

Best practices​

  1. Use the Responses API — it is the default and recommended path
  2. Leverage ReasoningSummary — get insight into model reasoning without full chain-of-thought overhead
  3. Check CostInUsdTicks — available in image and video provider metadata for cost tracking

See also​

May 2026 parity updates​

Non-image file URLs​

The xAI Responses provider accepts non-image file URLs by serializing URL file parts as input_file with file_url.

messages := []types.Message{{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Summarize this report."},
types.FileContent{FileData: types.FileData{
Type: types.FileDataTypeURL,
URL: "https://example.com/report.pdf",
MediaType: "application/pdf",
}},
},
}}

Cost and ZDR metadata​

xAI response, image, and video calls surface costInUsdTicks under provider metadata when the API returns it. Encrypted reasoning is preserved for zero-data-retention follow-up turns.

File uploads​

Use ai.UploadFile with xai.New(...). The returned ProviderReference is map[string]string{"xai": fileID}.

Workflow serialization​

xAI's image, speech, and transcription models can cross a workflow boundary with providerutils.SerializeModel / DeserializeModel (language models could already be serialized). See Provider Serialization for the mechanism; video models are not yet serializable.