# xAI Provider

xAI provides Grok models with real-time knowledge through X (Twitter) integration, image/video generation, and strong reasoning capabilities.

## Setup

### Installation

```go
import (
    "github.com/digitallysavvy/go-ai/pkg/ai"
    "github.com/digitallysavvy/go-ai/pkg/providers/xai"
)
```

### Configuration

```go
provider := xai.New(xai.Config{
    APIKey: os.Getenv("XAI_API_KEY"),
})

// Responses API (the only language model API xAI supports)
model, err := provider.LanguageModel("grok-3")
```

### Get API key

1. Sign up at [x.ai](https://x.ai)
2. Request API access
3. Set environment variable:

```bash
export XAI_API_KEY=xai-...
```

## Available models

### Language models

| Model ID | Description |
|----------|-------------|
| `grok-3` | Latest Grok 3 model |
| `grok-3-latest` | Alias for latest Grok 3 |
| `grok-3-mini` | Smaller, faster Grok 3 variant |
| `grok-4` | Latest Grok 4 model |
| `grok-4-latest` | Alias for latest Grok 4 |
| `grok-4.3` | Current Grok 4.3 model |
| `grok-4.20-multi-agent` | Multi-agent orchestration |
| `grok-4.20-reasoning` | Reasoning-optimized |
| `grok-4.20-non-reasoning` | Fast non-reasoning |

### Image models

| Model ID | Description |
|----------|-------------|
| `grok-imagine-image` | Image generation |
| `grok-imagine-image-pro` | Higher quality image generation |

### Video models

| Model ID | Description |
|----------|-------------|
| `grok-imagine-video` | Video generation |

### Audio models

xAI speech and transcription use provider factory methods rather than named
model IDs. The Go provider method keeps the shared provider interface shape,
but the model ID parameter is ignored to match the TypeScript provider default.

```go
speechModel, _ := provider.SpeechModel("")
transcriptionModel, _ := provider.TranscriptionModel("")
```

## Responses API (only)

`LanguageModel()` returns a Responses API model. Starting in v0.4.0 this was the default (with the legacy Chat Completions API available as an opt-in escape hatch); as of this release the Chat Completions API (`ChatCompletionsLanguageModel()`) has been removed entirely, matching the upstream TypeScript SDK. See the migration guide if you were still using `ChatCompletionsLanguageModel()`.

```go
model, _ := provider.LanguageModel("grok-3")

result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: "What's happening in AI today?",
})
```

## Provider options

### Provider-executed web search

Use `xai.WebSearch` with the Responses API. `EnableImageSearch` maps to xAI's `enable_image_search` request field.

```go
enableImageSearch := true
result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: "Show me images of SpaceX Starship on the launch pad.",
    Tools: []types.Tool{
        xai.WebSearch(xai.WebSearchConfig{
            EnableImageSearch: &enableImageSearch,
        }),
    },
})
```

The legacy chat-completions `SearchParameters` option was removed along with `ChatCompletionsLanguageModel()`. Use the provider-executed `WebSearch` or `XSearch` tools on the Responses API instead.

### Reasoning summary

Control reasoning output in Responses API:

```go
result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: "Explain quantum computing",
    ProviderOptions: map[string]interface{}{
        "xai": map[string]interface{}{
            "reasoningSummary": "detailed", // "auto", "concise", or "detailed"
        },
    },
})
```

### File URL inputs

The xAI Responses API path accepts public file URLs for non-image content, including `application/pdf` and `text/*` media types. The SDK serializes these as `input_file` with `file_url`, matching the TypeScript provider. Inline non-image bytes remain unsupported by xAI; use a public URL or provider file reference for PDFs and text files.

```go
result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model: model,
    Messages: []types.Message{{
        Role: types.RoleUser,
        Content: []types.ContentPart{
            types.TextContent{Text: "Summarize this report."},
            types.FileContent{
                MediaType: "application/pdf",
                FileData: types.FileData{
                    Type: types.FileDataTypeURL,
                    URL:  "https://example.com/report.pdf",
                },
            },
        },
    }},
})
```

### Cost metadata

xAI Responses usage metadata is exposed under `ProviderMetadata["xai"]`. `costInUsdTicks` is forwarded directly, and token costs are available in the nested `cost` map using TypeScript-compatible camelCase keys.

```go
meta := result.ProviderMetadata["xai"].(map[string]interface{})
fmt.Println(meta["costInUsdTicks"])

if cost, ok := meta["cost"].(map[string]interface{}); ok {
    fmt.Println(cost["inputTokensCost"])
    fmt.Println(cost["outputTokensCost"])
}
```

### Logprobs

```go
result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: "Hello",
    ProviderOptions: map[string]interface{}{
        "xai": map[string]interface{}{
            "logprobs":    true,
            "topLogprobs": 5,
        },
    },
})
```

## Image generation

```go
imageModel, _ := provider.ImageModel("grok-imagine-image")

result, _ := ai.GenerateImage(ctx, ai.GenerateImageOptions{
    Model:  imageModel,
    Prompt: "A futuristic cityscape at sunset",
    ProviderOptions: map[string]interface{}{
        "xai": map[string]interface{}{
            "quality": "high",  // "low", "medium", "high"
        },
    },
})
```

Supports multiple input images for editing, b64_json output format, and `CostInUsdTicks` in provider metadata.

## Video generation

```go
videoModel, _ := provider.VideoModel("grok-imagine-video")

result, _ := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{
    Model:  videoModel,
    Prompt: "A cat playing piano",
})
```

Returns `ModerationError` if content is filtered. `CostInUsdTicks` available in provider metadata.

### Async video

`xai.VideoModel` implements `provider.VideoModelStarter` / `VideoModelStatusChecker`, so it also works with the fire-and-forget `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus` calls instead of the polling `ai.GenerateVideo`:

```go
started, _ := ai.ExperimentalStartVideo(ctx, ai.StartVideoOptions{
    Model:  videoModel,
    Prompt: ai.VideoPrompt{Text: "A cat playing piano"},
})

// Persist started.Operation, then later (even from another process):
status, _ := ai.ExperimentalGetVideoStatus(ctx, videoModel, ai.GetVideoStatusOptions{
    Operation: started.Operation,
})
if status.Status == provider.VideoOperationStatusCompleted {
    fmt.Println(status.Videos)
}
```

## Speech and transcription

`SpeechModel("")` maps to xAI's `/v1/tts` endpoint. It defaults to voice
`"eve"`, language `"auto"`, and MP3 output, matching the TypeScript provider.
Use provider options under the `"xai"` key for xAI-specific audio controls.

```go
speechModel, _ := provider.SpeechModel("")

result, _ := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{
    Model:        speechModel,
    Text:         "Hello from Go-AI.",
    Voice:        "ara",
    OutputFormat: "wav",
    ProviderOptions: map[string]interface{}{
        "xai": map[string]interface{}{
            "sampleRate":               44100,
            "optimizeStreamingLatency": 1,
            "textNormalization":        true,
        },
    },
})
```

`TranscriptionModel("")` maps to `/v1/stt`. Multipart request fields match the
TypeScript provider, including repeated `keyterm` fields and the audio `file`
part last.

```go
transcriptionModel, _ := provider.TranscriptionModel("")

transcript, _ := ai.Transcribe(ctx, ai.TranscribeOptions{
    Model:     transcriptionModel,
    Audio:     audioBytes,
    MimeType:  "audio/wav",
    ProviderOptions: map[string]interface{}{
        "xai": map[string]interface{}{
            "language":    "en",
            "diarize":     true,
            "keyterm":     []string{"Go-AI", "Grok"},
            "fillerWords": true,
        },
    },
})

fmt.Println(transcript.Text)
```

## Best practices

1. **Use the Responses API** — it is the default and recommended path
2. **Leverage ReasoningSummary** — get insight into model reasoning without full chain-of-thought overhead
3. **Check CostInUsdTicks** — available in image and video provider metadata for cost tracking

## See also

- [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md)
- [xAI Documentation](https://docs.x.ai)

## May 2026 parity updates

### Non-image file URLs

The xAI Responses provider accepts non-image file URLs by serializing URL file parts as `input_file` with `file_url`.

```go
messages := []types.Message{{
    Role: types.RoleUser,
    Content: []types.ContentPart{
        types.TextContent{Text: "Summarize this report."},
        types.FileContent{FileData: types.FileData{
            Type:      types.FileDataTypeURL,
            URL:       "https://example.com/report.pdf",
            MediaType: "application/pdf",
        }},
    },
}}
```

### Cost and ZDR metadata

xAI response, image, and video calls surface `costInUsdTicks` under provider metadata when the API returns it. Encrypted reasoning is preserved for zero-data-retention follow-up turns.

### File uploads

Use `ai.UploadFile` with `xai.New(...)`. The returned `ProviderReference` is `map[string]string{"xai": fileID}`.

## Workflow serialization

xAI's image, speech, and transcription models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel` (language models could already be serialized). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; video models are not yet serializable.
