# Go AI SDK > The Go AI SDK is a powerful toolkit for building AI applications and agents with Go, featuring support for 49 providers. Canonical URL: https://goaisdk.com/docs/introduction/ Documentation index: https://goaisdk.com/llms.txt The Go AI SDK is a comprehensive toolkit designed to help developers build AI-powered applications and agents using Go. It provides a unified interface for working with multiple AI providers while leveraging Go's strengths in concurrency, performance, and type safety. ## Why use the Go AI SDK? Integrating large language models (LLMs) into applications is complicated and heavily dependent on the specific model provider you use. The Go AI SDK standardizes integrating artificial intelligence (AI) models across [49 supported providers](https://goaisdk.com/docs/providers.md). This enables developers to focus on building great AI applications, not waste time on technical details. For example, here's how you can generate text with various models using the Go AI SDK: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() // Use OpenAI openaiProvider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) openaiModel, _ := openaiProvider.LanguageModel("gpt-6-astra") // Use Anthropic anthropicProvider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) anthropicModel, _ := anthropicProvider.LanguageModel("claude-sonnet-4-5") // Same API for both! for _, model := range []provider.LanguageModel{openaiModel, anthropicModel} { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Why is Go great for AI applications?", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } } ``` ## What is the Go AI SDK? The Go AI SDK provides: - **Unified API**: Work with 49 AI providers through a consistent interface - **Type Safety**: Leverage Go's strong typing for safer AI applications - **Concurrency**: Built on Go's goroutines, channels, and contexts for efficient streaming - **Production Ready**: Comprehensive error handling, testing utilities, and telemetry - **Server-Side Focus**: Optimized for backend services, APIs, and agents ## Core Features ### Text Generation Generate text with any supported model: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing in simple terms.", }) ``` ### Streaming Stream responses in real-time with automatic backpressure: ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a story about...", }) for chunk := range stream.Chunks() { if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } } ``` ### Structured Data Generation Generate type-safe structured data: ```go type Recipe struct { Name string `json:"name"` Ingredients []string `json:"ingredients"` Instructions []string `json:"instructions"` } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate a recipe for chocolate chip cookies.", Output: ai.ObjectOutput[Recipe](ai.ObjectOutputOptions{ Schema: schema.NewSimpleStructSchema(reflect.TypeOf(Recipe{})), }), }) recipe := result.Output.(Recipe) ``` ### Tool Calling Integrate external tools and APIs: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in San Francisco?", Tools: []types.Tool{ { Name: "getWeather", Description: "Get current weather for a location", Parameters: weatherSchema, Execute: func(ctx context.Context, params map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return fetchWeather(params["location"].(string)) }, }, }, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) ``` ### Agent Framework Build autonomous agents with workflow support: ```go myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful assistant.", Tools: tools, MaxSteps: 10, }) result, err := myAgent.Execute(ctx, "Help me plan a trip to Japan") ``` ## Model Providers The Go AI SDK supports 49 model providers including: - **OpenAI** - GPT-6, GPT-5, o-series - **Anthropic** - Claude Opus, Sonnet, Haiku and Fable - **Google** - Gemini Pro, Gemini Flash - **Mistral** - Mistral Large, Medium, Small - **Cohere** - Command R+, Command R - **Amazon Bedrock** - Access multiple models through AWS - **Azure OpenAI** - Enterprise OpenAI models - **Groq** - Ultra-fast inference - **Together AI** - Open source models - **Fireworks AI** - Fast model serving - And 16+ more providers! ## Go-Specific Advantages ### Native Concurrency Go's goroutines and channels provide natural patterns for AI workloads: ```go // Process multiple requests concurrently results := make(chan string) for _, prompt := range prompts { go func(p string) { result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: p, }) results <- result.Text }(prompt) } ``` ### Context-Based Cancellation Clean cancellation and timeouts using Go's context: ```go ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Long generation...", }) // Automatically cancelled after 30 seconds ``` ### Automatic Backpressure Unbuffered channels provide natural flow control: ```go for chunk := range stream.Chunks() { // Processing automatically slows down generation time.Sleep(100 * time.Millisecond) processChunk(chunk) } ``` ## Getting Started Ready to start building? Choose your path: - **[Quick Start](https://goaisdk.com/docs/getting-started/golang.md)** - Build your first AI application in 5 minutes - **[Foundations](https://goaisdk.com/docs/foundations.md)** - Learn core concepts - **[AI SDK Core](https://goaisdk.com/docs/ai-sdk-core.md)** - Explore the complete API - **[Agents](https://goaisdk.com/docs/agents.md)** - Build autonomous agents - **[Build a chat app](https://goaisdk.com/docs/build-a-chat-app.md)** - A `useChat` frontend on a Go backend, with tool approval and coding agents - **[Recipes](https://goaisdk.com/docs/recipes.md)** - Short, complete programs for common tasks ### Reference app [Shipyard](https://github.com/digitallysavvy/go-ai-demo) is a Next.js chat UI backed by a Go server built with the Go AI SDK. It shows the stock `useChat` hook, an approval-gated tool and Claude Code or Codex running through the harness. Three guides walk through its code: [serve a useChat frontend from Go](https://goaisdk.com/docs/build-a-chat-app/serve-usechat-from-go.md), [tool approval end to end](https://goaisdk.com/docs/build-a-chat-app/tool-approval.md) and [coding agents with the harness](https://goaisdk.com/docs/build-a-chat-app/coding-agents-harness.md). ## Installation ```bash go get github.com/digitallysavvy/go-ai ``` ## Join the Community If you have questions about anything related to the Go AI SDK, you're welcome to: - Open issues on [GitHub](https://github.com/digitallysavvy/go-ai/issues) - Contribute to the project - Share your experiences and use cases ## Why Go for AI? While Python dominates AI/ML model training, Go excels at building production AI applications: - **Performance**: Fast execution and low memory overhead - **Concurrency**: Built-in support for parallel processing - **Deployment**: Single binary deployment, no dependencies - **Type Safety**: Catch errors at compile time - **Production Ready**: Excellent for APIs, microservices, and backend systems - **Cloud Native**: Perfect for Kubernetes, Docker, and cloud deployments The Go AI SDK brings the power of modern LLMs to Go's strengths in building scalable, production-ready systems. --- # Announcing the Go AI SDK > A comprehensive toolkit for building AI applications and agents with Go. Canonical URL: https://goaisdk.com/docs/announcing-go-ai/ Documentation index: https://goaisdk.com/llms.txt We're excited to introduce the **Go AI SDK** - a complete, production-ready toolkit for building AI-powered applications and agents using Go. ## Why Go AI SDK? While the AI ecosystem has been dominated by Python and TypeScript, Go developers have lacked a comprehensive, idiomatic solution for building AI applications. The Go AI SDK fills this gap by bringing the power of modern LLMs to Go's strengths: - **Performance**: Go's execution speed and low memory overhead - **Concurrency**: Native goroutines and channels for streaming - **Type Safety**: Compile-time error checking - **Production Ready**: Built for scalable backend systems - **Cloud Native**: Perfect for microservices and Kubernetes deployments ## What's Included? ### Unified Provider Interface Access 49 AI providers through a single, consistent API: ```go // Switch providers with just one line openaiModel, _ := openaiProvider.LanguageModel("gpt-6-astra") claudeModel, _ := anthropicProvider.LanguageModel("claude-sonnet-4-5") geminiModel, _ := googleProvider.LanguageModel("gemini-2.5-flash") // Same API for all for _, model := range []provider.LanguageModel{openaiModel, claudeModel, geminiModel} { result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) fmt.Println(result.Text) } ``` ### Comprehensive Feature Set The Go AI SDK provides everything you need: - **Text Generation**: `GenerateText` and `StreamText` - **Structured Data**: Type-safe object generation - **Tool Calling**: Integrate external APIs and functions - **Embeddings**: Semantic search and RAG - **Image Generation**: Create images with AI - **Speech & Transcription**: Audio capabilities - **Agent Framework**: Build autonomous agents - **Middleware**: Default settings/instructions, JSON/reasoning extraction, simulated streaming - **Testing Utilities**: Mock providers and responses ### Built for Go's Strengths #### Context-Based Cancellation Go's `context.Context` provides clean cancellation patterns: ```go ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Long generation...", }) // Automatically stops after 30 seconds or if cancelled ``` #### Natural Streaming with Channels Unbuffered channels provide automatic backpressure: ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a story...", }) for chunk := range stream.Chunks() { if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } // Producer automatically slows down to match consumer } ``` #### Concurrent Processing Process multiple requests efficiently: ```go var wg sync.WaitGroup results := make(chan string, len(prompts)) for _, prompt := range prompts { wg.Add(1) go func(p string) { defer wg.Done() result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: p, }) results <- result.Text }(prompt) } wg.Wait() close(results) ``` ## Key Features ### Type-Safe Structured Output Generate structured data with compile-time type safety: ```go type ProductReview struct { Rating int `json:"rating"` Pros []string `json:"pros"` Cons []string `json:"cons"` Summary string `json:"summary"` } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Analyze this product review...", Output: ai.ObjectOutput[ProductReview](ai.ObjectOutputOptions{ Schema: schema.NewSimpleStructSchema(reflect.TypeOf(ProductReview{})), }), }) // result.Output is fully typed! review := result.Output.(ProductReview) fmt.Printf("Rating: %d/5\n", review.Rating) ``` ### Powerful Agent Framework Build autonomous agents with full control: ```go myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a customer support agent.", Tools: []types.Tool{ searchDocsTool, getUserInfoTool, createTicketTool, }, MaxSteps: 10, }) result, err := myAgent.Execute(ctx, "Help me reset my password") ``` ### Production-Ready Features #### Middleware Support Wrap models with default settings, instructions, and other cross-cutting behavior: ```go wrappedModel := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: floatPtr(0.7), }), }, nil, nil, ) ``` #### Error Handling Comprehensive error types for robust applications: ```go result, err := ai.GenerateText(ctx, options) if err != nil { var rateLimitErr *providererrors.RateLimitError var validationErr *providererrors.ValidationError switch { case errors.As(err, &rateLimitErr): // Implement backoff case errors.As(err, &validationErr): // Handle invalid inputs default: // Handle other errors } } ``` #### Testing Utilities Mock providers for unit tests: ```go mockModel := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { return &types.GenerateResult{ Text: "Hello, I'm a mock response!", FinishReason: types.FinishReasonStop, }, nil }, } result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: mockModel, Prompt: "Test prompt", }) ``` ## Supported Providers The Go AI SDK supports 49 providers out of the box: ### Major Providers - **OpenAI** - GPT-6, GPT-5, o-series - **Anthropic** - Claude Opus, Sonnet, Haiku and Fable - **Google** - Gemini Pro, Gemini Flash, Gemini Ultra - **Mistral** - Mistral Large, Medium, Small, Codestral - **Cohere** - Command R+, Command R, Command Light ### Cloud Platforms - **Amazon Bedrock** - Access multiple models through AWS - **Azure OpenAI** - Enterprise OpenAI integration - **Google Vertex AI** - Google Cloud AI Platform ### Specialized Providers - **Groq** - Ultra-fast inference - **Together AI** - Open source models - **Fireworks AI** - High-performance serving - **Replicate** - Easy model deployment - **Hugging Face** - Access thousands of models ### And Many More - xAI Grok, Perplexity, Cerebras, Deepseek, Fal, Friendli, LlamaCpp, Novita, and more! ## Use Cases The Go AI SDK is perfect for: - **Backend APIs**: Build AI-powered REST/GraphQL APIs - **Microservices**: AI capabilities in your service mesh - **CLI Tools**: Interactive command-line AI assistants - **Batch Processing**: Process large datasets with AI - **Real-time Streaming**: WebSocket/SSE AI responses - **Kubernetes Workloads**: Cloud-native AI services - **Edge Computing**: Deploy to edge with minimal overhead ## Getting Started Install the SDK: ```bash go get github.com/digitallysavvy/go-ai ``` Create your first AI application: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Why is Go great for building AI applications?", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Next Steps - **[Quick Start Guide](https://goaisdk.com/docs/getting-started/golang.md)** - Build your first agent - **[Foundations](https://goaisdk.com/docs/foundations.md)** - Learn core concepts - **[AI SDK Core](https://goaisdk.com/docs/ai-sdk-core.md)** - Explore the complete API - **[Agents](https://goaisdk.com/docs/agents.md)** - Build autonomous agents - **[Advanced Topics](https://goaisdk.com/docs/advanced.md)** - Production patterns ## Community & Contribution The Go AI SDK is open source and welcomes contributions: - **GitHub**: [github.com/digitallysavvy/go-ai](https://github.com/digitallysavvy/go-ai) - **Issues**: Report bugs or request features - **Discussions**: Share ideas and get help - **Pull Requests**: Contribute code, docs, or examples ## Philosophy The Go AI SDK is built on these principles: 1. **Go Idioms First**: Use Go's natural patterns (contexts, channels, interfaces) 2. **Type Safety**: Leverage compile-time checking wherever possible 3. **Production Ready**: Built for real-world deployments 4. **Provider Agnostic**: Switch providers without changing code 5. **Performance**: Minimize overhead and maximize throughput 6. **Developer Experience**: Clear APIs and comprehensive documentation ## Comparison with Other SDKs | Feature | Go AI SDK | Python SDKs | TypeScript AI SDK | |---------|-----------|-------------|-------------------| | **Type Safety** | Compile-time | Runtime | TypeScript compile-time | | **Concurrency** | Native goroutines | asyncio/threads | async/await | | **Deployment** | Single binary | Python + deps | Node.js + deps | | **Performance** | Very fast | Fast | Fast | | **Use Case** | Backend services | Scripts, research | Full-stack apps | | **Memory** | Low | Medium-High | Medium | ## The Future We're committed to keeping the Go AI SDK up-to-date with: - New provider integrations - Latest AI capabilities - Performance optimizations - Community-requested features - Comprehensive documentation ## Get Started Today ```bash go get github.com/digitallysavvy/go-ai ``` Join us in bringing the power of AI to the Go ecosystem! --- # Overview > An overview of foundational concepts critical to understanding the Go AI SDK Canonical URL: https://goaisdk.com/docs/foundations/overview Documentation index: https://goaisdk.com/llms.txt > **Note:** This page is a beginner-friendly introduction to high-level artificial intelligence (AI) concepts. To dive right into implementing the Go AI SDK, feel free to skip ahead to the [Generating Text guide](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) or learn about our [supported models and providers](https://goaisdk.com/docs/foundations/providers-and-models.md). The Go AI SDK standardizes integrating artificial intelligence (AI) models across [supported providers](https://goaisdk.com/docs/foundations/providers-and-models.md). This enables developers to focus on building great AI applications in Go, not waste time on technical details. For example, here's how you can generate text with various models using the Go AI SDK: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/providers/google" ) func main() { ctx := context.Background() // OpenAI openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) gpt, _ := openaiProvider.LanguageModel("gpt-6-astra") // Anthropic anthropicProvider := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) claude, _ := anthropicProvider.LanguageModel("claude-sonnet-5-5") // Google googleProvider := google.New(google.Config{APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY")}) gemini, _ := googleProvider.LanguageModel("gemini-2.5-flash") // Same API for all providers for _, model := range []provider.LanguageModel{gpt, claude, gemini} { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What is love?", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } } ``` To effectively leverage the Go AI SDK, it helps to familiarize yourself with the following concepts: ## Generative Artificial Intelligence **Generative artificial intelligence** refers to models that predict and generate various types of outputs (such as text, images, or audio) based on what's statistically likely, pulling from patterns they've learned from their training data. For example: - Given a photo, a generative model can generate a caption. - Given an audio file, a generative model can generate a transcription. - Given a text description, a generative model can generate an image. ## Large Language Models A **large language model (LLM)** is a subset of generative models focused primarily on **text**. An LLM takes a sequence of words as input and aims to predict the most likely sequence to follow. It assigns probabilities to potential next sequences and then selects one. The model continues to generate sequences until it meets a specified stopping criterion. LLMs learn by training on massive collections of written text, which means they will be better suited to some use cases than others. For example, a model trained on GitHub data would understand the probabilities of sequences in source code particularly well. However, it's crucial to understand LLMs' limitations. When asked about less known or absent information, like the birthday of a personal relative, LLMs might "hallucinate" or make up information. It's essential to consider how well-represented the information you need is in the model. ## Embedding Models An **embedding model** is used to convert complex data (like words or images) into a dense vector (a list of numbers) representation, known as an embedding. Unlike generative models, embedding models do not generate new text or data. Instead, they provide representations of semantic and syntactic relationships between entities that can be used as input for other models or other natural language processing tasks. ## Reranking Models A **reranking model** is used to reorder a list of documents based on their relevance to a query. Unlike embedding models that generate vector representations, reranking models directly score document-query pairs for relevance. This is particularly useful for search and retrieval-augmented generation (RAG) applications where you need to rank multiple candidates by their relevance to a specific query. ## Go-Specific Advantages The Go AI SDK brings the power of AI to Go developers with several advantages: - **Idiomatic Go**: Uses Go conventions like `context.Context`, channels, and error returns - **Type Safety**: Strong typing with Go's type system - **Performance**: Compiled language performance for production workloads - **Concurrency**: Built-in support for concurrent operations with goroutines - **Easy Deployment**: Single binary deployment without runtime dependencies - **Standard Library**: Integrates seamlessly with Go's standard library --- In the next section, you will learn about the difference between model providers and models, and which ones are available in the Go AI SDK. ## See Also - [Providers and Models](https://goaisdk.com/docs/foundations/providers-and-models.md) - [Prompts](https://goaisdk.com/docs/foundations/prompts.md) - [Tools](https://goaisdk.com/docs/foundations/tools.md) - [Streaming](https://goaisdk.com/docs/foundations/streaming.md) --- # Providers and Models > Explains the Go AI SDK's unified provider.Provider interface, listing supported providers, model capabilities, self-hosted models, and the provider registry. Canonical URL: https://goaisdk.com/docs/foundations/providers-and-models Documentation index: https://goaisdk.com/llms.txt Companies such as OpenAI and Anthropic (providers) offer access to a range of large language models (LLMs) with differing strengths and capabilities through their own APIs. Each provider typically has its own unique method for interfacing with their models, complicating the process of switching providers and increasing the risk of vendor lock-in. To solve these challenges, the Go AI SDK offers a standardized approach to interacting with LLMs through a unified provider interface that abstracts differences between providers. This unified interface allows you to switch between providers with ease while using the same API for all providers. ## Unified Provider Architecture The Go AI SDK uses a consistent provider architecture where all providers implement the same interfaces: ```go // All providers implement these interfaces type Provider interface { LanguageModel(modelID string) (LanguageModel, error) EmbeddingModel(modelID string) (EmbeddingModel, error) ImageModel(modelID string) (ImageModel, error) SpeechModel(modelID string) (SpeechModel, error) TranscriptionModel(modelID string) (TranscriptionModel, error) RerankingModel(modelID string) (RerankingModel, error) } ``` This means switching providers is as simple as changing which provider you instantiate: ```go // OpenAI openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model1, _ := openaiProvider.LanguageModel("gpt-6-astra") // Anthropic anthropicProvider := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) model2, _ := anthropicProvider.LanguageModel("claude-sonnet-5-5") // Same API for both result1, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model1, Prompt: "Hello"}) result2, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model2, Prompt: "Hello"}) ``` ## Supported Providers The Go AI SDK includes built-in support for 49 providers: ### Language Model Providers - **[OpenAI](https://goaisdk.com/docs/providers/openai.md)** - GPT-6, GPT-5, o-series ```go ``` - **[Anthropic](https://goaisdk.com/docs/providers/anthropic.md)** - Claude Opus, Sonnet, Haiku and Fable ```go ``` - **[Google](https://goaisdk.com/docs/providers/google.md)** - Gemini Pro, Gemini Flash ```go ``` - **[Azure OpenAI](https://goaisdk.com/docs/providers/azure.md)** - All Azure-hosted models ```go ``` - **[Amazon Bedrock](https://goaisdk.com/docs/providers/bedrock.md)** - Claude, Titan, Llama, and more ```go ``` - **[Cohere](https://goaisdk.com/docs/providers/cohere.md)** - Command, Command-R series ```go ``` - **[Mistral](https://goaisdk.com/docs/providers/mistral.md)** - Mistral Large, Medium, Small ```go ``` - **[Groq](https://goaisdk.com/docs/providers/groq.md)** - Fast inference with Llama, Mixtral ```go ``` - **[DeepSeek](https://goaisdk.com/docs/providers/deepseek.md)** - DeepSeek Chat, Coder ```go ``` - **[xAI](https://goaisdk.com/docs/providers/xai.md)** - Grok models ```go ``` - **[Perplexity](https://goaisdk.com/docs/providers/perplexity.md)** - Sonar models ```go ``` - **[Together AI](https://goaisdk.com/docs/providers/together.md)** - Llama, Mixtral, Qwen ```go ``` - **[Fireworks](https://goaisdk.com/docs/providers/fireworks.md)** - Fast inference for various models ```go ``` - **[Replicate](https://goaisdk.com/docs/providers/replicate.md)** - All hosted models ```go ``` - **[Hugging Face](https://goaisdk.com/docs/providers/huggingface.md)** - Router Responses API models ```go ``` - **[Ollama](https://goaisdk.com/docs/providers/ollama.md)** - Local models ```go ``` ### Image Generation Providers - **[Stability AI](https://goaisdk.com/docs/providers/stability.md)** - Stable Diffusion ```go ``` - **[Black Forest Labs](https://goaisdk.com/docs/providers/bfl.md)** - FLUX models ```go ``` - **[Fal.ai](https://goaisdk.com/docs/providers/fal.md)** - Fast image generation ```go ``` ### Speech & Transcription Providers - **[ElevenLabs](https://goaisdk.com/docs/providers/elevenlabs.md)** - Text-to-speech ```go ``` - **[Deepgram](https://goaisdk.com/docs/providers/deepgram.md)** - Speech-to-text ```go ``` - **[AssemblyAI](https://goaisdk.com/docs/providers/assemblyai.md)** - Speech recognition ```go ``` ### Infrastructure Providers - **[Baseten](https://goaisdk.com/docs/providers/baseten.md)** - Model serving infrastructure ```go ``` - **[Cerebras](https://goaisdk.com/docs/providers/cerebras.md)** - Fast inference hardware ```go ``` - **[DeepInfra](https://goaisdk.com/docs/providers/deepinfra.md)** - Multi-model inference ```go ``` - **[Vercel AI](https://goaisdk.com/docs/providers/vercel.md)** - Vercel's hosted models ```go ``` ## Self-Hosted Models You can run models locally with: - **Ollama** - Run Llama, Mistral, and other models locally - **LM Studio** (via OpenAI-compatible interface) - Any OpenAI-compatible server Example with Ollama: ```go import "github.com/digitallysavvy/go-ai/pkg/providers/ollama" provider := ollama.New(ollama.Config{ BaseURL: "http://localhost:11434", }) model, _ := provider.LanguageModel("llama2") result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello from my local machine!", }) ``` ## Model Capabilities Different models support different capabilities. You can check what a model supports: ```go model, _ := provider.LanguageModel("gpt-6-astra") // Check capabilities if model.SupportsTools() { fmt.Println("Model supports tool calling") } if model.SupportsStructuredOutput() { fmt.Println("Model supports structured output") } if model.SupportsImageInput() { fmt.Println("Model can process images") } ``` ### Capability Matrix Here are the capabilities of popular models: | Provider | Model | Image Input | Structured Output | Tool Calling | Tool Streaming | |----------|-------|-------------|-------------------|--------------|----------------| | OpenAI | gpt-6-astra | ✓ | ✓ | ✓ | ✓ | | OpenAI | gpt-6-sol | ✓ | ✓ | ✓ | ✓ | | OpenAI | gpt-5.4-nano | ✗ | ✓ | ✓ | ✓ | | Anthropic | claude-opus-4 | ✓ | ✓ | ✓ | ✓ | | Anthropic | claude-sonnet-4 | ✓ | ✓ | ✓ | ✓ | | Anthropic | claude-haiku-4-5 | ✓ | ✓ | ✓ | ✓ | | Google | gemini-2.0-flash | ✓ | ✓ | ✓ | ✓ | | Google | gemini-2.5-pro | ✓ | ✓ | ✓ | ✓ | | Mistral | pixtral-large | ✓ | ✓ | ✓ | ✓ | | Mistral | mistral-large | ✗ | ✓ | ✓ | ✓ | | DeepSeek | deepseek-chat | ✗ | ✓ | ✓ | ✓ | | Groq | llama-3.3-70b | ✗ | ✓ | ✓ | ✓ | > **Note:** This table is not exhaustive. Additional models and their capabilities can be found in the individual provider documentation pages. ## Using Multiple Providers One advantage of the unified interface is that you can easily use multiple providers in the same application: ```go package main import ( "context" "fmt" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() // Use GPT-6 for initial response openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) gpt, _ := openaiProvider.LanguageModel("gpt-6-astra") result1, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: gpt, Prompt: "Generate a creative story idea", }) // Use Claude for critique anthropicProvider := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) claude, _ := anthropicProvider.LanguageModel("claude-sonnet-5-5") result2, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: claude, Prompt: fmt.Sprintf("Critique this story idea: %s", result1.Text), }) fmt.Println("Story Idea:", result1.Text) fmt.Println("Critique:", result2.Text) } ``` ## Provider Registry For applications that need dynamic provider/model selection, use the registry: ```go import "github.com/digitallysavvy/go-ai/pkg/registry" // Create registry and register providers reg := registry.NewRegistry() reg.RegisterProvider("openai", openaiProvider) reg.RegisterProvider("anthropic", anthropicProvider) reg.RegisterProvider("google", googleProvider) // Resolve models by string ID model, _ := reg.ResolveLanguageModel("openai:gpt-6-astra") model2, _ := reg.ResolveLanguageModel("anthropic:claude-opus-5-5") // Use consistently result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) ``` See the [Provider Registry documentation](https://goaisdk.com/docs/ai-sdk-core/provider-management.md) for more details. ## Next Steps - Learn about [Prompts](https://goaisdk.com/docs/foundations/prompts.md) - Explore [Tools and Tool Calling](https://goaisdk.com/docs/foundations/tools.md) - Understand [Streaming](https://goaisdk.com/docs/foundations/streaming.md) - Check out individual [Provider Documentation](https://goaisdk.com/docs/providers/overview.md) ## See Also - [Core API Overview](https://goaisdk.com/docs/ai-sdk-core/overview.md) - [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - [Provider Management](https://goaisdk.com/docs/ai-sdk-core/provider-management.md) - [Registry API Reference](https://goaisdk.com/docs/reference/registry/provider-registry.md) --- # Prompts > Covers text, system, and message prompt formats in the Go AI SDK, including provider-specific options and multi-turn conversation context. Canonical URL: https://goaisdk.com/docs/foundations/prompts Documentation index: https://goaisdk.com/llms.txt Prompts are instructions that you give a [large language model (LLM)](https://goaisdk.com/docs/foundations/overview.md#large-language-models) to tell it what to do. It's like when you ask someone for directions; the clearer your question, the better the directions you'll get. Many LLM providers offer complex interfaces for specifying prompts. They involve different roles and message types. While these interfaces are powerful, they can be hard to use and understand. In order to simplify prompting, the Go AI SDK supports text, message, and system prompts. ## Text Prompts Text prompts are strings. They are ideal for simple generation use cases, e.g. repeatedly generating content for variants of the same prompt text. You can set text prompts using the `Prompt` field available in functions like `ai.StreamText()` or `ai.GenerateObject()`. You can structure the text in any way and inject variables using Go's string formatting. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Invent a new holiday and describe its traditions.", }) ``` You can also use string formatting to provide dynamic data to your prompt: ```go destination := "Paris" lengthOfStay := 5 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("I am planning a trip to %s for %d days. "+ "Please suggest the best tourist activities for me to do.", destination, lengthOfStay), }) ``` ## System Prompts System prompts are the initial set of instructions given to models that help guide and constrain the models' behaviors and responses. You can set system prompts using the `System` field. System prompts work with both the `Prompt` and the `Messages` fields. ```go destination := "Paris" lengthOfStay := 5 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You help planning travel itineraries. " + "Respond to the users' request with a list " + "of the best stops to make in their destination.", Prompt: fmt.Sprintf("I am planning a trip to %s for %d days. "+ "Please suggest the best tourist activities for me to do.", destination, lengthOfStay), }) ``` > **Note:** When you use a message prompt, you can also use system messages instead of a system prompt. ## Message Prompts A message prompt is a slice of user, assistant, and tool messages. They are great for chat interfaces and more complex, multi-modal prompts. You can use the `Messages` field to set message prompts. Each message has a `Role` and a `Content` field. The content can either be text (for user and assistant messages), or a slice of relevant parts (data) for that message type. ```go import "github.com/digitallysavvy/go-ai/pkg/provider/types" result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ {Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "Hi!"}}}, {Role: types.RoleAssistant, Content: []types.ContentPart{types.TextContent{Text: "Hello, how can I help?"}}}, {Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "Where can I buy the best Currywurst in Berlin?"}}}, }, }) ``` Instead of sending text in the `Content` field, you can send a slice of parts that includes a mix of text and other content parts. > **Warning:** Not all language models support all message and content types. For example, some models might not be capable of handling multi-modal inputs or tool messages. [Learn more about the capabilities of select models](https://goaisdk.com/docs/foundations/providers-and-models.md#model-capabilities). ### User Messages #### Text Parts Text content is the most common type of content. It is a string that is passed to the model. If you only need to send text content in a message, `Content` can be a single-element slice of `TextContent`, but it can also hold multiple content parts. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{ Text: "Where can I buy the best Currywurst in Berlin?", }, }, }, }, }) ``` #### Image Parts User messages can include image parts. An image can be one of the following: - Base64-encoded image: `string` with base-64 encoded content or data URL - Binary image: `[]byte` - URL: `string` with HTTP(S) URL ##### Example: Binary image ([]byte) ```go import "os" imageData, _ := os.ReadFile("./data/comic-cat.png") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Describe the image in detail."}, types.ImageContent{ Image: imageData, MimeType: "image/png", }, }, }, }, }) ``` ##### Example: Base-64 encoded image (string) ```go imageData, _ := os.ReadFile("./data/comic-cat.png") base64Image := base64.StdEncoding.EncodeToString(imageData) decodedImage, _ := base64.StdEncoding.DecodeString(base64Image) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Describe the image in detail."}, types.ImageContent{ Image: decodedImage, MimeType: "image/png", }, }, }, }, }) ``` ##### Example: Image URL (string) ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Describe the image in detail."}, types.ImageContent{ URL: "https://github.com/vercel/ai/blob/main/examples/ai-core/data/comic-cat.png?raw=true", }, }, }, }, }) ``` #### File Parts > **Warning:** Only a few providers and models currently support file parts: Google Generative AI, Google Vertex AI, OpenAI (for `wav` and `mp3` audio with `gpt-4o-audio-preview`), Anthropic, and OpenAI (for `pdf`). User messages can include file parts. A file can be one of the following: - Base64-encoded file: `string` with base-64 encoded content or data URL - Binary data: `[]byte` - URL: `string` with HTTP(S) URL You need to specify the MIME type of the file you are sending. ##### Example: PDF file from []byte ```go import ( "github.com/digitallysavvy/go-ai/pkg/providers/google" "github.com/digitallysavvy/go-ai/pkg/ai" ) pdfData, _ := os.ReadFile("./data/example.pdf") model, err := googleProvider.LanguageModel("gemini-2.5-flash") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What is the file about?"}, types.FileContent{ MediaType: "application/pdf", Data: pdfData, Filename: "example.pdf", // optional, not used by all providers }, }, }, }, }) ``` ##### Example: mp3 audio file from []byte ```go import ( "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/ai" ) audioData, _ := os.ReadFile("./data/galileo.mp3") model, err := openaiProvider.LanguageModel("gpt-4o-audio-preview") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What is the audio saying?"}, types.FileContent{ MediaType: "audio/mpeg", Data: audioData, }, }, }, }, }) ``` ### Assistant Messages Assistant messages are messages that have a role of `assistant`. They are typically previous responses from the assistant and can contain text, reasoning, and tool call parts. #### Example: Assistant message with text content ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ {Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "Hi!"}}}, {Role: types.RoleAssistant, Content: []types.ContentPart{types.TextContent{Text: "Hello, how can I help?"}}}, }, }) ``` #### Example: Assistant message with tool call content ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "How many calories are in this block of cheese?"}}, }, { Role: types.RoleAssistant, Content: []types.ContentPart{ types.ToolCallContent{ ToolCallID: "12345", ToolName: "get-nutrition-data", Arguments: map[string]interface{}{ "cheese": "Roquefort", }, }, }, }, }, }) ``` ### Tool Messages > **Note:** [Tools](https://goaisdk.com/docs/foundations/tools.md) (also known as function calling) are programs that you can provide an LLM to extend its built-in functionality. This can be anything from calling an external API to calling functions within your application. Learn more about Tools in [the next section](https://goaisdk.com/docs/foundations/tools.md). For models that support [tool](https://goaisdk.com/docs/foundations/tools.md) calls, assistant messages can contain tool call parts, and tool messages can contain tool output parts. A single assistant message can call multiple tools, and a single tool message can contain multiple tool results. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "How many calories are in this block of cheese?"}, types.ImageContent{Image: imageData, MimeType: "image/jpeg"}, }, }, { Role: types.RoleAssistant, Content: []types.ContentPart{ types.ToolCallContent{ ToolCallID: "12345", ToolName: "get-nutrition-data", Arguments: map[string]interface{}{ "cheese": "Roquefort", }, }, // there could be more tool calls here (parallel calling) }, }, { Role: types.RoleTool, Content: []types.ContentPart{ types.ToolResultContent{ ToolCallID: "12345", // needs to match the tool call id ToolName: "get-nutrition-data", Output: &types.ToolResultOutput{ Type: types.ToolResultOutputJSON, Value: map[string]interface{}{ "name": "Cheese, roquefort", "calories": 369, "fat": 31, "protein": 22, }, }, }, // there could be more tool results here (parallel calling) }, }, }, }) ``` ### System Messages System messages are messages that are sent to the model before the user messages to guide the assistant's behavior. You can alternatively use the `System` field. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleSystem, Content: []types.ContentPart{types.TextContent{Text: "You help planning travel itineraries."}}, }, { Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{ Text: "I am planning a trip to Berlin for 3 days. " + "Please suggest the best tourist activities for me to do.", }}, }, }, }) ``` ## Provider-Specific Options Some providers support additional options that can be passed at different levels: ### Function Call Level Provider options can be passed at the function call level using the `ProviderOptions` field: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: azureModel, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "reasoningEffort": "low", }, }, }) ``` ### Message Level For granular control, you can pass `ProviderOptions` to individual messages: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleSystem, Content: []types.ContentPart{types.TextContent{Text: "Cached system message"}}, ProviderOptions: map[string]interface{}{ "anthropic": map[string]interface{}{ "cacheControl": map[string]interface{}{ "type": "ephemeral", }, }, }, }, }, }) ``` ### Message Part Level Certain provider-specific options require configuration at the message part level: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.ImageContent{ URL: imageURL, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "imageDetail": "low", }, }, }, }, }, }, }) ``` ## Working with Context All AI SDK functions in Go accept a `context.Context` as the first parameter. This enables: - **Cancellation**: Cancel long-running operations - **Timeouts**: Set deadlines for operations - **Values**: Pass request-scoped values ```go // Timeout example ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a long story...", }) if err != nil { if ctx.Err() == context.DeadlineExceeded { fmt.Println("Request timed out") } } ``` ## Next Steps - Learn about [Tools](https://goaisdk.com/docs/foundations/tools.md) - Understand [Streaming](https://goaisdk.com/docs/foundations/streaming.md) - Explore [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) ## See Also - [Generating Structured Data](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) - [Tools and Tool Calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - [Message Types Reference](https://goaisdk.com/docs/reference/types/messages.md) --- # Tools > Introduces tools in the Go AI SDK: defining schemas, multi-step tool calling, tool choice, execution callbacks, and handling tool errors in Go. Canonical URL: https://goaisdk.com/docs/foundations/tools Documentation index: https://goaisdk.com/llms.txt While [large language models (LLMs)](https://goaisdk.com/docs/foundations/overview.md#large-language-models) have incredible generation capabilities, they struggle with discrete tasks (e.g. mathematics) and interacting with the outside world (e.g. getting the weather). Tools are actions that an LLM can invoke. The results of these actions can be reported back to the LLM to be considered in the next response. For example, when you ask an LLM for the "weather in London", and there is a weather tool available, it could call a tool with London as the argument. The tool would then fetch the weather data and return it to the LLM. The LLM can then use this information in its response. ## What is a tool? A tool is an object that can be called by the model to perform a specific task. You can use tools with [`ai.GenerateText()`](https://goaisdk.com/docs/reference/ai/generate-text.md) and [`ai.StreamText()`](https://goaisdk.com/docs/reference/ai/stream-text.md) by passing one or more tools to the `Tools` field. A tool in Go consists of several properties: ```go type Tool struct { Name string // Required: name of the tool Description string // Optional: description to guide when tool is picked Parameters interface{} // Required: JSON schema for tool input Execute func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) // Optional: execution function } ``` If the LLM decides to use a tool, it will generate a tool call. Tools with an `Execute` function are run automatically when these calls are generated. The output of the tool calls are returned using tool result objects. You can automatically pass tool results back to the LLM using [multi-step calls](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md#multi-step-calls) with `StreamText` and `GenerateText`. ## Basic Tool Example Here's a simple weather tool: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-6-astra") weatherTool := types.Tool{ Name: "get_weather", Description: "Get current weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "City name", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location := input["location"].(string) // In a real application, you would call a weather API here return map[string]interface{}{ "temperature": 72, "condition": "sunny", "location": location, }, nil }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in San Francisco?", Tools: []types.Tool{weatherTool}, // Allow a follow-up step so the model can turn the tool result into // a final answer (GenerateText/StreamText default to a single step). StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Tool Schemas Schemas define the parameters for tools and validate the [tool calls](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md). The Go AI SDK uses JSON Schema format for tool parameters. ### JSON Schema You can define schemas using Go's `map[string]interface{}`: ```go recipeSchema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "recipe": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{ "type": "string", }, "ingredients": map[string]interface{}{ "type": "array", "items": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{ "type": "string", }, "amount": map[string]interface{}{ "type": "string", }, }, }, }, "steps": map[string]interface{}{ "type": "array", "items": map[string]interface{}{ "type": "string", }, }, }, }, }, } ``` ### Using Go Structs For better type safety, you can generate JSON schemas from Go structs: ```go type RecipeInput struct { Recipe struct { Name string `json:"name"` Ingredients []struct { Name string `json:"name"` Amount string `json:"amount"` } `json:"ingredients"` Steps []string `json:"steps"` } `json:"recipe"` } // Use a JSON schema generation library or manually create the schema // from your struct definition ``` ## Multi-Step Tool Calling The Go AI SDK automatically handles multi-step tool execution. When a model makes tool calls, the SDK: 1. Executes the tools 2. Adds the results to the conversation 3. Sends back to the model for the next response ```go maxSteps := 5 // Allow up to 5 steps of tool calling result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in London and Paris?", Tools: []types.Tool{weatherTool}, MaxSteps: &maxSteps, }) ``` ## Tool Choice You can control when and which tools the model should use: ```go // Auto: Let the model decide when to use tools (default) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather?", Tools: []types.Tool{weatherTool}, ToolChoice: types.AutoToolChoice(), }) // None: Prevent the model from using any tools result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather?", Tools: []types.Tool{weatherTool}, ToolChoice: types.NoneToolChoice(), }) // Required: Force the model to use at least one tool result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather?", Tools: []types.Tool{weatherTool}, ToolChoice: types.RequiredToolChoice(), }) // Specific: Force the model to use a specific tool result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather?", Tools: []types.Tool{weatherTool, calculatorTool}, ToolChoice: types.SpecificToolChoice("get_weather"), }) ``` ## Complex Tool Example Here's a more complex example with multiple tools: ```go package main import ( "context" "fmt" "log" "math" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() provider := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) model, _ := provider.LanguageModel("claude-sonnet-5-5") // Calculator tool calculatorTool := types.Tool{ Name: "calculator", Description: "Perform mathematical calculations", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "operation": map[string]interface{}{ "type": "string", "enum": []string{"add", "subtract", "multiply", "divide", "power"}, }, "a": map[string]interface{}{"type": "number"}, "b": map[string]interface{}{"type": "number"}, }, "required": []string{"operation", "a", "b"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { op := input["operation"].(string) a := input["a"].(float64) b := input["b"].(float64) var result float64 switch op { case "add": result = a + b case "subtract": result = a - b case "multiply": result = a * b case "divide": if b == 0 { return nil, fmt.Errorf("division by zero") } result = a / b case "power": result = math.Pow(a, b) default: return nil, fmt.Errorf("unknown operation: %s", op) } return map[string]interface{}{"result": result}, nil }, } // Web search tool searchTool := types.Tool{ Name: "web_search", Description: "Search the web for information", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{ "type": "string", "description": "The search query", }, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := input["query"].(string) // In a real application, call a search API return map[string]interface{}{ "results": []string{ "Result 1 for: " + query, "Result 2 for: " + query, }, }, nil }, } maxSteps := 5 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Search for the population of Tokyo, then calculate what percentage it is of Japan's total population (125 million)", Tools: []types.Tool{calculatorTool, searchTool}, MaxSteps: &maxSteps, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Tool Execution Callbacks You can monitor tool execution with callbacks: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in Tokyo?", Tools: []types.Tool{weatherTool}, OnStepFinish: func(ctx context.Context, step types.StepResult, userContext interface{}) { fmt.Printf("Step %d finished\n", step.StepNumber) for _, toolCall := range step.ToolCalls { fmt.Printf(" Tool: %s, Args: %v\n", toolCall.ToolName, toolCall.Arguments) } for _, toolResult := range step.ToolResults { fmt.Printf(" Result: %v\n", toolResult.Result) } }, }) ``` ## Error Handling Tools should handle errors gracefully: ```go weatherTool := types.Tool{ Name: "get_weather", Description: "Get current weather for a location", Parameters: weatherSchema, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location, ok := input["location"].(string) if !ok { return nil, fmt.Errorf("location must be a string") } // Call weather API resp, err := http.Get(fmt.Sprintf("https://api.weather.com/v1/%s", location)) if err != nil { return nil, fmt.Errorf("failed to fetch weather: %w", err) } defer resp.Body.Close() if resp.StatusCode != http.StatusOK { return nil, fmt.Errorf("weather API returned status: %d", resp.StatusCode) } var data map[string]interface{} if err := json.NewDecoder(resp.Body).Decode(&data); err != nil { return nil, fmt.Errorf("failed to decode response: %w", err) } return data, nil }, } ``` ## Context Cancellation Tools respect context cancellation: ```go weatherTool := types.Tool{ Name: "get_weather", Parameters: weatherSchema, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Check if context is cancelled select { case <-ctx.Done(): return nil, ctx.Err() default: } // Create request with context req, _ := http.NewRequestWithContext(ctx, "GET", weatherURL, nil) resp, err := http.DefaultClient.Do(req) if err != nil { return nil, err } defer resp.Body.Close() // ... rest of implementation }, } ``` ## Tool Packages Since tools are just Go structs, they can be packaged and distributed as Go modules. This makes it easy to share reusable tools across projects. ### Creating a Tool Package ```go // weathertools/weather.go package weathertools import ( "context" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func GetWeatherTool(apiKey string) types.Tool { return types.Tool{ Name: "get_weather", Description: "Get current weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Implementation return nil, nil }, } } ``` ### Using a Tool Package ```go skip-compile import "github.com/yourusername/weathertools" weatherTool := weathertools.GetWeatherTool(apiKey) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather?", Tools: []types.Tool{weatherTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) ``` ## Best Practices 1. **Clear Names**: Use descriptive tool names that clearly indicate their purpose 2. **Good Descriptions**: Write clear descriptions to help the model know when to use the tool 3. **Schema Validation**: Define comprehensive parameter schemas with descriptions 4. **Error Handling**: Return clear error messages when tools fail 5. **Context Support**: Respect context cancellation in long-running tools 6. **Idempotency**: Make tools idempotent when possible 7. **Logging**: Log tool executions for debugging 8. **Rate Limiting**: Implement rate limiting for external API calls ## Next Steps - Learn about [Streaming](https://goaisdk.com/docs/foundations/streaming.md) - Explore [Tools and Tool Calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - Build [Agents](https://goaisdk.com/docs/agents/overview.md) ## See Also - [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - [Tool Types Reference](https://goaisdk.com/docs/reference/types/tools.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [Agent Overview](https://goaisdk.com/docs/agents/overview.md) --- # Streaming > Compares blocking and streaming interfaces for AI apps, then shows basic Go streaming, callbacks, accumulating text, and cancelling streams with StreamText. Canonical URL: https://goaisdk.com/docs/foundations/streaming Documentation index: https://goaisdk.com/llms.txt Streaming conversational text UIs (like ChatGPT) have gained massive popularity over the past few months. This section explores the benefits and drawbacks of streaming and blocking interfaces. [Large language models (LLMs)](https://goaisdk.com/docs/foundations/overview.md#large-language-models) are extremely powerful. However, when generating long outputs, they can be very slow compared to the latency you're likely used to. If you try to build a traditional blocking UI, your users might easily find themselves waiting 5, 10, even up to 40 seconds for the entire LLM response to be generated. This can lead to a poor user experience, especially in conversational applications like chatbots. Streaming UIs can help mitigate this issue by **displaying parts of the response as they become available**. ## Blocking vs Streaming ### Blocking Response With a blocking response, you wait until the full response is available before displaying it: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a detailed essay about artificial intelligence.", }) if err != nil { log.Fatal(err) } // User waits 10-40 seconds... fmt.Println(result.Text) // Finally displays all at once ``` ### Streaming Response With a streaming response, parts of the response are transmitted as they become available: ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a detailed essay about artificial intelligence.", }) if err != nil { log.Fatal(err) } defer result.Close() // Display text as it arrives for chunk := range result.Chunks() { fmt.Print(chunk.Text) // Displays immediately } ``` ## Real-world Impact The difference between blocking and streaming can be dramatic: - **Blocking**: User sees nothing for 30 seconds, then sees the complete response - **Streaming**: User starts seeing the response within 1-2 seconds, with continuous updates While streaming interfaces can greatly enhance user experiences, especially with larger language models, they aren't always necessary or beneficial. If you can achieve your desired functionality using a smaller, faster model without resorting to streaming, this route can often lead to simpler and more manageable development processes. However, regardless of the speed of your model, the Go AI SDK is designed to make implementing streaming as simple as possible. ## Basic Streaming in Go The Go AI SDK uses Go channels to provide a natural streaming interface: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a poem about Go programming.", }) if err != nil { log.Fatal(err) } defer result.Close() // Stream text chunks as they arrive for chunk := range result.Chunks() { fmt.Print(chunk.Text) } fmt.Println("\n\nFinish reason:", result.FinishReason()) } ``` ## Streaming with Callbacks For more control, you can use callbacks to process each chunk: ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Explain quantum computing.", OnChunk: func(chunk provider.StreamChunk) { fmt.Print(chunk.Text) // You can also process other chunk data if chunk.Usage != nil && chunk.Usage.TotalTokens != nil { fmt.Printf("[Tokens: %d]", *chunk.Usage.TotalTokens) } }, OnFinish: func(result *ai.StreamTextResult) { fmt.Printf("\n\nTotal tokens: %d\n", result.Usage().GetTotalTokens()) fmt.Printf("Finish reason: %s\n", result.FinishReason()) }, }) if err != nil { log.Fatal(err) } defer result.Close() // Still need to consume the channel to trigger callbacks for range result.Chunks() { } ``` ## Accumulating Streamed Text If you need the complete text after streaming: ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate a story.", }) if err != nil { log.Fatal(err) } defer result.Close() var fullText strings.Builder for chunk := range result.Chunks() { fmt.Print(chunk.Text) // Display to user fullText.WriteString(chunk.Text) // Accumulate } // Now you have the full text completeStory := fullText.String() // Or use the built-in ReadAll method // fullText, err := result.ReadAll() ``` ## Streaming with Tool Calls Streaming also works with tool calls. Tool calls are streamed as they're generated: ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "What's the weather in Tokyo and Paris?", Tools: []types.Tool{weatherTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, OnChunk: func(chunk provider.StreamChunk) { if chunk.Text != "" { fmt.Print(chunk.Text) } // Handle tool calls as they arrive if chunk.ToolCall != nil { fmt.Printf("\n[Calling tool: %s]\n", chunk.ToolCall.ToolName) } }, }) if err != nil { log.Fatal(err) } defer result.Close() for range result.Chunks() { } ``` ## Tool Execution Timing Since v0.4.0, tool execution in `StreamText` is deferred until the stream is fully consumed. This means: - All text, reasoning, and custom content chunks arrive first - Tool calls are accumulated during streaming - Tool execution happens after the stream ends - Tool-result chunks are forwarded to `OnChunk` after execution **Chunk ordering:** 1. Provider chunks — text (`ChunkTypeText`), reasoning (`ChunkTypeReasoning`), and custom content (`ChunkTypeCustom`, `ChunkTypeReasoningFile`) — arrive in whatever order the provider streams them and are forwarded to `OnChunk` immediately 2. Stream end 3. Tool execution (deferred) 4. Tool-result chunks (`ChunkTypeToolResult`) forwarded to `OnChunk` after execution This ensures consumers receive all content before any tool side-effects fire. If your application previously relied on tool results arriving mid-stream, see the [v0.3 to v0.4 migration guide](https://goaisdk.com/docs/migration-guides/from-v0.3-to-v0.4.md) for details. ## Cancellation Streaming respects context cancellation, allowing you to stop generation early: ```go ctx, cancel := context.WithCancel(context.Background()) go func() { // Cancel after 5 seconds time.Sleep(5 * time.Second) cancel() }() result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a very long story...", }) if err != nil { log.Fatal(err) } defer result.Close() for chunk := range result.Chunks() { select { case <-ctx.Done(): fmt.Println("\n\nCancelled!") return default: fmt.Print(chunk.Text) } } ``` ## Streaming Objects You can also stream structured data: ```go result, err := ai.StreamObject(ctx, ai.StreamObjectOptions{ Model: model, Prompt: "Generate a recipe for chocolate chip cookies", Schema: recipeSchema, OnChunk: func(partialObject interface{}) { fmt.Printf("Partial object: %v\n", partialObject) }, }) if err != nil { log.Fatal(err) } // StreamObject is deprecated — use StreamText with ObjectOutput[T]() instead. // The result is a *GenerateObjectResult with the complete object. fmt.Println(result.Object) ``` ## Channel-Based Architecture The Go AI SDK leverages Go's channel-based concurrency model for streaming: ```go // Channels close automatically when done result, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Hello", }) defer result.Close() // Safe to range over channel - will exit when complete for chunk := range result.Chunks() { fmt.Print(chunk.Text) } // Channel is closed, loop exits naturally ``` ## Error Handling in Streams Handle errors during streaming: ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate text", }) if err != nil { log.Fatal(err) } defer result.Close() for chunk := range result.Chunks() { if chunk.Type == provider.ChunkTypeError { fmt.Printf("Error during streaming: %s\n", chunk.AbortReason) break } fmt.Print(chunk.Text) } // Check for final error if err := result.Err(); err != nil { log.Printf("Stream ended with error: %v", err) } ``` ## Concurrent Streaming Stream from multiple models concurrently: ```go package main import ( "context" "fmt" "os" "sync" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func streamFromModel(ctx context.Context, model provider.LanguageModel, prompt string, wg *sync.WaitGroup) { defer wg.Done() result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: prompt, }) if err != nil { fmt.Printf("Error: %v\n", err) return } defer result.Close() for chunk := range result.Chunks() { fmt.Printf("[%s]: %s", model.ModelID(), chunk.Text) } } func main() { ctx := context.Background() var wg sync.WaitGroup openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) gptModel, _ := openaiProvider.LanguageModel("gpt-6-astra") anthropicProvider := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) claudeModel, _ := anthropicProvider.LanguageModel("claude-sonnet-5-5") // Stream from GPT and Claude simultaneously wg.Add(2) go streamFromModel(ctx, gptModel, "What is AI?", &wg) go streamFromModel(ctx, claudeModel, "What is AI?", &wg) wg.Wait() } ``` ## Backpressure Handling Go's channels naturally handle backpressure. If your consumer is slow, the producer will wait: ```go result, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate text", }) defer result.Close() for chunk := range result.Chunks() { fmt.Print(chunk.Text) // Slow consumer - producer will wait time.Sleep(100 * time.Millisecond) } ``` ## When to Use Streaming **Use streaming when:** - Generating long-form content (essays, stories, reports) - Building interactive chat applications - User experience is critical - You need to show progress - Working with slower models **Use blocking when:** - Generating short responses - Batch processing - You need the complete response before proceeding - Building APIs that return complete responses - Working with fast models ## Performance Considerations Streaming adds minimal overhead: ```go // Blocking: Simple, slightly lower overhead result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Short response", }) // Streaming: Slightly more overhead, better UX streamResult, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Short response", }) ``` For short responses (< 100 tokens), the overhead is negligible. For long responses (> 1000 tokens), streaming significantly improves perceived latency. ## Next Steps - Start with [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - Learn about [Advanced Streaming](https://goaisdk.com/docs/advanced/backpressure.md) - Explore [Stream Helpers](https://goaisdk.com/docs/reference/ai/stream-text.md) ## See Also - [Core API Overview](https://goaisdk.com/docs/ai-sdk-core/overview.md) - [Generating Structured Data](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) - [Backpressure Handling](https://goaisdk.com/docs/advanced/backpressure.md) --- # Provider options > Documents ProviderOptions in the Go AI SDK: namespaced provider-specific settings for OpenAI, Anthropic, Google/Vertex, and xAI, with type-safe combining. Canonical URL: https://goaisdk.com/docs/foundations/provider-options Documentation index: https://goaisdk.com/llms.txt `ProviderOptions` lets you pass provider-specific configuration that goes beyond standard settings like `Temperature` and `MaxTokens`. Options are namespaced by provider name, so you can include settings for multiple providers in the same call — only the options matching the active provider are used, and the rest are ignored. This is useful when you want to access features that are unique to a specific provider, such as exact reasoning token budgets, reasoning summaries, or server-side persistence controls. > **Note:** For reasoning effort, prefer the top-level `Reasoning` parameter for portability across providers. Use provider-specific options only when you need exact token budgets or provider-unique features. See [Reasoning](https://goaisdk.com/docs/ai-sdk-core/reasoning.md) for details. ## Common provider options Provider options are passed as a `map[string]interface{}` where the top-level keys are provider names (e.g., `"openai"`, `"xai"`, `"google"`). Each provider key maps to another `map[string]interface{}` containing that provider's specific settings. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum entanglement.", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "reasoningEffort": "low", }, }, }) ``` If you switch from an OpenAI model to an XAI model, the `"openai"` options are silently ignored. You can include options for multiple providers to support model switching without code changes: ```go ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "reasoningEffort": "low", }, "xai": map[string]interface{}{ "reasoningSummary": "concise", }, } ``` The sections below cover the most frequently used provider options. For a complete reference, see the individual provider pages: - [OpenAI provider](https://goaisdk.com/docs/providers/openai.md) - [Anthropic provider](https://goaisdk.com/docs/providers/anthropic.md) - [Google provider](https://goaisdk.com/docs/providers/google.md) - [XAI provider](https://goaisdk.com/docs/providers/xai.md) ## OpenAI ### Reasoning effort Controls how much internal reasoning a model performs before responding. Available on reasoning models like `o3`, `o4-mini`, and `gpt-5.2`. Lower values are faster and use fewer tokens. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "How many r's are in strawberry?", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "reasoningEffort": "low", }, }, }) // Access reasoning tokens from the result if openaiMeta, ok := result.ProviderMetadata["openai"].(map[string]interface{}); ok { fmt.Println("Reasoning tokens:", openaiMeta["reasoningTokens"]) } ``` | Value | Description | |-------|-------------| | `"none"` | No reasoning (GPT-5.1 models only) | | `"minimal"` | Bare-minimum reasoning | | `"low"` | Fast, concise reasoning | | `"medium"` | Balanced reasoning (default) | | `"high"` | Thorough reasoning | | `"xhigh"` | Maximum reasoning (GPT-5.1-Codex-Max only) | > **Note:** `"none"` and `"xhigh"` are only supported on specific models. Using them with unsupported models will result in an error. ### Reasoning summary Surfaces the model's internal reasoning process in the response. Useful for debugging, auditing, or showing the user how the model arrived at an answer. | Value | Description | |-------|-------------| | `"auto"` | Condensed summary of reasoning | | `"detailed"` | Comprehensive reasoning output | **Non-streaming example:** ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What is the capital of France?", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "reasoningSummary": "detailed", }, }, }) // Access reasoning from the result fmt.Println("Reasoning:", result.ReasoningText) ``` **Streaming example:** ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Solve this step by step: what is 127 * 43?", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "reasoningSummary": "detailed", }, }, }) ``` When streaming, reasoning summary chunks arrive as reasoning content before the main text response. ### Text verbosity Controls the length and detail of the model's text response independently of reasoning: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a poem about a boy and his first pet dog.", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "textVerbosity": "low", }, }, }) ``` | Value | Description | |-------|-------------| | `"low"` | Terse, minimal responses | | `"medium"` | Balanced detail (default) | | `"high"` | Verbose, comprehensive responses | ### Store Controls whether OpenAI persists the request and response on their servers. When set to `false`, the SDK automatically strips unencrypted reasoning content from conversation history to comply with OpenAI's requirements. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Analyze this confidential data...", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "store": false, }, }, }) ``` ## Anthropic Anthropic options are configured via `ModelOptions` when creating the language model, rather than through the per-call `ProviderOptions` map. This is because Anthropic features like thinking and speed mode are typically set once per model instance. ### Thinking (extended reasoning) Enables Anthropic's extended thinking, which gives the model a dedicated thinking phase before responding. You control the budget in tokens — higher budgets allow deeper reasoning but increase latency and cost. For Opus 4.6 and newer models, use adaptive thinking: ```go model, err := provider.LanguageModelWithOptions("claude-opus-4-6", &anthropic.ModelOptions{ Thinking: &anthropic.ThinkingConfig{ Type: anthropic.ThinkingTypeAdaptive, }, }) ``` For earlier models, use enabled thinking with an explicit budget: ```go budget := 12000 model, err := provider.LanguageModelWithOptions("claude-sonnet-4-6", &anthropic.ModelOptions{ Thinking: &anthropic.ThinkingConfig{ Type: anthropic.ThinkingTypeEnabled, BudgetTokens: &budget, }, }) ``` > **Note:** Thinking is supported on `claude-opus-4-6`, `claude-sonnet-4-6`, and similar models. You can also use the top-level `Reasoning` parameter for a portable alternative that works across providers. ### Effort The `Effort` option provides a simpler way to control reasoning depth without specifying a token budget. It affects thinking, text responses, and function calls. ```go model, err := provider.LanguageModelWithOptions("claude-opus-4-6", &anthropic.ModelOptions{ Effort: anthropic.EffortLow, }) ``` | Value | Description | |-------|-------------| | `EffortLow` | Minimal reasoning, fastest responses | | `EffortMedium` | Balanced reasoning | | `EffortHigh` | Thorough reasoning (default) | | `EffortMax` | Maximum reasoning effort | ### Fast mode (speed) For `claude-opus-4-6`, the `Speed` option enables approximately 2.5x faster output token speeds: ```go model, err := provider.LanguageModelWithOptions("claude-opus-4-6", &anthropic.ModelOptions{ Speed: anthropic.SpeedFast, }) ``` | Value | Description | |-------|-------------| | `SpeedFast` | ~2.5x faster output token speeds | | `SpeedStandard` | Standard inference speed (default) | ### Send reasoning Controls whether reasoning content from previous turns is sent back to the model in multi-turn conversations. Set to `false` to reduce token usage when the model's reasoning history is not needed. ```go disabled := false model, err := provider.LanguageModelWithOptions("claude-sonnet-4-6", &anthropic.ModelOptions{ SendReasoning: &disabled, }) ``` ## Google / Vertex ### Thinking config Controls the thinking budget for Google and Vertex AI models that support reasoning (e.g., Gemini 2.5 Flash). The `thinkingBudget` is specified in tokens. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a proof that there are infinitely many primes.", ProviderOptions: map[string]interface{}{ "google": map[string]interface{}{ "thinkingConfig": map[string]interface{}{ "thinkingBudget": 10000, }, }, }, }) ``` For Vertex AI models, use the `"googleVertex"` key instead of `"google"`. Vertex checks provider options keys in this order: `"googleVertex"` (preferred), then `"vertex"` (legacy, still honored), then `"google"` (cross-namespace fallback) — never the hyphenated `"google-vertex"`, which is the provider's identity string, not a provider options key. ```go ProviderOptions: map[string]interface{}{ "googleVertex": map[string]interface{}{ "thinkingConfig": map[string]interface{}{ "thinkingBudget": 10000, }, }, } ``` ## XAI ### Reasoning summary XAI (Grok) models support reasoning summaries with three levels of detail. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What are the key differences between TCP and UDP?", ProviderOptions: map[string]interface{}{ "xai": map[string]interface{}{ "reasoningSummary": "concise", }, }, }) ``` | Value | Description | |-------|-------------| | `"auto"` | Let the model decide the summary level | | `"concise"` | Brief reasoning summary | | `"detailed"` | Comprehensive reasoning output | ## Combining options You can combine multiple provider options in a single call. For example, using both reasoning effort and reasoning summaries with OpenAI: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What are the implications of quantum computing for cryptography?", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "reasoningEffort": "high", "reasoningSummary": "detailed", }, }, }) ``` Or combining thinking with a specific effort level for Anthropic: ```go budget := 8000 model, err := provider.LanguageModelWithOptions("claude-opus-4-6", &anthropic.ModelOptions{ Thinking: &anthropic.ThinkingConfig{ Type: anthropic.ThinkingTypeEnabled, BudgetTokens: &budget, }, Effort: anthropic.EffortMedium, }) ``` You can also combine the top-level `Reasoning` parameter with provider-specific options. The top-level parameter sets portable reasoning effort, while provider options handle provider-unique features: ```go reasoning := types.ReasoningLow result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain the Riemann hypothesis.", Reasoning: &reasoning, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "reasoningSummary": "detailed", }, }, }) ``` When both `Reasoning` and a provider-specific reasoning effort are set, the provider-specific setting takes precedence. For Anthropic, a call-level `Reasoning` value overrides the model-level `Thinking` configuration. ## Type safety Anthropic options use typed Go structs for compile-time safety: ```go import "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" // Typed structs — typos caught at compile time opts := &anthropic.ModelOptions{ Thinking: &anthropic.ThinkingConfig{Type: anthropic.ThinkingTypeAdaptive}, Effort: anthropic.EffortHigh, Speed: anthropic.SpeedFast, } ``` OpenAI, Google, XAI, and other providers use `map[string]interface{}` with documented string keys. Refer to the individual provider pages for the full list of supported keys: - [OpenAI provider](https://goaisdk.com/docs/providers/openai.md) - [Anthropic provider](https://goaisdk.com/docs/providers/anthropic.md) ## See also - [Providers and models](https://goaisdk.com/docs/foundations/providers-and-models.md) — overview of available providers - [Streaming](https://goaisdk.com/docs/foundations/streaming.md) — how streaming works with provider options - [OpenAI provider](https://goaisdk.com/docs/providers/openai.md) - [Anthropic provider](https://goaisdk.com/docs/providers/anthropic.md) - [Google provider](https://goaisdk.com/docs/providers/google.md) - [Google Vertex provider](https://goaisdk.com/docs/providers/google-vertex.md) - [XAI provider](https://goaisdk.com/docs/providers/xai.md) --- # Foundations > A section that covers foundational knowledge around LLMs and concepts crucial to the Go AI SDK Canonical URL: https://goaisdk.com/docs/foundations/ Documentation index: https://goaisdk.com/llms.txt This section covers the foundational concepts you need to understand when working with Large Language Models and the Go AI SDK. ## Topics - **[Overview](https://goaisdk.com/docs/foundations/overview.md)** - Learn about foundational concepts around AI and LLMs. - **[Providers and Models](https://goaisdk.com/docs/foundations/providers-and-models.md)** - Learn about the providers and models that you can use with the Go AI SDK. - **[Prompts](https://goaisdk.com/docs/foundations/prompts.md)** - Learn about how prompts are used and defined in the Go AI SDK. - **[Tools](https://goaisdk.com/docs/foundations/tools.md)** - Learn about tools in the Go AI SDK and how to integrate external functionality. - **[Streaming](https://goaisdk.com/docs/foundations/streaming.md)** - Learn why streaming is used for AI applications and how it works in Go. - **[Provider Options](https://goaisdk.com/docs/foundations/provider-options.md)** - Learn how to pass provider-specific configuration beyond standard settings. ## Next Steps After understanding these foundational concepts, explore: - [AI SDK Core](https://goaisdk.com/docs/ai-sdk-core.md) for detailed API documentation - [Agents](https://goaisdk.com/docs/agents.md) to build autonomous AI agents - [Advanced Topics](https://goaisdk.com/docs/advanced.md) for production-ready patterns --- # Navigating the Library > Learn how to navigate the Go AI SDK and choose the right tools for your needs. Canonical URL: https://goaisdk.com/docs/getting-started/navigating-the-library Documentation index: https://goaisdk.com/llms.txt The Go AI SDK is a powerful toolkit for building AI applications in Go. This page will help you pick the right tools for your requirements. ## SDK Structure The Go AI SDK is organized into several key packages: ### Core Packages - **`pkg/ai`** - Main package with high-level functions (`GenerateText`, `StreamText`, `GenerateObject`, `Embed`, `GenerateImage`, `GenerateSpeech`, `Transcribe`, `Rerank`, etc.) - **`pkg/providers`** - All 49 provider implementations (OpenAI, Anthropic, Google, etc.) - **`pkg/agent`** - Agent framework for building autonomous workflows (`ToolLoopAgent`) - **`pkg/middleware`** - Middleware for wrapping models with default settings, instructions, JSON/reasoning extraction, and simulated streaming - **`pkg/testutil`** - Testing utilities and mock providers ### Optional Packages - **`pkg/schema`** - JSON Schema utilities for structured output - **`pkg/telemetry`** - OpenTelemetry integration ## Choosing the Right Tool When deciding which part of the SDK to use, consider your use case: | Package | Purpose | When to Use | |---------|---------|-------------| | **`ai.GenerateText`** | Generate text completions | Single-shot text generation, no streaming needed | | **`ai.StreamText`** | Stream text responses | Real-time responses, user interfaces, long generations | | **`ai.GenerateObject`** | Generate structured data | Type-safe JSON output, data extraction | | **`ai.StreamObject`** | Stream structured data | Real-time structured data with progress updates | | **`agent`** | Build autonomous agents | Multi-step workflows, tool calling, decision making | | **`middleware`** | Wrap provider calls | Default settings/instructions, JSON extraction, observability | ## Use Case Guide ### Simple Text Generation For straightforward text generation without streaming: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing.", }) ``` **Use when:** - You don't need real-time updates - The response is short - You're building batch processes ### Streaming Responses For real-time streaming to users: ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long essay...", }) for chunk := range stream.Chunks() { if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } } ``` **Use when:** - Building user-facing applications - Responses are long - You want to show progress - Implementing SSE or WebSocket endpoints ### Structured Data Extraction For extracting structured information: ```go type EmailSummary struct { Sender string `json:"sender"` Subject string `json:"subject"` KeyPoints []string `json:"key_points"` Priority string `json:"priority"` } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Summarize this email: ...", Output: ai.ObjectOutput[EmailSummary](ai.ObjectOutputOptions{ Schema: schema.NewSimpleStructSchema(reflect.TypeOf(EmailSummary{})), }), }) summary := result.Output.(EmailSummary) fmt.Printf("Priority: %s\n", summary.Priority) ``` **Use when:** - You need structured JSON output - Type safety is important - Building data pipelines - Extracting information from text ### Building Agents For complex, multi-step workflows: ```go myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a customer support agent.", Tools: []types.Tool{ searchKnowledgeBase, createSupportTicket, checkOrderStatus, }, MaxSteps: 10, }) result, err := myAgent.Execute(ctx, "Help me track my order #12345") ``` **Use when:** - You need multi-step reasoning - Agents should use external tools - Building autonomous workflows - Implementing complex business logic ## Environment Compatibility The Go AI SDK works in any Go environment: | Environment | AI SDK Core | Agents | Middleware | All Features | |-------------|-------------|--------|------------|--------------| | **HTTP Server** (net/http) | ✓ | ✓ | ✓ | ✓ | | **Gin** | ✓ | ✓ | ✓ | ✓ | | **Echo** | ✓ | ✓ | ✓ | ✓ | | **Fiber** | ✓ | ✓ | ✓ | ✓ | | **gRPC** | ✓ | ✓ | ✓ | ✓ | | **CLI Applications** | ✓ | ✓ | ✓ | ✓ | | **Background Jobs** | ✓ | ✓ | ✓ | ✓ | | **Serverless** (Lambda, Cloud Functions) | ✓ | ✓ | ✓ | ✓ | ## Provider Selection Choose the right provider for your needs: ### Development For quick prototyping and development: - **OpenAI** - Easy to use, well-documented - **Anthropic** - Excellent for complex reasoning - **Groq** - Ultra-fast for testing ### Production For production deployments: - **Azure OpenAI** - Enterprise SLAs and compliance - **Amazon Bedrock** - AWS integration and security - **Google Vertex AI** - GCP integration ### Cost Optimization For cost-conscious applications: - **Groq** - Fast and affordable - **Together AI** - Competitive pricing - **Mistral** - Open models, good value ### Specialized Use Cases - **Long Context**: Anthropic Claude (200K+ tokens) - **Code Generation**: OpenAI GPT-6, Codestral - **Fast Inference**: Groq, Fireworks - **Open Source**: Together AI, Replicate, Hugging Face ## Package Import Guide ### Minimum Imports For basic text generation: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) ``` ### With Structured Output Add schema support: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) ``` ### With Agents Include the agent framework: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) ``` ### With Middleware Add observability: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) ``` ## Feature Matrix | Feature | GenerateText | StreamText | GenerateObject | StreamObject | Agents | |---------|--------------|------------|----------------|--------------|--------| | **Text Output** | ✓ | ✓ | - | - | ✓ | | **Streaming** | - | ✓ | - | ✓ | ✓ | | **Structured Output** | - | - | ✓ | ✓ | ✓ | | **Tool Calling** | ✓ | ✓ | - | - | ✓ | | **Multi-Step** | Manual | Manual | - | - | Automatic | | **Conversation Memory** | Manual | Manual | - | - | Automatic | ## Decision Tree **Need real-time updates?** - Yes → Use `StreamText` or `StreamObject` - No → Continue **Need structured data?** - Yes → Use `GenerateObject` or `StreamObject` - No → Continue **Need multi-step workflows?** - Yes → Use the `agent` package - No → Use `GenerateText` ## Best Practices ### 1. Always Use Context ```go // Good - with timeout ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() result, err := ai.GenerateText(ctx, options) ``` ### 2. Handle Errors Properly ```go import providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" result, err := ai.GenerateText(ctx, options) if err != nil { var rateLimitErr *providererrors.RateLimitError switch { case errors.As(err, &rateLimitErr): // Wait and retry, respecting rateLimitErr.RetryAfterSeconds if set default: return err } } ``` ### 3. Use Appropriate Concurrency ```go // Process multiple prompts concurrently var wg sync.WaitGroup semaphore := make(chan struct{}, 5) // Limit to 5 concurrent requests for _, prompt := range prompts { wg.Add(1) go func(p string) { defer wg.Done() semaphore <- struct{}{} defer func() { <-semaphore }() result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: p, }) // Process result }(prompt) } wg.Wait() ``` ### 4. Choose the Right Model ```go // Fast, cheap model for simple tasks quickModel, _ := provider.LanguageModel("gpt-5.4-nano") // Powerful model for complex reasoning smartModel, _ := provider.LanguageModel("gpt-6-astra") // Use based on task complexity model := quickModel if isComplexTask { model = smartModel } ``` ## Next Steps - **[Go Quick Start](https://goaisdk.com/docs/getting-started/golang.md)** - Build your first AI application - **[Foundations](https://goaisdk.com/docs/foundations.md)** - Learn core concepts - **[AI SDK Core](https://goaisdk.com/docs/ai-sdk-core.md)** - Explore the complete API - **[Agents](https://goaisdk.com/docs/agents.md)** - Build autonomous agents --- # Go Quick Start > Step-by-step quick start for building a first AI agent in Go, covering setup, provider choice, adding tools, and enabling multi-step tool calls. Canonical URL: https://goaisdk.com/docs/getting-started/golang Documentation index: https://goaisdk.com/llms.txt In this quickstart, you'll make a first text generation call, stream the response, give the model a tool, and wrap it all in a chat loop. Each step is a complete program you can run. ## Prerequisites - **Go 1.26+** installed on your local development machine - An **API key** from a supported provider (OpenAI, Anthropic, Google, etc.) If you haven't obtained your API key, sign up at your chosen provider's website. ## Set up your application Create a new directory and initialize a Go module: ```bash mkdir my-ai-app cd my-ai-app go mod init my-ai-app go get github.com/digitallysavvy/go-ai ``` > **Note:** The Go AI SDK is a unified interface to large language models. You can change > model and provider with a couple of lines of code. Learn more about > [available providers](https://goaisdk.com/docs/providers.md). ### Set your API key The OpenAI provider reads `OPENAI_API_KEY` from the environment. Export it in the shell you will run the program from: ```bash export OPENAI_API_KEY=sk-... ``` Other providers read their own variable, for example `ANTHROPIC_API_KEY` or `GOOGLE_GENERATIVE_AI_API_KEY`. If you prefer a `.env` file, Go does not load it on its own. Put `OPENAI_API_KEY=sk-...` in `.env`, install the loader with `go get github.com/joho/godotenv`, and add a blank import of its `autoload` package to `main.go`: ```go skip-compile import _ "github.com/joho/godotenv/autoload" ``` The loader reads `.env` from the working directory when the program starts. Don't commit the file. ## Generate text Create `main.go`: ```go filename="main.go" package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { model, err := openai.New(openai.Config{}).LanguageModel(openai.ModelGPT6Astra) if err != nil { log.Fatal(err) } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Explain what a goroutine is in one sentence.", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` Run it: ```bash go run main.go ``` You should see one sentence about goroutines printed to your terminal. `openai.ModelGPT6Astra` is a constant for the model ID `gpt-6-astra`. Every provider package has constants like it in `model_ids.go`. ## Stream the response `ai.GenerateText` waits for the whole answer. For a chat interface you want text as it arrives. Replace `main.go` with: ```go filename="main.go" package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { model, err := openai.New(openai.Config{}).LanguageModel(openai.ModelGPT6Astra) if err != nil { log.Fatal(err) } stream, err := ai.StreamText(context.Background(), ai.StreamTextOptions{ Model: model, Prompt: "Write a haiku about concurrency.", }) if err != nil { log.Fatal(err) } for chunk := range stream.Chunks() { if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } } fmt.Println() if err := stream.Err(); err != nil { log.Fatal(err) } } ``` `stream.Chunks()` returns a channel of chunks. Text chunks carry the next piece of the answer. After the channel closes, `stream.Err()` reports any error that ended the stream early. ## Switch providers To change providers, change the provider and model lines. The rest of your code stays the same. ### Anthropic ```go import "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" p := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, err := p.LanguageModel(anthropic.ClaudeSonnet5_5) ``` ### Google ```go import "github.com/digitallysavvy/go-ai/pkg/providers/google" p := google.New(google.Config{ APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"), }) model, err := p.LanguageModel(google.ModelGemini31FlashLitePreview) ``` ## Add a tool Language models are poor at exact tasks like arithmetic and cannot see the outside world. [Tools](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) are Go functions the model can ask you to run. The SDK runs the function and sends the result back to the model. This program gives the model a weather tool. `StopWhen` lets the model take several steps: call the tool, read the result, then answer. ```go filename="main.go" package main import ( "context" "fmt" "log" "math/rand" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { model, err := openai.New(openai.Config{}).LanguageModel(openai.ModelGPT6Astra) if err != nil { log.Fatal(err) } weather := types.Tool{ Name: "getWeather", Description: "Get the weather in a location (fahrenheit)", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "The location to get the weather for", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, params map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location, _ := params["location"].(string) return map[string]interface{}{ "location": location, "temperature": rand.Intn(59) + 32, // random 32-90°F }, nil }, } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in New York?", Tools: []types.Tool{weather}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` The tool returns a random temperature, so the answer changes on each run. Replace the body of `Execute` with a call to a real weather API when you are ready. The pieces of a tool: 1. **`Description`** helps the model decide when to use the tool. 2. **`Parameters`** is a JSON Schema describing the input. Here it requires a `location` string. 3. **`Execute`** runs your code and returns any JSON-serializable value. ## Build a chat loop Now put the pieces together: a terminal chat that streams answers, keeps history, and can call the tool. ```go filename="main.go" package main import ( "bufio" "context" "fmt" "log" "math/rand" "os" "strings" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() model, err := openai.New(openai.Config{}).LanguageModel(openai.ModelGPT6Astra) if err != nil { log.Fatal(err) } tools := []types.Tool{ { Name: "getWeather", Description: "Get the weather in a location (fahrenheit)", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "The location to get the weather for", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, params map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location, _ := params["location"].(string) return map[string]interface{}{ "location": location, "temperature": rand.Intn(59) + 32, }, nil }, }, } messages := []types.Message{} reader := bufio.NewReader(os.Stdin) fmt.Println("AI chat (type 'exit' to quit)") for { fmt.Print("You: ") userInput, err := reader.ReadString('\n') userInput = strings.TrimSpace(userInput) if err != nil || userInput == "exit" { break } if userInput == "" { continue } messages = append(messages, types.Message{ Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: userInput}}, }) stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Messages: messages, Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Printf("Error: %v\n", err) messages = messages[:len(messages)-1] continue } fmt.Print("\nAssistant: ") var answer strings.Builder for chunk := range stream.Chunks() { if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) answer.WriteString(chunk.Text) } } fmt.Print("\n\n") if err := stream.Err(); err != nil { log.Printf("Stream error: %v\n", err) messages = messages[:len(messages)-1] continue } messages = append(messages, types.Message{ Role: types.RoleAssistant, Content: []types.ContentPart{types.TextContent{Text: answer.String()}}, }) } } ``` Run it with `go run main.go` and ask "What's the weather in New York in celsius?". The model calls `getWeather`, reads the result, converts it, and streams the answer. What the loop does: 1. **Keeps history** in the `messages` slice. Each turn appends the user message, then the assistant's answer. 2. **Streams** each answer with `ai.StreamText` and prints text chunks as they arrive. 3. **Allows several steps** with `StopWhen`, so tool calls resolve before the final answer. 4. **Rolls back** the last user message if a call fails, so history stays consistent. To add more tools, append them to the `tools` slice. A model can chain tools, for example one that fetches a temperature and one that converts it. ## Where to Next? You've built an AI agent using the Go AI SDK! From here, you have several paths to explore: - **[Build a chat app](https://goaisdk.com/docs/build-a-chat-app.md)** - Serve a `useChat` frontend from Go, add tool approval and run coding agents, based on the [Shipyard demo](https://github.com/digitallysavvy/go-ai-demo) - **[Recipes](https://goaisdk.com/docs/recipes.md)** - Short, complete programs for common tasks - **[Foundations](https://goaisdk.com/docs/foundations.md)** - Learn about core concepts like providers, prompts, and streaming - **[AI SDK Core](https://goaisdk.com/docs/ai-sdk-core.md)** - Explore the complete API reference - **[Agents](https://goaisdk.com/docs/agents.md)** - Build more sophisticated autonomous agents - **[Advanced Topics](https://goaisdk.com/docs/advanced.md)** - Learn about production patterns like caching, rate limiting, and backpressure - **[Examples](https://github.com/digitallysavvy/go-ai/tree/main/examples)** - Check out more example applications ## Production Tips When moving to production, consider: 1. **Error Handling** - Implement proper error handling and retry logic 2. **Context Timeouts** - Set appropriate timeouts for your use case 3. **Rate Limiting** - Respect provider rate limits 4. **Monitoring** - Add telemetry and logging 5. **Cost Management** - Monitor token usage and implement caching Example with timeout and better error handling: ```go import providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() result, err := ai.GenerateText(ctx, options) if err != nil { var rateLimitErr *providererrors.RateLimitError var providerErr *providererrors.ProviderError switch { case errors.As(err, &rateLimitErr): log.Printf("Rate limited: %v", rateLimitErr) // Wait and retry, respecting rateLimitErr.RetryAfterSeconds if set case errors.As(err, &providerErr): log.Printf("Provider error (%d): %v", providerErr.StatusCode, providerErr) // Retry on 5xx, fail fast on 4xx case providererrors.IsValidationError(err): log.Printf("Invalid input: %v", err) // Fix the request default: log.Printf("Unknown error: %v", err) } return } ``` --- # Use Go AI SDK with coding agents > Point Claude Code, Cursor, Codex and other coding agents at the Go AI SDK docs, with a snippet for AGENTS.md or CLAUDE.md, an agent skill, and a docs MCP server. Canonical URL: https://goaisdk.com/docs/getting-started/using-go-ai-with-coding-agents Documentation index: https://goaisdk.com/llms.txt Coding agents write better Go AI SDK code when they read the docs instead of guessing from training data. The site publishes the docs in forms an agent can use directly. ## Agent entry points | URL | What it is | |---|---| | [`/agents.md`](https://goaisdk.com/agents.md) | A 3 KB orientation file: module path, install, the six most common calls and the usual mistakes. Start here. | | [`/llms.txt`](https://goaisdk.com/llms.txt) | An index of every page with a "Start here" block. Providers, migration guides and troubleshooting sit under "Optional". | | [`/sitemap.md`](https://goaisdk.com/sitemap.md) | Every page with its type, a one-line summary and prerequisites. | | [`/llms-core.txt`](https://goaisdk.com/llms-core.txt) | The core, agents and reference docs in one file, under 300 KB. | | `/docs/
/llms.txt` | An index for one section, for example `/docs/agents/llms.txt`. | | [`/llms-full.txt`](https://goaisdk.com/llms-full.txt) | All docs in one file. It is large; prefer the files above. | | `.md` | That page as markdown, with its title, description and canonical URL at the top. | ## Add the SDK to your project instructions Paste this into your project's `AGENTS.md` or `CLAUDE.md`: ```md ## Go AI SDK This project uses the Go AI SDK (`github.com/digitallysavvy/go-ai`, Go 1.26 or later). - Read https://goaisdk.com/agents.md before writing SDK code. It lists the common calls and mistakes. - Find a page with https://goaisdk.com/llms.txt, or search https://goaisdk.com/sitemap.md. Fetch any docs page as markdown by appending `.md` to its URL. - `LanguageModel(id)` returns `(model, error)`. Streams are `TextStream` or `Chunks()`, not `io.Reader`. - Provider option keys are camelCase. Use `ToolApproval`, not `NeedsApproval`. - `GenerateText` and `StreamText` run one step unless you set `StopWhen`. - For a useChat endpoint, use `ai.PipeUIMessageStreamToResponse` or `agent.PipeAgentUIStreamFromUIMessagesToResponse`. They set the stream headers. - Take model IDs from the provider package constants (for example `anthropic.ClaudeSonnet5_5`). - Reference app: https://github.com/digitallysavvy/go-ai-demo ``` ## Install the agent skill The repository ships the same guidance as an agent skill at [`skills/go-ai/SKILL.md`](https://github.com/digitallysavvy/go-ai/blob/main/skills/go-ai/SKILL.md). To use it in Claude Code, copy the folder into your project or your home directory: ```bash mkdir -p .claude/skills/go-ai curl -fsSL https://raw.githubusercontent.com/digitallysavvy/go-ai/main/skills/go-ai/SKILL.md \ -o .claude/skills/go-ai/SKILL.md ``` Use `~/.claude/skills/go-ai/` instead to make it available in every project. The skill loads when you work on code that imports the SDK. ## Search the docs from your editor The `goai-docs-mcp` server gives an agent two tools, `search_docs` and `read_doc`, over the full docs. It runs locally over stdio, and the docs are compiled into the binary, so it needs no network access and no API key. ```bash go install github.com/digitallysavvy/go-ai/cmd/goai-docs-mcp@latest ``` Add it to Claude Code: ```bash claude mcp add go-ai-docs -- goai-docs-mcp ``` Add it to Cursor in `.cursor/mcp.json` (or `~/.cursor/mcp.json` for every project): ```json { "mcpServers": { "go-ai-docs": { "command": "goai-docs-mcp" } } } ``` Make sure `$(go env GOPATH)/bin` is on your `PATH`, or use the absolute path to the binary as the command. Then ask your agent to "search the Go AI SDK docs for tool approval". The docs in the binary match the version you installed, so update it with the SDK. ## Next steps - [Quick start](https://goaisdk.com/docs/getting-started/golang.md) builds a first chat agent. - The [Shipyard demo](https://github.com/digitallysavvy/go-ai-demo) is a complete app with a useChat frontend, a Go backend, an approval-gated tool and a coding-agent harness. --- # Getting Started > Entry point for Go AI SDK docs, pointing to the quick start, installation, a first example application, and guidance for API servers and frameworks. Canonical URL: https://goaisdk.com/docs/getting-started/ Documentation index: https://goaisdk.com/llms.txt The following guides are intended to provide you with an introduction to some of the core features provided by the Go AI SDK. ## Quick Start Get started quickly with our comprehensive tutorial: - **[Go Quick Start](https://goaisdk.com/docs/getting-started/golang.md)** - Build your first AI agent with the Go AI SDK in 5 minutes ## Learn the Library - **[Navigating the Library](https://goaisdk.com/docs/getting-started/navigating-the-library.md)** - Understand the structure and choose the right tools for your needs ## What You'll Learn These guides will walk you through: 1. **Setting up your environment** - Install the SDK and configure your API keys 2. **Core concepts** - Understand providers, models, and streaming 3. **Streaming** - Print text as the model produces it 4. **Using tools** - Extend your agent with external capabilities and build a chat loop 5. **Production patterns** - Best practices for real-world applications ## Prerequisites To follow these guides, you'll need: - **Go 1.26+** installed on your development machine - **An API key** from at least one AI provider (OpenAI, Anthropic, Google, etc.) - **Basic Go knowledge** - Understanding of packages, error handling, and goroutines ## Quick Installation ```bash go get github.com/digitallysavvy/go-ai ``` ## Example: Your First AI Application Set your API key in the shell, then run a minimal program: ```bash export OPENAI_API_KEY=sk-... ``` ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { model, err := openai.New(openai.Config{}).LanguageModel(openai.ModelGPT6Astra) if err != nil { log.Fatal(err) } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Explain what makes Go great for AI applications.", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` The [Go Quick Start](https://goaisdk.com/docs/getting-started/golang.md) builds on this program with streaming, a tool, and a chat loop. ## API Servers and Frameworks You can use the Go AI SDK with any Go web framework or server: - **net/http** - Standard library HTTP server - **Gin** - High-performance web framework - **Echo** - Minimalist web framework - **Fiber** - Express-inspired framework - **Chi** - Lightweight router - **Gorilla** - Web toolkit - **gRPC** - For high-performance RPC services Example with net/http: ```go func streamHandler(w http.ResponseWriter, r *http.Request) { ctx := r.Context() stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a story...", }) if err != nil { http.Error(w, err.Error(), http.StatusInternalServerError) return } w.Header().Set("Content-Type", "text/event-stream") w.Header().Set("Cache-Control", "no-cache") flusher, _ := w.(http.Flusher) for chunk := range stream.Chunks() { fmt.Fprintf(w, "data: %s\n\n", chunk.Text) flusher.Flush() } } ``` ## Next Steps After completing the getting started guide, explore: - **[Build a chat app](https://goaisdk.com/docs/build-a-chat-app.md)** - Serve a `useChat` frontend from Go, add tool approval and run coding agents, based on the [Shipyard demo](https://github.com/digitallysavvy/go-ai-demo) - **[Recipes](https://goaisdk.com/docs/recipes.md)** - Short, complete programs for common tasks - **[Foundations](https://goaisdk.com/docs/foundations.md)** - Deep dive into core concepts - **[AI SDK Core](https://goaisdk.com/docs/ai-sdk-core.md)** - Complete API reference - **[Agents](https://goaisdk.com/docs/agents.md)** - Build autonomous agents - **[Advanced Topics](https://goaisdk.com/docs/advanced.md)** - Production-ready patterns ## Need Help? - **Documentation**: You're reading it! - **GitHub Issues**: [Report bugs or ask questions](https://github.com/digitallysavvy/go-ai/issues) - **Examples**: Check the `examples/` directory in the repository --- # Overview > Explains how agents combine LLMs and tools in a loop using ToolLoopAgent, covering execution flow, stopping conditions, callbacks, and structured workflows. Canonical URL: https://goaisdk.com/docs/agents/overview Documentation index: https://goaisdk.com/llms.txt Agents are **large language models (LLMs)** that use **tools** in a **loop** to accomplish tasks. These components work together: - **LLMs** process input and decide the next action - **Tools** extend capabilities beyond text generation (reading files, calling APIs, writing to databases) - **Loop** orchestrates execution through: - **Context management** - Maintaining conversation history and deciding what the model sees (input) at each step - **Stopping conditions** - Determining when the loop (task) is complete ## ToolLoopAgent The `ToolLoopAgent` handles these three components. Here's an agent that uses multiple tools in a loop to accomplish a task: ```go package main import ( "context" "fmt" "log" "math/rand" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() // Set up provider and model provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-6-astra") // Create agent with tools weatherAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: []types.Tool{ { Name: "weather", Description: "Get the weather in a location (in Fahrenheit)", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "The location to get the weather for", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location := input["location"].(string) return map[string]interface{}{ "location": location, "temperature": 72 + rand.Intn(21) - 10, }, nil }, }, { Name: "convertFahrenheitToCelsius", Description: "Convert temperature from Fahrenheit to Celsius", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "temperature": map[string]interface{}{ "type": "number", "description": "Temperature in Fahrenheit", }, }, "required": []string{"temperature"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { temp := input["temperature"].(float64) celsius := int((temp - 32) * (5.0 / 9.0)) return map[string]interface{}{ "celsius": celsius, }, nil }, }, }, StopWhen: []ai.StopCondition{ai.IsStepCount(20)}, // Agent stops after maximum of 20 steps }) // Execute agent result, err := weatherAgent.Execute(ctx, "What is the weather in San Francisco in celsius?") if err != nil { log.Fatal(err) } fmt.Println("Final Answer:", result.Text) fmt.Printf("Steps taken: %d\n", len(result.Steps)) } ``` The agent automatically: 1. Calls the `weather` tool to get the temperature in Fahrenheit 2. Calls `convertFahrenheitToCelsius` to convert it 3. Generates a final text response with the result The `ToolLoopAgent` handles the loop, context management, and stopping conditions. ## Why Use Agents? Using agents provides several benefits: - **Reduces boilerplate** - Manages loops and message arrays automatically - **Improves reusability** - Define once, use throughout your application - **Simplifies maintenance** - Single place to update agent configuration - **Handles complexity** - Manages tool execution, context, and termination logic For most use cases, start with agents. Use core functions (`ai.GenerateText`, `ai.StreamText`) when you need explicit control over each step for complex structured workflows. ## Basic Agent Structure Every agent needs: ### 1. Model The language model that powers the agent: ```go provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-6-astra") ``` ### 2. Tools Functions the agent can call to perform actions: ```go tools := []types.Tool{ { Name: "search", Description: "Search the web for information", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{"type": "string"}, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Implementation return map[string]interface{}{"results": []string{}}, nil }, }, } ``` ### 3. Configuration Settings that control agent behavior: ```go config := agent.AgentConfig{ Model: model, Tools: tools, MaxSteps: 15, // Maximum iterations System: "You are a helpful assistant", // System prompt } ``` ## Agent Execution Flow When you call `agent.Execute()`, the following happens: 1. **Initial Call**: Agent receives user prompt and converts it to messages 2. **Generation**: Model generates response (text and/or tool calls) 3. **Tool Execution**: If tool calls are present, execute them 4. **Context Update**: Add assistant response and tool results to conversation 5. **Loop**: Repeat steps 2-4 until stopping condition is met 6. **Return**: Final result with text, steps, and metadata ``` ┌─────────────┐ │ User Prompt │ └──────┬──────┘ │ ▼ ┌─────────────────┐ │ Model Generate │◄────┐ └────────┬────────┘ │ │ │ ▼ │ ┌────────┐ │ │ Tools? │──No───►Stop └───┬────┘ │ │Yes │ ▼ │ ┌──────────────┐ │ │ Execute Tools│ │ └──────┬───────┘ │ │ │ └────────────────┘ ``` ## Stopping Conditions Agents stop when: 1. **StopWhen reached**: Agent hits the configured step limit (default: `ai.IsStepCount(20)`) 2. **No tool calls**: Model returns text without requesting tool calls 3. **Error occurs**: Tool execution or generation fails 4. **Context canceled**: Context deadline or cancellation ```go config := agent.AgentConfig{ Model: model, Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(20)}, // Stop after 20 iterations } ``` ## Callbacks Monitor agent execution with callbacks: ```go config := agent.AgentConfig{ Model: model, Tools: tools, OnStepStart: func(stepNum int) { fmt.Printf("Starting step %d\n", stepNum) }, OnStepFinish: func(step types.StepResult) { fmt.Printf("Step %d complete: %s\n", step.StepNumber, step.Text) }, OnToolCall: func(toolCall types.ToolCall) { fmt.Printf("Calling tool: %s\n", toolCall.ToolName) }, OnToolResult: func(toolResult types.ToolResult) { fmt.Printf("Tool %s returned: %v\n", toolResult.ToolName, toolResult.Result) }, OnFinish: func(result *agent.AgentResult) { fmt.Printf("Agent complete after %d steps\n", len(result.Steps)) }, } ``` ## Agent vs Core Functions ### When to use Agents Use agents when: - Task requires multiple tool calls - You want automatic loop management - You need reusable agent configurations - You want simplified context management ```go // Agent handles everything automatically agent := agent.NewToolLoopAgent(config) result, _ := agent.Execute(ctx, "Complex multi-step task") ``` ### When to use Core Functions Use core functions when: - You need explicit control over each step - Building complex structured workflows - Implementing custom loop logic - Requiring specific error handling ```go // Manual control over each step for step := 0; step < maxSteps; step++ { result, _ := ai.GenerateText(ctx, options) // Custom logic for each step if shouldStop(result) { break } } ``` ## Structured Workflows Agents are flexible and powerful, but non-deterministic. When you need reliable, repeatable outcomes with explicit control flow, use core functions with structured workflow patterns combining: - Conditional statements for explicit branching - Standard functions for reusable logic - Error handling for robustness - Explicit control flow for predictability [Explore workflow patterns](https://goaisdk.com/docs/agents/workflows.md) to learn more about building structured, reliable systems. ## Real-World Example Here's a practical agent that can search the web and summarize results: ```go package main import ( "context" "fmt" "log" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() provider := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) model, _ := provider.LanguageModel("claude-sonnet-4-5") // Create research agent researchAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a research assistant. Use the available tools to find and summarize information.", Tools: []types.Tool{ { Name: "search", Description: "Search the web for information", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{ "type": "string", "description": "The search query", }, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := input["query"].(string) // Actual implementation would call a search API return fmt.Sprintf("Search results for: %s", query), nil }, }, { Name: "fetchWebpage", Description: "Fetch and return the content of a webpage", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "url": map[string]interface{}{ "type": "string", "description": "The URL to fetch", }, }, "required": []string{"url"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { url := input["url"].(string) resp, err := http.Get(url) if err != nil { return nil, err } defer resp.Body.Close() // Actual implementation would parse and clean HTML return fmt.Sprintf("Content from %s", url), nil }, }, }, MaxSteps: 15, OnStepFinish: func(step types.StepResult) { fmt.Printf("[Step %d] %s\n", step.StepNumber, step.Text) }, }) result, err := researchAgent.Execute(ctx, "Research the latest developments in quantum computing and provide a summary") if err != nil { log.Fatal(err) } fmt.Println("\n=== Final Answer ===") fmt.Println(result.Text) fmt.Printf("\nCompleted in %d steps\n", len(result.Steps)) } ``` ## Next Steps - **[Building Agents](https://goaisdk.com/docs/agents/building-agents.md)** - Detailed guide to creating agents - **[Workflow Patterns](https://goaisdk.com/docs/agents/workflows.md)** - Structured patterns using core functions - **[Loop Control](https://goaisdk.com/docs/agents/loop-control.md)** - Advanced execution control - **[Configuring Call Options](https://goaisdk.com/docs/agents/configuring-call-options.md)** - Fine-tune agent behavior ## See Also - [Tools and Tool Calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) --- # Building Agents > Complete guide to creating agents with ToolLoopAgent in Go: configuration options, system instructions, callbacks, and human-in-the-loop tool approval. Canonical URL: https://goaisdk.com/docs/agents/building-agents Documentation index: https://goaisdk.com/llms.txt The `ToolLoopAgent` provides a structured way to encapsulate LLM configuration, tools, and behavior into reusable components. It handles the agent loop for you, allowing the LLM to call tools multiple times in sequence to accomplish complex tasks. Define agents once and use them across your application. ## Why Use ToolLoopAgent? When building AI applications, you often need to: - **Reuse configurations** - Same model settings, tools, and prompts across different parts of your application - **Maintain consistency** - Ensure the same behavior and capabilities throughout your codebase - **Simplify API routes** - Reduce boilerplate in your endpoints - **Type safety** - Get full compile-time type checking for your agent's configuration The `ToolLoopAgent` provides a single place to define your agent's behavior. ## Creating an Agent Define an agent by calling `NewToolLoopAgent` with your desired configuration: ```go package main import ( "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful assistant.", Tools: []types.Tool{ // Your tools here }, }) _ = myAgent } ``` ## Configuration Options The `AgentConfig` struct accepts all the same settings as `GenerateText` and `StreamText`. Configure: ### Model and System Instructions ```go package main import ( "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are an expert software engineer.", }) _ = myAgent } ``` ### Tools Provide tools that the agent can use to accomplish tasks: ```go package main import ( "context" "fmt" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") codeAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: []types.Tool{ { Name: "runCode", Description: "Execute Python code", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "code": map[string]interface{}{ "type": "string", "description": "The Python code to execute", }, }, "required": []string{"code"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { code := input["code"].(string) // Execute code and return result return map[string]interface{}{ "output": fmt.Sprintf("Executed: %s", code), }, nil }, }, }, }) _ = codeAgent } ``` ### Loop Control By default, agents stop after 20 completed steps (`StopWhen: []ai.StopCondition{ai.IsStepCount(20)}`). In each step, the model either generates text or calls a tool. If it generates text, the agent completes. If it calls a tool, the SDK executes that tool. To let agents call multiple tools in sequence, configure `StopWhen` with `ai.IsStepCount`. After each tool execution, the agent triggers a new generation where the model can call another tool or generate text: ```go package main import ( "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, StopWhen: []ai.StopCondition{ai.IsStepCount(20)}, // Allow up to 20 steps }) _ = myAgent } ``` Each step represents one generation (which results in either text or a tool call). The loop continues until: - A finish reason other than `FinishReasonToolCalls` is returned, or - A tool execution fails, or - A tool call needs approval (see [tool approval](#tool-approval-and-human-in-the-loop)), or - The step limit configured with `StopWhen` is reached Learn more about [loop control](https://goaisdk.com/docs/agents/loop-control.md). ### Generation Settings Control temperature, max tokens, and other generation parameters: ```go package main import ( "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") temperature := 0.7 maxTokens := 2000 myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a creative writer.", Temperature: &temperature, MaxTokens: &maxTokens, }) _ = myAgent } ``` ## Define Agent Behavior with System Instructions System instructions define your agent's behavior, personality, and constraints. They set the context for all interactions and guide how the agent responds to user queries and uses tools. ### Basic System Instructions Set the agent's role and expertise: ```go myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are an expert data analyst. You provide clear insights from complex data.", }) ``` ### Detailed Behavioral Instructions Provide specific guidelines for agent behavior: ```go codeReviewAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: `You are a senior software engineer conducting code reviews. Your approach: - Focus on security vulnerabilities first - Identify performance bottlenecks - Suggest improvements for readability and maintainability - Be constructive and educational in your feedback - Always explain why something is an issue and how to fix it`, }) ``` ### Constrain Agent Behavior Set boundaries and ensure consistent behavior: ```go customerSupportAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: `You are a customer support specialist for an e-commerce platform. Rules: - Never make promises about refunds without checking the policy - Always be empathetic and professional - If you don't know something, say so and offer to escalate - Keep responses concise and actionable - Never share internal company information`, Tools: []types.Tool{ checkOrderStatusTool, lookupPolicyTool, createTicketTool, }, }) ``` ### Tool Usage Instructions Guide how the agent should use available tools: ```go researchAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: `You are a research assistant with access to search and document tools. When researching: 1. Always start with a broad search to understand the topic 2. Use document analysis for detailed information 3. Cross-reference multiple sources before drawing conclusions 4. Cite your sources when presenting information 5. If information conflicts, present both viewpoints`, Tools: []types.Tool{ webSearchTool, analyzeDocumentTool, extractQuotesTool, }, }) ``` ### Format and Style Instructions Control the output format and communication style: ```go technicalWriterAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: `You are a technical documentation writer. Writing style: - Use clear, simple language - Avoid jargon unless necessary - Structure information with headers and bullet points - Include code examples where relevant - Write in second person ("you" instead of "the user") Always format responses in Markdown.`, }) ``` ## Using an Agent Once defined, you can use your agent with its execution methods: ### Execute with Simple Prompt Use `Execute()` for one-time text generation with a simple prompt: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful assistant.", }) result, err := myAgent.Execute(ctx, "What is the weather like?") if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Execute with Message History Use `ExecuteWithMessages()` for conversations with history: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful assistant.", }) // Create conversation history messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What is 2+2?"}, }, }, { Role: types.RoleAssistant, Content: []types.ContentPart{ types.TextContent{Text: "2+2 equals 4."}, }, }, { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What about 3+3?"}, }, }, } result, err := myAgent.ExecuteWithMessages(ctx, messages) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Monitoring Agent Execution with Callbacks Monitor and debug agent behavior with lifecycle callbacks. Callbacks fire at key points during agent execution, providing visibility into the agent's decision-making process and allowing you to implement custom logic like logging, monitoring, or budget controls. ### Available Callbacks - **OnStepStart** - Called before each agent step begins - **OnStepFinish** - Called after each step completes (most useful for monitoring) - **OnToolCall** - Called when a tool is about to be executed - **OnToolResult** - Called after a tool execution completes - **OnFinish** - Called when the entire agent execution completes ### Basic Callback Example ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful research assistant.", MaxSteps: 15, OnStepStart: func(stepNum int) { fmt.Printf("Starting step %d\n", stepNum) }, OnStepFinish: func(step types.StepResult) { fmt.Printf("Step %d complete: %s\n", step.StepNumber, step.Text) if step.Usage.GetTotalTokens() > 0 { fmt.Printf(" Tokens used: %d\n", step.Usage.GetTotalTokens()) } }, OnToolCall: func(toolCall types.ToolCall) { fmt.Printf("Calling tool: %s\n", toolCall.ToolName) fmt.Printf(" Arguments: %v\n", toolCall.Arguments) }, OnToolResult: func(toolResult types.ToolResult) { if toolResult.Error != nil { fmt.Printf("Tool %s failed: %v\n", toolResult.ToolName, toolResult.Error) } else { fmt.Printf("Tool %s returned: %v\n", toolResult.ToolName, toolResult.Result) } }, OnFinish: func(result *agent.AgentResult) { fmt.Printf("Agent complete after %d steps\n", len(result.Steps)) if result.Usage.GetTotalTokens() > 0 { fmt.Printf("Total tokens used: %d\n", result.Usage.GetTotalTokens()) } }, }) result, err := myAgent.Execute(ctx, "Research the latest developments in quantum computing") if err != nil { log.Fatal(err) } fmt.Println("\n=== Final Answer ===") fmt.Println(result.Text) } ``` ### OnStepFinish Callback The `OnStepFinish` callback is called after each agent step completes. It provides detailed information about what happened during the step, making it the most useful callback for monitoring, logging, and implementing custom control logic. #### StepResult Structure The `StepResult` passed to `OnStepFinish` contains: - `StepNumber` (int) - The step number (0-indexed) - `Text` (string) - Text generated in this step - `ToolCalls` ([]ToolCall) - Tool calls made in this step - `ToolResults` ([]ToolResult) - Results from tool executions - `FinishReason` (FinishReason) - Why this step ended ("stop", "tool-calls", "length", etc.) - `RawFinishReason` (string) - Raw finish reason from the provider - `Usage` (Usage) - Token usage for this step - `Warnings` ([]Warning) - Warnings generated during this step - `ResponseMessages` ([]Message) - The assistant message generated in this step #### Common Use Cases **Logging and Debugging:** ```go OnStepFinish: func(step types.StepResult) { log.Printf("[Step %d] Generated: %s", step.StepNumber, step.Text) log.Printf("[Step %d] Tool calls: %d, Tokens: %d", step.StepNumber, len(step.ToolCalls), getTokens(step.Usage.TotalTokens)) } ``` **Token Usage Monitoring:** ```go totalTokens := int64(0) OnStepFinish: func(step types.StepResult) { if step.Usage.GetTotalTokens() > 0 { totalTokens += step.Usage.GetTotalTokens() } fmt.Printf("Step %d: %d tokens (total: %d)\n", step.StepNumber, getTokens(step.Usage.TotalTokens), totalTokens) if totalTokens > 10000 { log.Println("WARNING: Approaching token limit") } } ``` **Warning Detection:** ```go OnStepFinish: func(step types.StepResult) { if len(step.Warnings) > 0 { for _, warning := range step.Warnings { log.Printf("WARNING [%s]: %s", warning.Type, warning.Message) } } } ``` **Pattern Analysis:** ```go OnStepFinish: func(step types.StepResult) { // Detect inefficient tool usage if len(step.ToolCalls) > 3 { log.Printf("Multiple tool calls in step %d - consider optimization", step.StepNumber) } // Track tool usage patterns for _, tc := range step.ToolCalls { metrics.RecordToolUsage(tc.ToolName) } } ``` See the [OnStepFinish examples](https://github.com/digitallysavvy/go-ai/tree/main/examples/agents/callbacks) for complete working examples. **Helper function for safely accessing token counts:** ```go func getTokens(tokens *int64) int64 { if tokens == nil { return 0 } return *tokens } ``` ## Tool approval and human-in-the-loop Mark a tool as requiring approval with `Tool.ToolApproval`. Set it to `true`, or to a `types.ToolNeedsApprovalFunc` that decides per call: ```go deleteFile := types.Tool{ Name: "deleteFile", Description: "Delete a file from the filesystem", Parameters: map[string]interface{}{"type": "object"}, ToolApproval: true, // always ask Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return "deleted", nil }, } ``` To set a policy for the whole agent instead, use `AgentConfig.ToolApproval`. It takes a function or a map keyed by tool name: ```go agent.AgentConfig{ Model: model, Tools: []types.Tool{deleteFile}, ToolApproval: map[string]types.ToolApprovalValue{ "deleteFile": types.ToolApprovalStatusUserApproval, }, } ``` When a call needs approval, the loop stops with a pending approval request instead of running the tool. Your application collects the decision and resumes the conversation. Set `ExperimentalToolApprovalSecret` to sign requests, so a client cannot forge an approval. ### Blocking approval for command-line programs If your program can block and ask the user inline, the legacy `ToolApprovalRequired` and `ToolApprover` fields do that. `ToolApprover` runs for each tool call and returns true to approve or false to reject: ```go package main import ( "bufio" "context" "fmt" "log" "os" "strings" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Create agent with tool approval myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful assistant with access to various tools.", Tools: []types.Tool{ { Name: "deleteFile", Description: "Delete a file from the filesystem", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "path": map[string]interface{}{ "type": "string", "description": "The file path to delete", }, }, "required": []string{"path"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { path := input["path"].(string) err := os.Remove(path) if err != nil { return nil, err } return map[string]interface{}{ "success": true, "message": fmt.Sprintf("Deleted file: %s", path), }, nil }, }, }, MaxSteps: 10, ToolApprovalRequired: true, ToolApprover: func(toolCall types.ToolCall) bool { fmt.Printf("\n=== Tool Approval Required ===\n") fmt.Printf("Tool: %s\n", toolCall.ToolName) fmt.Printf("Arguments: %v\n", toolCall.Arguments) fmt.Print("Approve this tool call? (yes/no): ") reader := bufio.NewReader(os.Stdin) response, _ := reader.ReadString('\n') response = strings.TrimSpace(strings.ToLower(response)) approved := response == "yes" || response == "y" if approved { fmt.Println("✓ Tool call approved") } else { fmt.Println("✗ Tool call rejected") } return approved }, }) result, err := myAgent.Execute(ctx, "Please clean up the temp directory") if err != nil { log.Fatal(err) } fmt.Println("\n=== Final Answer ===") fmt.Println(result.Text) } ``` ## Updating Agent Configuration Modify agent configuration at runtime: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful assistant.", Tools: []types.Tool{}, }) // Update system prompt myAgent.SetSystem("You are now an expert mathematician.") // Add a tool myAgent.AddTool(types.Tool{ Name: "calculate", Description: "Perform a calculation", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "expression": map[string]interface{}{ "type": "string", "description": "The mathematical expression to evaluate", }, }, "required": []string{"expression"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { expr := input["expression"].(string) // Evaluate expression (simplified example) return map[string]interface{}{ "expression": expr, "result": 42, }, nil }, }) // Update max steps myAgent.SetMaxSteps(20) // Remove a tool myAgent.RemoveTool("calculate") result, err := myAgent.Execute(ctx, "Solve a complex problem") if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Accessing Agent Results The `AgentResult` provides detailed information about the agent's execution: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful assistant.", MaxSteps: 15, }) result, err := myAgent.Execute(ctx, "What's the weather in Paris and how many hours until sunset?") if err != nil { log.Fatal(err) } // Access final text fmt.Println("Final Answer:", result.Text) // Access all steps taken fmt.Printf("\nSteps taken: %d\n", len(result.Steps)) for _, step := range result.Steps { fmt.Printf(" Step %d: %s\n", step.StepNumber, step.Text) if len(step.ToolCalls) > 0 { fmt.Printf(" Tool calls: %d\n", len(step.ToolCalls)) } } // Access tool results fmt.Printf("\nTools used: %d\n", len(result.ToolResults)) for _, tr := range result.ToolResults { if tr.Error != nil { fmt.Printf(" %s: Error - %v\n", tr.ToolName, tr.Error) } else { fmt.Printf(" %s: Success\n", tr.ToolName) } } // Access usage information fmt.Printf("\nToken Usage:\n") if result.Usage.InputTokens != nil { fmt.Printf(" Input tokens: %d\n", *result.Usage.InputTokens) } if result.Usage.OutputTokens != nil { fmt.Printf(" Output tokens: %d\n", *result.Usage.OutputTokens) } if result.Usage.GetTotalTokens() > 0 { fmt.Printf(" Total tokens: %d\n", result.Usage.GetTotalTokens()) } // Check finish reason fmt.Printf("\nFinish Reason: %s\n", result.FinishReason) // Check warnings if len(result.Warnings) > 0 { fmt.Printf("\nWarnings: %d\n", len(result.Warnings)) for _, warning := range result.Warnings { fmt.Printf(" %s: %s\n", warning.Type, warning.Message) } } } ``` ## Complete Example: Research Agent Here's a complete example of a research agent with tools, callbacks, and proper error handling: ```go package main import ( "context" "fmt" "io" "log" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() // Create provider and model provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, err := provider.LanguageModel("claude-sonnet-4-5") if err != nil { log.Fatal(err) } // Create research agent with tools researchAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: `You are a research assistant. Use the available tools to find and analyze information. Guidelines: - Start with broad searches to understand the topic - Use webpage fetching for detailed information - Synthesize information from multiple sources - Cite your sources in your final answer`, Tools: []types.Tool{ { Name: "search", Description: "Search the web for information", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{ "type": "string", "description": "The search query", }, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := input["query"].(string) fmt.Printf("[Search] Query: %s\n", query) // Actual implementation would call a search API return map[string]interface{}{ "results": []map[string]interface{}{ { "title": "Example Result 1", "url": "https://example.com/1", "snippet": "Sample search result...", }, }, }, nil }, }, { Name: "fetchWebpage", Description: "Fetch and return the content of a webpage", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "url": map[string]interface{}{ "type": "string", "description": "The URL to fetch", }, }, "required": []string{"url"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { url := input["url"].(string) fmt.Printf("[Fetch] URL: %s\n", url) resp, err := http.Get(url) if err != nil { return nil, fmt.Errorf("failed to fetch URL: %w", err) } defer resp.Body.Close() body, err := io.ReadAll(resp.Body) if err != nil { return nil, fmt.Errorf("failed to read response: %w", err) } // Actual implementation would parse and clean HTML return map[string]interface{}{ "url": url, "content": string(body), }, nil }, }, }, MaxSteps: 15, // Add callbacks for monitoring OnStepStart: func(stepNum int) { fmt.Printf("\n[Agent] Starting step %d\n", stepNum) }, OnStepFinish: func(step types.StepResult) { if step.Text != "" { fmt.Printf("[Agent] Step %d: %s\n", step.StepNumber, step.Text) } }, OnToolCall: func(toolCall types.ToolCall) { fmt.Printf("[Agent] Calling tool: %s\n", toolCall.ToolName) }, OnFinish: func(result *agent.AgentResult) { fmt.Printf("\n[Agent] Completed in %d steps\n", len(result.Steps)) if result.Usage.GetTotalTokens() > 0 { fmt.Printf("[Agent] Total tokens: %d\n", result.Usage.GetTotalTokens()) } }, }) // Execute agent result, err := researchAgent.Execute( ctx, "Research the latest developments in quantum computing and provide a summary with key points", ) if err != nil { log.Fatalf("Agent execution failed: %v", err) } // Display results fmt.Println("\n=== Final Answer ===") fmt.Println(result.Text) fmt.Println("\n=== Agent Statistics ===") fmt.Printf("Steps: %d\n", len(result.Steps)) fmt.Printf("Tools used: %d\n", len(result.ToolResults)) if result.Usage.GetTotalTokens() > 0 { fmt.Printf("Total tokens: %d\n", result.Usage.GetTotalTokens()) } fmt.Printf("Finish reason: %s\n", result.FinishReason) } ``` ## Best Practices ### 1. Clear System Instructions Provide clear, specific instructions about the agent's role and behavior: ```go // Good - Clear and specific agent.AgentConfig{ System: `You are a code review assistant. Focus on security, performance, and maintainability. Rules: - Identify potential bugs and security issues - Suggest specific improvements with code examples - Explain why changes are needed - Be constructive and educational`, } // Bad - Too vague agent.AgentConfig{ System: "You help with code", } ``` ### 2. Appropriate MaxSteps Set `MaxSteps` based on task complexity: ```go // Simple tasks - fewer steps agent.AgentConfig{ MaxSteps: 5, // Quick Q&A, simple calculations } // Complex tasks - more steps agent.AgentConfig{ MaxSteps: 20, // Research, multi-step analysis } ``` ### 3. Tool Descriptions Write clear tool descriptions that help the model understand when to use each tool: ```go // Good - Clear purpose and usage types.Tool{ Name: "getWeather", Description: "Get current weather for a specific city. Use when user asks about weather conditions, temperature, or forecast.", } // Bad - Unclear types.Tool{ Name: "getWeather", Description: "Weather", } ``` ### 4. Error Handling in Tools Always handle errors properly in tool execution functions: ```go types.Tool{ Name: "fetchData", Description: "Fetch data from API", Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Good - Proper error handling data, err := fetchFromAPI(ctx, input) if err != nil { return nil, fmt.Errorf("failed to fetch data: %w", err) } return data, nil }, } ``` ### 5. Use Callbacks for Monitoring Add callbacks in development to understand agent behavior: ```go agent.AgentConfig{ OnToolCall: func(toolCall types.ToolCall) { log.Printf("Tool: %s, Args: %v", toolCall.ToolName, toolCall.Arguments) }, OnToolResult: func(toolResult types.ToolResult) { if toolResult.Error != nil { log.Printf("Tool %s failed: %v", toolResult.ToolName, toolResult.Error) } }, } ``` ### 6. Context Usage Always pass and respect context for cancellation and timeouts: ```go ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute) defer cancel() result, err := myAgent.Execute(ctx, prompt) if err != nil { if ctx.Err() == context.DeadlineExceeded { log.Println("Agent execution timed out") } return err } ``` ### 7. Tool Approval for Dangerous Operations Require approval for operations that modify state or access sensitive data: ```go agent.AgentConfig{ Tools: []types.Tool{deleteFile}, // deleteFile sets ToolApproval: true ToolApproval: types.GenericToolApprovalFunc(func(opts types.ToolApprovalOptions) types.ToolApprovalResult { // Implement your approval policy if opts.ToolCall.ToolName == "deleteFile" { return types.ToolApprovalResult{Status: types.ToolApprovalStatusUserApproval} } return types.ToolApprovalResult{} }), } ``` For a blocking prompt in a command-line program, use the legacy `ToolApprover` callback described above. ## Next Steps Now that you understand building agents, you can: - Explore [workflow patterns](https://goaisdk.com/docs/agents/workflows.md) for structured patterns using core functions - Learn about [loop control](https://goaisdk.com/docs/agents/loop-control.md) for advanced execution control - See [configuring call options](https://goaisdk.com/docs/agents/configuring-call-options.md) for fine-tuning behavior --- # Workflow Patterns > Learn workflow patterns for building reliable agents with the Go AI SDK. Canonical URL: https://goaisdk.com/docs/agents/workflows Documentation index: https://goaisdk.com/llms.txt Combine the building blocks from the [overview](https://goaisdk.com/docs/agents/overview.md) with these patterns to add structure and reliability to your agents: - [Sequential Processing](#sequential-processing-chains) - Steps executed in order - [Parallel Processing](#parallel-processing) - Independent tasks run simultaneously - [Evaluation/Feedback Loops](#evaluator-optimizer) - Results checked and improved iteratively - [Orchestration](#orchestrator-worker) - Coordinating multiple components - [Routing](#routing) - Directing work based on context ## Choose Your Approach Consider these key factors: - **Flexibility vs Control** - How much freedom does the LLM need vs how tightly you must constrain its actions? - **Error Tolerance** - What are the consequences of mistakes in your use case? - **Cost Considerations** - More complex systems typically mean more LLM calls and higher costs - **Maintenance** - Simpler architectures are easier to debug and modify **Start with the simplest approach that meets your needs**. Add complexity only when required by: 1. Breaking down tasks into clear steps 2. Adding tools for specific capabilities 3. Implementing feedback loops for quality control 4. Introducing multiple agents for complex workflows Let's look at examples of these patterns in action. ## Patterns with Examples These patterns, adapted from [Anthropic's guide on building effective agents](https://www.anthropic.com/research/building-effective-agents), serve as building blocks you can combine to create comprehensive workflows. Each pattern addresses specific aspects of task execution. Combine them thoughtfully to build reliable solutions for complex problems. ## Sequential Processing (Chains) The simplest workflow pattern executes steps in a predefined order. Each step's output becomes input for the next step, creating a clear chain of operations. Use this pattern for tasks with well-defined sequences, like content generation pipelines or data transformation processes. ```go package main import ( "context" "encoding/json" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/schema" ) type QualityMetrics struct { HasCallToAction bool `json:"hasCallToAction"` EmotionalAppeal float64 `json:"emotionalAppeal"` Clarity float64 `json:"clarity"` } func generateMarketingCopy(ctx context.Context, input string) (string, QualityMetrics, error) { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Step 1: Generate marketing copy copyResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Write persuasive marketing copy for: %s. Focus on benefits and emotional appeal.", input), }) if err != nil { return "", QualityMetrics{}, err } copy := copyResult.Text // Step 2: Perform quality check on copy qualityResult, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "hasCallToAction": map[string]interface{}{ "type": "boolean", "description": "Whether the copy has a clear call to action", }, "emotionalAppeal": map[string]interface{}{ "type": "number", "description": "Emotional appeal rating from 1-10", "minimum": 1, "maximum": 10, }, "clarity": map[string]interface{}{ "type": "number", "description": "Clarity rating from 1-10", "minimum": 1, "maximum": 10, }, }, "required": []string{"hasCallToAction", "emotionalAppeal", "clarity"}, }), Prompt: fmt.Sprintf(`Evaluate this marketing copy for: 1. Presence of call to action (true/false) 2. Emotional appeal (1-10) 3. Clarity (1-10) Copy to evaluate: %s`, copy), }) if err != nil { return "", QualityMetrics{}, err } var metrics QualityMetrics objBytes, err := json.Marshal(qualityResult.Object) if err != nil { return "", QualityMetrics{}, err } if err := json.Unmarshal(objBytes, &metrics); err != nil { return "", QualityMetrics{}, err } // Step 3: If quality check fails, regenerate with more specific instructions if !metrics.HasCallToAction || metrics.EmotionalAppeal < 7 || metrics.Clarity < 7 { improvements := "" if !metrics.HasCallToAction { improvements += "- A clear call to action\n" } if metrics.EmotionalAppeal < 7 { improvements += "- Stronger emotional appeal\n" } if metrics.Clarity < 7 { improvements += "- Improved clarity and directness\n" } improvedResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf(`Rewrite this marketing copy with: %s Original copy: %s`, improvements, copy), }) if err != nil { return "", QualityMetrics{}, err } return improvedResult.Text, metrics, nil } return copy, metrics, nil } func main() { ctx := context.Background() copy, metrics, err := generateMarketingCopy(ctx, "Revolutionary new AI-powered productivity app") if err != nil { log.Fatal(err) } fmt.Println("Final Copy:", copy) fmt.Printf("Quality Metrics: CTA=%v, Emotional=%v, Clarity=%v\n", metrics.HasCallToAction, metrics.EmotionalAppeal, metrics.Clarity) } ``` ## Routing This pattern lets the model decide which path to take through a workflow based on context and intermediate results. The model acts as an intelligent router, directing the flow of execution between different branches of your workflow. Use this when handling varied inputs that require different processing approaches. In the example below, the first LLM call's results determine the second call's model size and system prompt. ```go package main import ( "context" "encoding/json" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/schema" ) type Classification struct { Reasoning string `json:"reasoning"` Type string `json:"type"` // "general", "refund", or "technical" Complexity string `json:"complexity"` // "simple" or "complex" } func handleCustomerQuery(ctx context.Context, query string) (string, Classification, error) { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Step 1: Classify the query type classificationResult, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "reasoning": map[string]interface{}{ "type": "string", "description": "Brief reasoning for classification", }, "type": map[string]interface{}{ "type": "string", "enum": []string{"general", "refund", "technical"}, "description": "The type of customer query", }, "complexity": map[string]interface{}{ "type": "string", "enum": []string{"simple", "complex"}, "description": "The complexity level of the query", }, }, "required": []string{"reasoning", "type", "complexity"}, }), Prompt: fmt.Sprintf(`Classify this customer query: %s Determine: 1. Query type (general, refund, or technical) 2. Complexity (simple or complex) 3. Brief reasoning for classification`, query), }) if err != nil { return "", Classification{}, err } var classification Classification classificationBytes, err := json.Marshal(classificationResult.Object) if err != nil { return "", Classification{}, err } if err := json.Unmarshal(classificationBytes, &classification); err != nil { return "", Classification{}, err } // Step 2: Route based on classification // Set model based on complexity var responseModel string if classification.Complexity == "simple" { responseModel = "gpt-5.4-mini" } else { responseModel = "gpt-6-astra" } // Set system prompt based on query type systemPrompts := map[string]string{ "general": "You are an expert customer service agent handling general inquiries.", "refund": "You are a customer service agent specializing in refund requests. Follow company policy and collect necessary information.", "technical": "You are a technical support specialist with deep product knowledge. Focus on clear step-by-step troubleshooting.", } routedModel, _ := provider.LanguageModel(responseModel) responseResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: routedModel, System: systemPrompts[classification.Type], Prompt: query, }) if err != nil { return "", Classification{}, err } return responseResult.Text, classification, nil } func main() { ctx := context.Background() query := "I need help resetting my password and it's not working" response, classification, err := handleCustomerQuery(ctx, query) if err != nil { log.Fatal(err) } fmt.Printf("Classification: %s (%s)\n", classification.Type, classification.Complexity) fmt.Printf("Reasoning: %s\n", classification.Reasoning) fmt.Printf("\nResponse:\n%s\n", response) } ``` ## Parallel Processing Break down tasks into independent subtasks that execute simultaneously. This pattern uses parallel execution to improve efficiency while maintaining the benefits of structured workflows. For example, analyze multiple documents or process different aspects of a single input concurrently (like code review). ```go package main import ( "context" "encoding/json" "fmt" "log" "os" "sync" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/schema" ) type SecurityReview struct { Vulnerabilities []string `json:"vulnerabilities"` RiskLevel string `json:"riskLevel"` // "low", "medium", or "high" Suggestions []string `json:"suggestions"` } type PerformanceReview struct { Issues []string `json:"issues"` Impact string `json:"impact"` // "low", "medium", or "high" Optimizations []string `json:"optimizations"` } type MaintainabilityReview struct { Concerns []string `json:"concerns"` QualityScore float64 `json:"qualityScore"` // 1-10 Recommendations []string `json:"recommendations"` } func parallelCodeReview(ctx context.Context, code string) (string, error) { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Run parallel reviews using goroutines var wg sync.WaitGroup var securityReview SecurityReview var performanceReview PerformanceReview var maintainabilityReview MaintainabilityReview errChan := make(chan error, 3) // Security review wg.Add(1) go func() { defer wg.Done() result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, System: "You are an expert in code security. Focus on identifying security vulnerabilities, injection risks, and authentication issues.", Schema: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "vulnerabilities": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, }, "riskLevel": map[string]interface{}{ "type": "string", "enum": []string{"low", "medium", "high"}, }, "suggestions": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, }, }, "required": []string{"vulnerabilities", "riskLevel", "suggestions"}, }), Prompt: fmt.Sprintf("Review this code:\n%s", code), }) if err != nil { errChan <- err return } securityBytes, err := json.Marshal(result.Object) if err != nil { errChan <- err return } if err := json.Unmarshal(securityBytes, &securityReview); err != nil { errChan <- err } }() // Performance review wg.Add(1) go func() { defer wg.Done() result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, System: "You are an expert in code performance. Focus on identifying performance bottlenecks, memory leaks, and optimization opportunities.", Schema: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "issues": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, }, "impact": map[string]interface{}{ "type": "string", "enum": []string{"low", "medium", "high"}, }, "optimizations": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, }, }, "required": []string{"issues", "impact", "optimizations"}, }), Prompt: fmt.Sprintf("Review this code:\n%s", code), }) if err != nil { errChan <- err return } performanceBytes, err := json.Marshal(result.Object) if err != nil { errChan <- err return } if err := json.Unmarshal(performanceBytes, &performanceReview); err != nil { errChan <- err } }() // Maintainability review wg.Add(1) go func() { defer wg.Done() result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, System: "You are an expert in code quality. Focus on code structure, readability, and adherence to best practices.", Schema: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "concerns": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, }, "qualityScore": map[string]interface{}{ "type": "number", "minimum": 1, "maximum": 10, }, "recommendations": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, }, }, "required": []string{"concerns", "qualityScore", "recommendations"}, }), Prompt: fmt.Sprintf("Review this code:\n%s", code), }) if err != nil { errChan <- err return } maintainabilityBytes, err := json.Marshal(result.Object) if err != nil { errChan <- err return } if err := json.Unmarshal(maintainabilityBytes, &maintainabilityReview); err != nil { errChan <- err } }() // Wait for all reviews to complete wg.Wait() close(errChan) // Check for errors for err := range errChan { if err != nil { return "", err } } // Aggregate results reviews := fmt.Sprintf(`Security Review: - Risk Level: %s - Vulnerabilities: %v - Suggestions: %v Performance Review: - Impact: %s - Issues: %v - Optimizations: %v Maintainability Review: - Quality Score: %.1f/10 - Concerns: %v - Recommendations: %v`, securityReview.RiskLevel, securityReview.Vulnerabilities, securityReview.Suggestions, performanceReview.Impact, performanceReview.Issues, performanceReview.Optimizations, maintainabilityReview.QualityScore, maintainabilityReview.Concerns, maintainabilityReview.Recommendations) // Generate summary summaryResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a technical lead summarizing multiple code reviews.", Prompt: fmt.Sprintf("Synthesize these code review results into a concise summary with key actions:\n%s", reviews), }) if err != nil { return "", err } return summaryResult.Text, nil } func main() { ctx := context.Background() code := ` func processPayment(amount float64, cardNumber string) error { query := fmt.Sprintf("SELECT * FROM payments WHERE card='%s'", cardNumber) db.Exec(query) return nil } ` summary, err := parallelCodeReview(ctx, code) if err != nil { log.Fatal(err) } fmt.Println("=== Code Review Summary ===") fmt.Println(summary) } ``` ## Orchestrator-Worker A primary model (orchestrator) coordinates the execution of specialized workers. Each worker optimizes for a specific subtask, while the orchestrator maintains overall context and ensures coherent results. This pattern excels at complex tasks requiring different types of expertise or processing. ```go package main import ( "context" "encoding/json" "fmt" "log" "os" "sync" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/schema" ) type FileChange struct { Purpose string `json:"purpose"` FilePath string `json:"filePath"` ChangeType string `json:"changeType"` // "create", "modify", or "delete" } type ImplementationPlan struct { Files []FileChange `json:"files"` EstimatedComplexity string `json:"estimatedComplexity"` // "low", "medium", or "high" } type CodeImplementation struct { Explanation string `json:"explanation"` Code string `json:"code"` } type FileImplementation struct { File FileChange Implementation CodeImplementation } func implementFeature(ctx context.Context, featureRequest string) (ImplementationPlan, []FileImplementation, error) { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Orchestrator: Plan the implementation planResult, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, System: "You are a senior software architect planning feature implementations.", Schema: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "files": map[string]interface{}{ "type": "array", "items": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "purpose": map[string]interface{}{ "type": "string", "description": "The purpose of this file change", }, "filePath": map[string]interface{}{ "type": "string", "description": "Path to the file", }, "changeType": map[string]interface{}{ "type": "string", "enum": []string{"create", "modify", "delete"}, }, }, "required": []string{"purpose", "filePath", "changeType"}, }, }, "estimatedComplexity": map[string]interface{}{ "type": "string", "enum": []string{"low", "medium", "high"}, }, }, "required": []string{"files", "estimatedComplexity"}, }), Prompt: fmt.Sprintf("Analyze this feature request and create an implementation plan:\n%s", featureRequest), }) if err != nil { return ImplementationPlan{}, nil, err } var plan ImplementationPlan planBytes, err := json.Marshal(planResult.Object) if err != nil { return ImplementationPlan{}, nil, err } if err := json.Unmarshal(planBytes, &plan); err != nil { return ImplementationPlan{}, nil, err } fmt.Printf("Plan created: %d files to change (Complexity: %s)\n", len(plan.Files), plan.EstimatedComplexity) // Workers: Execute the planned changes in parallel var wg sync.WaitGroup fileChanges := make([]FileImplementation, len(plan.Files)) errChan := make(chan error, len(plan.Files)) for i, file := range plan.Files { wg.Add(1) go func(idx int, f FileChange) { defer wg.Done() // Each worker is specialized for the type of change systemPrompts := map[string]string{ "create": "You are an expert at implementing new files following best practices and project patterns.", "modify": "You are an expert at modifying existing code while maintaining consistency and avoiding regressions.", "delete": "You are an expert at safely removing code while ensuring no breaking changes.", } changeResult, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, System: systemPrompts[f.ChangeType], Schema: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "explanation": map[string]interface{}{ "type": "string", "description": "Explanation of the changes", }, "code": map[string]interface{}{ "type": "string", "description": "The implementation code", }, }, "required": []string{"explanation", "code"}, }), Prompt: fmt.Sprintf(`Implement the changes for %s to support: %s Consider the overall feature context: %s`, f.FilePath, f.Purpose, featureRequest), }) if err != nil { errChan <- err return } var impl CodeImplementation implBytes, err := json.Marshal(changeResult.Object) if err != nil { errChan <- err return } if err := json.Unmarshal(implBytes, &impl); err != nil { errChan <- err return } fileChanges[idx] = FileImplementation{ File: f, Implementation: impl, } fmt.Printf("✓ Completed %s: %s\n", f.ChangeType, f.FilePath) }(i, file) } wg.Wait() close(errChan) // Check for errors for err := range errChan { if err != nil { return ImplementationPlan{}, nil, err } } return plan, fileChanges, nil } func main() { ctx := context.Background() featureRequest := "Add user authentication with JWT tokens, including login, logout, and token refresh endpoints" plan, changes, err := implementFeature(ctx, featureRequest) if err != nil { log.Fatal(err) } fmt.Printf("\n=== Implementation Plan ===\n") fmt.Printf("Complexity: %s\n", plan.EstimatedComplexity) fmt.Printf("Files to change: %d\n\n", len(changes)) for _, change := range changes { fmt.Printf("=== %s: %s ===\n", change.File.ChangeType, change.File.FilePath) fmt.Printf("Purpose: %s\n", change.File.Purpose) fmt.Printf("Explanation: %s\n", change.Implementation.Explanation) fmt.Printf("Code:\n%s\n\n", change.Implementation.Code) } } ``` ## Evaluator-Optimizer Add quality control to workflows with dedicated evaluation steps that assess intermediate results. Based on the evaluation, the workflow proceeds, retries with adjusted parameters, or takes corrective action. This creates robust workflows capable of self-improvement and error recovery. ```go package main import ( "context" "encoding/json" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/schema" ) type TranslationEvaluation struct { QualityScore float64 `json:"qualityScore"` // 1-10 PreservesTone bool `json:"preservesTone"` PreservesNuance bool `json:"preservesNuance"` CulturallyAccurate bool `json:"culturallyAccurate"` SpecificIssues []string `json:"specificIssues"` ImprovementSuggestions []string `json:"improvementSuggestions"` } func translateWithFeedback(ctx context.Context, text, targetLanguage string) (string, int, error) { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") const maxIterations = 3 iterations := 0 // Initial translation translationResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are an expert literary translator.", Prompt: fmt.Sprintf("Translate this text to %s, preserving tone and cultural nuances:\n%s", targetLanguage, text), }) if err != nil { return "", 0, err } currentTranslation := translationResult.Text // Evaluation-optimization loop for iterations < maxIterations { fmt.Printf("Iteration %d: Evaluating translation...\n", iterations+1) // Evaluate current translation evalResult, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, System: "You are an expert in evaluating literary translations.", Schema: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "qualityScore": map[string]interface{}{ "type": "number", "minimum": 1, "maximum": 10, "description": "Overall quality score", }, "preservesTone": map[string]interface{}{ "type": "boolean", "description": "Whether the translation preserves the original tone", }, "preservesNuance": map[string]interface{}{ "type": "boolean", "description": "Whether the translation preserves subtle nuances", }, "culturallyAccurate": map[string]interface{}{ "type": "boolean", "description": "Whether the translation is culturally accurate", }, "specificIssues": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, "description": "Specific issues found in the translation", }, "improvementSuggestions": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, "description": "Suggestions for improving the translation", }, }, "required": []string{"qualityScore", "preservesTone", "preservesNuance", "culturallyAccurate", "specificIssues", "improvementSuggestions"}, }), Prompt: fmt.Sprintf(`Evaluate this translation: Original: %s Translation: %s Consider: 1. Overall quality 2. Preservation of tone 3. Preservation of nuance 4. Cultural accuracy`, text, currentTranslation), }) if err != nil { return "", 0, err } var evaluation TranslationEvaluation evalBytes, err := json.Marshal(evalResult.Object) if err != nil { return "", 0, err } if err := json.Unmarshal(evalBytes, &evaluation); err != nil { return "", 0, err } fmt.Printf("Quality Score: %.1f/10\n", evaluation.QualityScore) // Check if quality meets threshold if evaluation.QualityScore >= 8 && evaluation.PreservesTone && evaluation.PreservesNuance && evaluation.CulturallyAccurate { fmt.Println("Translation meets quality standards!") break } // Generate feedback for improvement feedback := "" for _, issue := range evaluation.SpecificIssues { feedback += fmt.Sprintf("- Issue: %s\n", issue) } for _, suggestion := range evaluation.ImprovementSuggestions { feedback += fmt.Sprintf("- Suggestion: %s\n", suggestion) } fmt.Printf("Issues found, improving translation...\n%s\n", feedback) // Generate improved translation based on feedback improvedResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are an expert literary translator.", Prompt: fmt.Sprintf(`Improve this translation based on the following feedback: %s Original: %s Current Translation: %s`, feedback, text, currentTranslation), }) if err != nil { return "", 0, err } currentTranslation = improvedResult.Text iterations++ } return currentTranslation, iterations, nil } func main() { ctx := context.Background() text := "The ancient temple stood silently, a testament to forgotten wisdom and the passage of time." targetLanguage := "Japanese" translation, iterations, err := translateWithFeedback(ctx, text, targetLanguage) if err != nil { log.Fatal(err) } fmt.Printf("\n=== Final Translation ===\n") fmt.Printf("Original: %s\n", text) fmt.Printf("Translation: %s\n", translation) fmt.Printf("Iterations required: %d\n", iterations) } ``` ## Combining Patterns Real-world applications often combine multiple patterns. Here's an example that uses routing, parallel processing, and sequential steps: ```go package main import ( "context" "encoding/json" "fmt" "log" "os" "sync" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/schema" ) type TaskType struct { Type string `json:"type"` // "research", "code", or "content" Complexity string `json:"complexity"` // "simple" or "complex" } func intelligentTaskProcessor(ctx context.Context, task string) (string, error) { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Step 1: Route - Classify the task classificationResult, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "type": map[string]interface{}{ "type": "string", "enum": []string{"research", "code", "content"}, }, "complexity": map[string]interface{}{ "type": "string", "enum": []string{"simple", "complex"}, }, }, "required": []string{"type", "complexity"}, }), Prompt: fmt.Sprintf("Classify this task:\n%s", task), }) if err != nil { return "", err } var taskType TaskType taskTypeBytes, err := json.Marshal(classificationResult.Object) if err != nil { return "", err } if err := json.Unmarshal(taskTypeBytes, &taskType); err != nil { return "", err } fmt.Printf("Task classified as: %s (%s)\n", taskType.Type, taskType.Complexity) // Step 2: Parallel - Process based on type var result string switch taskType.Type { case "research": // Parallel research from multiple angles var wg sync.WaitGroup results := make([]string, 3) errChan := make(chan error, 3) angles := []string{"technical", "business", "user"} for i, angle := range angles { wg.Add(1) go func(idx int, perspective string) { defer wg.Done() r, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: fmt.Sprintf("You are a %s expert.", perspective), Prompt: fmt.Sprintf("Research this from a %s perspective:\n%s", perspective, task), }) if err != nil { errChan <- err return } results[idx] = r.Text }(i, angle) } wg.Wait() close(errChan) for err := range errChan { if err != nil { return "", err } } // Synthesize synthResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Synthesize these research findings:\n%s\n%s\n%s", results[0], results[1], results[2]), }) if err != nil { return "", err } result = synthResult.Text case "code": // Sequential code generation and review codeResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are an expert programmer.", Prompt: fmt.Sprintf("Implement this:\n%s", task), }) if err != nil { return "", err } reviewResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a code reviewer.", Prompt: fmt.Sprintf("Review this code and suggest improvements:\n%s", codeResult.Text), }) if err != nil { return "", err } result = fmt.Sprintf("Code:\n%s\n\nReview:\n%s", codeResult.Text, reviewResult.Text) case "content": // Sequential content creation with quality check contentResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a professional content writer.", Prompt: task, }) if err != nil { return "", err } result = contentResult.Text } return result, nil } func main() { ctx := context.Background() tasks := []string{ "Research the impact of quantum computing on cryptography", "Write a function to validate email addresses in Go", "Create engaging social media content for a new productivity app", } for _, task := range tasks { fmt.Printf("\n=== Processing Task ===\n%s\n\n", task) result, err := intelligentTaskProcessor(ctx, task) if err != nil { log.Printf("Error processing task: %v", err) continue } fmt.Printf("Result:\n%s\n", result) } } ``` ## Best Practices ### 1. Start Simple Begin with the simplest pattern that solves your problem: ```go // Good - Simple sequential for straightforward tasks func processDocument(ctx context.Context, doc string) (string, error) { summary, err := generateSummary(ctx, doc) if err != nil { return "", err } return enhanceSummary(ctx, summary) } ``` ### 2. Use Parallel Processing for Independent Tasks Run independent operations concurrently: ```go // Good - Parallel for independent analyses var wg sync.WaitGroup results := make([]string, 3) for i, task := range independentTasks { wg.Add(1) go func(idx int, t string) { defer wg.Done() results[idx], _ = processTask(ctx, t) }(i, task) } wg.Wait() ``` ### 3. Add Error Handling Always handle errors properly in workflows: ```go // Good - Proper error handling with context errChan := make(chan error, len(tasks)) for _, task := range tasks { go func(t string) { if err := process(ctx, t); err != nil { errChan <- err } }(task) } for err := range errChan { if err != nil { return fmt.Errorf("workflow failed: %w", err) } } ``` ### 4. Use Context for Cancellation Respect context cancellation in all steps: ```go ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute) defer cancel() // All operations respect this context result, err := multiStepWorkflow(ctx, input) ``` ### 5. Monitor and Debug Add logging for complex workflows: ```go func processWithLogging(ctx context.Context, input string) (string, error) { log.Printf("Starting workflow for input: %s", input) result, err := step1(ctx, input) if err != nil { log.Printf("Step 1 failed: %v", err) return "", err } log.Printf("Step 1 completed") // Continue with more steps... return result, nil } ``` ## Next Steps - Learn about [loop control](https://goaisdk.com/docs/agents/loop-control.md) for advanced agent execution control - See [configuring call options](https://goaisdk.com/docs/agents/configuring-call-options.md) for fine-tuning agent behavior - Explore [error handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) for robust workflows --- # Loop Control > Control agent execution with stop conditions, PrepareCall, and manual loop patterns Canonical URL: https://goaisdk.com/docs/agents/loop-control Documentation index: https://goaisdk.com/llms.txt The Go AI SDK provides several ways to control how long an agent's tool-calling loop runs. The recommended approach is `StopWhen` — a declarative list of stop conditions evaluated after each step that contains tool results. For lower-level control, `PrepareCall` lets you modify each generation call dynamically. Manual loops using `ai.GenerateText` directly give you full control for advanced use cases. --- ## Stop Conditions `StopWhen` accepts a slice of `StopCondition` functions. After every step that produces tool results, the SDK evaluates the conditions and stops the loop as soon as one returns a non-empty reason string. The reason is available as `result.StopReason`. ### IsStepCount The most common condition: stop after *n* steps. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Analyze this dataset and create a summary report", Tools: tools, StopWhen: []ai.StopCondition{ ai.IsStepCount(5), }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) fmt.Printf("Stopped because: %s\n", result.StopReason) // → "Stopped because: maximum number of steps (5) reached" ``` `IsStepCount(n)` stops the loop once exactly *n* steps have completed. Because it is evaluated after every step, it is checked against every step count from 1 up to *n* and never skips the threshold. ### HasToolCall Stop when the model calls a specific tool — useful for semantic completion signals where the model is expected to invoke a "finish" or "submit" tool when it's done. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Research the topic and call the finish tool when done", Tools: []types.Tool{ researchTool, finishTool, // model calls this to signal completion }, StopWhen: []ai.StopCondition{ ai.HasToolCall("finish"), ai.IsStepCount(10), // safety ceiling }, }) ``` ### Combining Conditions Conditions are evaluated OR — the loop stops on the **first match**. Place conditions you care about most, or that have side effects, **before** safety ceilings like `IsStepCount`. ```go StopWhen: []ai.StopCondition{ ai.HasToolCall("submit"), // semantic completion ai.IsStepCount(20), // hard limit }, ``` ### Custom Closure Any function with signature `func(ai.StopConditionState) string` qualifies as a `StopCondition`. Return a non-empty string to stop, or an empty string to continue. ```go // Token-budget guard — placed BEFORE IsStepCount so it always runs. // Because evaluation is eager (all conditions run before the first match is // returned), this condition fires even if IsStepCount would also match. tokenBudget := func(state ai.StopConditionState) string { if state.Usage.GetTotalTokens() > 5000 { return fmt.Sprintf("token budget exceeded (%d tokens)", state.Usage.GetTotalTokens()) } return "" } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Perform a deep analysis", Tools: tools, StopWhen: []ai.StopCondition{ tokenBudget, // side-effectful or priority condition first ai.IsStepCount(10), // safety ceiling last }, }) ``` `StopConditionState` gives you full context: ```go type StopConditionState struct { // Steps completed so far; the step that just finished is the last element. Steps []types.StepResult // Full message history including the latest tool-result messages. Messages []types.Message // Accumulated token usage across all steps. Usage types.Usage } ``` ### Condition Evaluation Order and Side Effects > **Important:** All conditions are always evaluated before the first match is returned. > This matches the TypeScript SDK's `Promise.all` behavior and ensures side-effectful > conditions (e.g. recording metrics, sending alerts) always execute regardless of their > position in the slice. Place conditions you want to guarantee run **before** conditions that might pre-empt them, and rely on the evaluation guarantee rather than ordering for correctness. ### StopReason `result.StopReason` (on `GenerateTextResult`) and `agentResult.StopReason` (on `AgentResult`) contain the reason string returned by whichever condition fired. Empty if the loop ended naturally (no tool calls) or no `StopWhen` was set. ```go result, err := myAgent.Execute(ctx, "Complete the task") if err != nil { log.Fatal(err) } switch { case result.StopReason != "": fmt.Printf("Stopped by condition: %s\n", result.StopReason) case result.FinishReason == types.FinishReasonStop: fmt.Println("Agent finished naturally") } ``` ### Defaults | Scenario | Behavior | |----------|----------| | Neither `StopWhen` nor `MaxSteps` set | `ai.GenerateText` / `ai.StreamText` default to `IsStepCount(1)`: one model call, whose tool calls still execute (as in TS). `agent.NewToolLoopAgent` defaults to `IsStepCount(20)`. | | `MaxSteps` set, `StopWhen` not set | Converts to `StopWhen{IsStepCount(MaxSteps)}` | | Both set | `StopWhen` takes precedence; `MaxSteps` is ignored | ### MaxSteps (Deprecated) `MaxSteps` is shorthand for `StopWhen: []ai.StopCondition{ai.IsStepCount(n)}`. It is still supported for backwards compatibility but **new code should use `StopWhen` directly**. ```go // Deprecated — works but prefer StopWhen maxSteps := 5 ai.GenerateTextOptions{MaxSteps: &maxSteps} // Preferred ai.GenerateTextOptions{StopWhen: []ai.StopCondition{ai.IsStepCount(5)}} ``` --- ## PrepareCall `PrepareCall` is called immediately before each generation step. Use it to adjust the system prompt, modify available tools, or change model parameters based on how many steps have already run or how many tokens have been consumed. ```go myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(10)}, PrepareCall: func(ctx context.Context, config agent.PrepareCallConfig) agent.PrepareCallConfig { // Switch to a more concise prompt after the first few steps if config.StepNumber > 3 { config.System = "Be concise. Finalize your answer." } // Restrict available tools in later steps if config.StepNumber > 7 { config.Tools = finalizeTools } return config }, }) ``` `PrepareCallConfig` exposes: | Field | Type | Description | |-------|------|-------------| | `StepNumber` | `int` | Zero-based index of the upcoming step | | `System` | `string` | System prompt for this call | | `Messages` | `[]types.Message` | Conversation history so far | | `Tools` | `[]types.Tool` | Tool set for this call | | `Temperature` | `*float64` | Sampling temperature override | | `MaxTokens` | `*int` | Token limit override | | `AccumulatedUsage` | `types.Usage` | Total tokens used so far | | `CustomData` | `interface{}` | Pass-through state between invocations | --- ## Manual Loop For cases where you need complete control outside the agent abstraction — dynamic model selection, custom context trimming, interleaved human input — you can implement the loop yourself using `ai.GenerateText`. ```go func manualAgentLoop(ctx context.Context, prompt string, tools []types.Tool) (string, error) { provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-6-astra") messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: prompt}}, }, } for step := 0; step < 10; step++ { select { case <-ctx.Done(): return "", ctx.Err() default: } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: messages, Tools: tools, }) if err != nil { return "", fmt.Errorf("step %d failed: %w", step, err) } // Natural completion — no tool calls if len(result.ToolCalls) == 0 { return result.Text, nil } // Append assistant turn messages = append(messages, types.Message{ Role: types.RoleAssistant, Content: []types.ContentPart{types.TextContent{Text: result.Text}}, }) // Execute tools and append results for _, toolCall := range result.ToolCalls { var toolResult interface{} var toolErr error for _, tool := range tools { if tool.Name == toolCall.ToolName { toolResult, toolErr = tool.Execute(ctx, toolCall.Arguments, types.ToolExecutionOptions{}) break } } resultContent := map[string]interface{}{"result": toolResult} if toolErr != nil { resultContent = map[string]interface{}{"error": toolErr.Error()} } messages = append(messages, types.Message{ Role: types.RoleTool, Content: []types.ContentPart{ types.ToolResultContent{ ToolCallID: toolCall.ID, ToolName: toolCall.ToolName, Result: resultContent, }, }, }) } } return "", fmt.Errorf("reached maximum steps without completion") } ``` ### When to Use a Manual Loop Use a manual loop when you need: - **Dynamic model selection** — swap models between steps based on complexity - **Context trimming** — prune the message history to stay within token limits - **Tool phase control** — expose different tool sets in different phases - **Interleaved human input** — pause for approval between steps For standard tool-calling agents, `StopWhen` + `PrepareCall` cover most needs without the boilerplate. --- ## See Also - [Stop-When Example](https://github.com/digitallysavvy/go-ai/blob/main/examples/agents/callbacks/stop-when/main.go) — focused example showing all three stop patterns - [Configuring Call Options](https://goaisdk.com/docs/agents/configuring-call-options.md) — temperature, max tokens, and other options - [Workflow Patterns](https://goaisdk.com/docs/agents/workflows.md) — structured agent designs - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) — robust error handling --- # Configuring Call Options > Shows how to pass runtime inputs to agents in Go using factory functions, so agent settings can be dynamically configured per request in HTTP handlers. Canonical URL: https://goaisdk.com/docs/agents/configuring-call-options Documentation index: https://goaisdk.com/llms.txt In Go, you can dynamically configure agent behavior by creating agent factory functions that accept runtime parameters. This pattern allows you to modify agent settings based on specific request context. ## Why Use Dynamic Configuration? When you need agent behavior to change based on runtime context: - **Add dynamic context** - Inject retrieved documents, user preferences, or session data into prompts - **Select models dynamically** - Choose faster or more capable models based on request complexity - **Configure tools per request** - Pass user location to search tools or adjust tool behavior - **Customize generation settings** - Set temperature, max tokens, or other provider-specific settings Without dynamic configuration, you'd need to create multiple agents or handle configuration logic outside the agent. ## How It Works Dynamic configuration in Go uses factory functions that: 1. **Accept runtime parameters** - Take configuration options as function parameters 2. **Build agent configuration** - Construct `AgentConfig` based on those parameters 3. **Return configured agent** - Create and return the agent ready for use ## Basic Example Create an agent factory that injects user context at runtime: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type UserContext struct { UserID string AccountType string // "free", "pro", or "enterprise" } func newSupportAgent(ctx UserContext) *agent.ToolLoopAgent { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") systemPrompt := fmt.Sprintf(`You are a helpful customer support agent. User context: - Account type: %s - User ID: %s Adjust your response based on the user's account level.`, ctx.AccountType, ctx.UserID) return agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: systemPrompt, Tools: []types.Tool{}, // Your tools here }) } func main() { ctx := context.Background() // Create agent with specific user context supportAgent := newSupportAgent(UserContext{ UserID: "user_123", AccountType: "free", }) result, err := supportAgent.Execute(ctx, "How do I upgrade my account?") if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Modifying Agent Settings Create factory functions that modify different agent settings based on runtime parameters. ### Dynamic Model Selection Choose models based on request characteristics: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type QueryComplexity string const ( ComplexitySimple QueryComplexity = "simple" ComplexityComplex QueryComplexity = "complex" ) func newAdaptiveAgent(complexity QueryComplexity) (*agent.ToolLoopAgent, error) { openaiProvider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) // Choose model based on complexity var model provider.LanguageModel var err error if complexity == ComplexitySimple { model, err = openaiProvider.LanguageModel("gpt-5.4-mini") fmt.Println("Using fast model for simple query") } else { model, err = openaiProvider.LanguageModel("gpt-6-astra") fmt.Println("Using powerful model for complex query") } if err != nil { return nil, err } return agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, MaxSteps: 15, }), nil } func main() { ctx := context.Background() // Use faster model for simple queries simpleAgent, _ := newAdaptiveAgent(ComplexitySimple) result1, err := simpleAgent.Execute(ctx, "What is 2+2?") if err != nil { log.Fatal(err) } fmt.Println("Simple answer:", result1.Text) // Use more capable model for complex reasoning complexAgent, _ := newAdaptiveAgent(ComplexityComplex) result2, err := complexAgent.Execute(ctx, "Explain quantum entanglement and its implications for computing") if err != nil { log.Fatal(err) } fmt.Println("Complex answer:", result2.Text) } ``` ### Dynamic Tool Configuration Configure tools based on runtime context using closures: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type Location struct { City string Region string Country string } func newNewsAgent(userLocation Location) *agent.ToolLoopAgent { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Create tool with user location captured in closure searchTool := types.Tool{ Name: "webSearch", Description: fmt.Sprintf("Search for news in %s, %s", userLocation.City, userLocation.Region), Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{ "type": "string", "description": "The search query", }, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := input["query"].(string) // Use location in search results := fmt.Sprintf("Searching for '%s' in %s, %s, %s", query, userLocation.City, userLocation.Region, userLocation.Country) // Actual implementation would call search API with location return map[string]interface{}{ "results": results, }, nil }, } return agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: fmt.Sprintf("You are a news assistant providing information for users in %s, %s.", userLocation.City, userLocation.Region), Tools: []types.Tool{searchTool}, }) } func main() { ctx := context.Background() newsAgent := newNewsAgent(Location{ City: "San Francisco", Region: "California", Country: "US", }) result, err := newsAgent.Execute(ctx, "What are the top local news stories?") if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Dynamic Generation Settings Configure temperature, max tokens, and other settings: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type GenerationConfig struct { Temperature *float64 MaxTokens *int MaxSteps int } func newConfigurableAgent(config GenerationConfig) *agent.ToolLoopAgent { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") return agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Temperature: config.Temperature, MaxTokens: config.MaxTokens, MaxSteps: config.MaxSteps, }) } func main() { ctx := context.Background() // Creative writing configuration creativeTemp := 0.9 creativeTokens := 1000 creativeAgent := newConfigurableAgent(GenerationConfig{ Temperature: &creativeTemp, MaxTokens: &creativeTokens, MaxSteps: 10, }) result1, err := creativeAgent.Execute(ctx, "Write a creative story") if err != nil { log.Fatal(err) } fmt.Println("Creative output:", result1.Text) // Precise, deterministic configuration preciseTemp := 0.1 preciseTokens := 500 preciseAgent := newConfigurableAgent(GenerationConfig{ Temperature: &preciseTemp, MaxTokens: &preciseTokens, MaxSteps: 5, }) result2, err := preciseAgent.Execute(ctx, "Calculate the answer precisely") if err != nil { log.Fatal(err) } fmt.Println("Precise output:", result2.Text) } ``` ## Advanced Patterns ### Retrieval Augmented Generation (RAG) Fetch relevant context and inject it into your agent: ```go package main import ( "context" "fmt" "log" "os" "strings" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) type Document struct { ID string Content string } func vectorSearch(ctx context.Context, query string) ([]Document, error) { // Simulated vector search - in production, use a vector database docs := []Document{ { ID: "doc1", Content: "Our refund policy allows returns within 30 days of purchase for a full refund.", }, { ID: "doc2", Content: "Refunds are processed within 5-7 business days to the original payment method.", }, { ID: "doc3", Content: "To request a refund, contact support@example.com with your order number.", }, } return docs, nil } func newRAGAgent(ctx context.Context, query string) (*agent.ToolLoopAgent, error) { provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") // Fetch relevant documents documents, err := vectorSearch(ctx, query) if err != nil { return nil, err } // Build context from documents var contextParts []string for _, doc := range documents { contextParts = append(contextParts, doc.Content) } context := strings.Join(contextParts, "\n\n") systemPrompt := fmt.Sprintf(`Answer questions using the following context: %s If the answer is not in the context, say so clearly.`, context) return agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: systemPrompt, }), nil } func main() { ctx := context.Background() query := "What is our refund policy?" ragAgent, err := newRAGAgent(ctx, query) if err != nil { log.Fatal(err) } result, err := ragAgent.Execute(ctx, query) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Combining Multiple Modifications Modify multiple settings together based on user role and urgency: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type RequestConfig struct { UserRole string // "admin" or "user" Urgency string // "low" or "high" } func readDatabaseTool() types.Tool { return types.Tool{ Name: "readDatabase", Description: "Read data from the database", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{ "type": "string", }, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{"data": "..."}, nil }, } } func writeDatabaseTool() types.Tool { return types.Tool{ Name: "writeDatabase", Description: "Write data to the database", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "data": map[string]interface{}{ "type": "object", }, }, "required": []string{"data"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{"success": true}, nil }, } } func newConfiguredAgent(config RequestConfig) *agent.ToolLoopAgent { openaiProvider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) // Select model based on urgency var model provider.LanguageModel if config.Urgency == "high" { model, _ = openaiProvider.LanguageModel("gpt-6-astra") // More capable model } else { model, _ = openaiProvider.LanguageModel("gpt-5.4-mini") // Faster, cheaper model } // Configure tools based on user role var tools []types.Tool if config.UserRole == "admin" { tools = []types.Tool{ readDatabaseTool(), writeDatabaseTool(), } } else { tools = []types.Tool{ readDatabaseTool(), } } // Build system prompt systemPrompt := fmt.Sprintf("You are a %s assistant.\n", config.UserRole) if config.UserRole == "admin" { systemPrompt += "You have full database access." } else { systemPrompt += "You have read-only access." } return agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: systemPrompt, Tools: tools, }) } func main() { ctx := context.Background() // Admin with high urgency adminAgent := newConfiguredAgent(RequestConfig{ UserRole: "admin", Urgency: "high", }) result, err := adminAgent.Execute(ctx, "Update the user record") if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Session-Based Configuration Create agents with session-specific configuration: ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type SessionContext struct { SessionID string UserPreferences map[string]string ConversationHistory []types.Message CreatedAt time.Time } func newSessionAgent(session SessionContext) *agent.ToolLoopAgent { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Build personalized system prompt systemPrompt := fmt.Sprintf(`You are a helpful assistant. Session ID: %s User preferences: %v Session started: %s Tailor your responses to the user's preferences.`, session.SessionID, session.UserPreferences, session.CreatedAt.Format("2006-01-02 15:04:05")) return agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: systemPrompt, }) } func main() { ctx := context.Background() session := SessionContext{ SessionID: "sess_123", UserPreferences: map[string]string{ "language": "English", "tone": "professional", "detail": "concise", }, ConversationHistory: []types.Message{}, CreatedAt: time.Now(), } sessionAgent := newSessionAgent(session) result, err := sessionAgent.Execute(ctx, "Tell me about your capabilities") if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Multi-Tenant Configuration Configure agents for different tenants or organizations: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) type TenantConfig struct { TenantID string OrganizationName string CustomBranding string AllowedFeatures []string MaxSteps int } func newTenantAgent(tenant TenantConfig) *agent.ToolLoopAgent { provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") systemPrompt := fmt.Sprintf(`You are an AI assistant for %s. %s Available features: %v Always maintain professionalism and follow the organization's guidelines.`, tenant.OrganizationName, tenant.CustomBranding, tenant.AllowedFeatures) // Filter tools based on allowed features var tools []types.Tool for _, feature := range tenant.AllowedFeatures { switch feature { case "search": tools = append(tools, searchTool()) case "analytics": tools = append(tools, analyticsTool()) case "reporting": tools = append(tools, reportingTool()) } } return agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: systemPrompt, Tools: tools, MaxSteps: tenant.MaxSteps, }) } func searchTool() types.Tool { return types.Tool{ Name: "search", Description: "Search for information", // ... tool definition } } func analyticsTool() types.Tool { return types.Tool{ Name: "analytics", Description: "Analyze data", // ... tool definition } } func reportingTool() types.Tool { return types.Tool{ Name: "reporting", Description: "Generate reports", // ... tool definition } } func main() { ctx := context.Background() tenantConfig := TenantConfig{ TenantID: "tenant_xyz", OrganizationName: "Acme Corporation", CustomBranding: "Acme values efficiency and innovation.", AllowedFeatures: []string{"search", "analytics"}, MaxSteps: 15, } tenantAgent := newTenantAgent(tenantConfig) result, err := tenantAgent.Execute(ctx, "Analyze our sales data") if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Using in HTTP Handlers Integrate dynamic agent configuration in web applications: ```go package main import ( "encoding/json" "fmt" "log" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type ChatRequest struct { Prompt string `json:"prompt"` UserID string `json:"userId"` AccountType string `json:"accountType"` } type ChatResponse struct { Text string `json:"text"` Steps int `json:"steps"` } func chatHandler(w http.ResponseWriter, r *http.Request) { var req ChatRequest if err := json.NewDecoder(r.Body).Decode(&req); err != nil { http.Error(w, err.Error(), http.StatusBadRequest) return } // Create agent with user context supportAgent := newSupportAgentFromRequest(req) // Execute agent result, err := supportAgent.Execute(r.Context(), req.Prompt) if err != nil { http.Error(w, err.Error(), http.StatusInternalServerError) return } // Return response response := ChatResponse{ Text: result.Text, Steps: len(result.Steps), } w.Header().Set("Content-Type", "application/json") json.NewEncoder(w).Encode(response) } func newSupportAgentFromRequest(req ChatRequest) *agent.ToolLoopAgent { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") systemPrompt := fmt.Sprintf(`You are a customer support agent. User: %s (Account: %s) Adjust responses based on the account level.`, req.UserID, req.AccountType) return agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: systemPrompt, }) } func main() { http.HandleFunc("/api/chat", chatHandler) fmt.Println("Server listening on :8080") log.Fatal(http.ListenAndServe(":8080", nil)) } ``` ## Best Practices ### 1. Use Factory Functions Create factory functions for reusable agent configurations: ```go // Good - Reusable factory function func newAgentWithContext(ctx RequestContext) *agent.ToolLoopAgent { // Configuration logic return agent.NewToolLoopAgent(config) } // Bad - Inline configuration everywhere myAgent := agent.NewToolLoopAgent(agent.AgentConfig{Model: model}) ``` ### 2. Validate Configuration Validate configuration parameters before creating agents: ```go func newValidatedAgent(config Config) (*agent.ToolLoopAgent, error) { if config.UserRole != "admin" && config.UserRole != "user" { return nil, fmt.Errorf("invalid user role: %s", config.UserRole) } if config.MaxSteps < 1 || config.MaxSteps > 50 { return nil, fmt.Errorf("maxSteps must be between 1 and 50") } // Create agent with validated config return agent.NewToolLoopAgent(agent.AgentConfig{Model: model}), nil } ``` ### 3. Use Closures for Tool Configuration Capture runtime context in tool closures: ```go func newAgentWithUserTools(userID string) *agent.ToolLoopAgent { tool := types.Tool{ Name: "getUserData", Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // userID is captured from outer scope return fetchUserData(ctx, userID) }, } return agent.NewToolLoopAgent(agent.AgentConfig{ Tools: []types.Tool{tool}, }) } ``` ### 4. Handle Errors Properly Always check for errors when creating agents: ```go func newAgent(complexity string) (*agent.ToolLoopAgent, error) { model, err := getModel(complexity) if err != nil { return nil, fmt.Errorf("failed to get model: %w", err) } return agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, }), nil } ``` ### 5. Document Configuration Options Document what configuration options are available: ```go // AgentOptions configures agent behavior type AgentOptions struct { // Model complexity: "simple" uses fast models, "complex" uses powerful models Complexity string // Maximum number of reasoning steps (1-50) MaxSteps int // User context for personalization UserContext UserContext } ``` ## Next Steps - Learn about [loop control](https://goaisdk.com/docs/agents/loop-control.md) for execution management - Explore [workflow patterns](https://goaisdk.com/docs/agents/workflows.md) for complex multi-step processes - See [building agents](https://goaisdk.com/docs/agents/building-agents.md) for agent fundamentals --- # Agent Callbacks > Reference for Go AI SDK agent lifecycle callbacks, covering both legacy and LangChain-style callback types, common use cases, data types, and run tracking. Canonical URL: https://goaisdk.com/docs/agents/callbacks Documentation index: https://goaisdk.com/llms.txt Callbacks provide hooks into agent execution, allowing you to monitor progress, log activities, handle errors, and integrate with external systems. The Go AI SDK supports both legacy callbacks and LangChain-style callbacks for enhanced interoperability. ## Overview Agent callbacks enable you to: - **Monitor execution**: Track agent progress and decisions - **Log activities**: Record tool calls, actions, and results - **Handle errors**: Implement custom error recovery logic - **Measure performance**: Track token usage and execution time - **Integrate with observability**: Connect to logging and monitoring platforms - **Debug agent behavior**: Understand decision-making processes ## Callback Types ### Legacy Callbacks These callbacks provide basic step-level tracking: | Callback | Invoked When | Parameters | |----------|-------------|------------| | `OnStepStart` | Step begins | `stepNum int` | | `OnStepFinish` | Step completes | `step types.StepResult` | | `OnToolCall` | Tool is called | `toolCall types.ToolCall` | | `OnToolResult` | Tool execution completes | `toolResult types.ToolResult` | | `OnFinish` | Agent completes | `result *AgentResult` | ### LangChain-Style Callbacks (v6.0.60+) These callbacks provide fine-grained control and align with LangChain's callback system: #### Chain Lifecycle | Callback | Invoked When | Parameters | |----------|-------------|------------| | `OnChainStart` | Agent begins execution | `input string, messages []types.Message` | | `OnChainEnd` | Agent completes successfully | `result *AgentResult` | | `OnChainError` | Agent encounters an error | `err error` | #### Agent Decisions | Callback | Invoked When | Parameters | |----------|-------------|------------| | `OnAgentAction` | Agent decides to take an action | `action AgentAction` | | `OnAgentFinish` | Agent reaches a final answer | `finish AgentFinish` | #### Tool Lifecycle | Callback | Invoked When | Parameters | |----------|-------------|------------| | `OnToolStart` | Tool execution begins | `toolCall types.ToolCall` | | `OnToolEnd` | Tool completes successfully | `toolResult types.ToolResult` | | `OnToolError` | Tool execution fails | `toolCall types.ToolCall, err error` | ## Basic Usage ```go import ( "context" "log" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful assistant.", Tools: tools, // Legacy callbacks OnStepStart: func(stepNum int) { log.Printf("Starting step %d", stepNum) }, OnStepFinish: func(step types.StepResult) { log.Printf("Step %d finished: %s", step.StepNumber, step.FinishReason) }, // LangChain-style callbacks OnChainStart: func(input string, messages []types.Message) { log.Printf("Agent starting with input: %s", input) }, OnChainEnd: func(result *agent.AgentResult) { log.Printf("Agent completed in %d steps", len(result.Steps)) }, OnChainError: func(err error) { log.Printf("Agent error: %v", err) }, }) ``` ## Common Use Cases ### Logging Agent Execution Track the complete execution flow: ```go import ( "fmt" "time" ) var startTime time.Time config := agent.AgentConfig{ Model: model, Tools: tools, OnChainStart: func(input string, messages []types.Message) { startTime = time.Now() fmt.Printf("\n=== Agent Started ===\n") fmt.Printf("Input: %s\n", input) }, OnAgentAction: func(action agent.AgentAction) { fmt.Printf("\n[Step %d] Agent Action\n", action.StepNumber) fmt.Printf(" Tool: %s\n", action.ToolCall.ToolName) if action.Reasoning != "" { fmt.Printf(" Reasoning: %s\n", action.Reasoning) } }, OnAgentFinish: func(finish agent.AgentFinish) { fmt.Printf("\n[Step %d] Agent Finished\n", finish.StepNumber) fmt.Printf(" Output: %s\n", finish.Output) fmt.Printf(" Reason: %s\n", finish.FinishReason) }, OnChainEnd: func(result *agent.AgentResult) { duration := time.Since(startTime) fmt.Printf("\n=== Agent Completed ===\n") fmt.Printf("Duration: %s\n", duration) fmt.Printf("Steps: %d\n", len(result.Steps)) fmt.Printf("Tool calls: %d\n", len(result.ToolResults)) fmt.Printf("Total tokens: %d\n", result.Usage.GetTotalTokens()) }, } ``` ### Tracking Token Usage and Costs Monitor API usage for cost tracking: ```go type UsageTracker struct { totalTokens int64 totalCost float64 toolCallCount int costPerToken float64 // Cost per 1000 tokens } func (t *UsageTracker) addUsage(tokens int64) { t.totalTokens += tokens t.totalCost += float64(tokens) * t.costPerToken / 1000.0 } tracker := &UsageTracker{ costPerToken: 0.002, // $0.002 per 1K tokens (example rate) } config := agent.AgentConfig{ Model: model, Tools: tools, OnStepFinish: func(step types.StepResult) { if step.Usage.GetTotalTokens() > 0 { tracker.addUsage(step.Usage.GetTotalTokens()) } }, OnToolEnd: func(toolResult types.ToolResult) { tracker.toolCallCount++ }, OnChainEnd: func(result *agent.AgentResult) { fmt.Printf("\n=== Usage Summary ===\n") fmt.Printf("Total tokens: %d\n", tracker.totalTokens) fmt.Printf("Estimated cost: $%.4f\n", tracker.totalCost) fmt.Printf("Tool calls: %d\n", tracker.toolCallCount) fmt.Printf("Average tokens per step: %d\n", tracker.totalTokens/int64(len(result.Steps))) }, } ``` ### Error Handling and Recovery Implement custom error handling: ```go config := agent.AgentConfig{ Model: model, Tools: tools, OnChainError: func(err error) { log.Printf("Chain error occurred: %v", err) // Send alert sendAlert("Agent execution failed", err.Error()) // Log to external service logToObservability("agent.error", map[string]interface{}{ "error": err.Error(), "timestamp": time.Now(), }) }, OnToolError: func(toolCall types.ToolCall, err error) { log.Printf("Tool %s failed: %v", toolCall.ToolName, err) // Track tool failures metrics.IncrementCounter("tool.failures", map[string]string{ "tool_name": toolCall.ToolName, }) // Implement retry logic elsewhere if needed // (Note: callbacks should not modify execution flow directly) }, } ``` ### Performance Monitoring Track execution time for each component: ```go type PerformanceMonitor struct { toolTimes map[string]time.Duration toolStarts map[string]time.Time stepTimes []time.Duration stepStart time.Time } func NewPerformanceMonitor() *PerformanceMonitor { return &PerformanceMonitor{ toolTimes: make(map[string]time.Duration), toolStarts: make(map[string]time.Time), stepTimes: make([]time.Duration, 0), } } monitor := NewPerformanceMonitor() config := agent.AgentConfig{ Model: model, Tools: tools, OnStepStart: func(stepNum int) { monitor.stepStart = time.Now() }, OnStepFinish: func(step types.StepResult) { duration := time.Since(monitor.stepStart) monitor.stepTimes = append(monitor.stepTimes, duration) }, OnToolStart: func(toolCall types.ToolCall) { monitor.toolStarts[toolCall.ID] = time.Now() }, OnToolEnd: func(toolResult types.ToolResult) { if startTime, ok := monitor.toolStarts[toolResult.ToolCallID]; ok { duration := time.Since(startTime) monitor.toolTimes[toolResult.ToolName] += duration delete(monitor.toolStarts, toolResult.ToolCallID) } }, OnChainEnd: func(result *agent.AgentResult) { fmt.Printf("\n=== Performance Report ===\n") // Step timings var totalStepTime time.Duration for i, duration := range monitor.stepTimes { fmt.Printf("Step %d: %s\n", i+1, duration) totalStepTime += duration } fmt.Printf("Average step time: %s\n", totalStepTime/time.Duration(len(monitor.stepTimes))) // Tool timings fmt.Printf("\nTool execution times:\n") for toolName, duration := range monitor.toolTimes { fmt.Printf(" %s: %s\n", toolName, duration) } }, } ``` ### Integration with Observability Platforms Connect to external monitoring systems: ```go import ( "go.opentelemetry.io/otel" "go.opentelemetry.io/otel/trace" ) tracer := otel.Tracer("agent-execution") config := agent.AgentConfig{ Model: model, Tools: tools, OnChainStart: func(input string, messages []types.Message) { ctx, span := tracer.Start(context.Background(), "agent.execute") span.SetAttributes(attribute.String("input", input)) // Store ctx for later use }, OnAgentAction: func(action agent.AgentAction) { _, span := tracer.Start(context.Background(), "agent.action") span.SetAttributes( attribute.String("tool", action.ToolCall.ToolName), attribute.Int("step", action.StepNumber), ) defer span.End() }, OnToolStart: func(toolCall types.ToolCall) { _, span := tracer.Start(context.Background(), "tool.execute") span.SetAttributes(attribute.String("tool", toolCall.ToolName)) // Store span for completion in OnToolEnd }, OnChainEnd: func(result *agent.AgentResult) { // Complete the main span // span.SetAttributes(...) // span.End() }, } ``` ## Data Types ### AgentAction Represents an action the agent has decided to take: ```go type AgentAction struct { ToolCall types.ToolCall // The tool being called StepNumber int // Step when action was decided Reasoning string // Agent's reasoning for the action RunID string // Unique identifier for this agent execution chain ParentRunID string // ID of the parent run (empty for top-level runs) Tags []string // User-defined labels for categorizing runs } ``` ### AgentFinish Represents the agent's final decision: ```go type AgentFinish struct { Output string // Final output text StepNumber int // Step when agent finished FinishReason types.FinishReason // Why the agent finished Metadata map[string]interface{} // Additional metadata RunID string // Unique identifier for this agent execution chain ParentRunID string // ID of the parent run (empty for top-level runs) Tags []string // User-defined labels for categorizing runs } ``` ## Best Practices ### 1. Keep Callbacks Fast Callbacks are invoked during execution. Long-running operations can slow down the agent: ```go // ❌ Bad: Slow synchronous operation OnToolEnd: func(toolResult types.ToolResult) { saveToDatabaseSync(toolResult) // Blocks execution }, // ✅ Good: Async logging OnToolEnd: func(toolResult types.ToolResult) { go saveToDatabaseAsync(toolResult) // Non-blocking }, ``` ### 2. Handle Errors in Callbacks Panics in callbacks can crash your application: ```go OnChainEnd: func(result *agent.AgentResult) { defer func() { if r := recover(); r != nil { log.Printf("Callback panic: %v", r) } }() // Your callback code }, ``` ### 3. Don't Modify Agent State Callbacks should observe, not modify: ```go // ❌ Bad: Trying to modify execution OnAgentAction: func(action agent.AgentAction) { // Don't try to cancel or modify the action }, // ✅ Good: Just observe and log OnAgentAction: func(action agent.AgentAction) { log.Printf("Agent is calling: %s", action.ToolCall.ToolName) }, ``` ### 4. Use Appropriate Callback Level Choose between legacy and LangChain-style callbacks based on your needs: - **Legacy callbacks**: Simple step tracking, basic monitoring - **LangChain-style callbacks**: Detailed lifecycle management, LangChain compatibility, fine-grained control ### 5. Thread Safety If callbacks access shared state, ensure thread safety: ```go type SafeCounter struct { mu sync.Mutex count int } func (c *SafeCounter) Increment() { c.mu.Lock() defer c.mu.Unlock() c.count++ } counter := &SafeCounter{} config := agent.AgentConfig{ OnToolEnd: func(toolResult types.ToolResult) { counter.Increment() // Thread-safe }, } ``` ## Combining Callbacks You can use both legacy and LangChain-style callbacks together: ```go config := agent.AgentConfig{ Model: model, Tools: tools, // Legacy - for backward compatibility OnStepFinish: func(step types.StepResult) { // Your existing logic }, // LangChain-style - for new features OnChainStart: func(input string, messages []types.Message) { // Enhanced tracking }, OnAgentAction: func(action agent.AgentAction) { // Fine-grained action monitoring }, } ``` ## Run Tracking (v6.0.61+) Run tracking enables you to correlate all callbacks, actions, and events from a single agent execution using unique identifiers. This is particularly useful for observability, debugging, and distributed tracing. ### Run Tracking Fields Each `AgentAction` and `AgentFinish` includes: - **`RunID`**: Unique identifier for this agent execution chain - **`ParentRunID`**: ID of the parent run (for subagents/nested executions) - **`Tags`**: User-defined labels for categorizing runs ### Automatic RunID Generation Run IDs are automatically generated when an agent starts: ```go agent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, OnAgentAction: func(action agent.AgentAction) { // RunID is automatically populated log.Printf("Action in run %s: %s", action.RunID, action.ToolCall.ToolName) }, }) result, err := agent.Execute(context.Background(), "question") // RunID is automatically generated and propagated ``` ### Custom Run IDs and Tags Provide your own run ID and tags using context helpers: ```go import "github.com/digitallysavvy/go-ai/pkg/agent" ctx := context.Background() // Option 1: Add tags only (RunID auto-generated) ctx = agent.WithTags(ctx, []string{"production", "user:123", "session:abc"}) // Option 2: Custom RunID ctx = agent.WithRunID(ctx, "my-custom-run-id-123") ctx = agent.WithTags(ctx, []string{"experiment", "variant:A"}) // Option 3: Nested execution (subagent tracking) parentCtx := agent.WithRunID(ctx, "parent-run-456") subagentCtx := agent.WithParentRunID(parentCtx, "parent-run-456") subagentCtx = agent.WithRunID(subagentCtx, "subagent-run-789") result, err := myAgent.Generate(ctx, agent.AgentGenerateOptions{Prompt: "question"}) ``` ### Retrieving Run Tracking Info Extract run tracking information from context: ```go runID := agent.GetRunID(ctx) parentRunID := agent.GetParentRunID(ctx) tags := agent.GetTags(ctx) log.Printf("Executing run %s with tags: %v", runID, tags) ``` ### Use Cases for Run Tracking #### 1. Distributed Tracing Correlate agent execution with other services: ```go import "go.opentelemetry.io/otel/trace" // Get trace ID from OpenTelemetry span := trace.SpanFromContext(ctx) traceID := span.SpanContext().TraceID().String() // Use trace ID as run ID for correlation ctx = agent.WithRunID(ctx, traceID) ctx = agent.WithTags(ctx, []string{"service:agent", "env:production"}) config := agent.AgentConfig{ Model: model, OnAgentAction: func(action agent.AgentAction) { // Log with trace ID for correlation log.Printf("[trace=%s] Agent action: %s", action.RunID, action.ToolCall.ToolName) }, } ``` #### 2. Multi-Tenant Observability Track executions by user or session: ```go // Tag runs with user/session info ctx = agent.WithTags(ctx, []string{ fmt.Sprintf("user:%s", userID), fmt.Sprintf("session:%s", sessionID), "plan:premium", }) config := agent.AgentConfig{ OnAgentFinish: func(finish agent.AgentFinish) { // Store metrics tagged by user metrics.RecordAgentCompletion( finish.RunID, finish.Tags, len(finish.Metadata), ) }, } ``` #### 3. A/B Testing and Experiments Track different agent configurations: ```go experimentID := "exp-123-variant-A" ctx = agent.WithRunID(ctx, experimentID) ctx = agent.WithTags(ctx, []string{ "experiment:prompt-variations", "variant:A", "temperature:0.7", }) config := agent.AgentConfig{ OnChainEnd: func(result *agent.AgentResult) { // Log results for experiment analysis experimentTracker.LogResult(experimentID, result) }, } ``` #### 4. Subagent Hierarchies Track parent-child relationships in delegations: ```go func executeWithSubagent(parentCtx context.Context) { parentRunID := agent.GetRunID(parentCtx) // Create subagent context with parent tracking subagentCtx := agent.WithParentRunID(context.Background(), parentRunID) subagentCtx = agent.WithTags(subagentCtx, []string{"type:subagent", "role:researcher"}) subagent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, OnChainStart: func(input string, messages []types.Message) { runID := agent.GetRunID(subagentCtx) log.Printf("Subagent %s starting (parent: %s)", runID, parentRunID) }, }) result, err := subagent.Execute(subagentCtx, "research question") } ``` #### 5. Performance Analysis Group and analyze runs by tags: ```go ctx = agent.WithTags(ctx, []string{ "tool:search", "complexity:high", "priority:p0", }) config := agent.AgentConfig{ OnChainEnd: func(result *agent.AgentResult) { // Store metrics for later analysis analytics.RecordExecution(analytics.Execution{ RunID: agent.GetRunID(ctx), Tags: agent.GetTags(ctx), Duration: time.Since(startTime), TokensUsed: result.Usage.GetTotalTokens(), StepCount: len(result.Steps), }) }, } // Later: Query executions by tags highComplexityRuns := analytics.FindByTag("complexity:high") ``` ## Examples See complete examples in `examples/agents/callbacks/`: - **onstepfinish**: Basic step tracking with OnStepFinish - **early-stopping**: Token limit monitoring with callbacks - **langchain-style**: Comprehensive example of all LangChain-style callbacks (includes run tracking demo) ## Next Steps - Explore [Agent Workflows](https://goaisdk.com/docs/agents/workflows.md) for complex agent patterns - Learn about [Loop Control](https://goaisdk.com/docs/agents/loop-control.md) for execution management - See [Agent Configuration](https://goaisdk.com/docs/agents/configuring-call-options.md) for more options --- *Introduced in v6.0.60: LangChain-style callbacks (OnChainStart, OnChainEnd, OnChainError, OnAgentAction, OnAgentFinish, OnToolStart, OnToolEnd, OnToolError)* *Introduced in v6.0.61: Run tracking with RunID, ParentRunID, and Tags fields in AgentAction and AgentFinish* --- # WorkflowAgent > Documents pkg/workflow's WorkflowAgent for durable, serializable Go workflows, covering the constructor, Generate, Stream, and chat transport. Canonical URL: https://goaisdk.com/docs/agents/workflow-agent Documentation index: https://goaisdk.com/llms.txt `pkg/workflow` provides a workflow-oriented agent surface for durable execution scenarios. It mirrors the TypeScript `@ai-sdk/workflow` package with Go option structs instead of JavaScript overloads. ## Key Types - `workflow.WorkflowAgent` - `workflow.WorkflowResult` - `workflow.WorkflowStreamResult` - `workflow.WorkflowRunMultiplexer` - `workflow.WorkflowChatTransport` - `workflow.StreamTextIterator` - `workflow.SerializableToolDef` ## Constructor ```go agent, err := workflow.NewWorkflowAgent(workflow.WorkflowAgent{ ID: "support-agent", Model: model, Instructions: "You are a support assistant.", Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(20)}, ActiveTools: []string{"search"}, Telemetry: &ai.TelemetrySettings{FunctionID: "support-agent"}, }) ``` `System` is still accepted for compatibility, but `Instructions` is the preferred name. `Instructions` accepts a string, a system `types.Message`, or a slice of system messages. Tools can be supplied as either a `[]types.Tool` or a TypeScript-style keyed `ToolSet` using `map[string]types.Tool`; missing tool names are filled from the map key. `MaxSteps` is intentionally not part of `WorkflowAgent`; use `StopWhen` with `ai.IsStepCount(n)`. ## Generate ```go result, err := agent.GenerateWithOptions(ctx, workflow.WorkflowGenerateOptions{ Prompt: "Help me debug this issue", OnError: func(ctx context.Context, err error) { log.Printf("workflow error: %v", err) }, }) if err != nil { return err } _ = result.IsLoopFinished() _ = result.Output ``` When `Output` is an `ai.Output` specification, `GenerateWithOptions` forwards its response format to the model and parses the final text into `WorkflowResult.Output`. ## Stream ```go streamResult, err := agent.StreamWithOptions(ctx, workflow.WorkflowStreamOptions{ Prompt: "Walk me through this fix", ActiveTools: []string{"search", "final_answer"}, OnAbort: func(ctx context.Context, steps []types.StepResult) { log.Printf("aborted after %d steps", len(steps)) }, }) if err != nil { return err } ``` `ExperimentalTransform` applies an ordered list of `ai.StreamTransformFunc` to raw model stream chunks before they reach `OnChunk` or the returned stream, letting you rewrite or drop chunks (e.g. smoothing text, or replacing an oversized provider tool result): ```go redactLargeResults := func(ctx context.Context, chunk provider.StreamChunk) []provider.StreamChunk { if chunk.Type == provider.ChunkTypeToolResult && chunk.ToolResult != nil { if s, ok := chunk.ToolResult.Result.(string); ok && len(s) > 2048 { replaced := *chunk.ToolResult replaced.Result = map[string]interface{}{"reason": "too_large_for_stream"} chunk.ToolResult = &replaced } } return []provider.StreamChunk{chunk} } streamResult, err := agent.StreamWithOptions(ctx, workflow.WorkflowStreamOptions{ Prompt: "Walk me through this fix", ExperimentalTransform: []ai.StreamTransformFunc{redactLargeResults}, }) if err != nil { return err } ``` ## Chat Transport `WorkflowRunMultiplexer` is the server-side `http.Handler` that emits run-scoped SSE and supports reattaching to a run: ```go transport := workflow.NewWorkflowRunMultiplexer() http.Handle("/api/workflow", transport) http.Handle("/api/workflow/resume", transport.Resume(runID)) ``` `Resume` supports a `startIndex` query parameter (including negative tail-relative offsets). `SendMessages`/`ReconnectToStream` on `WorkflowRunMultiplexer` return `([]*streaming.SSEEvent, error)` and take `workflow.SendMessagesOptions`/`workflow.ReconnectToStreamOptions` respectively: ```go events, err := transport.SendMessages(ctx, workflow.SendMessagesOptions{ ChatID: "chat-1", Trigger: "submit-message", Messages: messages, }) events, err = transport.ReconnectToStream(ctx, workflow.ReconnectToStreamOptions{ RunID: "run-1", ChatID: "chat-1", StartIndex: -10, }) ``` `WorkflowChatTransport` is the separate Go client-side equivalent of the TypeScript `WorkflowChatTransport`: it implements `ai.ChatTransport`, POSTing a chat's message history to a workflow endpoint and streaming the response as `ai.UIMessageChunk` values over channels: ```go transport := workflow.NewWorkflowChatTransport(workflow.WorkflowChatTransportOptions{ API: "https://example.com/api/chat", InitialStartIndex: -10, // tail-relative offset used on reconnect }) chunks, errs := transport.SendMessages(ctx, ai.ChatTransportSendMessagesRequest{ ChatID: "chat-1", Trigger: "submit-message", Messages: messages, }) chunks, errs = transport.ReconnectToStream(ctx, ai.ChatTransportReconnectToStreamRequest{ ChatID: "chat-1", }) ``` ## Serializable Schemas And Tools Use `workflow.MarshalSchema` and `workflow.UnmarshalSchema` for JSON Schema round trips. Use `workflow.SerializeToolSet` to strip function fields from tool definitions before crossing a workflow boundary, then `workflow.ResolveSerializableTools` to restore non-executable tool descriptors in the resumed step. Use `workflow.ValidateSerializableToolInput` to validate reconstructed tool inputs against the serialized JSON Schema before executing an attached Go function. ## Provider Serialization `providerutils.SerializeModel` returns a JSON-safe map containing `provider`, `modelId`, and `config`. Provider auth fields and function-valued fields are omitted. Use `providerutils.DeserializeModel(providerID, modelID, config)` after importing the target provider package so its deserializer is registered. `providerutils.SerializeModel`/`DeserializeModel` are convenience wrappers around `provider.SerializeModel(provider.LanguageModel) (provider.SerializedModel, error)` and `provider.DeserializeModel(provider.SerializedModel) (provider.LanguageModel, error)` for the language model case specifically. Every other model kind that a provider can serialize has the same pair of functions and a matching `Register*Deserializer`, one set per kind — for every provider model TS marks `WORKFLOW_SERIALIZE`: ```go provider.SerializeImageModel / DeserializeImageModel / RegisterImageModelDeserializer provider.SerializeVideoModel / DeserializeVideoModel / RegisterVideoModelDeserializer provider.SerializeSpeechModel / DeserializeSpeechModel / RegisterSpeechModelDeserializer provider.SerializeTranscriptionModel / DeserializeTranscriptionModel / RegisterTranscriptionModelDeserializer provider.SerializeEmbeddingModel / DeserializeEmbeddingModel / RegisterEmbeddingModelDeserializer provider.SerializeEvaluationModel / DeserializeEvaluationModel / RegisterEvaluationModelDeserializer ``` None of this is wired into the workflow runtime automatically — there is no implicit "serialize every model field on a workflow struct" step. Your application calls `Serialize*Model` explicitly before persisting workflow state or crossing a boundary, and calls `Deserialize*Model` (after importing the target provider package, so its deserializer registers itself via `init()`) explicitly when resuming. A model that implements `provider.SerializableModelStrict` instead of the plain `SerializableModel` interface can refuse to serialize — for example, an Open Responses model with registered extension codecs, since the codecs are functions and cannot be reconstructed from JSON. --- # Agent Callbacks > Legacy reference for Go AI SDK agent callback support, covering the OnStepFinish callback, other callbacks, a complete example, and TypeScript comparison. Canonical URL: https://goaisdk.com/docs/agents/agent-callbacks Documentation index: https://goaisdk.com/llms.txt The Go AI SDK provides comprehensive callback support for monitoring and controlling agent execution. Callbacks allow you to track progress, log intermediate results, monitor resource usage, and build real-time UI updates. ## OnStepFinish Callback The `OnStepFinish` callback is triggered after each step in a multi-step agent execution. This is particularly useful for: - **Progress tracking**: Monitor how many steps have been completed - **Token usage monitoring**: Track costs in real-time - **Debugging**: Inspect tool calls and results at each step - **UI updates**: Build progress indicators and status displays - **Logging**: Record detailed execution traces ### Signature ```go OnStepFinish func(step types.StepResult) ``` ### StepResult Structure The `StepResult` passed to the callback contains complete information about the step: ```go type StepResult struct { // Step number (0-indexed) StepNumber int // Text generated in this step Text string // Tool calls made in this step ToolCalls []ToolCall // Tool results from this step ToolResults []ToolResult // Finish reason for this step FinishReason FinishReason // Raw finish reason from the provider RawFinishReason string // Usage for this step (token counts) Usage Usage // Context management information (Anthropic-specific) ContextManagement interface{} // Warnings from this step Warnings []Warning // Response messages generated in this step // Deprecated: use Response.Messages instead. ResponseMessages []Message } ``` ### Basic Example ```go agent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful assistant.", Tools: tools, OnStepFinish: func(step types.StepResult) { fmt.Printf("Step %d complete: %s\n", step.StepNumber, step.Text) if len(step.ToolCalls) > 0 { fmt.Printf("Called %d tools\n", len(step.ToolCalls)) } if step.Usage.InputTokens != nil { fmt.Printf("Used %d input tokens\n", *step.Usage.InputTokens) } }, }) ``` ### Advanced Example: Token Usage Tracking ```go var totalCost float64 const INPUT_COST_PER_1K = 0.01 // $0.01 per 1K input tokens const OUTPUT_COST_PER_1K = 0.03 // $0.03 per 1K output tokens agent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: tools, OnStepFinish: func(step types.StepResult) { // Calculate cost for this step var stepCost float64 if step.Usage.InputTokens != nil { inputCost := float64(*step.Usage.InputTokens) / 1000.0 * INPUT_COST_PER_1K stepCost += inputCost } if step.Usage.OutputTokens != nil { outputCost := float64(*step.Usage.OutputTokens) / 1000.0 * OUTPUT_COST_PER_1K stepCost += outputCost } totalCost += stepCost fmt.Printf("Step %d cost: $%.4f (Total: $%.4f)\n", step.StepNumber, stepCost, totalCost) }, }) ``` ### Example: Building a Progress UI ```go type ProgressTracker struct { steps []types.StepResult currentStep int totalTokens int64 mu sync.Mutex } func (p *ProgressTracker) OnStepFinish(step types.StepResult) { p.mu.Lock() defer p.mu.Unlock() p.steps = append(p.steps, step) p.currentStep = step.StepNumber if step.Usage.GetTotalTokens() > 0 { p.totalTokens += step.Usage.GetTotalTokens() } // Update UI (simplified example) fmt.Printf("\r[%d steps] Processing... %d tokens used", p.currentStep, p.totalTokens) } tracker := &ProgressTracker{} agent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, OnStepFinish: tracker.OnStepFinish, }) ``` ### Example: Logging to File ```go logFile, _ := os.Create("agent-execution.log") defer logFile.Close() logger := log.New(logFile, "", log.LstdFlags) agent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: tools, OnStepFinish: func(step types.StepResult) { // Log structured data as JSON data, _ := json.MarshalIndent(step, "", " ") logger.Printf("Step %d:\n%s\n", step.StepNumber, string(data)) }, }) ``` ### Example: Early Termination Based on Steps ```go const MAX_ALLOWED_STEPS = 5 stepCount := 0 agent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, MaxSteps: 10, // hard limit for this agent (plain int, no SDK-imposed ceiling) OnStepFinish: func(step types.StepResult) { stepCount++ if stepCount >= MAX_ALLOWED_STEPS { fmt.Println("Warning: Approaching step limit") } // Note: In current implementation, OnStepFinish cannot stop execution // Use MaxSteps in AgentConfig for hard limits }, }) ``` ## Other Callbacks ### OnStepStart Called before each step begins: ```go OnStepStart func(stepNum int) ``` ### OnToolCall Called when the agent decides to call a tool: ```go OnToolCall func(toolCall types.ToolCall) ``` ### OnToolResult Called when a tool execution completes: ```go OnToolResult func(toolResult types.ToolResult) ``` ### OnFinish Called when the agent completes execution: ```go OnFinish func(result *agent.AgentResult) ``` ### LangChain-Style Callbacks For compatibility with LangChain patterns: ```go OnChainStart func(input string, messages []types.Message) OnChainEnd func(result *agent.AgentResult) OnChainError func(err error) OnAgentAction func(action agent.AgentAction) OnAgentFinish func(finish agent.AgentFinish) ``` ## Complete Example See `/examples/agent/on-step-finish/main.go` for a complete working example demonstrating all features of the `OnStepFinish` callback. ## Best Practices 1. **Keep callbacks fast**: Callbacks are called synchronously during agent execution. Avoid blocking operations. 2. **Handle errors gracefully**: Don't panic in callbacks - log errors instead. 3. **Use goroutines for expensive operations**: If you need to do expensive work, spawn a goroutine: ```go OnStepFinish: func(step types.StepResult) { go func() { // Expensive operation like uploading logs uploadStepData(step) }() } ``` 4. **Thread safety**: If callbacks share state, use proper synchronization: ```go var mu sync.Mutex var sharedState []types.StepResult OnStepFinish: func(step types.StepResult) { mu.Lock() defer mu.Unlock() sharedState = append(sharedState, step) } ``` 5. **Monitor token usage**: Use `OnStepFinish` to track costs and set budgets. 6. **Debug with context**: `StepResult` includes warnings, raw responses, and context management info. ## Comparison with TypeScript AI SDK The Go implementation closely follows the TypeScript AI SDK's callback patterns: | TypeScript | Go | Notes | |------------|-----|-------| | `onStepFinish` | `OnStepFinish` | Same functionality | | `onFinish` | `OnFinish` | Called at completion | | `experimental_onToolCall` | `OnToolCall` | Stable in Go | The `StepResult` structure matches the TypeScript `StreamStep` with all the same fields and behavior. ## See Also - [Agent Guide](https://goaisdk.com/docs/agents.md) - [Tool Integration](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - [Example: OnStepFinish](https://github.com/digitallysavvy/go-ai/blob/main/examples/agent/on-step-finish/main.go) --- # Agent Skills > Explains agent skills in the Go AI SDK: reusable behaviors registered with agents, covering skill structure, the skill registry, and skills vs. tools. Canonical URL: https://goaisdk.com/docs/agents/agent-skills Documentation index: https://goaisdk.com/llms.txt Agent skills are reusable behaviors that can be registered with agents to extend their capabilities. Skills provide a clean way to encapsulate common operations and share them across multiple agents. ## Overview Skills in the Go AI SDK allow you to: - **Encapsulate reusable behaviors**: Define skills once and use them across multiple agents - **Extend agent capabilities**: Add custom functionality beyond what tools provide - **Organize agent logic**: Group related behaviors into named skills - **Share instructions**: Include guidance on when and how to use each skill ## Basic Usage ### Creating a Skill ```go skill := &agent.Skill{ Name: "weather", Description: "Get weather information for a location", Instructions: "Use this skill when the user asks about weather or climate", Handler: func(ctx context.Context, input string) (string, error) { // Implement weather fetching logic return fmt.Sprintf("Weather for %s: Sunny, 72°F", input), nil }, } ``` ### Adding Skills to an Agent ```go agentInstance := agent.NewToolLoopAgent(config) if err := agentInstance.AddSkill(skill); err != nil { log.Fatalf("Failed to add skill: %v", err) } ``` ### Executing Skills ```go result, err := agentInstance.ExecuteSkill(ctx, "weather", "San Francisco") if err != nil { log.Fatalf("Skill execution failed: %v", err) } fmt.Println(result) // "Weather for San Francisco: Sunny, 72°F" ``` ## Skill Structure ### Required Fields - **Name** (string): Unique identifier for the skill - **Description** (string): Brief description of what the skill does - **Handler** (SkillHandler): Function that executes the skill logic ### Optional Fields - **Instructions** (string): Guidance for agents on when to use the skill - **Metadata** (map[string]interface{}): Additional information about the skill ## Skill Registry The `SkillRegistry` manages a collection of skills and provides methods for registration, retrieval, and execution. ### Creating a Registry ```go registry := agent.NewSkillRegistry() ``` ### Registry Operations ```go // Register a skill err := registry.Register(skill) // Check if a skill exists exists := registry.Has("weather") // Get a skill skill, found := registry.Get("weather") // Execute a skill result, err := registry.Execute(ctx, "weather", "New York") // List all skills skills := registry.List() // Get skill names names := registry.Names() // Remove a skill registry.Unregister("weather") // Count skills count := registry.Count() // Clear all skills registry.Clear() ``` ## Complete Example ```go package main import ( "context" "fmt" "os" "strings" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { // Create agent openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := openaiProvider.LanguageModel("gpt-5.4-mini") config := agent.AgentConfig{ Model: model, System: "You are a helpful assistant with text processing skills.", MaxSteps: 5, } agentInstance := agent.NewToolLoopAgent(config) // Define skills uppercaseSkill := &agent.Skill{ Name: "uppercase", Description: "Converts text to uppercase", Instructions: "Use when the user wants text in all caps", Handler: func(ctx context.Context, input string) (string, error) { return strings.ToUpper(input), nil }, } wordCountSkill := &agent.Skill{ Name: "word_count", Description: "Counts words in text", Instructions: "Use when the user wants to know word count", Handler: func(ctx context.Context, input string) (string, error) { words := strings.Fields(input) return fmt.Sprintf("Word count: %d", len(words)), nil }, Metadata: map[string]interface{}{ "category": "text-analysis", "version": "1.0", }, } // Add skills agentInstance.AddSkill(uppercaseSkill) agentInstance.AddSkill(wordCountSkill) // Execute skills result1, _ := agentInstance.ExecuteSkill(context.Background(), "uppercase", "hello world") fmt.Println(result1) // "HELLO WORLD" result2, _ := agentInstance.ExecuteSkill(context.Background(), "word_count", "hello world") fmt.Println(result2) // "Word count: 2" // List skills for _, skill := range agentInstance.ListSkills() { fmt.Printf("Skill: %s - %s\n", skill.Name, skill.Description) } } ``` ## Best Practices ### 1. Choose Clear Names Use descriptive, action-oriented names for skills: ```go // Good "format_text" "analyze_sentiment" "extract_keywords" // Avoid "skill1" "helper" "util" ``` ### 2. Provide Detailed Instructions Help agents understand when to use each skill: ```go skill := &agent.Skill{ Name: "summarize", Description: "Creates a concise summary of text", Instructions: "Use this skill when the user asks for a summary, synopsis, or brief overview of long text. Works best with text over 100 words.", Handler: summarizeHandler, } ``` ### 3. Handle Errors Gracefully Return descriptive errors from skill handlers: ```go Handler: func(ctx context.Context, input string) (string, error) { if input == "" { return "", fmt.Errorf("input cannot be empty") } result, err := processData(input) if err != nil { return "", fmt.Errorf("processing failed: %w", err) } return result, nil } ``` ### 4. Use Metadata for Organization Add metadata to categorize and version skills: ```go skill := &agent.Skill{ Name: "analyze_sentiment", Description: "Analyzes text sentiment", Handler: sentimentHandler, Metadata: map[string]interface{}{ "category": "text-analysis", "version": "1.0.0", "author": "team-ai", "model": "sentiment-v2", }, } ``` ### 5. Keep Skills Focused Each skill should do one thing well: ```go // Good: Focused skills formatSkill // Just formats text validateSkill // Just validates transformSkill // Just transforms // Avoid: Overly complex skills formatAndValidateAndTransformSkill ``` ## Skills vs Tools ### When to Use Skills - **Reusable agent behaviors**: When multiple agents need the same capability - **Simple operations**: When the operation doesn't need LLM-style tool calling - **Organizational clarity**: When you want to group related agent behaviors ### When to Use Tools - **LLM-driven operations**: When the LLM should decide when to use it - **Complex parameters**: When you need structured input schemas - **Tool calling**: When using the model's native tool-calling capabilities ### Skills + Tools You can use both together: ```go // Tool for LLM to call weatherTool := types.Tool{ Name: "get_weather", Description: "Get current weather", Execute: func(ctx context.Context, args map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location := args["location"].(string) // Fetch weather... return weatherData, nil }, } // Skill for direct agent use weatherSkill := &agent.Skill{ Name: "format_weather", Description: "Formats weather data for display", Handler: func(ctx context.Context, input string) (string, error) { // Format weather... return formatted, nil }, } config := agent.AgentConfig{ Tools: []types.Tool{weatherTool}, Skills: registry, } ``` ## Advanced Patterns ### Skill Composition Combine multiple skills: ```go compositeSkill := &agent.Skill{ Name: "process_and_analyze", Handler: func(ctx context.Context, input string) (string, error) { // Use other skills processed, err := agentInstance.ExecuteSkill(ctx, "process", input) if err != nil { return "", err } analyzed, err := agentInstance.ExecuteSkill(ctx, "analyze", processed) return analyzed, err }, } ``` ### Conditional Skills Add/remove skills dynamically: ```go if userHasPremium { agentInstance.AddSkill(premiumSkill) } if !userNeedsBasicFeatures { agentInstance.RemoveSkill("basic_feature") } ``` ### Skill Middleware Wrap skill execution with common logic: ```go func withLogging(skill *agent.Skill) *agent.Skill { originalHandler := skill.Handler skill.Handler = func(ctx context.Context, input string) (string, error) { log.Printf("Executing skill: %s with input: %s", skill.Name, input) result, err := originalHandler(ctx, input) if err != nil { log.Printf("Skill %s failed: %v", skill.Name, err) } else { log.Printf("Skill %s succeeded", skill.Name) } return result, err } return skill } // Usage agentInstance.AddSkill(withLogging(mySkill)) ``` ## See Also - [Agent Subagents](https://goaisdk.com/docs/agents/agent-subagents.md) - Hierarchical agent delegation - [Tool Loop Agent](https://goaisdk.com/docs/reference/ai/tool-loop-agent.md) - Core agent implementation - [Agent Configuration](https://goaisdk.com/docs/agents/configuring-call-options.md) - Configuring agents ## Examples See the [agent-skills example](https://github.com/digitallysavvy/go-ai/tree/main/examples/agent-skills) for a complete working example. --- # Agent Subagents > Covers hierarchical agent systems in the Go AI SDK where a main agent delegates to subagents, including the subagent registry and delegation tracking. Canonical URL: https://goaisdk.com/docs/agents/agent-subagents Documentation index: https://goaisdk.com/llms.txt Subagents enable hierarchical agent systems where a main agent can delegate tasks to specialized subagents. This allows for complex workflows with division of labor and specialized expertise. ## Overview Subagents in the Go AI SDK allow you to: - **Delegate specialized tasks**: Route work to agents optimized for specific domains - **Build hierarchies**: Create multi-level agent structures with subagents having their own subagents - **Separate concerns**: Keep agents focused on their areas of expertise - **Coordinate workflows**: Orchestrate complex tasks across multiple specialized agents ## Basic Usage ### Creating Subagents ```go // Create main agent mainConfig := agent.AgentConfig{ Model: model, System: "You are a coordinator agent", MaxSteps: 5, } mainAgent := agent.NewToolLoopAgent(mainConfig) // Create research subagent researchConfig := agent.AgentConfig{ Model: model, System: "You are a research specialist", MaxSteps: 3, } researchAgent := agent.NewToolLoopAgent(researchConfig) // Register subagent if err := mainAgent.AddSubagent("research", researchAgent); err != nil { log.Fatalf("Failed to add subagent: %v", err) } ``` ### Delegating to Subagents ```go // Delegate with a prompt result, err := mainAgent.DelegateToSubagent( ctx, "research", "Find information about Go concurrency patterns", ) // Delegate with messages messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Analyze this data"}, }, }, } result, err := mainAgent.DelegateToSubagentWithMessages( ctx, "research", messages, ) ``` ## Subagent Registry The `SubagentRegistry` manages a collection of subagents and provides methods for registration, retrieval, and delegation. ### Creating a Registry ```go registry := agent.NewSubagentRegistry() ``` ### Registry Operations ```go // Register a subagent err := registry.Register("research", researchAgent) // Check if a subagent exists exists := registry.Has("research") // Get a subagent subagent, found := registry.Get("research") // Execute delegation result, err := registry.Execute(ctx, "research", "find data") // Execute with messages result, err := registry.ExecuteWithMessages(ctx, "research", messages) // List all subagents names := registry.List() // Remove a subagent registry.Unregister("research") // Count subagents count := registry.Count() // Clear all subagents registry.Clear() ``` ## Complete Example ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := openaiProvider.LanguageModel("gpt-5.4-mini") // Create main coordinator agent mainConfig := agent.AgentConfig{ Model: model, System: "You coordinate tasks between specialized agents", MaxSteps: 5, } mainAgent := agent.NewToolLoopAgent(mainConfig) // Create research subagent researchConfig := agent.AgentConfig{ Model: model, System: "You are a research specialist who finds and summarizes information", MaxSteps: 3, } researchAgent := agent.NewToolLoopAgent(researchConfig) // Create analysis subagent analysisConfig := agent.AgentConfig{ Model: model, System: "You are an analysis specialist who provides insights from data", MaxSteps: 3, } analysisAgent := agent.NewToolLoopAgent(analysisConfig) // Register subagents mainAgent.AddSubagent("research", researchAgent) mainAgent.AddSubagent("analysis", analysisAgent) // Delegate to research subagent researchResult, err := mainAgent.DelegateToSubagent( context.Background(), "research", "Find information about Go interfaces", ) if err != nil { log.Fatal(err) } fmt.Println("Research result:", researchResult.Text) // Delegate to analysis subagent analysisResult, err := mainAgent.DelegateToSubagent( context.Background(), "analysis", "Analyze the benefits of using interfaces", ) if err != nil { log.Fatal(err) } fmt.Println("Analysis result:", analysisResult.Text) } ``` ## Hierarchical Agents Subagents can have their own subagents, creating multi-level hierarchies: ```go // Main agent mainAgent := agent.NewToolLoopAgent(mainConfig) // Research agent (subagent of main) researchAgent := agent.NewToolLoopAgent(researchConfig) mainAgent.AddSubagent("research", researchAgent) // Deep research agent (subagent of research) deepResearchAgent := agent.NewToolLoopAgent(deepResearchConfig) // researchAgent is already a *agent.ToolLoopAgent, so its subagent methods // are available directly -- no type assertion needed. researchAgent.AddSubagent("deep_research", deepResearchAgent) // Delegate through hierarchy // Main -> Research result1, _ := mainAgent.DelegateToSubagent(ctx, "research", "broad research task") // Research -> Deep Research result2, _ := researchAgent.DelegateToSubagent(ctx, "deep_research", "detailed research task") ``` ## Delegation Tracking Track delegation history with `DelegationTracker`: ```go tracker := agent.NewDelegationTracker() // Track a delegation delegation := agent.SubagentDelegation{ SubagentName: "research", Prompt: "Find data", Result: result, } tracker.Track(delegation) // Get all delegations delegations := tracker.GetDelegations() for _, d := range delegations { fmt.Printf("Delegated to %s: %s\n", d.SubagentName, d.Prompt) if d.Result != nil { fmt.Printf("Result: %s\n", d.Result.Text) } } ``` ## Best Practices ### 1. Clear Specialization Give each subagent a clear, focused role: ```go // Good: Specialized subagents researchAgent // Finds information analysisAgent // Analyzes data summaryAgent // Creates summaries // Avoid: Generic subagents helperAgent // Too vague utilityAgent // Unclear purpose ``` ### 2. Appropriate System Prompts Tailor system prompts to each subagent's specialty: ```go researchConfig := agent.AgentConfig{ System: `You are a research specialist. Your role is to: - Find relevant information on requested topics - Verify sources and accuracy - Provide comprehensive summaries - Focus on factual, well-sourced content`, } analysisConfig := agent.AgentConfig{ System: `You are a data analysis specialist. Your role is to: - Analyze patterns and trends in data - Provide insights and interpretations - Identify key findings - Present clear, actionable conclusions`, } ``` ### 3. Manage Subagent Scope Limit subagent steps to prevent runaway execution: ```go subagentConfig := agent.AgentConfig{ Model: model, StopWhen: []ai.StopCondition{ai.IsStepCount(3)}, // Limit subagent steps } mainConfig := agent.AgentConfig{ Model: model, StopWhen: []ai.StopCondition{ai.IsStepCount(10)}, // Main agent can take more steps } ``` ### 4. Error Handling Handle delegation failures gracefully: ```go result, err := mainAgent.DelegateToSubagent(ctx, "research", prompt) if err != nil { // Log the error log.Printf("Delegation to research failed: %v", err) // Try alternative approach result, err = mainAgent.DelegateToSubagent(ctx, "backup_research", prompt) // Or handle in main agent if err != nil { return mainAgent.Execute(ctx, "Handle this yourself: "+prompt) } } ``` ### 5. Combine with Skills Subagents can have their own skills: ```go // Create subagent with skills analysisAgent := agent.NewToolLoopAgent(analysisConfig) // Add skills to subagent statisticsSkill := &agent.Skill{ Name: "calculate_stats", Description: "Calculates statistics", Handler: statsHandler, } analysisAgent.AddSkill(statisticsSkill) // Register as subagent mainAgent.AddSubagent("analysis", analysisAgent) // Access subagent's skills subagent, _ := mainAgent.GetSubagent("analysis") toolLoopSubagent := subagent.(*agent.ToolLoopAgent) result, _ := toolLoopSubagent.ExecuteSkill(ctx, "calculate_stats", data) ``` ## Design Patterns ### 1. Coordinator Pattern Main agent coordinates work across subagents: ```go mainAgent := agent.NewToolLoopAgent(coordinatorConfig) mainAgent.AddSubagent("research", researchAgent) mainAgent.AddSubagent("analysis", analysisAgent) mainAgent.AddSubagent("writer", writerAgent) // Main agent decides which subagent to use result, _ := mainAgent.Execute(ctx, "Create a report on AI trends") ``` ### 2. Pipeline Pattern Subagents form a processing pipeline: ```go // Step 1: Research researchResult, _ := mainAgent.DelegateToSubagent(ctx, "research", topic) // Step 2: Analyze research analysisResult, _ := mainAgent.DelegateToSubagent(ctx, "analysis", researchResult.Text) // Step 3: Summarize analysis summaryResult, _ := mainAgent.DelegateToSubagent(ctx, "writer", analysisResult.Text) ``` ### 3. Expert Panel Pattern Multiple subagents provide different perspectives: ```go // Get opinions from multiple experts technicalView, _ := mainAgent.DelegateToSubagent(ctx, "technical_expert", question) businessView, _ := mainAgent.DelegateToSubagent(ctx, "business_expert", question) legalView, _ := mainAgent.DelegateToSubagent(ctx, "legal_expert", question) // Main agent synthesizes synthesis := fmt.Sprintf( "Technical: %s\nBusiness: %s\nLegal: %s", technicalView.Text, businessView.Text, legalView.Text, ) ``` ### 4. Hierarchical Delegation Multi-level delegation for complex tasks: ```go // Level 1: Main agent mainAgent.AddSubagent("operations", operationsAgent) // Level 2: Operations has sub-specialists operationsAgent.AddSubagent("data_processing", dataAgent) operationsAgent.AddSubagent("quality_control", qcAgent) // Level 3: Data processing has its own subagents dataAgent.AddSubagent("cleaning", cleaningAgent) dataAgent.AddSubagent("transformation", transformAgent) ``` ## Subagents vs Tools ### When to Use Subagents - **Complex specialized tasks**: When a task requires its own agent loop - **Domain expertise**: When you need focused, specialized behavior - **Workflow steps**: When breaking a task into distinct phases - **Resource management**: When you want separate token/cost tracking ### When to Use Tools - **Simple operations**: When a single function call suffices - **LLM-driven decisions**: When the model should choose when to call - **Structured I/O**: When you need parameter schemas - **Native integration**: When using model's tool-calling features ### Subagents + Tools Use both for maximum flexibility: ```go // Subagent with tools researchAgent := agent.NewToolLoopAgent(researchConfig) // Add tools to subagent (AgentConfig.Tools is unexported on ToolLoopAgent; // use AddTool to add tools after construction) researchAgent.AddTool(searchTool) researchAgent.AddTool(summarizerTool) // Register as subagent mainAgent.AddSubagent("research", researchAgent) // Subagent can use its tools when delegated to result, _ := mainAgent.DelegateToSubagent(ctx, "research", "find information") ``` ## Performance Considerations ### Token Usage Subagents add overhead: ```go // Each delegation is a full agent execution result, _ := mainAgent.DelegateToSubagent(ctx, "research", prompt) // Uses tokens for: main agent + research agent + context passing ``` ### Limiting Costs Control subagent resource usage: ```go subagentConfig := agent.AgentConfig{ Model: model, MaxSteps: 2, // Limit steps MaxTokens: &maxTokens, // Limit tokens per step } ``` ### Caching Reuse subagents instead of recreating: ```go // Good: Create once, reuse researchAgent := agent.NewToolLoopAgent(config) mainAgent.AddSubagent("research", researchAgent) // Multiple delegations reuse the same subagent result1, _ := mainAgent.DelegateToSubagent(ctx, "research", prompt1) result2, _ := mainAgent.DelegateToSubagent(ctx, "research", prompt2) ``` ## See Also - [Agent Skills](https://goaisdk.com/docs/agents/agent-skills.md) - Reusable agent behaviors - [Tool Loop Agent](https://goaisdk.com/docs/reference/ai/tool-loop-agent.md) - Core agent implementation - [Agent Configuration](https://goaisdk.com/docs/agents/configuring-call-options.md) - Configuring agents ## Examples See these examples for complete working code: - [agent-subagents example](https://github.com/digitallysavvy/go-ai/tree/main/examples/agent-subagents) - Basic subagent usage - [agent-skills-subagents example](https://github.com/digitallysavvy/go-ai/tree/main/examples/agent-skills-subagents) - Skills and subagents together --- # Agents > Landing page for the agents section of the Go AI SDK docs, linking guides on building agents, workflows, callbacks, and call-option configuration. Canonical URL: https://goaisdk.com/docs/agents/ Documentation index: https://goaisdk.com/llms.txt The following section shows you how to build agents with the Go AI SDK - systems where large language models (LLMs) use tools in a loop to accomplish tasks. ## Contents ### [Overview](https://goaisdk.com/docs/agents/overview.md) Learn what agents are and why to use the ToolLoopAgent. ### [Building Agents](https://goaisdk.com/docs/agents/building-agents.md) Complete guide to creating agents with the ToolLoopAgent. ### [Workflow Patterns](https://goaisdk.com/docs/agents/workflows.md) Structured patterns using core functions for complex workflows. ### [Loop Control](https://goaisdk.com/docs/agents/loop-control.md) Advanced execution control with MaxSteps and custom loop patterns. ### [Configuring Call Options](https://goaisdk.com/docs/agents/configuring-call-options.md) Pass runtime inputs to dynamically configure agent behavior. ### [WorkflowAgent](https://goaisdk.com/docs/agents/workflow-agent.md) Durable and serializable workflow agent execution with resumable SSE streams. --- # Core API Overview > Overview of the Go AI SDK Core API: core generation functions, context-based operations, error handling, middleware, telemetry, and Go-specific features. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/overview Documentation index: https://goaisdk.com/llms.txt Large Language Models (LLMs) are advanced programs that can understand, create, and engage with human language on a large scale. They are trained on vast amounts of written material to recognize patterns in language and predict what might come next in a given piece of text. The **Go AI SDK Core** simplifies working with LLMs by offering a standardized way of integrating them into your Go applications - so you can focus on building great AI applications for your users, not waste time on technical details. For example, here's how you can generate text with various models using the Go AI SDK: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/providers/google" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() // Pick a provider and model. The provider constants hold current model IDs. openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) gpt, err := openaiProvider.LanguageModel(openai.ModelGPT6Astra) if err != nil { log.Fatal(err) } anthropicProvider := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) claude, err := anthropicProvider.LanguageModel(anthropic.ClaudeSonnet5_5) if err != nil { log.Fatal(err) } googleProvider := google.New(google.Config{APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY")}) gemini, err := googleProvider.LanguageModel(google.ModelGemini31FlashLitePreview) if err != nil { log.Fatal(err) } // Same API for all providers for _, model := range []provider.LanguageModel{gpt, claude, gemini} { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing in simple terms.", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } } ``` ## Core Functions The Go AI SDK Core has various functions designed for [text generation](https://goaisdk.com/docs/ai-sdk-core/generating-text.md), [structured data generation](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md), and [tool usage](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md). These functions take a standardized approach to setting up [prompts](https://goaisdk.com/docs/foundations/prompts.md) and [settings](https://goaisdk.com/docs/ai-sdk-core/settings.md), making it easier to work with different models. ### Text Generation - **[`ai.GenerateText()`](https://goaisdk.com/docs/ai-sdk-core/generating-text.md)** - Generates text and [tool calls](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md). This function is ideal for non-interactive use cases such as automation tasks where you need to write text (e.g. drafting email or summarizing web pages) and for agents that use tools. - **[`ai.StreamText()`](https://goaisdk.com/docs/ai-sdk-core/generating-text.md#streamtext)** - Stream text and tool calls. You can use the `StreamText` function for interactive use cases such as chat bots and content streaming. ```go // Generate complete text at once result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a haiku about Go programming.", }) fmt.Println(result.Text) // Stream text as it's generated stream, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a story about AI.", }) defer stream.Close() for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } ``` ### Structured Data Generation - **[`ai.GenerateObject()`](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md)** - Generates a typed, structured object that matches a JSON schema. You can use this function to force the language model to return structured data, e.g. for information extraction, synthetic data generation, or classification tasks. - **[`ai.StreamObject()`](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md#stream-object)** - Stream a structured object that matches a JSON schema. You can use this function to stream generated structured data. ```go personSchema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "age": map[string]interface{}{"type": "number"}, }, "required": []string{"name", "age"}, } result, _ := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema.NewSimpleJSONSchema(personSchema), Prompt: "Generate a person profile.", }) fmt.Printf("Object: %v\n", result.Object) ``` ### Embeddings & Semantic Search - **[`ai.Embed()`](https://goaisdk.com/docs/ai-sdk-core/embeddings.md)** - Generate a single embedding vector for text. - **[`ai.EmbedMany()`](https://goaisdk.com/docs/ai-sdk-core/embeddings.md)** - Generate multiple embedding vectors in batch. - **Similarity Functions** - `CosineSimilarity`, `EuclideanDistance`, `DotProduct`, and more. ```go // Single embedding result, _ := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "Go is a statically typed, compiled language", }) // Batch embeddings results, _ := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: []string{"text1", "text2", "text3"}, }) // Find most similar idx, score, _ := ai.FindMostSimilar(queryEmbedding, results.Embeddings) ``` ### Document Reranking - **[`ai.Rerank()`](https://goaisdk.com/docs/ai-sdk-core/reranking.md)** - Rerank documents based on relevance to a query. ```go topN := 3 result, _ := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: "What is machine learning?", Documents: []string{"doc1", "doc2", "doc3"}, TopN: &topN, }) ``` ### Multi-Modal Generation - **[`ai.GenerateImage()`](https://goaisdk.com/docs/ai-sdk-core/image-generation.md)** - Generate images from text prompts. - **[`ai.GenerateSpeech()`](https://goaisdk.com/docs/ai-sdk-core/speech.md)** - Convert text to speech. - **[`ai.Transcribe()`](https://goaisdk.com/docs/ai-sdk-core/transcription.md)** - Convert speech to text. ```go // Text to image imageResult, _ := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "A serene mountain landscape", }) // Text to speech speechResult, _ := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "alloy", }) // Speech to text transcription, _ := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioBytes, }) ``` ## Context-Based Operations All Go AI SDK functions use `context.Context` as the first parameter, enabling: - **Cancellation**: Cancel operations in progress - **Timeouts**: Set deadlines for operations - **Request-scoped values**: Pass data through the call stack ```go // With timeout ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Long generation...", }) if err != nil { if ctx.Err() == context.DeadlineExceeded { fmt.Println("Operation timed out") } } ``` ## Error Handling The Go AI SDK uses Go's standard error handling patterns. All functions return an error as the last return value: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello", }) if err != nil { // Handle error // providererrors is github.com/digitallysavvy/go-ai/pkg/provider/errors if providererrors.IsRateLimitError(err) { fmt.Println("Rate limit exceeded") } log.Fatal(err) } ``` See [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) for comprehensive error handling strategies. ## Configuration & Settings Control model behavior with settings: ```go temperature := 0.8 maxTokens := 500 topP := 0.9 topK := 40 result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate creative text", Temperature: &temperature, // Higher = more creative MaxTokens: &maxTokens, // Limit response length TopP: &topP, // Nucleus sampling TopK: &topK, // Top-k sampling StopSequences: []string{"\n\n"}, // Stop at double newline }) ``` See [Settings](https://goaisdk.com/docs/ai-sdk-core/settings.md) for all available options. ## Middleware Add cross-cutting concerns like logging, caching, or rate limiting: ```go import "github.com/digitallysavvy/go-ai/pkg/middleware" wrappedModel := middleware.WrapLanguageModel(model, []*middleware.LanguageModelMiddleware{ loggingMiddleware, cachingMiddleware, rateLimitingMiddleware, }, nil, nil) ``` See [Middleware](https://goaisdk.com/docs/ai-sdk-core/middleware.md) for details. ## Telemetry Built-in OpenTelemetry support for observability: ```go import "github.com/digitallysavvy/go-ai/pkg/telemetry" // Register an OpenTelemetry-backed integration once at startup. telemetry.RegisterTelemetryIntegration(telemetry.NewLegacyOpenTelemetry(telemetry.LegacyOpenTelemetryOptions{})) // Calls are traced when telemetry is enabled or an integration is registered. result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello", Telemetry: &telemetry.Options{}, }) ``` See [Telemetry](https://goaisdk.com/docs/ai-sdk-core/telemetry.md) for configuration options. ## Go-Specific Features The Go AI SDK leverages Go's strengths: ### Channels for Streaming ```go stream, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate text", }) defer stream.Close() // Channels close automatically when done for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } ``` ### Goroutines for Concurrency ```go var wg sync.WaitGroup for _, prompt := range prompts { wg.Add(1) go func(p string) { defer wg.Done() result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: p, }) fmt.Println(result.Text) }(prompt) } wg.Wait() ``` ### defer for Resource Cleanup ```go stream, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate text", }) defer stream.Close() // Always closes, even on error for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } ``` ## API Reference For detailed API documentation, see the [API Reference](https://goaisdk.com/docs/reference.md) section: - [AI Package Functions](https://goaisdk.com/docs/reference.md) - [Provider Interfaces](https://goaisdk.com/docs/reference.md) - [Type Definitions](https://goaisdk.com/docs/reference.md) ## Next Steps - Learn about [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - Explore [Generating Structured Data](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) - Understand [Tools and Tool Calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - Build [Agents](https://goaisdk.com/docs/agents/overview.md) ## See Also - [Foundations](https://goaisdk.com/docs/foundations/overview.md) - [Providers and Models](https://goaisdk.com/docs/foundations/providers-and-models.md) - [Examples](https://github.com/digitallysavvy/go-ai/tree/main/examples) --- # Generating Text > Shows how to generate and stream text with GenerateText and StreamText in Go, covering message-based generation, settings, and error-handling patterns. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/generating-text Documentation index: https://goaisdk.com/llms.txt Large language models (LLMs) can generate text in response to a prompt, which can contain instructions and information to process. For example, you can ask a model to come up with a recipe, draft an email, or summarize a document. The Go AI SDK Core provides two functions to generate text from LLMs: - [`ai.GenerateText()`](#generatetext) - Generates text for a given prompt and model. - [`ai.StreamText()`](#streamtext) - Streams text from a given prompt and model. Advanced LLM features such as [tool calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) and [structured data generation](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) are built on top of text generation. ## GenerateText You can generate text using the [`ai.GenerateText()`](https://goaisdk.com/docs/reference/ai/generate-text.md) function. This function is ideal for non-interactive use cases where you need to write text (e.g. drafting email or summarizing web pages) and for agents that use tools. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-5.2") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a vegetarian lasagna recipe for 4 people.", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` You can use more [advanced prompts](https://goaisdk.com/docs/foundations/prompts.md) to generate text with more complex instructions and content: ```go article := "..." // your article text result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a professional writer. " + "You write simple, clear, and concise content.", Prompt: fmt.Sprintf("Summarize the following article in 3-5 sentences: %s", article), }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` ### Result Object The result object of `GenerateText` contains several fields that provide information about the generation: ```go type GenerateTextResult struct { // The generated text Text string // Ordered output content from all completed steps. Includes text, // reasoning, tool calls, tool results, files, sources, and provider // metadata-bearing parts in the same order emitted by the model/tool loop. Content []types.ContentPart // Output contains the parsed output when a WithOutput option was provided. // Type-assert to the concrete type, e.g.: recipe := result.Output.(Recipe) // Nil when no Output option was set. Output any // Tool calls made during generation ToolCalls []types.ToolCall // Tool results from executed tools ToolResults []types.ToolResult // Steps taken during generation (for multi-step tool calling) Steps []types.StepResult // Reason the model finished generating FinishReason types.FinishReason // Token usage information Usage types.Usage // Warnings from the model provider Warnings []types.Warning // Provider-specific metadata ProviderMetadata map[string]interface{} // Raw request/response (for debugging) RawRequest interface{} RawResponse interface{} } ``` ### Accessing Response Headers & Body Sometimes you need access to the full response from the model provider, e.g. to access some provider-specific headers or body content. You can access the raw request and response using the `RawRequest` and `RawResponse` fields: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello", }) if err != nil { log.Fatal(err) } fmt.Printf("Raw request: %+v\n", result.RawRequest) fmt.Printf("Raw response: %+v\n", result.RawResponse) ``` ### OnFinish Callback When using `GenerateText`, you can provide an `OnFinish` callback that is triggered after the last step is finished. It contains the text, usage information, finish reason, and more: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Invent a new holiday and describe its traditions.", OnFinish: func(ctx context.Context, result *ai.GenerateTextResult, userContext interface{}) { // Your own logic, e.g. for saving the chat history or recording usage fmt.Printf("Text: %s\n", result.Text) fmt.Printf("Finish reason: %s\n", result.FinishReason) fmt.Printf("Usage: %+v\n", result.Usage) fmt.Printf("Steps: %d\n", len(result.Steps)) }, }) ``` ### ToolOrder `ToolOrder` controls the order tools are sent to providers after active-tool filtering, middleware transforms, and `PrepareStep` updates. Listed names are sent first in the provided order; unlisted tools follow alphabetically. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Inspect the user account.", Tools: tools, ToolOrder: []string{"readProfile", "listOrders"}, }) ``` ### OnStepEnd `OnStepEnd` is the canonical callback name for step completion. `OnStepFinish` remains as a deprecated compatibility alias; when both are set, `OnStepEnd` wins. The same migration applies to typed event callbacks: use `OnStepEndEvent` instead of `OnStepFinishEvent`. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Use a tool, then summarize.", Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, OnStepEnd: func(ctx context.Context, step types.StepResult, userContext interface{}) { fmt.Println(step.StepNumber, step.FinishReason, step.Usage) }, }) ``` ## StreamText Depending on your model and prompt, it can take a large language model (LLM) up to a minute to finish generating its response. This delay can be unacceptable for interactive use cases such as chatbots or real-time applications, where users expect immediate responses. The Go AI SDK Core provides the [`ai.StreamText()`](https://goaisdk.com/docs/reference/ai/stream-text.md) function which simplifies streaming text from LLMs: ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Invent a new holiday and describe its traditions.", }) if err != nil { log.Fatal(err) } defer result.Close() // Stream text chunks as they arrive for chunk := range result.Chunks() { fmt.Print(chunk.Text) } ``` > **Note:** `StreamText` immediately starts streaming and errors become part of the stream. Use the `OnError` callback to log errors. > **Note:** `StreamText` uses backpressure and only generates tokens as they are requested. You need to consume the channel for the stream to finish. ### Stream Result Object The stream result object provides access to the streamed data: ```go // StreamTextResult provides these methods to access streamed data. type streamTextResult interface { // Streaming methods (available immediately) Chunks() <-chan provider.StreamChunk // Channel for receiving chunks Close() error // Close the stream ReadAll() (string, error) // Read all text (blocks until complete) // Accessor methods (available after stream completes) Text() string // Complete generated text Content() []types.ContentPart // Ordered content across all steps FinishReason() types.FinishReason // Reason the model finished Usage() types.Usage // Token usage information ToolCalls() []types.ToolCall // Tool calls made during streaming ToolResults() []types.ToolResult // Tool results from executed tools ContextManagement() interface{} // Context management stats (Anthropic) Err() error // Error if stream failed } ``` ### Standalone Stream Helpers For new code, use `result.Stream()` with the package-level helpers. The result-bound helpers are retained only as deprecated compatibility wrappers. ```go textStream, errStream := ai.ToTextStream(ctx, result.Stream()) for text := range textStream { fmt.Print(text) } if err := <-errStream; err != nil { log.Fatal(err) } ``` Use `ai.ToUIMessageStream(ctx, result.Stream(), ...)` to convert provider stream chunks into UI message chunks. `result.FullStream()` remains as a deprecated alias for `result.Stream()`. ### Channel-Based Streaming The Go AI SDK uses Go channels for streaming, providing a natural and idiomatic way to handle streaming data: ```go result, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a story", }) defer result.Close() // Channel closes automatically when done for chunk := range result.Chunks() { if chunk.Type == provider.ChunkTypeError { fmt.Printf("Error: %s\n", chunk.AbortReason) break } fmt.Print(chunk.Text) } ``` ### OnError Callback `StreamText` immediately starts streaming to enable sending data without waiting for the model. Errors become part of the stream and are not returned to prevent servers from crashing. To log errors, you can provide an `OnError` callback that is triggered when an error occurs: ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate text", OnError: func(ctx context.Context, err error) { log.Printf("Stream error: %v", err) // Your error logging logic here }, }) ``` ### OnChunk Callback When using `StreamText`, you can provide an `OnChunk` callback that is triggered for each chunk of the stream. It receives the following chunk types: - `text` - Text content - `tool-call` - Tool call - `tool-result` - Tool result - `finish` - Stream finish ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate text", OnChunk: func(chunk provider.StreamChunk) { // Implement your own logic here if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } else if chunk.Type == provider.ChunkTypeToolCall { fmt.Printf("\n[Tool call: %s]\n", chunk.ToolCall.ToolName) } }, }) ``` ### OnFinish Callback When using `StreamText`, you can provide an `OnFinish` callback that is triggered when the stream is finished. It contains the text, usage information, finish reason, and more: ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate text", OnFinish: func(result *ai.StreamTextResult) { // Your own logic, e.g. for saving the chat history or recording usage fmt.Printf("\n\nText: %s\n", result.Text()) fmt.Printf("Finish reason: %s\n", result.FinishReason()) fmt.Printf("Usage: %+v\n", result.Usage()) }, }) ``` ### Accumulating Streamed Text If you need the complete text after streaming, you can use `ReadAll()`: ```go result, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a story", }) defer result.Close() // Option 1: Use ReadAll (blocks until complete) fullText, err := result.ReadAll() if err != nil { log.Fatal(err) } fmt.Println(fullText) // Option 2: Manually accumulate var builder strings.Builder for chunk := range result.Chunks() { fmt.Print(chunk.Text) // Display to user builder.WriteString(chunk.Text) // Accumulate } completeText := builder.String() ``` ### Context Cancellation Streaming respects context cancellation, allowing you to stop generation early: ```go ctx, cancel := context.WithCancel(context.Background()) go func() { // Cancel after 5 seconds time.Sleep(5 * time.Second) cancel() }() result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a very long story...", }) if err != nil { log.Fatal(err) } defer result.Close() for chunk := range result.Chunks() { select { case <-ctx.Done(): fmt.Println("\n\nCancelled!") return default: fmt.Print(chunk.Text) } } ``` ## Advanced Features ### Multi-Step Generation with Tools Both `GenerateText` and `StreamText` support multi-step generation with automatic tool execution: ```go weatherTool := types.Tool{ Name: "get_weather", Description: "Get weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{"type": "string"}, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location := input["location"].(string) return map[string]interface{}{ "temperature": 72, "condition": "sunny", "location": location, }, nil }, } maxSteps := 5 // Allow up to 5 steps of tool calling result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in London?", Tools: []types.Tool{weatherTool}, MaxSteps: &maxSteps, }) fmt.Println(result.Text) // Output: The weather in London is sunny with a temperature of 72°F. ``` See [Tools and Tool Calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) for more details. ### Step-by-Step Processing Access intermediate steps in multi-step generations: ```go maxSteps := 10 result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in Tokyo and Paris?", Tools: []types.Tool{weatherTool}, MaxSteps: &maxSteps, OnStepFinish: func(ctx context.Context, step types.StepResult, userContext interface{}) { fmt.Printf("Step %d finished\n", step.StepNumber) for _, toolCall := range step.ToolCalls { fmt.Printf(" Tool: %s\n", toolCall.ToolName) } for _, toolResult := range step.ToolResults { fmt.Printf(" Result: %v\n", toolResult.Result) } }, }) // Access all steps after completion for i, step := range result.Steps { fmt.Printf("Step %d: %d tool calls, %d results\n", i+1, len(step.ToolCalls), len(step.ToolResults)) } ``` ### Sources Some providers such as Perplexity and Google Generative AI include sources in the response. These are web pages that ground the response. You can access them using the `Sources` field of the result: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "List the top 5 San Francisco news from the past week.", }) if err != nil { log.Fatal(err) } for _, source := range result.Sources { if source.SourceType == "url" { fmt.Printf("ID: %s\n", source.ID) fmt.Printf("Title: %s\n", source.Title) fmt.Printf("URL: %s\n", source.URL) fmt.Println() } } ``` When streaming, sources are available both as `ChunkTypeSource` chunks and via `result.Sources()` after the stream completes: ```go result, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Latest news", }) defer result.Close() for chunk := range result.Chunks() { if chunk.Type == provider.ChunkTypeSource && chunk.SourceContent != nil { if chunk.SourceContent.SourceType == "url" { fmt.Printf("Source: %s - %s\n", chunk.SourceContent.Title, chunk.SourceContent.URL) } } } // Or access all sources after streaming completes for _, source := range result.Sources() { fmt.Printf("Source: %s - %s\n", source.Title, source.URL) } ``` ## Message-Based Generation For chat applications, use message-based generation: ```go import "github.com/digitallysavvy/go-ai/pkg/provider/types" messages := []types.Message{ {Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "Hi!"}}}, {Role: types.RoleAssistant, Content: []types.ContentPart{types.TextContent{Text: "Hello! How can I help?"}}}, {Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "Tell me a joke."}}}, } result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: messages, }) fmt.Println(result.Text) ``` ## Settings Control generation behavior with various settings: ```go temperature := 0.8 maxTokens := 500 topP := 0.9 topK := 40 result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate creative text", Temperature: &temperature, // Higher = more creative MaxTokens: &maxTokens, // Limit response length TopP: &topP, // Nucleus sampling TopK: &topK, // Top-k sampling StopSequences: []string{"\n\n"}, // Stop at double newline }) ``` See [Settings](https://goaisdk.com/docs/ai-sdk-core/settings.md) for all available options. ## Error Handling Handle errors appropriately: ```go import providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", }) if err != nil { // Check for specific error types var rateLimitErr *providererrors.RateLimitError if errors.As(err, &rateLimitErr) { fmt.Println("Rate limit exceeded, retry later") return } if providererrors.IsValidationError(err) { fmt.Println("Invalid request parameters") return } log.Fatal(err) } ``` See [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) for comprehensive error handling strategies. ## Best Practices 1. **Use Context**: Always pass context for cancellation and timeouts 2. **Close Streams**: Use `defer result.Close()` for streaming 3. **Handle Errors**: Check errors from both function calls and stream chunks 4. **Choose Wisely**: Use `GenerateText` for short responses, `StreamText` for long ones 5. **Set Limits**: Use `MaxTokens` to control costs and response length 6. **Monitor Usage**: Track token usage for cost management 7. **Cache Results**: Consider caching for repeated queries ## Examples ### Basic Text Generation ```go result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum entanglement in simple terms.", }) fmt.Println(result.Text) ``` ### Streaming with Progress ```go result, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long article about AI", }) defer result.Close() tokenCount := 0 for chunk := range result.Chunks() { fmt.Print(chunk.Text) tokenCount += len(strings.Fields(chunk.Text)) fmt.Printf("\r[Tokens: ~%d]", tokenCount) } fmt.Println("\nComplete!") ``` ### Concurrent Generation ```go prompts := []string{ "Explain photosynthesis", "Explain gravity", "Explain evolution", } var wg sync.WaitGroup results := make([]string, len(prompts)) for i, prompt := range prompts { wg.Add(1) go func(idx int, p string) { defer wg.Done() result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: p, }) results[idx] = result.Text }(i, prompt) } wg.Wait() for i, result := range results { fmt.Printf("%d: %s\n\n", i+1, result) } ``` ## Custom Content Types Some providers return ordered content that carries provider-specific metadata. The Go AI SDK preserves that metadata on output content as `ProviderMetadata`; when the content is replayed to a provider as response messages, it is converted to provider-facing `ProviderOptions`, matching the TypeScript SDK. This applies to standard content such as text, reasoning, tool-call, tool-result, and tool-error parts, and to provider-specific content that does not map to standard text, reasoning, or tool call types. The Go AI SDK represents provider-specific parts as `CustomContent` and `ReasoningFileContent`. ### CustomContent `CustomContent` wraps provider-specific response parts that have no standard mapping (e.g., XAI citations, OpenAI compaction events): ```go import "github.com/digitallysavvy/go-ai/pkg/provider/types" // CustomContent is emitted by providers for unknown/provider-specific parts type CustomContent struct { Kind string // e.g. "xai.citation", "openai.compaction" ProviderOptions map[string]interface{} // input: provider-specific options to forward ProviderMetadata json.RawMessage // output: raw JSON from provider } ``` - `Kind` identifies the content type using the format `"{provider}.{type}"` (e.g., `"xai.citation"`) - `ProviderMetadata` carries the raw JSON returned by the provider (output direction) - `ProviderOptions` carries provider-specific options keyed by provider name (input direction) ### ReasoningFileContent `ReasoningFileContent` carries reasoning output as a binary file (e.g., PDFs from Google models): ```go // ReasoningFileContent carries reasoning output as a binary file type ReasoningFileContent struct { MediaType string // e.g. "application/pdf", "image/png" Data []byte // auto base64-encodes in JSON marshaling ProviderOptions map[string]interface{} // input: provider-specific options to forward ProviderMetadata json.RawMessage // output: raw JSON from provider } ``` - `MediaType` is the IANA media type of the file - `Data` holds raw bytes; Go's `encoding/json` automatically handles base64 encoding/decoding - `ProviderOptions` and `ProviderMetadata` serve the same input/output roles as on `CustomContent` ### Stream Chunk Handling Both content types flow through stream chunks via `OnChunk`. Use `ChunkTypeCustom` and `ChunkTypeReasoningFile` to detect them: ```go import "github.com/digitallysavvy/go-ai/pkg/provider" result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Analyze this data with reasoning.", OnChunk: func(chunk provider.StreamChunk) { switch chunk.Type { case provider.ChunkTypeText: fmt.Print(chunk.Text) case provider.ChunkTypeCustom: // Provider-specific content fmt.Printf("Custom content [%s]: %s\n", chunk.CustomContent.Kind, string(chunk.CustomContent.ProviderMetadata)) case provider.ChunkTypeReasoningFile: // Reasoning output as file fmt.Printf("Reasoning file: %s (%d bytes)\n", chunk.ReasoningFileContent.MediaType, len(chunk.ReasoningFileContent.Data)) } }, }) if err != nil { log.Fatal(err) } defer result.Close() for range result.Chunks() { } ``` ### Multi-Turn Behavior In multi-turn conversations, `CustomContent` and `ReasoningFileContent` appear in assistant messages within the conversation history: - **`ProviderOptions`** is forwarded back to the provider in subsequent requests. This allows provider-specific state (e.g., cached references) to round-trip correctly. - **`ProviderMetadata`** is output-only and is **not** re-sent to the provider. It is preserved in conversation history for logging and debugging. ## Next Steps - Learn about [Generating Structured Data](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) - Explore [Tools and Tool Calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - Build [Agents](https://goaisdk.com/docs/agents/overview.md) ## See Also - [Prompts](https://goaisdk.com/docs/foundations/prompts.md) - [Streaming](https://goaisdk.com/docs/foundations/streaming.md) - [Settings](https://goaisdk.com/docs/ai-sdk-core/settings.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [API Reference: StreamText](https://goaisdk.com/docs/reference/ai/stream-text.md) --- # Generating Structured Data > Explains generating structured JSON output with GenerateObject and StreamObject in Go, including the SchemaFor helper, output strategies, and schema naming. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/generating-structured-data Documentation index: https://goaisdk.com/llms.txt While text generation can be useful, your use case will likely call for generating structured data. For example, you might want to extract information from text, classify data, or generate synthetic data. Many language models are capable of generating structured data, often defined as using "JSON modes" or "tools". However, you need to manually provide schemas and then validate the generated data as LLMs can produce incorrect or incomplete structured data. The Go AI SDK standardizes structured output via the `Output` option on [`ai.GenerateText()`](https://goaisdk.com/docs/reference/ai/generate-text.md) and [`ai.StreamText()`](https://goaisdk.com/docs/reference/ai/stream-text.md). Five output factories cover all common patterns: | Factory | Use case | |---|---| | `ai.TextOutput()` | Plain text (default) | | `ai.ObjectOutput[T]()` | Typed struct from JSON | | `ai.ArrayOutput[T]()` | Slice of typed structs from JSON | | `ai.ChoiceOutput[T]()` | Enum/classification | | `ai.JSONOutput()` | Untyped `interface{}` JSON | > **Note:** The older `ai.GenerateObject()` and `ai.StreamObject()` functions are **deprecated**. Use the `Output` option instead. ## Structured Output with GenerateText Use `GenerateText` with an `Output` factory to get a fully typed, validated result in a single call. The SDK automatically sets the correct `ResponseFormat` and parses the model's response. ### Object Output ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type Recipe struct { Name string `json:"name"` Ingredients []string `json:"ingredients"` Steps []string `json:"steps"` } func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate a lasagna recipe.", Output: ai.ObjectOutput[Recipe](ai.ObjectOutputOptions{ Schema: ai.SchemaFor[Recipe](), Name: "recipe", Description: "A complete recipe with name, ingredients, and steps", }), }) if err != nil { log.Fatal(err) } // Type-assert the Output field to get your typed struct recipe := result.Output.(Recipe) fmt.Printf("Recipe: %s\n", recipe.Name) fmt.Printf("Steps: %v\n", recipe.Steps) } ``` ### Array Output ```go type Hero struct { Name string `json:"name"` Class string `json:"class"` Description string `json:"description"` } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate 3 hero descriptions for a fantasy RPG.", Output: ai.ArrayOutput[Hero](ai.ArrayOutputOptions[Hero]{ ElementSchema: ai.SchemaFor[Hero](), Name: "heroes", }), }) if err != nil { log.Fatal(err) } heroes := result.Output.([]Hero) for _, h := range heroes { fmt.Printf("%s (%s): %s\n", h.Name, h.Class, h.Description) } ``` ### Choice Output ```go type Sentiment string const ( Positive Sentiment = "positive" Negative Sentiment = "negative" Neutral Sentiment = "neutral" ) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: `Classify the sentiment of: "I love this product!"`, Output: ai.ChoiceOutput[Sentiment](ai.ChoiceOutputOptions[Sentiment]{ Options: []Sentiment{Positive, Negative, Neutral}, }), }) if err != nil { log.Fatal(err) } sentiment := result.Output.(Sentiment) fmt.Printf("Sentiment: %s\n", sentiment) // positive ``` ### JSON Output (Untyped) ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Return some arbitrary structured data about the Go language.", Output: ai.JSONOutput(ai.JSONOutputOptions{}), }) if err != nil { log.Fatal(err) } data := result.Output.(map[string]interface{}) fmt.Printf("Data: %v\n", data) ``` ## Structured Streaming with StreamText Use `StreamText` with an `Output` factory to receive structured data as it streams. Call `result.PartialOutput()` at any time to get the most recently parsed partial value. ```go type Notification struct { Name string `json:"name"` Message string `json:"message"` } result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate a system notification.", Output: ai.ObjectOutput[Notification](ai.ObjectOutputOptions{ Schema: ai.SchemaFor[Notification](), }), }) if err != nil { log.Fatal(err) } defer result.Close() for chunk := range result.Chunks() { if chunk.Type == provider.ChunkTypeText { // Read incremental parsed value during streaming partial := result.PartialOutput() if n, ok := partial.(Notification); ok { fmt.Printf("Partial name: %s\n", n.Name) } } } // Full text after stream completes fmt.Printf("Raw JSON: %s\n", result.Text()) ``` ### Element Streaming for Arrays When using `ArrayOutput`, you can receive fully validated array elements one at a time as they stream: ```go streamResult, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate 5 notification objects.", Output: arrayOut, // ArrayOutput[Notification] }) if err != nil { log.Fatal(err) } elements := ai.ElementStreamWithOutput(streamResult, arrayOut) for elem := range elements { fmt.Printf("Notification %d: %s\n", elem.Index, elem.Element.Name) } // Check for a stream-level error after consuming all elements if err := streamResult.Err(); err != nil { log.Printf("Error: %v", err) } ``` ## SchemaFor Helper `ai.SchemaFor[T]()` generates a JSON Schema from a Go struct's field types and `json` tags. Use it instead of writing schemas by hand: ```go type Address struct { Street string `json:"street"` City string `json:"city"` Country string `json:"country"` } schema := ai.SchemaFor[Address]() // Result: {"type":"object","properties":{"street":{"type":"string"},...},"required":["street","city","country"]} ``` Fields with `omitempty` in their json tag become optional (not included in `required`). Fields tagged `json:"-"` are excluded. --- ## Legacy API (Deprecated) > **Deprecated:** The functions below still work but are deprecated. Prefer `GenerateText` / `StreamText` with `Output` instead. ## Generate Object The `ai.GenerateObject()` function generates structured data from a prompt. The schema is also used to validate the generated data, ensuring type safety and correctness. ```go package main import ( "context" "encoding/json" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/schema" ) type Recipe struct { Name string `json:"name"` Ingredients []struct { Name string `json:"name"` Amount string `json:"amount"` } `json:"ingredients"` Steps []string `json:"steps"` } func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-6-astra") // Define JSON Schema recipeSchema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "recipe": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "ingredients": map[string]interface{}{ "type": "array", "items": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "amount": map[string]interface{}{"type": "string"}, }, }, }, "steps": map[string]interface{}{ "type": "array", "items": map[string]interface{}{"type": "string"}, }, }, }, }, "required": []string{"recipe"}, } result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema.NewSimpleJSONSchema(recipeSchema), Prompt: "Generate a lasagna recipe.", }) if err != nil { log.Fatal(err) } // result.Object is already unmarshaled; re-marshal it to decode into a typed struct objBytes, err := json.Marshal(result.Object) if err != nil { log.Fatal(err) } var recipe Recipe if err := json.Unmarshal(objBytes, &recipe); err != nil { log.Fatal(err) } fmt.Printf("Recipe: %s\n", recipe.Name) fmt.Printf("Ingredients: %+v\n", recipe.Ingredients) fmt.Printf("Steps: %+v\n", recipe.Steps) } ``` ### Accessing Response Headers & Body Sometimes you need access to the full response from the model provider, e.g. to access some provider-specific headers or body content. You can access the raw response headers and body using the `Response` field: ```go result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema, Prompt: "Generate a recipe.", }) if err != nil { log.Fatal(err) } fmt.Printf("Headers: %+v\n", result.Response.Headers) fmt.Printf("Body: %+v\n", result.Response.Body) ``` ## Stream Object Given the added complexity of returning structured data, model response time can be unacceptable for your interactive use case. With the [`ai.StreamObject()`](https://goaisdk.com/docs/reference/ai/stream-object.md) function, you can receive partial objects as they are parsed during generation via the `OnChunk` callback. Unlike `StreamText`, `ai.StreamObject()` blocks until generation completes and returns the final `*GenerateObjectResult` directly — it does not return a channel-based stream. ```go result, err := ai.StreamObject(ctx, ai.StreamObjectOptions{ Model: model, Schema: schema, Prompt: "Generate a lasagna recipe.", OnChunk: func(partialObject interface{}) { fmt.Printf("Partial object: %v\n", partialObject) }, }) if err != nil { log.Fatal(err) } // Get final complete object finalObject := result.Object fmt.Printf("Final object: %v\n", finalObject) ``` ### OnError Callback `StreamObject` surfaces stream-level errors (distinct from schema parse/validation errors) via the `OnError` callback, which is called before the function returns its final error: ```go result, err := ai.StreamObject(ctx, ai.StreamObjectOptions{ Model: model, Schema: schema, Prompt: "Generate data", OnError: func(ctx context.Context, err error) { log.Printf("Stream error: %v", err) // Your error logging logic here }, }) ``` ## Output Strategy You can use both functions with different output strategies: `object`, `array`, `enum`, or `no-schema`. ### Object The default output strategy is `object`, which returns the generated data as an object. You don't need to specify the output strategy if you want to use the default. ```go result, _ := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema, Prompt: "Generate a person profile.", // OutputMode: "object", // default, can be omitted }) ``` ### Array If you want to generate an array of objects, you can set the output strategy to `array`. When you use the `array` output strategy, the schema specifies the shape of an array element. With `StreamObject`, you can receive the partial array through `OnChunk` as it is generated: ```go // Schema for a single hero heroSchema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "class": map[string]interface{}{ "type": "string", "description": "Character class, e.g. warrior, mage, or thief.", }, "description": map[string]interface{}{"type": "string"}, }, "required": []string{"name", "class", "description"}, } result, _ := ai.StreamObject(ctx, ai.StreamObjectOptions{ Model: model, OutputMode: "array", Schema: schema.NewSimpleJSONSchema(heroSchema), Prompt: "Generate 3 hero descriptions for a fantasy role playing game.", OnChunk: func(partialObject interface{}) { fmt.Printf("Partial array so far: %v\n", partialObject) }, }) // The final array is available once StreamObject returns for _, hero := range result.Array { fmt.Printf("Hero: %v\n", hero) } ``` > **Note:** To receive fully validated elements one at a time as they complete, use the non-deprecated `StreamText` with `ArrayOutput[T]()` and `ai.ElementStreamWithOutput()` shown above under [Element Streaming for Arrays](#element-streaming-for-arrays). ### Enum If you want to generate a specific enum value, e.g. for classification tasks, you can set the output strategy to `enum` and provide a list of possible values in the `EnumValues` parameter. > **Note:** Enum output is only available with `GenerateObject`. ```go result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, OutputMode: "enum", EnumValues: []string{"action", "comedy", "drama", "horror", "sci-fi"}, Prompt: "Classify the genre of this movie plot: " + "\"A group of astronauts travel through a wormhole in search of a " + "new habitable planet for humanity.\"", }) if err != nil { log.Fatal(err) } // result.EnumValue contains the selected enum value fmt.Printf("Genre: %s\n", result.EnumValue) // Output: sci-fi ``` ### No Schema In some cases, you might not want to use a schema, for example when the data is a dynamic user request. You can use the `OutputMode` setting to set the output format to `no-schema` in those cases and omit the schema parameter. ```go result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, OutputMode: "no-schema", Prompt: "Generate a lasagna recipe.", }) if err != nil { log.Fatal(err) } // result.Object contains the parsed value without schema validation fmt.Printf("Object: %v\n", result.Object) ``` ## Schema Name and Description You can optionally specify a name and description for the schema. These are used by some providers for additional LLM guidance, e.g. via tool or schema name. ```go recipeSchema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "ingredients": map[string]interface{}{ "type": "array", "items": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "amount": map[string]interface{}{"type": "string"}, }, }, }, "steps": map[string]interface{}{ "type": "array", "items": map[string]interface{}{"type": "string"}, }, }, "required": []string{"name", "ingredients", "steps"}, } result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, SchemaName: "Recipe", SchemaDescription: "A recipe for a dish.", Schema: schema.NewSimpleJSONSchema(recipeSchema), Prompt: "Generate a lasagna recipe.", }) ``` ## Error Handling When `GenerateObject` cannot generate a valid object, it returns an error. This error occurs when the AI provider fails to generate a parsable object that conforms to the schema. It can arise due to the following reasons: - The model failed to generate a response - The model generated a response that could not be parsed - The model generated a response that could not be validated against the schema ```go import providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema, Prompt: "Generate data", }) if err != nil { // Check for specific error types var noObjectErr *ai.NoObjectGeneratedError if errors.As(err, &noObjectErr) { fmt.Println("No object was generated:", noObjectErr.Message) } else if providererrors.IsValidationError(err) { fmt.Println("Object validation failed") } else { fmt.Printf("Error: %v\n", err) } return } ``` ## Advanced Usage ### Using Go Structs for Type Safety For better type safety, define Go structs and generate schemas from them: ```go type Person struct { Name string `json:"name"` Age int `json:"age"` Email string `json:"email"` Address Address `json:"address"` } type Address struct { Street string `json:"street"` City string `json:"city"` Country string `json:"country"` } // Generate a JSON schema directly from the struct's json tags result, _ := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: ai.SchemaFor[Person](), Prompt: "Generate a person profile.", }) // result.Object is already unmarshaled; re-marshal it to decode into a typed struct objBytes, _ := json.Marshal(result.Object) var person Person json.Unmarshal(objBytes, &person) fmt.Printf("Person: %+v\n", person) ``` ### Information Extraction Extract structured information from unstructured text: ```go articleSchema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "title": map[string]interface{}{"type": "string"}, "author": map[string]interface{}{"type": "string"}, "date": map[string]interface{}{"type": "string"}, "summary": map[string]interface{}{"type": "string"}, "keywords": map[string]interface{}{ "type": "array", "items": map[string]interface{}{"type": "string"}, }, }, } article := "..." // your article text result, _ := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema.NewSimpleJSONSchema(articleSchema), Prompt: fmt.Sprintf("Extract key information from this article: %s", article), }) // result.Object is already unmarshaled; re-marshal it to decode into a typed value objBytes, _ := json.Marshal(result.Object) var extractedInfo map[string]interface{} json.Unmarshal(objBytes, &extractedInfo) fmt.Printf("Extracted: %+v\n", extractedInfo) ``` ### Classification Use enum mode for classification tasks: ```go result, _ := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, OutputMode: "enum", EnumValues: []string{"positive", "negative", "neutral"}, Prompt: "Classify the sentiment of this review: \"This product is amazing!\"", }) fmt.Printf("Sentiment: %s\n", result.EnumValue) // Output: positive ``` ### Synthetic Data Generation Generate synthetic test data: ```go userSchema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "username": map[string]interface{}{"type": "string"}, "email": map[string]interface{}{"type": "string"}, "age": map[string]interface{}{"type": "integer", "minimum": 18, "maximum": 80}, "role": map[string]interface{}{"type": "string", "enum": []string{"admin", "user", "moderator"}}, }, } result, _ := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, OutputMode: "array", Schema: schema.NewSimpleJSONSchema(userSchema), Prompt: "Generate 5 realistic test users for a web application.", }) // For array mode, elements are available in result.Array for i, user := range result.Array { fmt.Printf("User %d: %+v\n", i+1, user) } ``` ## Callbacks ### OnFinish Callback Monitor when object generation completes: ```go result, _ := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema, Prompt: "Generate data", OnFinish: func(ctx context.Context, result *ai.GenerateObjectResult, userContext interface{}) { fmt.Printf("Object generated\n") fmt.Printf("Usage: %+v\n", result.Usage) fmt.Printf("Finish reason: %s\n", result.FinishReason) }, }) ``` ### OnChunk Callback (Streaming) Process partial objects as they arrive during streaming (note: `StreamObject` blocks until generation completes; `OnChunk` is a callback, not a channel): ```go result, _ := ai.StreamObject(ctx, ai.StreamObjectOptions{ Model: model, Schema: schema, Prompt: "Generate data", OnChunk: func(partialObject interface{}) { fmt.Printf("Partial: %v\n", partialObject) }, }) ``` ## Validation The Go AI SDK automatically validates generated objects against your schema. If validation fails, an error is returned: ```go import providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema, Prompt: "Generate data", }) if err != nil { if providererrors.IsValidationError(err) { fmt.Println("Generated object does not match schema") // Handle validation failure } } ``` ## Best Practices 1. **Define Clear Schemas**: Provide detailed schemas with descriptions for better results 2. **Use Enums for Classification**: Use enum mode for fixed-choice classification tasks 3. **Stream Long Objects**: Use `StreamObject` for large or complex objects 4. **Handle Errors**: Always check for validation and generation errors 5. **Use Type-Safe Structs**: Define Go structs for compile-time type safety 6. **Add Descriptions**: Include field descriptions in your schema to guide the model 7. **Test Schemas**: Validate your schemas work correctly before production use ## Common Patterns ### Multi-Entity Extraction Extract multiple entities from text: ```go entitiesSchema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "people": map[string]interface{}{ "type": "array", "items": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "role": map[string]interface{}{"type": "string"}, }, }, }, "organizations": map[string]interface{}{ "type": "array", "items": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "type": map[string]interface{}{"type": "string"}, }, }, }, "locations": map[string]interface{}{ "type": "array", "items": map[string]interface{}{"type": "string"}, }, }, } text := "Apple CEO Tim Cook announced a new initiative in Cupertino..." result, _ := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema.NewSimpleJSONSchema(entitiesSchema), Prompt: fmt.Sprintf("Extract all people, organizations, and locations from: %s", text), }) ``` ### Form Data Extraction Extract form data from images or documents: ```go formSchema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "email": map[string]interface{}{"type": "string"}, "phone": map[string]interface{}{"type": "string"}, "address": map[string]interface{}{"type": "string"}, "dob": map[string]interface{}{"type": "string"}, }, } result, _ := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema.NewSimpleJSONSchema(formSchema), Prompt: "Extract form data from this document: ...", }) ``` ## Next Steps - Learn about [Tools and Tool Calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - Explore [Embeddings](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) - See [Examples](https://github.com/digitallysavvy/go-ai/tree/main/examples) ## See Also - [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - [Settings](https://goaisdk.com/docs/ai-sdk-core/settings.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [API Reference: GenerateObject](https://goaisdk.com/docs/reference/ai/generate-object.md) - [API Reference: StreamObject](https://goaisdk.com/docs/reference/ai/stream-object.md) --- # Tool Calling > Covers multi-step tool calling in the Go AI SDK: strict mode, response messages, tool choice, execution options, timeouts, and packaging complex tools. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling Documentation index: https://goaisdk.com/llms.txt As covered under [Foundations](https://goaisdk.com/docs/foundations/tools.md), [tools](https://goaisdk.com/docs/foundations/tools.md) are objects that can be called by the model to perform a specific task. Go AI SDK Core tools contain several core elements: - **`Name`**: The name of the tool - **`Description`**: An optional description of the tool that can influence when the tool is picked - **`Parameters`**: A JSON schema that defines the input parameters. The schema is consumed by the LLM and also used to validate the LLM tool calls - **`Execute`**: An optional function that is called with the inputs from the tool call. It is optional because you might want to forward tool calls to the client or to a queue instead of executing them in the same process - **`Strict`**: _(optional)_ Enables strict tool calling when supported by the provider The `Tools` parameter of `GenerateText` and `StreamText` is a slice of tools: ```go package main import ( "context" "fmt" "log" "math/rand" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-5.2") weatherTool := types.Tool{ Name: "weather", Description: "Get the weather in a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "The location to get the weather for", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location := input["location"].(string) return map[string]interface{}{ "location": location, "temperature": 72 + rand.Intn(21) - 10, }, nil }, } maxSteps := 5 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: []types.Tool{weatherTool}, MaxSteps: &maxSteps, Prompt: "What is the weather in San Francisco?", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` > **Note:** When a model uses a tool, it is called a "tool call" and the output of the tool is called a "tool result". ## Strict Mode When enabled, language model providers that support strict tool calling will only generate tool calls that are valid according to your defined `Parameters` schema. This increases the reliability of tool calling. However, not all schemas may be supported in strict mode, and what is supported depends on the specific provider. By default, strict mode is disabled. You can enable it per-tool by setting `Strict: types.BoolPtr(true)`: ```go weatherTool := types.Tool{ Name: "weather", Description: "Get the weather in a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{"type": "string"}, }, "required": []string{"location"}, }, Strict: types.BoolPtr(true), // Enable strict validation for this tool Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // ... return nil, nil }, } ``` > **Note:** Not all providers or models support strict mode. For those that do not, this option is ignored. ## Multi-Step Calls With the `MaxSteps` setting, you can enable multi-step calls in `GenerateText` and `StreamText`. When `MaxSteps` is set and the model generates a tool call, the AI SDK will trigger a new generation passing in the tool result until there are no further tool calls or the maximum steps is reached. > **Note:** The step limit is only evaluated when the last step contains tool results. By default, when you use `GenerateText` or `StreamText`, it triggers a single generation. This works well for many use cases where you can rely on the model's training data to generate a response. However, when you provide tools, the model now has the choice to either generate a normal text response, or generate a tool call. If the model generates a tool call, its generation is complete and that step is finished. You may want the model to generate text after the tool has been executed, either to summarize the tool results in the context of the user's query. In many cases, you may also want the model to use multiple tools in a single response. This is where multi-step calls come in. ### Example In the following example, there are two steps: 1. **Step 1** 1. The prompt `"What is the weather in San Francisco?"` is sent to the model 2. The model generates a tool call 3. The tool call is executed 2. **Step 2** 1. The tool result is sent to the model 2. The model generates a response considering the tool result ```go maxSteps := 5 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: []types.Tool{ { Name: "weather", Description: "Get the weather in a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "The location to get the weather for", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location := input["location"].(string) return map[string]interface{}{ "location": location, "temperature": 72 + rand.Intn(21) - 10, }, nil }, }, }, MaxSteps: &maxSteps, // Stop after a maximum of 5 steps if tools were called Prompt: "What is the weather in San Francisco?", }) ``` > **Note:** You can use `StreamText` in a similar way. ### Steps To access intermediate tool calls and results, you can use the `Steps` field in the result object or the `OnFinish` callback. It contains all the text, tool calls, tool results, and more from each step. #### Example: Extract tool results from all steps ```go maxSteps := 10 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, MaxSteps: &maxSteps, Tools: tools, Prompt: "What's the weather in Tokyo and Paris?", }) if err != nil { log.Fatal(err) } // Extract all tool calls from the steps var allToolCalls []types.ToolCall for _, step := range result.Steps { allToolCalls = append(allToolCalls, step.ToolCalls...) } fmt.Printf("Total tool calls: %d\n", len(allToolCalls)) ``` ### OnStepFinish Callback When using `GenerateText` or `StreamText`, you can provide an `OnStepFinish` callback that is triggered when a step is finished, i.e. all text deltas, tool calls, and tool results for the step are available. When you have multiple steps, the callback is triggered for each step. ```go maxSteps := 5 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: tools, MaxSteps: &maxSteps, Prompt: "What's the weather in multiple cities?", OnStepFinish: func(ctx context.Context, step types.StepResult, userContext interface{}) { // Your own logic, e.g. for saving the chat history or recording usage fmt.Printf("Step %d finished\n", step.StepNumber) fmt.Printf(" Text: %s\n", step.Text) fmt.Printf(" Tool calls: %d\n", len(step.ToolCalls)) fmt.Printf(" Tool results: %d\n", len(step.ToolResults)) fmt.Printf(" Finish reason: %s\n", step.FinishReason) fmt.Printf(" Usage: %+v\n", step.Usage) }, }) ``` ## Response Messages Adding the generated assistant and tool messages to your conversation history is a common task, especially if you are using multi-step tool calls. Both `GenerateText` and `StreamText` have a `Response.Messages` property that you can use to add the assistant and tool messages to your conversation history. It is also available in the `OnFinish` callback of `StreamText`. The `Response.Messages` field contains a slice of `Message` objects that you can add to your conversation history: ```go import "github.com/digitallysavvy/go-ai/pkg/provider/types" maxSteps := 5 messages := []types.Message{ {Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "What's the weather?"}}}, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: messages, Tools: tools, MaxSteps: &maxSteps, }) if err != nil { log.Fatal(err) } // Add the response messages from the last step to your conversation history if len(result.Steps) > 0 { lastStep := result.Steps[len(result.Steps)-1] messages = append(messages, lastStep.ResponseMessages...) } // Continue the conversation result2, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: messages, Tools: tools, MaxSteps: &maxSteps, }) ``` ## Tool Choice You can use the `ToolChoice` setting to influence when a tool is selected. It supports the following settings: - `auto` (default): the model can choose whether and which tools to call - `required`: the model must call a tool. It can choose which tool to call - `none`: the model must not call tools - `tool`: the model must call the specified tool ```go // Auto: Let the model decide (default) result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: []types.Tool{weatherTool}, ToolChoice: types.ToolChoice{Type: "auto"}, Prompt: "What's the weather?", }) // Required: Force the model to call a tool result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: []types.Tool{weatherTool}, ToolChoice: types.ToolChoice{Type: "required"}, Prompt: "What's the weather in San Francisco?", }) // None: Prevent the model from using tools result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: []types.Tool{weatherTool}, ToolChoice: types.ToolChoice{Type: "none"}, Prompt: "Tell me about San Francisco.", }) // Specific: Force the model to use a specific tool result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: []types.Tool{weatherTool, calculatorTool}, ToolChoice: types.ToolChoice{ Type: "tool", ToolName: "weather", }, Prompt: "What's the weather?", }) ``` ## Tool Execution Options When tools are called, they receive the context and input parameters. You can use the context for cancellation, timeouts, and passing request-scoped values. ### Context Cancellation The abort signals from `GenerateText` and `StreamText` are forwarded to the tool execution via the context: ```go // Create a cancellable context ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() weatherTool := types.Tool{ Name: "weather", Description: "Get the weather in a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{"type": "string"}, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location := input["location"].(string) // Create HTTP request with context for cancellation req, err := http.NewRequestWithContext( ctx, "GET", fmt.Sprintf("https://api.weatherapi.com/v1/current.json?q=%s", location), nil, ) if err != nil { return nil, err } resp, err := http.DefaultClient.Do(req) if err != nil { return nil, err } defer resp.Body.Close() // Parse and return weather data var data map[string]interface{} json.NewDecoder(resp.Body).Decode(&data) return data, nil }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: []types.Tool{weatherTool}, Prompt: "What's the weather in San Francisco?", }) ``` ## Tool Timeouts Set timeouts for tool execution to prevent long-running tools from blocking the step loop. The Go AI SDK provides both global and per-tool timeout configuration. ### Global Tool Timeout Set a default timeout for all tools using `ToolMs`: ```go import ( "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) fiveSeconds := 5 * time.Second result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather?", Tools: []types.Tool{weatherTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, Timeout: &ai.TimeoutConfig{ ToolMs: &fiveSeconds, // 5 second timeout for all tools }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` ### Per-Tool Timeout Override timeouts for individual tools using the `Tools` map. Per-tool timeouts take precedence over the global `ToolMs`: ```go import "time" fiveSeconds := 5 * time.Second result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Search for weather and news", Tools: []types.Tool{weatherTool, searchTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, Timeout: &ai.TimeoutConfig{ ToolMs: &fiveSeconds, // Default: 5s for all tools Tools: map[string]time.Duration{ "get_weather": 5 * time.Second, // 5s for weather "search_web": 15 * time.Second, // 15s for web search }, }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` ### Timeout Resolution Order The `GetToolTimeout()` helper resolves the effective timeout for a tool name using this priority: 1. **Per-tool override** (`Tools[toolName]`) — highest priority 2. **Global default** (`ToolMs`) — used if no per-tool entry exists 3. **No timeout** — if neither is set, the tool runs without a timeout ```go timeout := &ai.TimeoutConfig{ ToolMs: &fiveSeconds, Tools: map[string]time.Duration{ "slow_tool": 30 * time.Second, }, } timeout.GetToolTimeout("slow_tool") // 30s (per-tool override) timeout.GetToolTimeout("fast_tool") // 5s (global default) timeout.GetToolTimeout("another_tool") // 5s (global default) ``` ### Timeout Behavior When a tool exceeds its timeout, the tool's context is cancelled and the `Execute` function receives a `context.DeadlineExceeded` error. Tools that respect context cancellation will stop immediately: ```go Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Long-running operation that respects context req, err := http.NewRequestWithContext(ctx, "GET", apiURL, nil) if err != nil { return nil, err } resp, err := http.DefaultClient.Do(req) if err != nil { // Returns context.DeadlineExceeded if the tool timeout fires return nil, fmt.Errorf("request failed: %w", err) } defer resp.Body.Close() // Process response... return result, nil }, ``` > **Tip:** Always use `http.NewRequestWithContext(ctx, ...)` and pass the context to external calls so they respect tool timeouts and cancellation. ## Complex Tool Examples ### Multiple Tools Using multiple tools together: ```go calculatorTool := types.Tool{ Name: "calculator", Description: "Perform mathematical calculations", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "operation": map[string]interface{}{ "type": "string", "enum": []string{"add", "subtract", "multiply", "divide"}, }, "a": map[string]interface{}{"type": "number"}, "b": map[string]interface{}{"type": "number"}, }, "required": []string{"operation", "a", "b"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { op := input["operation"].(string) a := input["a"].(float64) b := input["b"].(float64) var result float64 switch op { case "add": result = a + b case "subtract": result = a - b case "multiply": result = a * b case "divide": if b == 0 { return nil, fmt.Errorf("division by zero") } result = a / b } return map[string]interface{}{"result": result}, nil }, } searchTool := types.Tool{ Name: "web_search", Description: "Search the web for information", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{ "type": "string", "description": "The search query", }, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := input["query"].(string) // Call search API return map[string]interface{}{ "results": []string{ "Result 1 for: " + query, "Result 2 for: " + query, }, }, nil }, } maxSteps := 5 result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: []types.Tool{calculatorTool, searchTool, weatherTool}, Prompt: "Search for the population of Tokyo, then calculate what " + "percentage it is of Japan's total population (125 million)", MaxSteps: &maxSteps, }) fmt.Println(result.Text) ``` ### Tool with Error Handling Robust tool implementation with comprehensive error handling: ```go weatherTool := types.Tool{ Name: "get_weather", Description: "Get current weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "City name", }, "units": map[string]interface{}{ "type": "string", "enum": []string{"celsius", "fahrenheit"}, "default": "celsius", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Validate input location, ok := input["location"].(string) if !ok || location == "" { return nil, fmt.Errorf("location must be a non-empty string") } units := "celsius" if u, ok := input["units"].(string); ok { units = u } // Check for context cancellation select { case <-ctx.Done(): return nil, ctx.Err() default: } // Call weather API apiKey := os.Getenv("WEATHER_API_KEY") url := fmt.Sprintf( "https://api.weatherapi.com/v1/current.json?key=%s&q=%s", apiKey, location, ) req, err := http.NewRequestWithContext(ctx, "GET", url, nil) if err != nil { return nil, fmt.Errorf("failed to create request: %w", err) } resp, err := http.DefaultClient.Do(req) if err != nil { return nil, fmt.Errorf("failed to fetch weather: %w", err) } defer resp.Body.Close() if resp.StatusCode != http.StatusOK { return nil, fmt.Errorf("weather API returned status: %d", resp.StatusCode) } var data map[string]interface{} if err := json.NewDecoder(resp.Body).Decode(&data); err != nil { return nil, fmt.Errorf("failed to decode response: %w", err) } // Extract and format weather data current := data["current"].(map[string]interface{}) temp := current["temp_c"].(float64) if units == "fahrenheit" { temp = temp*9/5 + 32 } return map[string]interface{}{ "location": location, "temperature": temp, "units": units, "condition": current["condition"].(map[string]interface{})["text"], }, nil }, } ``` ### Tool with Streaming Progress For long-running operations, you can provide progress updates: ```go dataProcessingTool := types.Tool{ Name: "process_data", Description: "Process a large dataset", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "dataset_id": map[string]interface{}{"type": "string"}, }, "required": []string{"dataset_id"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { datasetID := input["dataset_id"].(string) // Simulate long-running processing with progress updates total := 100 for i := 0; i < total; i++ { // Check for cancellation select { case <-ctx.Done(): return nil, ctx.Err() default: } // Simulate work time.Sleep(100 * time.Millisecond) // In a real implementation, you might send progress // via a channel or callback mechanism if i%10 == 0 { log.Printf("Processing: %d%% complete", i) } } return map[string]interface{}{ "status": "completed", "dataset_id": datasetID, "records": 1000, }, nil }, } ``` ## Tool Packaging Create reusable tool packages: ```go // weathertools/weather.go package weathertools import ( "context" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func NewWeatherTool(apiKey string) types.Tool { return types.Tool{ Name: "get_weather", Description: "Get current weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Implementation using apiKey return nil, nil }, } } func NewCalculatorTool() types.Tool { return types.Tool{ Name: "calculator", Description: "Perform calculations", Parameters: map[string]interface{}{ // ... }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Implementation return nil, nil }, } } ``` Usage: ```go skip-compile import "myapp/weathertools" weatherTool := weathertools.NewWeatherTool(apiKey) calcTool := weathertools.NewCalculatorTool() result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: []types.Tool{weatherTool, calcTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, Prompt: "What's the weather and calculate something", }) ``` ## Error Handling Handle tool execution errors: ```go maxSteps := 5 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: tools, MaxSteps: &maxSteps, Prompt: "Use tools", }) if err != nil { // Check for specific error types var noSuchToolErr *ai.NoSuchToolError var invalidInputErr *ai.InvalidToolInputError if errors.As(err, &noSuchToolErr) { fmt.Println("Model tried to call unknown tool") } else if errors.As(err, &invalidInputErr) { fmt.Println("Model called tool with invalid inputs") } else { fmt.Printf("Error: %v\n", err) } return } // Check for tool errors in the steps for _, step := range result.Steps { for _, toolResult := range step.ToolResults { if toolResult.Error != nil { fmt.Printf("Tool %s failed: %v\n", toolResult.ToolName, toolResult.Error) } } } ``` ## Best Practices 1. **Clear Descriptions**: Write clear tool descriptions to help the model know when to use them 2. **Schema Validation**: Define comprehensive parameter schemas with descriptions 3. **Error Handling**: Return clear error messages when tools fail 4. **Context Support**: Respect context cancellation in long-running tools 5. **Idempotency**: Make tools idempotent when possible 6. **Logging**: Log tool executions for debugging 7. **Rate Limiting**: Implement rate limiting for external API calls 8. **Type Safety**: Use Go structs for input validation 9. **Timeouts**: Set appropriate timeouts for external calls 10. **Testing**: Write unit tests for tool execution logic ## Concurrency and Parallel Tool Execution The Go AI SDK executes tools sequentially by default, but you can implement parallel execution patterns: ```go type ParallelToolExecutor struct { maxConcurrent int } func (e *ParallelToolExecutor) ExecuteTools( ctx context.Context, toolCalls []types.ToolCall, tools map[string]types.Tool, ) []types.ToolResult { results := make([]types.ToolResult, len(toolCalls)) var wg sync.WaitGroup sem := make(chan struct{}, e.maxConcurrent) for i, toolCall := range toolCalls { wg.Add(1) go func(idx int, tc types.ToolCall) { defer wg.Done() // Acquire semaphore sem <- struct{}{} defer func() { <-sem }() tool, ok := tools[tc.ToolName] if !ok { results[idx] = types.ToolResult{ ToolCallID: tc.ID, ToolName: tc.ToolName, Error: fmt.Errorf("tool not found: %s", tc.ToolName), } return } output, err := tool.Execute(ctx, tc.Arguments, types.ToolExecutionOptions{ ToolCallID: tc.ID, }) results[idx] = types.ToolResult{ ToolCallID: tc.ID, ToolName: tc.ToolName, Result: output, Error: err, } }(i, toolCall) } wg.Wait() return results } ``` ## Advanced Patterns ### Conditional Tool Selection Dynamically select tools based on context: ```go func getAvailableTools(userRole string) []types.Tool { basicTools := []types.Tool{weatherTool, calculatorTool} if userRole == "admin" { return append(basicTools, adminTools...) } return basicTools } result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: getAvailableTools(user.Role), StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, Prompt: prompt, }) ``` ### Tool Chaining Chain tool results: ```go maxSteps := 10 result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: []types.Tool{searchTool, summarizeTool, translateTool}, Prompt: "Search for AI news, summarize it, then translate to Spanish", MaxSteps: &maxSteps, OnStepFinish: func(ctx context.Context, step types.StepResult, userContext interface{}) { fmt.Printf("Step %d: Used %d tools\n", step.StepNumber, len(step.ToolCalls)) }, }) ``` ### Deferred Tool Search For large tool registries, mark tools `DeferLoading: true` and add `ai.ToolSearch()` so the model discovers only the tools it needs, instead of seeing every definition up front. By default, `ai.ToolSearch()` ranks deferred tools with built-in keyword scoring over their names and descriptions, returning up to five matches per call: ```go tools := []types.Tool{ ai.ToolSearch(), { Name: "getWeather", DeferLoading: true, Description: "Get the current weather for a city", // ... }, } ``` Use `ToolSearchConfig.MaxResults` to change how many matches a search returns (must be a positive integer; omitted or zero defaults to five), and `ToolSearchConfig.Search` to replace the built-in keyword scoring with your own ranking function: ```go tools := []types.Tool{ ai.ToolSearch(ai.ToolSearchConfig{ MaxResults: 10, Search: func(ctx context.Context, query string, candidates []ai.ToolSearchCandidate) ([]string, error) { // Return matching tool names in ranked order, e.g. from a // vector search or external index. Unknown names and // duplicates are ignored, and the result is capped at // MaxResults. return rankToolsByEmbedding(ctx, query, candidates) }, }), // ... deferred tools } ``` ## Next Steps - Learn about [Embeddings](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) - Explore [Agents](https://goaisdk.com/docs/agents/overview.md) - See [Tool Examples](https://github.com/digitallysavvy/go-ai/tree/main/examples) ## See Also - [Foundations: Tools](https://goaisdk.com/docs/foundations/tools.md) - [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - [Settings](https://goaisdk.com/docs/ai-sdk-core/settings.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [API Reference: Tool Types](https://goaisdk.com/docs/reference/types/tools.md) --- # Model Context Protocol (MCP) > Learn how to connect to Model Context Protocol (MCP) servers and use their tools with Go AI SDK. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/mcp-tools Documentation index: https://goaisdk.com/llms.txt The Go AI SDK supports connecting to [Model Context Protocol (MCP)](https://modelcontextprotocol.io/) servers to access their tools, resources, and prompts. This enables your AI applications to discover and use capabilities across various services through a standardized interface. ## Overview MCP (Model Context Protocol) is an open protocol that allows AI applications to: - **Discover Tools**: Automatically detect available functions from MCP servers - **Execute Tools**: Call remote functions and get structured results - **Access Resources**: Read data from various sources (files, databases, APIs) - **Use Prompts**: Retrieve reusable prompt templates ## Installation The MCP package is included in the Go AI SDK: ```bash go get github.com/digitallysavvy/go-ai ``` ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/mcp" ) ``` ## Initializing an MCP Client We recommend using HTTP transport for production deployments. The stdio transport should only be used for connecting to local servers as it cannot be deployed to production environments. ### HTTP Transport (Recommended) For production deployments, use the HTTP transport: ```go import ( "context" "github.com/digitallysavvy/go-ai/pkg/mcp" ) ctx := context.Background() // Create HTTP transport transport := mcp.NewHTTPTransport(mcp.HTTPTransportConfig{ URL: "https://your-server.com/mcp", TimeoutMS: 30000, // Optional: Add custom headers (set on the base transport config) Config: mcp.TransportConfig{ Headers: map[string]string{ "Authorization": "Bearer my-api-key", }, }, }) // Create client client := mcp.NewMCPClient(transport, mcp.MCPClientConfig{ ClientName: "my-go-app", ClientVersion: "1.0.0", }) // Connect err := client.Connect(ctx) if err != nil { log.Fatal(err) } defer client.Close() ``` ### HTTP with OAuth Authentication For authenticated MCP servers: ```go transport := mcp.NewHTTPTransport(mcp.HTTPTransportConfig{ URL: "https://mcp.example.com", TimeoutMS: 30000, OAuth: &mcp.OAuthConfig{ ClientID: "your-client-id", ClientSecret: "your-client-secret", TokenURL: "https://auth.example.com/token", }, }) client := mcp.NewMCPClient(transport, mcp.MCPClientConfig{ ClientName: "my-go-app", ClientVersion: "1.0.0", }) err := client.Connect(ctx) if err != nil { log.Fatal(err) } defer client.Close() ``` ### Custom OAuth Client Providers For full control over credential storage (e.g. a database-backed provider shared by multiple processes), implement `mcp.OAuthClientProvider` and drive the flow yourself with `mcp.Auth`. If the server rotates its OAuth issuer, rediscovered authorization server metadata may no longer match the metadata that issued your stored credentials. `mcp.Auth` surfaces this as a permanent failure, detectable with `mcp.IsAuthorizationServerMismatchError`: ```go result, err := mcp.Auth(ctx, provider, mcp.AuthOptions{ServerURL: serverURL}) if mcp.IsAuthorizationServerMismatchError(err) { // Treat as permanent: drop credentials and start a fresh authorization // flow rather than retrying with the stale pin. } ``` If your provider also implements `mcp.OAuthTokenInvalidator`, `mcp.Auth` identifies the exact token generation a failed refresh attempted, letting providers that share storage across clients delete only that stale generation instead of unconditionally clearing storage — so a concurrent refresh's newer tokens survive: ```go func (p *myOAuthProvider) InvalidateCredentialsForTokens(ctx context.Context, tokens mcp.OAuthTokens) error { if p.tokens != nil && p.tokens.AccessToken == tokens.AccessToken && p.tokens.RefreshToken == tokens.RefreshToken { p.tokens = nil } return nil } ``` ### HTTP Error Handling HTTP transport failures return `*mcp.MCPClientError` with structured transport fields. `Code` remains the JSON-RPC application error code; `StatusCode`, `URL`, and `ResponseBody` describe the HTTP failure. ```go err := client.Connect(ctx) if err != nil { var clientErr *mcp.MCPClientError if errors.As(err, &clientErr) && clientErr.StatusCode == http.StatusUnauthorized { log.Printf("MCP auth failed for %s: %s", clientErr.URL, clientErr.ResponseBody) return } log.Fatal(err) } ``` ### Stdio Transport (Local Servers Only) > **Warning:** The stdio transport should only be used for local development with MCP servers running as child processes. ```go // Create stdio transport for local Node.js MCP server transport := mcp.NewStdioTransport(mcp.StdioTransportConfig{ Command: "node", Args: []string{"server.mjs"}, // Optional: Set working directory WorkingDir: "/path/to/mcp/server", // Optional: Set environment variables Env: []string{ "API_KEY=your-api-key", }, }) client := mcp.NewMCPClient(transport, mcp.MCPClientConfig{ ClientName: "my-go-app", ClientVersion: "1.0.0", }) err := client.Connect(ctx) if err != nil { log.Fatal(err) } defer client.Close() ``` ### Closing the MCP Client After initialization, you should close the MCP client based on your usage pattern: - **Short-lived usage**: Close the client when the response is finished - **Long-running clients**: Keep the client open but ensure it's closed when the application terminates ```go // For short-lived usage client := setupMCPClient() defer client.Close() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: tools, Prompt: "Your prompt here", }) // Client closes automatically via defer ``` ```go // For long-running applications client := setupMCPClient() // Register cleanup handler c := make(chan os.Signal, 1) signal.Notify(c, os.Interrupt, syscall.SIGTERM) go func() { <-c client.Close() os.Exit(0) }() // Use client throughout application lifecycle ``` ## Using MCP Tools The Go AI SDK provides two approaches for working with MCP tools: ### Approach 1: Schema Discovery (Recommended) With schema discovery, all tools offered by the server are automatically listed and converted to Go AI SDK tools: ```go import ( "context" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/mcp" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) ctx := context.Background() // Set up MCP client transport := mcp.NewHTTPTransport(mcp.HTTPTransportConfig{ URL: "https://mcp.example.com", }) client := mcp.NewMCPClient(transport, mcp.MCPClientConfig{ ClientName: "my-app", ClientVersion: "1.0.0", }) err := client.Connect(ctx) if err != nil { log.Fatal(err) } defer client.Close() // Convert MCP tools to Go AI SDK tools converter := mcp.NewMCPToolConverter(client) tools, err := converter.ConvertToGoAITools(ctx) if err != nil { log.Fatal(err) } // Use tools with AI model openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := openaiProvider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in Brooklyn, NY?", Tools: tools, // All MCP tools are available // Allow a follow-up step so the model can turn the tool result into a // final answer (GenerateText/StreamText default to a single step). StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` This approach is simpler and automatically stays in sync with server changes. However, you won't have compile-time type safety for tool parameters. ### Approach 2: Schema Definition For better type safety and control, you can define specific tools and their schemas: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) // Define tool schema weatherTool := types.Tool{ Name: "get-weather", Description: "Get weather information for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]string{ "type": "string", "description": "City name", }, "units": map[string]interface{}{ "type": "string", "enum": []string{"celsius", "fahrenheit"}, "description": "Temperature units", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, args map[string]interface{}, _ types.ToolExecutionOptions) (interface{}, error) { // Call MCP tool manually result, err := client.CallTool(ctx, "get-weather", args) if err != nil { return nil, err } // Convert result content, err := mcp.ConvertMCPContentToAISDK(result.Content) if err != nil { return nil, err } return content, nil }, } // Use with AI model result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the temperature in Paris?", Tools: []types.Tool{weatherTool}, // Allow a follow-up step so the model can turn the tool result into a // final answer (GenerateText/StreamText default to a single step). StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) ``` This approach provides full type safety and lets you catch parameter mismatches during development. ## Image Content Handling > **Important:** The Go AI SDK automatically handles image content from MCP tools to prevent token explosions (200K+ tokens per image). When MCP tools return images (screenshots, charts, diagrams), the SDK properly converts them to the AI SDK's native image format: ```go // MCP server returns tool result with image // { // "type": "image", // "data": "iVBORw0KGgo...", // base64 data // "mimeType": "image/png" // } // Automatically converted to ImageContent // - Base64 decoded to bytes // - Token usage: ~1K instead of 200K+ // - Cost savings: ~99% reduction // Use tools that return images normally: tools, _ := converter.ConvertToGoAITools(ctx) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Create a sales chart and analyze it", Tools: tools, // Tool may return image StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) // Images in results are properly handled // No token explosion ``` ### Supported Image Formats The SDK handles all standard image formats: 1. **Base64 data**: Raw base64-encoded image 2. **Data URLs**: `data:image/png;base64,...` 3. **HTTP/HTTPS URLs**: Remote image URLs ### Performance Impact - **Without image handling**: 200K+ tokens per image - **With proper handling**: ~1K tokens per image - **Cost savings**: large reduction in image tokens sent to the model ## Using MCP Resources Resources are application-driven data sources that provide context to the model. Unlike tools (which are model-controlled), your application decides when to fetch and pass resources as context. ### Listing Resources List all available resources from the MCP server: ```go resources, err := client.ListResources(ctx) if err != nil { log.Fatal(err) } for _, resource := range resources { fmt.Printf("Resource: %s (%s)\n", resource.Name, resource.URI) fmt.Printf(" Description: %s\n", resource.Description) fmt.Printf(" MIME Type: %s\n", resource.MimeType) } ``` ### Reading Resource Contents Read the contents of a specific resource by its URI: ```go resourceData, err := client.ReadResource(ctx, "file:///example/document.txt") if err != nil { log.Fatal(err) } // Use resource content in prompt result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Summarize this document:\n\n%s", resourceData.Contents[0].Text), }) ``` ### Listing Resource Templates Resource templates are dynamic URI patterns that allow flexible queries: ```go templates, err := client.ListResourceTemplates(ctx) if err != nil { log.Fatal(err) } for _, template := range templates.ResourceTemplates { fmt.Printf("Template: %s\n", template.Name) fmt.Printf(" URI Template: %s\n", template.URITemplate) fmt.Printf(" Description: %s\n", template.Description) } ``` ## Handling Elicitation Requests Some MCP servers pause a tool call to ask the client (your application) for additional user input mid-execution, using the `elicitation/create` request. Register a handler with `OnElicitationRequest` before calling `Connect`: ```go client.OnElicitationRequest(func(ctx context.Context, req mcp.ElicitationRequest) (mcp.ElicitResult, error) { fmt.Println("Server asks:", req.Message) // req.RequestedSchema describes the JSON Schema for the expected input. answer := promptUserSomehow(req) return mcp.ElicitResult{ Action: "accept", // or "decline" / "cancel" Content: map[string]interface{}{"answer": answer}, }, nil }) err := client.Connect(ctx) ``` If no handler is registered, the client rejects `elicitation/create` requests automatically. A handler error is reported through the client's `OnError` callback rather than propagating to the tool call. ## Using MCP Prompts > **Warning:** MCP Prompts is an experimental feature and may change in the future. Prompts are user-controlled templates that servers expose for clients to list and retrieve with optional arguments. ### Listing Prompts ```go prompts, err := client.ListPrompts(ctx) if err != nil { log.Fatal(err) } for _, prompt := range prompts { fmt.Printf("Prompt: %s\n", prompt.Name) fmt.Printf(" Description: %s\n", prompt.Description) if len(prompt.Arguments) > 0 { fmt.Println(" Arguments:") for _, arg := range prompt.Arguments { fmt.Printf(" - %s (required: %v): %s\n", arg.Name, arg.Required, arg.Description) } } } ``` ### Getting a Prompt Retrieve prompt messages, optionally passing arguments: ```go prompt, err := client.GetPrompt(ctx, "code_review", map[string]interface{}{ "code": "func add(a, b int) int { return a + b }", "language": "go", }) if err != nil { log.Fatal(err) } // Convert prompt messages to AI SDK messages messages := make([]types.Message, len(prompt.Messages)) for i, msg := range prompt.Messages { messages[i] = types.Message{ Role: types.MessageRole(msg.Role), Content: []types.ContentPart{ types.TextContent{Text: msg.Content.Text}, }, } } // Use prompt messages with AI model result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: messages, }) ``` ## Complete Example Here's a complete example that demonstrates all MCP features: ```go package main import ( "context" "fmt" "log" "os" "os/signal" "syscall" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/mcp" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() // 1. Set up MCP client with HTTP transport transport := mcp.NewHTTPTransport(mcp.HTTPTransportConfig{ URL: os.Getenv("MCP_SERVER_URL"), TimeoutMS: 30000, Config: mcp.TransportConfig{ Headers: map[string]string{ "Authorization": "Bearer " + os.Getenv("MCP_API_KEY"), }, }, }) client := mcp.NewMCPClient(transport, mcp.MCPClientConfig{ ClientName: "go-ai-example", ClientVersion: "1.0.0", }) err := client.Connect(ctx) if err != nil { log.Fatal(err) } // Ensure cleanup defer client.Close() // Handle interrupt signal c := make(chan os.Signal, 1) signal.Notify(c, os.Interrupt, syscall.SIGTERM) go func() { <-c client.Close() os.Exit(0) }() // 2. List available resources resources, err := client.ListResources(ctx) if err != nil { log.Fatal(err) } fmt.Println("Available resources:") for _, r := range resources { fmt.Printf(" - %s: %s\n", r.Name, r.Description) } // 3. Convert MCP tools to AI SDK tools converter := mcp.NewMCPToolConverter(client) tools, err := converter.ConvertToGoAITools(ctx) if err != nil { log.Fatal(err) } fmt.Printf("\nConverted %d MCP tools\n\n", len(tools)) // 4. Use tools with AI model provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, err := provider.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } maxSteps := 5 // Allow multiple tool calls result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Use the available tools to get the current weather " + "in Brooklyn, New York and create a chart showing " + "temperature trends", Tools: tools, MaxSteps: &maxSteps, }) if err != nil { log.Fatal(err) } // 5. Display results fmt.Println("AI Response:") fmt.Println(result.Text) // Show tool calls made if len(result.Steps) > 0 { fmt.Println("\nTool Calls:") for i, step := range result.Steps { fmt.Printf(" Step %d:\n", i+1) for _, call := range step.ToolCalls { fmt.Printf(" - %s(%v)\n", call.ToolName, call.Arguments) } } } // Display usage statistics fmt.Printf("\nUsage: %d input tokens, %d output tokens\n", result.Usage.GetInputTokens(), result.Usage.GetOutputTokens()) } ``` ## Error Handling Handle common MCP errors gracefully: ```go import ( "errors" "github.com/digitallysavvy/go-ai/pkg/mcp" ) // Connect with error handling err := client.Connect(ctx) if err != nil { var transportErr *mcp.TransportError if errors.As(err, &transportErr) { log.Printf("Failed to connect to MCP server: %v", err) // Retry or fallback logic } else { log.Fatal(err) } } // Tool execution with error handling result, err := client.CallTool(ctx, "get-weather", map[string]interface{}{ "location": "New York", }) if err != nil { var clientErr *mcp.MCPClientError var timeoutErr *mcp.TimeoutError if errors.As(err, &timeoutErr) { log.Printf("Tool execution timed out: %v", err) // Retry with longer timeout } else if errors.As(err, &clientErr) { log.Printf("Tool execution failed: %v", err) // Handle tool failure } else { log.Fatal(err) } } else { fmt.Printf("Tool result: %+v\n", result.Content) } ``` ## Best Practices ### 1. Use HTTP Transport in Production ```go // ✅ Good: HTTP transport for production httpTransport := mcp.NewHTTPTransport(mcp.HTTPTransportConfig{ URL: "https://mcp.production.com", }) // ❌ Avoid: Stdio transport in production (won't work) stdioTransport := mcp.NewStdioTransport(mcp.StdioTransportConfig{ Command: "node", // Cannot deploy to cloud }) ``` ### 2. Always Close the Client ```go // ✅ Good: Use defer to ensure cleanup client := setupMCPClient() defer client.Close() // ✅ Good: Handle signals for long-running apps c := make(chan os.Signal, 1) signal.Notify(c, os.Interrupt) go func() { <-c client.Close() os.Exit(0) }() ``` ### 3. Handle Context Cancellation ```go // ✅ Good: Respect context cancellation ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: tools, Prompt: prompt, }) if err != nil { if ctx.Err() == context.DeadlineExceeded { log.Println("Request timed out") } return err } ``` ### 4. Use Schema Discovery for Simplicity ```go // ✅ Good: Simple schema discovery converter := mcp.NewMCPToolConverter(client) tools, err := converter.ConvertToGoAITools(ctx) // Use all available tools result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, Prompt: "Use whatever tools you need", }) ``` ### 5. Monitor Token Usage ```go // ✅ Good: Monitor usage to catch image token explosions result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, Prompt: prompt, }) if err != nil { return err } // Check if image handling is working correctly if result.Usage.GetInputTokens() > 100000 { log.Printf("WARNING: High token usage (%d) - possible image handling issue", result.Usage.GetInputTokens()) } ``` ## Testing The MCP package includes comprehensive tests: ```bash # Run all MCP tests go test ./pkg/mcp/... -v # Run tests with coverage go test ./pkg/mcp/... -cover # Run specific test go test ./pkg/mcp -run TestContentConversion -v ``` ## Troubleshooting ### Connection Failures If you cannot connect to an MCP server: ```go // Enable debug logging transport := mcp.NewHTTPTransport(mcp.HTTPTransportConfig{ URL: "https://mcp.example.com", Config: mcp.TransportConfig{ EnableLogging: true, // Enable verbose logging }, }) // Check network connectivity // - Verify the URL is correct // - Check firewall settings // - Test with curl: curl https://mcp.example.com/mcp ``` ### Tool Execution Failures If tools fail to execute: ```go // Check tool availability tools, err := converter.ConvertToGoAITools(ctx) if err != nil { log.Fatal("Failed to convert tools:", err) } fmt.Printf("Available tools: %d\n", len(tools)) for _, tool := range tools { fmt.Printf(" - %s: %s\n", tool.Name, tool.Description) } // Test tool manually result, err := client.CallTool(ctx, "problem-tool", map[string]interface{}{}) if err != nil { fmt.Printf("Tool error: %v\n", err) } else { fmt.Printf("Tool result: %+v\n", result.Content) } ``` ### High Token Usage with Images If you're seeing unexpectedly high token usage: ```go // This should NOT happen with proper image handling // but if you see 200K+ token usage, check: // 1. Verify image conversion is working result, err := client.CallTool(ctx, "create-chart", args) if err != nil { log.Fatal(err) } // 2. Inspect the result content for _, content := range result.Content { if content.Type == "image" { fmt.Println("Image found - should be converted to bytes") // If you see base64 text here, there's a problem } } ``` ## See Also - [MCP Specification](https://spec.modelcontextprotocol.io/) - [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - [Tools](https://goaisdk.com/docs/foundations/tools.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [Advanced: Tool Calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) ## May 2026 parity updates MCP JSON parsing uses secure parsing to avoid prototype-pollution style payloads when consuming remote server data. MCP clients expose server instructions and attach TypeScript-compatible MCP provider metadata to converted tools under the `"mcp"` key: `clientName`, `toolName`, optional `title`, and optional MCP Apps `app` metadata. The default client name is `ai-sdk-mcp-client`; `ClientName` overrides it, and `Name` remains as a deprecated alias. MCP Apps hosts can advertise app-rendering support with `mcp.MCPAppClientCapabilities()`, inspect `_meta.ui` using `GetMCPAppToolMeta`, split tools with `SplitMCPAppTools`, and read `ui://` app resources with `ReadMCPAppResource`. MCP tool and prompt content also accept `resource_link` parts, matching the current MCP schema. Custom transports remain supported through the transport interfaces in `pkg/mcp`. Use explicit redirect handling for HTTP transports and prefer the default redirect error mode unless the server is trusted. ```go transport := mcp.NewHTTPTransport(mcp.HTTPTransportConfig{ URL: "https://mcp.example.com", Config: mcp.TransportConfig{ Redirect: mcp.MCPRedirectError, }, }) ``` --- # Prompt Engineering > Teaches prompt engineering for LLMs in Go, from a slogan-generator quick start to advanced techniques, Go-specific schema patterns, and debugging prompts. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/prompt-engineering Documentation index: https://goaisdk.com/llms.txt Prompt engineering is the practice of designing and refining prompts to get the best results from Large Language Models (LLMs). This guide covers fundamental concepts, advanced techniques, and Go-specific patterns for working with the Go AI SDK. ## What is a Large Language Model (LLM)? A Large Language Model is a prediction engine that takes a sequence of words as input and predicts the most likely sequence to follow. It assigns probabilities to potential next sequences and selects one, continuing until a stopping criterion is met. These models are trained on massive text corpuses, making them better suited to some use cases than others. For example, a model trained on GitHub data understands source code patterns particularly well. However, generated sequences, while often plausible, can sometimes be random and not grounded in reality. ## Quick Start: Build a Slogan Generator ### Start with an Instruction ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Create a slogan for an organic coffee shop.", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) // Output: "Naturally Brewed, Ethically Sourced" } ``` ### Include Examples (Few-Shot Learning) ```go prompt := `Create three slogans for a business with unique features. Business: Bookstore with cats Slogans: "Purr-fect Pages", "Books and Whiskers", "Novels and Nuzzles" Business: Gym with rock climbing Slogans: "Peak Performance", "Reach New Heights", "Climb Your Way Fit" Business: Coffee shop with live music Slogans:` result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) // Output: "Sip, Savor, and Swing", "Beans and Beats", "Melody in Every Cup" ``` ### Adjust Temperature Temperature controls output randomness (0 = deterministic, 1 = creative): ```go // Deterministic output (same result every time) temperature := 0.0 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Temperature: &temperature, Prompt: "Create a slogan for a coffee shop.", }) // Creative, varied output (different each time) temperature := 0.8 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Temperature: &temperature, Prompt: "Create a slogan for a coffee shop.", }) ``` **Temperature Guide:** - **0.0-0.3**: Focused, deterministic, factual (good for tool calls, data extraction) - **0.4-0.7**: Balanced creativity and consistency (good for general tasks) - **0.8-1.0**: Creative, varied, exploratory (good for brainstorming, creative writing) ## Advanced Prompting Techniques ### System Prompts Define the model's behavior and personality: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: `You are a professional marketing copywriter with 10 years of experience. You specialize in creating memorable, punchy slogans. Always think about: - Brand voice and personality - Target audience - Emotional impact - Memorability`, Prompt: "Create three slogans for an eco-friendly coffee shop.", }) ``` ### Chain-of-Thought Prompting Encourage step-by-step reasoning: ```go prompt := `Create a slogan for a sustainable fashion brand. Think through this step by step: 1. First, identify the key values of sustainable fashion 2. Consider the target audience and their motivations 3. Think about what makes this brand unique 4. Combine these elements into a memorable slogan 5. Present your final slogan Let's begin:` result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) ``` ### Iterative Refinement Use multiple calls to refine output: ```go // Step 1: Generate initial slogans result1, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Create 5 slogans for a tech startup making AI productivity tools.", }) // Step 2: Refine the best ones refinePrompt := fmt.Sprintf(`Here are some slogans: %s Pick the best 2 and refine them to be: - More punchy and memorable - Focused on user benefits - Under 6 words each`, result1.Text) result2, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: refinePrompt, }) ``` ## Prompts for Tools When using tools, getting good results becomes more challenging as tool complexity increases. ### Tips for Effective Tool Prompts 1. **Use Strong Models**: Use models proficient at tool calling (GPT-6, GPT-5, Claude Sonnet) 2. **Limit Tool Count**: Keep tools to 5 or less for best results 3. **Simplify Schemas**: Avoid deeply nested or overly complex schemas 4. **Use Semantic Names**: Choose descriptive names for tools and parameters 5. **Add Descriptions**: Use field tags to provide context 6. **Show Examples**: Include example tool calls in your prompt 7. **Explain Tool Outputs**: Document what each tool returns ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() // Define tool with clear descriptions weatherTool := types.Tool{ Name: "get_weather", Description: "Get current weather for a location. Returns temperature in Celsius, conditions (e.g. sunny, rainy), and humidity percentage.", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]string{ "type": "string", "description": "City name (e.g., 'New York' or 'London, UK')", }, "units": map[string]interface{}{ "type": "string", "enum": []string{"celsius", "fahrenheit"}, "description": "Temperature units (defaults to celsius)", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, args map[string]interface{}, _ types.ToolExecutionOptions) (interface{}, error) { // Implementation return map[string]interface{}{ "temperature": 22, "conditions": "sunny", "humidity": 65, }, nil }, } provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Use temperature 0 for deterministic tool calls temperature := 0.0 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Temperature: &temperature, Tools: []types.Tool{weatherTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, Prompt: "What's the weather in Brooklyn, New York?", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Tool Prompt with Examples ```go prompt := `You have access to weather and flight tools. Here are examples of how to use them: Example 1: User: "What's the weather in Paris?" Action: Call get_weather({"location": "Paris, France"}) Example 2: User: "Find flights from NYC to LAX" Action: Call search_flights({"from": "JFK", "to": "LAX"}) Now answer: What's the weather like in Tokyo and are there flights from San Francisco?` result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, Prompt: prompt, }) ``` ## Go-Specific Schema Patterns ### Using Struct Tags When working with structured outputs in Go, leverage struct tags: ```go type MarketingCampaign struct { Slogan string `json:"slogan" description:"Catchy slogan under 10 words"` Tagline string `json:"tagline" description:"Supporting tagline explaining the value"` Target string `json:"target" description:"Target audience (e.g., 'young professionals')"` Channels []string `json:"channels" description:"Marketing channels (e.g., 'social', 'email')"` Budget int `json:"budget" description:"Estimated monthly budget in USD"` Tone string `json:"tone" description:"Brand tone (e.g., 'professional', 'playful')"` } // Build a JSON schema from the struct's json tags schema := ai.SchemaFor[MarketingCampaign]() result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema, Prompt: "Create a marketing campaign for an eco-friendly water bottle company.", }) ``` ### JSON Schema Considerations When defining schemas for tool parameters or structured output: ```go // ✅ Good: Simple, clear schema schema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]string{ "type": "string", "description": "Person's full name", }, "age": map[string]string{ "type": "number", "description": "Age in years", }, "email": map[string]string{ "type": "string", "description": "Email address", "pattern": "^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}$", }, }, "required": []string{"name", "age"}, } // ❌ Avoid: Overly complex nested schemas schema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "person": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "first": map[string]string{"type": "string"}, "middle": map[string]string{"type": "string"}, "last": map[string]string{"type": "string"}, // ... too deep }, }, }, }, }, } ``` ### Optional vs Required Parameters Be explicit about required fields: ```go // Clear required fields schema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "command": map[string]string{ "type": "string", "description": "Shell command to execute", }, "workdir": map[string]string{ "type": "string", "description": "Working directory (optional, defaults to current dir)", }, "timeout": map[string]string{ "type": "number", "description": "Timeout in seconds (optional, defaults to 30)", }, }, "required": []string{"command"}, // Only command is required } ``` ## Structured Output Prompting Guide the model to return structured information: ```go prompt := `Create marketing slogans for a vegan restaurant. Return your response in this exact format: PRIMARY SLOGAN: [main slogan] TAGLINE: [supporting tagline] SOCIAL MEDIA: [shorter version for social media] EMOTIONAL APPEAL: [one-word emotion] Example: PRIMARY SLOGAN: "Plant Power, Every Hour" TAGLINE: "Deliciously sustainable dining" SOCIAL MEDIA: "🌱 Power Up Plant-Based" EMOTIONAL APPEAL: Energetic` result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) ``` For guaranteed structured output, use `GenerateObject`: ```go type SloganSet struct { Primary string `json:"primary" description:"Main slogan under 8 words"` Tagline string `json:"tagline" description:"Supporting tagline"` Social string `json:"social" description:"Social media version with emoji"` Emotion string `json:"emotion" description:"Single word describing emotion"` } schema := ai.SchemaFor[SloganSet]() result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema, Prompt: "Create marketing slogans for a vegan restaurant.", }) // result.Object is guaranteed to match SloganSet structure slogans := result.Object.(SloganSet) fmt.Printf("Primary: %s\n", slogans.Primary) ``` ## Debugging Techniques ### Inspect HTTP Requests Pass an `http.Client` with a logging `http.RoundTripper` to see full API requests. See [Development tools](https://goaisdk.com/docs/ai-sdk-core/devtools.md#intercepting-http-requests) for a complete `LoggingTransport`: ```go provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), HTTPClient: &http.Client{ Transport: &LoggingTransport{Transport: http.DefaultTransport}, // Logs requests and responses }, }) ``` ### Check Token Usage Monitor token consumption: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { log.Fatal(err) } fmt.Printf("Input tokens: %d\n", result.Usage.GetInputTokens()) fmt.Printf("Output tokens: %d\n", result.Usage.GetOutputTokens()) fmt.Printf("Total tokens: %d\n", result.Usage.GetTotalTokens()) // Check for unexpectedly high usage if result.Usage.GetTotalTokens() > 10000 { log.Printf("WARNING: High token usage detected") } ``` ### Inspect Tool Calls Debug tool execution: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: tools, Prompt: "What's the weather in New York?", OnStepEnd: func(ctx context.Context, step types.StepResult, userContext interface{}) { // Called after each step fmt.Printf("Step %d:\n", step.StepNumber) for _, call := range step.ToolCalls { fmt.Printf(" Tool: %s\n", call.ToolName) fmt.Printf(" Args: %v\n", call.Arguments) } for _, res := range step.ToolResults { fmt.Printf(" Result (%s): %v\n", res.ToolName, res.Result) } }, }) ``` ### Log Provider Warnings Some providers return warnings about your prompts: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { log.Fatal(err) } // Check for warnings if len(result.Warnings) > 0 { for _, warning := range result.Warnings { log.Printf("Warning: %s", warning) } } ``` ### Test with Different Models Compare results across models: ```go models := []string{"gpt-6-astra", "gpt-6-sol", "gpt-5.4-nano"} for _, modelID := range models { model, _ := provider.LanguageModel(modelID) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { log.Printf("Model %s failed: %v", modelID, err) continue } fmt.Printf("\n%s:\n%s\n", modelID, result.Text) } ``` ## Best Practices ### 1. Be Specific ```text // ❌ Bad: Vague "Write about coffee" // ✅ Good: Specific "Write a 3-sentence product description for organic Ethiopian coffee beans emphasizing fruity notes and sustainability" ``` ### 2. Use Examples (Few-Shot Learning) ```go // Show the pattern you want prompt := `Convert casual text to professional: Casual: "Hey, wanna meet up?" Professional: "Would you be available for a meeting?" Casual: "Got your email, thx!" Professional: "Thank you for your email." Casual: "Can't make it tomorrow" Professional:` ``` ### 3. Set Constraints ```go prompt := `Create a slogan for a coffee shop. Requirements: - Maximum 5 words - Include the word "brew" - Sound energetic - Avoid cliches` ``` ### 4. Use System Prompts for Consistent Behavior ```go system := `You are a professional copywriter. For every slogan: - Keep it under 6 words - Make it memorable - Avoid cliches - Explain your rationale` result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: system, Prompt: "Create a slogan for an artisan bakery", }) ``` ### 5. Test Different Temperatures ```go prompts := []float64{0.0, 0.5, 1.0} for _, temp := range prompts { temperature := temp result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Temperature: &temperature, Prompt: prompt, }) fmt.Printf("Temperature %.1f: %s\n", temp, result.Text) } ``` ### 6. Use Temperature 0 for Tool Calls ```go // ✅ Good: Deterministic tool calls temperature := 0.0 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Temperature: &temperature, Tools: tools, Prompt: "Get the weather in Tokyo", }) // ❌ Avoid: High temperature with tools temperature := 0.8 // Can cause inconsistent tool usage ``` ### 7. Handle Context Limits ```go // Check if prompt fits in context maxTokens := 128000 // example context limit promptTokens := len(prompt) / 4 // Rough estimate (1 token ≈ 4 chars) if promptTokens > maxTokens-2000 { log.Printf("Warning: Prompt may exceed context limit") // Truncate or summarize prompt } ``` ## Common Pitfalls ### 1. Ambiguous Instructions ```go // ❌ Bad "Make it better" // ✅ Good "Rewrite this slogan to be more energetic and under 5 words" ``` ### 2. Conflicting Instructions ```go // ❌ Bad: Contradictory "Be creative and unique, but follow these 10 exact examples" // ✅ Good: Clear expectations "Use these examples as style inspiration, but create original slogans" ``` ### 3. Assuming Context ```text // ❌ Bad: No context "What should I do?" // ✅ Good: Full context "I'm creating a marketing campaign for a vegan restaurant targeting young professionals. What social media platforms should I focus on?" ``` ### 4. Ignoring Token Limits ```go // ❌ Bad: No MaxTokens limit result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a comprehensive guide to Go programming", // Could generate 10,000+ tokens }) // ✅ Good: Set reasonable limits maxTokens := 500 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, MaxTokens: &maxTokens, // Limit output length Prompt: "Write a brief introduction to Go programming", }) ``` ### 5. Not Handling Errors ```go // ✅ Good: Handle provider-specific errors result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { // Check for specific error types if strings.Contains(err.Error(), "rate_limit") { log.Println("Rate limit hit, retrying after delay...") time.Sleep(60 * time.Second) // Retry logic } else if strings.Contains(err.Error(), "context_length") { log.Println("Prompt too long, truncating...") // Truncate and retry } else { log.Fatal(err) } } ``` ## Provider-Specific Considerations ### OpenAI ```go // A larger model is better at complex tasks model, _ := provider.LanguageModel("gpt-6-astra") // A smaller model is faster and cheaper model, _ := provider.LanguageModel("gpt-5.4-nano") // o1 models use extended thinking model, _ := provider.LanguageModel("o1") // Note: o1 models have different parameters (no temperature, system prompts, etc.) ``` ### Anthropic ```go // Claude excels at long-form content and analysis model, _ := provider.LanguageModel("claude-sonnet-4-5") // Use system prompts effectively system := `You are Claude, an AI assistant created by Anthropic.` ``` ### Google ```go // Gemini models support multimodal input model, _ := provider.LanguageModel("gemini-2.0-flash") // Long context windows (up to 1M tokens) model, _ := provider.LanguageModel("gemini-2.5-pro") ``` ## Resources - [Anthropic Prompt Engineering Guide](https://docs.anthropic.com/claude/docs/prompt-engineering) - [OpenAI Prompt Engineering Guide](https://platform.openai.com/docs/guides/prompt-engineering) - [Prompt Engineering Guide](https://www.promptingguide.ai/) - [Google Gemini Prompting Guide](https://ai.google.dev/docs/prompt_best_practices) ## See Also - [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - [Tools](https://goaisdk.com/docs/foundations/tools.md) - [Generating Structured Data](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) - [Advanced: Caching](https://goaisdk.com/docs/advanced/caching.md) - [Advanced: Rate Limiting](https://goaisdk.com/docs/advanced/rate-limiting.md) --- # Settings > Reference for configuring the Go AI SDK: common settings, provider configuration headers, provider-specific options, and checking generation warnings. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/settings Documentation index: https://goaisdk.com/llms.txt Large language models (LLMs) typically provide settings to augment their output. All AI SDK functions support the following common settings in addition to the model, the [prompt](https://goaisdk.com/docs/foundations/prompts.md), and additional provider-specific settings: ```go import ( "context" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) maxTokens := 512 temperature := 0.3 maxRetries := 5 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Invent a new holiday and describe its traditions.", MaxTokens: &maxTokens, Temperature: &temperature, MaxRetries: &maxRetries, }) ``` > **Note:** Some providers do not support all common settings. If you use a setting with a provider that does not support it, a warning will be generated. You can check the `Warnings` field in the result to see if any warnings were generated. ## Common Settings ### MaxTokens Maximum number of tokens to generate. ```go maxTokens := 512 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing", MaxTokens: &maxTokens, }) ``` The pointer allows you to distinguish between "not set" (nil) and "set to 0" (pointer to 0). ### Temperature Temperature setting controls randomness in the output. The value is passed through to the provider. The range depends on the provider and model. For most providers, `0` means almost deterministic results, and higher values mean more randomness. It is recommended to set either `Temperature` or `TopP`, but not both. ```go temperature := 0.7 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a creative story", Temperature: &temperature, }) ``` > **Note:** In Go AI SDK, temperature is not set by default (nil), allowing the provider's default to be used. ### TopP Nucleus sampling - only sample from tokens that comprise the top P probability mass. The value is passed through to the provider. The range depends on the provider and model. For most providers, nucleus sampling is a number between 0 and 1. For example, 0.1 would mean that only tokens with the top 10% probability mass are considered. It is recommended to set either `Temperature` or `TopP`, but not both. ```go topP := 0.9 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate diverse text", TopP: &topP, }) ``` ### TopK Only sample from the top K options for each subsequent token. Used to remove "long tail" low probability responses. Recommended for advanced use cases only. You usually only need to use `Temperature`. ```go topK := 40 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", TopK: &topK, }) ``` ### PresencePenalty The presence penalty affects the likelihood of the model to repeat information that is already in the prompt. The value is passed through to the provider. The range depends on the provider and model. For most providers, `0` means no penalty. ```go presencePenalty := 0.6 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Discuss different topics", PresencePenalty: &presencePenalty, }) ``` ### FrequencyPenalty The frequency penalty affects the likelihood of the model to repeatedly use the same words or phrases. The value is passed through to the provider. The range depends on the provider and model. For most providers, `0` means no penalty. ```go frequencyPenalty := 0.8 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write varied text", FrequencyPenalty: &frequencyPenalty, }) ``` ### StopSequences The stop sequences to use for stopping the text generation. If set, the model will stop generating text when one of the stop sequences is generated. Providers may have limits on the number of stop sequences. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "List three items", StopSequences: []string{"\n\n", "END", "---"}, }) ``` ### Seed The seed (integer) to use for random sampling. If set and supported by the model, calls will generate deterministic results. ```go seed := 42 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate reproducible text", Seed: &seed, }) ``` Using the same seed with the same model and prompt should produce identical results (when supported by the provider). ### MaxRetries Maximum number of retries for failed requests. Set to 0 to disable retries. Default varies by provider (typically 2). ```go maxRetries := 5 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", MaxRetries: &maxRetries, // Retry up to 5 times on failure }) ``` Retries use exponential backoff to avoid overwhelming the provider. ### Context (Timeouts and Cancellation) Go uses `context.Context` for timeouts and cancellation instead of abort signals. The context can be forwarded from a user interface to cancel the call, or to define a timeout. #### Example: Timeout ```go import ( "context" "time" ) // Create context with 5 second timeout ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Invent a new holiday and describe its traditions.", }) if err != nil { if ctx.Err() == context.DeadlineExceeded { fmt.Println("Request timed out") } log.Fatal(err) } ``` #### Example: Manual Cancellation ```go ctx, cancel := context.WithCancel(context.Background()) // Start generation in goroutine go func() { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Long running task", }) // Handle result... }() // Cancel after some condition time.Sleep(3 * time.Second) cancel() // Cancels the generation ``` ### Headers Additional HTTP headers to be sent with the request. Only applicable for HTTP-based providers. You can use the request headers to provide additional information to the provider, depending on what the provider supports. For example, some observability providers support headers such as `Prompt-Id`. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Invent a new holiday and describe its traditions.", Headers: map[string]string{ "Prompt-Id": "my-prompt-id", "X-Custom-Field": "custom-value", }, }) ``` > **Note:** The `Headers` setting is for request-specific headers. You can also set headers in the provider configuration. These headers will be sent with every request made by the provider. ## Provider Configuration Headers Set headers globally for all requests from a provider: ```go import "github.com/digitallysavvy/go-ai/pkg/providers/openai" provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), Headers: map[string]string{ "X-Organization": "my-org-id", "X-Project": "my-project-id", }, }) ``` ## Provider-Specific Settings Each provider may support additional settings that are not part of the common API. Use the `ProviderOptions` field to pass provider-specific settings: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "user": "user-12345", "logit_bias": map[string]int{"50256": -100}, "response_format": map[string]string{"type": "json_object"}, }, }, }) ``` Provider-specific options are documented in each provider's documentation. ## Checking Warnings When using settings that aren't supported by a provider, warnings are returned: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", TopK: &topK, // May not be supported by all providers }) if err != nil { log.Fatal(err) } // Check for warnings for _, warning := range result.Warnings { fmt.Printf("Warning: %s - %s\n", warning.Type, warning.Message) } ``` ## Settings Across Different Functions These settings are supported across all core AI SDK functions with appropriate variations: ### GenerateText / StreamText ```go maxTokens := 512 temperature := 0.7 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a story", MaxTokens: &maxTokens, Temperature: &temperature, Seed: &seed, }) ``` ### GenerateObject / StreamObject ```go maxTokens := 256 temperature := 0.3 result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema, Prompt: "Generate structured data", MaxTokens: &maxTokens, Temperature: &temperature, }) ``` ### Embed / EmbedMany Embedding models typically don't use generation settings, but may support: ```go result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "Text to embed", Headers: map[string]string{"X-User-ID": "user-123"}, }) ``` ### GenerateImage Image generation uses different settings: ```go n := 4 // Number of images result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "A beautiful landscape", Size: "1024x1024", Seed: &seed, N: &n, }) ``` ### GenerateSpeech Speech generation settings: ```go speed := 1.5 result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "alloy", Speed: &speed, }) ``` ### Transcribe Transcription settings: ```go maxRetries := 2 result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, MimeType: "audio/mpeg", MaxRetries: &maxRetries, }) // The detected language is available on the result, not as an input option fmt.Println(result.Language) ``` ## Best Practices ### Use Pointers for Optional Settings Go uses pointers to distinguish between "not set" and "set to zero value": ```go // Not set (uses provider default) result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", // Temperature: nil (not set) }) ``` ```go // Explicitly set to 0 temperature := 0.0 result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", Temperature: &temperature, // Set to 0 }) ``` ### Balance Temperature and TopP Don't set both `Temperature` and `TopP` - use one or the other: ```go // Good: Use temperature temperature := 0.7 result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate creative text", Temperature: &temperature, }) ``` ```go // Bad: Don't use both temperature := 0.7 topP := 0.9 result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", Temperature: &temperature, // Don't use both TopP: &topP, // simultaneously }) ``` ### Always Set Timeouts Use context timeouts to prevent hanging requests: ```go import "time" ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", }) ``` ### Handle Warnings Always check for and log warnings: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", TopK: &topK, }) if err != nil { log.Fatal(err) } if len(result.Warnings) > 0 { for _, w := range result.Warnings { log.Printf("Warning: %s - %s", w.Type, w.Message) } } ``` ### Use Seeds for Reproducibility When you need consistent results (testing, demos), use seeds: ```go seed := 42 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate reproducible output", Seed: &seed, }) ``` ## Common Setting Combinations ### Deterministic Output ```go temperature := 0.0 seed := 42 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain how photosynthesis works", Temperature: &temperature, Seed: &seed, }) ``` ### Creative Writing ```go temperature := 0.9 presencePenalty := 0.6 frequencyPenalty := 0.3 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a creative story", Temperature: &temperature, PresencePenalty: &presencePenalty, FrequencyPenalty: &frequencyPenalty, }) ``` ### Factual/Technical Output ```go temperature := 0.2 topP := 0.1 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain the TCP/IP protocol", Temperature: &temperature, }) ``` ### Concise Output ```go maxTokens := 100 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Summarize this article", MaxTokens: &maxTokens, StopSequences: []string{"\n\n"}, }) ``` ## See Also - [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - [Prompts](https://goaisdk.com/docs/foundations/prompts.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [Provider Documentation](https://goaisdk.com/docs/providers/overview.md) --- # Reasoning > Learn how to control reasoning across providers with the top-level Reasoning parameter. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/reasoning Documentation index: https://goaisdk.com/llms.txt Many language models support an internal "reasoning" phase (sometimes also called "thinking") before producing a response. The Go AI SDK provides a top-level `Reasoning` parameter on `GenerateText` and `StreamText` that controls this behavior across providers with a single, portable setting. ## Basic usage ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() provider := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) model, err := provider.LanguageModel("claude-sonnet-4-6") if err != nil { log.Fatalf("Failed to create model: %v", err) } reasoning := types.ReasoningMedium result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "How many people will live in the world in 2040?", Reasoning: &reasoning, }) if err != nil { log.Fatalf("Generation failed: %v", err) } fmt.Println("Answer:", result.Text) // Access reasoning content from the result's Content parts. // Reasoning is returned as types.ReasoningContent blocks within the // step results, not as a top-level field. if len(result.Steps) > 0 { lastStep := result.Steps[len(result.Steps)-1] for _, part := range lastStep.Content { if rc, ok := part.(types.ReasoningContent); ok { fmt.Println("Reasoning:", rc.Text) } } } } ``` The `Reasoning` field is a `*types.ReasoningLevel`. Pass a pointer to one of the level constants to set the desired effort. When `nil` (the default), the provider's own default behavior is used. ## Streaming The `Reasoning` parameter works the same way with `StreamText`. Use the `OnChunk` callback to receive reasoning and text deltas as they arrive: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/google" ) func main() { ctx := context.Background() p := google.New(google.Config{APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY")}) model, err := p.LanguageModel("gemini-3-flash-preview") if err != nil { log.Fatalf("Failed to create model: %v", err) } reasoning := types.ReasoningHigh result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Explain the Riemann hypothesis in simple terms.", Reasoning: &reasoning, OnChunk: func(chunk provider.StreamChunk) { switch chunk.Type { case provider.ChunkTypeReasoning: fmt.Fprint(os.Stderr, chunk.Reasoning) case provider.ChunkTypeText: fmt.Print(chunk.Text) } }, }) if err != nil { log.Fatalf("Streaming failed: %v", err) } // ReadAll blocks until the stream completes. if _, err := result.ReadAll(); err != nil { log.Fatalf("ReadAll failed: %v", err) } } ``` The stream also emits `ChunkTypeReasoningStart` and `ChunkTypeReasoningEnd` to mark reasoning block boundaries. These carry an `ID` field that links them to their corresponding `ChunkTypeReasoning` deltas, but contain no text content themselves. ## Reasoning levels | Constant | Value | Behavior | | --- | --- | --- | | `types.ReasoningDefault` | `"provider-default"` | Use the provider's default reasoning behavior (default when `nil`) | | `types.ReasoningNone` | `"none"` | Disable reasoning | | `types.ReasoningMinimal` | `"minimal"` | Bare-minimum reasoning | | `types.ReasoningLow` | `"low"` | Fast, concise reasoning | | `types.ReasoningMedium` | `"medium"` | Balanced reasoning | | `types.ReasoningHigh` | `"high"` | Thorough reasoning | | `types.ReasoningXHigh` | `"xhigh"` | Maximum reasoning | ## Provider support Each provider translates the level to its native reasoning API. Some providers support all levels natively, while others coerce to fewer levels (a warning is emitted when coercion occurs). Some providers use a numeric token budget instead of an enum; in those cases the level is mapped to a budget calculated as a percentage of the model's maximum output tokens. ### Level mapping | Level | Anthropic | OpenAI | Bedrock (Claude) | Bedrock (Nova) | Google / Vertex | | --- | --- | --- | --- | --- | --- | | `none` | `disabled` | `disabled` | `disabled` | `disabled` | `thinkingBudget: 0` | | `minimal` | 2% of max | `low` | 2% of max | `low` | 2% of max | | `low` | 10% | `low` | 10% | `low` | 10% | | `medium` | 30% | `medium` | 30% | `medium` | 30% | | `high` | 60% | `high` | 60% | `high` | 60% | | `xhigh` | 90% | `high` | 90% | `high` | 90% | | `provider-default` | omitted | omitted | omitted | omitted | omitted | **Anthropic dynamic budgets:** Budget tokens are calculated dynamically per model using the formula `percentage * maxOutputTokens`, with a floor of 1,024 tokens. For example, `claude-sonnet-4-6` has a 128,000 max output token limit, so `medium` (30%) maps to 38,400 tokens. Model-specific limits: 128,000 for claude-4-6, 64,000 for claude-4-5, 32,000 for claude-4-1, 4,096 default. **Google / Vertex dynamic budgets:** Budget tokens are calculated as 2/10/30/60/90% of the model's max thinking tokens. Gemini 3 models use a `thinkingLevel` string instead of a numeric budget. ### Other supported providers | Provider | Native mapping | Notes | | --- | --- | --- | | **DeepSeek** | Provider-specific | Enables reasoning for DeepSeek-R1 and similar models. | | **Alibaba** | Provider-specific | Enables reasoning for Qwen models with thinking support. | | **Fireworks** | Provider-specific | Forwards reasoning configuration to supported models. | | **Groq** | Provider-specific | Supports reasoning on compatible models. | | **XAI** | `reasoning_effort` | OpenAI-compatible mapping. Reasoning deltas emitted via `OnReasoningDelta`. | | **Open Responses** | `reasoning.effort` | OpenAI-compatible Responses object mapping; `ProviderOptions["open-responses"].reasoningSummary` maps to `reasoning.summary`. | ### Providers that warn | Provider | Behavior | | --- | --- | | **Perplexity** | Emits an `unsupported-setting` warning and ignores the parameter. | | **Cohere** | Emits an `unsupported-setting` warning and ignores the parameter. | | **Mistral** | Only `mistral-small-latest` supports reasoning. Other models emit a warning. | ## Precedence rules The top-level `Reasoning` parameter and provider-specific `ProviderOptions` are **never merged**. If you set reasoning-related options in `ProviderOptions`, they take full precedence and the top-level `Reasoning` parameter is ignored. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() p := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, err := p.ResponsesModel("gpt-5.2") if err != nil { log.Fatalf("Failed to create model: %v", err) } reasoning := types.ReasoningLow // ignored because ProviderOptions is set result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum entanglement.", Reasoning: &reasoning, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "reasoningEffort": "high", // this wins }, }, }) if err != nil { log.Fatalf("Generation failed: %v", err) } fmt.Println(result.Text) } ``` This design lets you use the portable `Reasoning` parameter by default and fall back to `ProviderOptions` only when you need provider-specific features like exact token budgets. ## Migrating from ProviderOptions If you currently control reasoning via `ProviderOptions`, you can migrate to the top-level `Reasoning` parameter for portability across providers. ### Before (Anthropic) ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: anthropicModel, Prompt: "How many people will live in the world in 2040?", ProviderOptions: map[string]interface{}{ "anthropic": map[string]interface{}{ "thinking": map[string]interface{}{ "type": "adaptive", "effort": "high", }, }, }, }) ``` ### After (Anthropic) ```go reasoning := types.ReasoningHigh result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: anthropicModel, Prompt: "How many people will live in the world in 2040?", Reasoning: &reasoning, }) ``` ### Before (Anthropic with exact budget) ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: anthropicModel, // e.g. claude-sonnet-4-20250514 Prompt: "How many people will live in the world in 2040?", ProviderOptions: map[string]interface{}{ "anthropic": map[string]interface{}{ "thinking": map[string]interface{}{ "type": "enabled", "budgetTokens": 12000, }, }, }, }) ``` ### After (Anthropic with exact budget) ```go reasoning := types.ReasoningMedium result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: anthropicModel, // e.g. claude-sonnet-4-20250514 Prompt: "How many people will live in the world in 2040?", Reasoning: &reasoning, }) ``` If you need to enforce an exact token budget (e.g., exactly 12,000 tokens), keep using `ProviderOptions` instead of the top-level `Reasoning` parameter. ### Before (OpenAI with reasoningSummary) ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: openaiModel, // e.g. o3 via ResponsesModel Prompt: "Explain quantum entanglement.", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "reasoningEffort": "high", "reasoningSummary": "auto", }, }, }) ``` ### After (OpenAI with reasoningSummary) ```go reasoning := types.ReasoningHigh result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: openaiModel, // e.g. o3 via ResponsesModel Prompt: "Explain quantum entanglement.", Reasoning: &reasoning, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "reasoningSummary": "auto", }, }, }) ``` ### Before (Google with thinkingConfig) ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: googleModel, Prompt: "Explain the Riemann hypothesis in simple terms.", ProviderOptions: map[string]interface{}{ "google": map[string]interface{}{ "thinkingConfig": map[string]interface{}{ "thinkingBudget": 4096, "includeThoughts": true, }, }, }, }) ``` ### After (Google with thinkingConfig) ```go reasoning := types.ReasoningMedium result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: googleModel, Prompt: "Explain the Riemann hypothesis in simple terms.", Reasoning: &reasoning, ProviderOptions: map[string]interface{}{ "google": map[string]interface{}{ "thinkingConfig": map[string]interface{}{ "includeThoughts": true, }, }, }, }) ``` Note that `ProviderOptions` can still be used alongside `Reasoning` for provider-specific features unrelated to reasoning effort. However, if `ProviderOptions` includes reasoning effort or budget settings (e.g., `reasoningEffort`, `thinking`, `thinkingConfig.thinkingBudget`), those take full precedence and the top-level `Reasoning` parameter is ignored. --- # Upload File and Upload Skill > Upload provider files and multi-file skills for later reference in model calls. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/upload-file-and-skill Documentation index: https://goaisdk.com/llms.txt The Go SDK provides top-level helpers for uploading files and skills through provider upload APIs: - `ai.UploadFile(...)` - `ai.UploadSkill(...)` Both helpers accept either: 1. A provider upload API directly (`provider.FilesAPI` / `provider.SkillsAPI`) 2. A provider implementing `Files()` / `Skills()` ## UploadFile ```go result, err := ai.UploadFile(ctx, ai.UploadFileOptions{ API: openaiProvider, Data: []byte("hello world"), Filename: "hello.txt", MediaType: "text/plain", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{"purpose": "assistants"}, }, }) if err != nil { panic(err) } fileID := result.ProviderReference["openai"] ``` `UploadFile` normalizes shorthand data inputs (`[]byte`, `string`) into `types.FileData`, resolves the target files API, and auto-detects media type when omitted. String data is treated as a TypeScript-compatible base64 data string and decoded by the provider upload implementation. Use the returned provider reference in later prompts without downloading or inlining the file: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Summarize this file."}, types.FileContent{ MediaType: upload.MediaType, Filename: upload.Filename, FileData: types.FileData{ Type: types.FileDataTypeReference, Reference: upload.ProviderReference, MediaType: upload.MediaType, }, }, }, }, }, }) ``` See `examples/upload-file/provider-reference` for a complete upload-then-reference flow. ## UploadSkill ```go result, err := ai.UploadSkill(ctx, ai.UploadSkillOptions{ API: openaiProvider, DisplayTitle: "My Skill", Files: []ai.UploadSkillFile{ {Path: "index.ts", Data: []byte("export default {}")}, {Path: "README.md", Data: "# Skill"}, }, }) if err != nil { panic(err) } skillID := result.ProviderReference["openai"] ``` `UploadSkill` normalizes shorthand file data and delegates to the provider skills API. --- # Embeddings > Explains embedding single and multiple values with the Go AI SDK's Embed and EmbedMany functions, plus similarity, token usage, and embedding middleware. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/embeddings Documentation index: https://goaisdk.com/llms.txt Embeddings are a way to represent words, phrases, or images as vectors in a high-dimensional space. In this space, similar words are close to each other, and the distance between words can be used to measure their similarity. ## Embedding a Single Value The Go AI SDK provides the [`ai.Embed()`](https://goaisdk.com/docs/reference/ai/embed.md) function to embed single values, which is useful for tasks such as finding similar words or phrases or clustering text. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) embeddingModel, _ := provider.EmbeddingModel("text-embedding-3-small") // 'Embedding' is a single embedding vector ([]float64) result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "sunny day at the beach", }) if err != nil { log.Fatal(err) } fmt.Printf("Embedding dimensions: %d\n", len(result.Embedding)) fmt.Printf("First 5 values: %v\n", result.Embedding[:5]) } ``` ## Embedding Many Values When loading data, e.g. when preparing a data store for retrieval-augmented generation (RAG), it is often useful to embed many values at once (batch embedding). The Go AI SDK provides the [`ai.EmbedMany()`](https://goaisdk.com/docs/reference/ai/embed-many.md) function for this purpose. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) embeddingModel, _ := provider.EmbeddingModel("text-embedding-3-small") // 'Embeddings' is a slice of embedding vectors ([][]float64). // It is sorted in the same order as the input values. result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: []string{ "sunny day at the beach", "rainy afternoon in the city", "snowy night in the mountains", }, }) if err != nil { log.Fatal(err) } fmt.Printf("Generated %d embeddings\n", len(result.Embeddings)) for i, embedding := range result.Embeddings { fmt.Printf("Embedding %d: %d dimensions\n", i+1, len(embedding)) } } ``` ## Embedding Similarity After embedding values, you can calculate the similarity between them using the similarity functions provided by the Go AI SDK. This is useful to e.g. find similar words or phrases in a dataset. You can also rank and filter related items based on their similarity. ### Cosine Similarity ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) embeddingModel, _ := provider.EmbeddingModel("text-embedding-3-small") result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: []string{ "sunny day at the beach", "rainy afternoon in the city", }, }) if err != nil { log.Fatal(err) } // Calculate cosine similarity similarity, err := ai.CosineSimilarity(result.Embeddings[0], result.Embeddings[1]) if err != nil { log.Fatal(err) } fmt.Printf("Cosine similarity: %.4f\n", similarity) } ``` ### Available Similarity Functions The Go AI SDK provides several similarity functions: ```go // Cosine Similarity (range: -1 to 1, higher is more similar) similarity, err := ai.CosineSimilarity(embedding1, embedding2) // Euclidean Distance (lower is more similar) distance, err := ai.EuclideanDistance(embedding1, embedding2) // Dot Product (higher is more similar) dotProduct, err := ai.DotProduct(embedding1, embedding2) ``` ### Finding Most Similar Find the most similar embedding from a list: ```go queryResult, _ := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "ocean waves crashing", }) candidatesResult, _ := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: []string{ "sunny day at the beach", "rainy afternoon in the city", "snowy night in the mountains", "tropical island paradise", }, }) // Find most similar mostSimilarIdx, similarity, _ := ai.FindMostSimilar( queryResult.Embedding, candidatesResult.Embeddings, ) fmt.Printf("Most similar: index %d with similarity %.4f\n", mostSimilarIdx, similarity) ``` ### Ranking by Similarity Rank all embeddings by similarity: ```go indices, similarities, err := ai.RankBySimilarity(queryResult.Embedding, candidatesResult.Embeddings) if err != nil { log.Fatal(err) } fmt.Println("Rankings (most to least similar):") for i, idx := range indices { fmt.Printf("%d. Index %d - Similarity: %.4f\n", i+1, idx, similarities[i]) } ``` ## Token Usage Many providers charge based on the number of tokens used to generate embeddings. Both `Embed` and `EmbedMany` provide token usage information in the `Usage` field of the result: ```go result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "sunny day at the beach", }) if err != nil { log.Fatal(err) } fmt.Printf("Tokens used: %d\n", result.Usage.TotalTokens) ``` For `EmbedMany`: ```go result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: []string{"text1", "text2", "text3"}, }) if err != nil { log.Fatal(err) } fmt.Printf("Total tokens used: %d\n", result.Usage.TotalTokens) ``` ## Settings ### Provider Options Embedding model settings can be configured using `ProviderOptions` for provider-specific parameters: ```go result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "sunny day at the beach", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "dimensions": 512, // Reduce embedding dimensions }, }, }) ``` ### Google Multimodal Embeddings Google embedding models accept AI SDK-compatible provider options under the `"google"` key. `google.GoogleEmbeddingProviderOptions` supports `TaskType`, `OutputDimensionality`, and multimodal `Content` entries for each embedded value. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/google" ) func main() { ctx := context.Background() provider := google.New(google.Config{APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY")}) embeddingModel, _ := provider.EmbeddingModel(google.EmbeddingModelGeminiEmbedding001) dimensions := 768 result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "retrieval document about mountain weather", ProviderOptions: map[string]interface{}{ "google": google.GoogleEmbeddingProviderOptions{ TaskType: "RETRIEVAL_DOCUMENT", OutputDimensionality: &dimensions, }, }, }) if err != nil { log.Fatal(err) } fmt.Printf("Embedding dimensions: %d\n", len(result.Embedding)) } ``` Use `Content` for per-value multimodal input. This matches the TypeScript SDK `providerOptions.google.content` behavior and is validated to have the same length as the embedded values. ```go result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "summarize the attached document for retrieval", ProviderOptions: map[string]interface{}{ "google": google.GoogleEmbeddingProviderOptions{ TaskType: "RETRIEVAL_DOCUMENT", Content: [][]google.EmbeddingPart{ { google.ImageEmbeddingPart{ MimeType: "application/pdf", Data: pdfData, }, }, }, }, }, }) ``` For batch embeddings, provide one `Content` entry for each input value. Nil entries are text-only: ```go result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: []string{"first document", "second document"}, ProviderOptions: map[string]interface{}{ "google": google.GoogleEmbeddingProviderOptions{ Content: [][]google.EmbeddingPart{ nil, { google.ImageEmbeddingPart{ MimeType: "application/pdf", Data: pdfData, }, }, }, }, }, }) ``` ### Parallel Requests The `EmbedMany` function supports parallel processing with configurable `MaxParallelCalls` to optimize performance: ```go result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, MaxParallelCalls: 2, // Limit parallel requests Inputs: []string{ "sunny day at the beach", "rainy afternoon in the city", "snowy night in the mountains", }, }) ``` ### Retries Both `Embed` and `EmbedMany` accept an optional `MaxRetries` parameter that you can use to set the maximum number of retries for the embedding process. It defaults to `2` retries (3 attempts in total). You can set it to `0` to disable retries. ```go maxRetries := 0 result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "sunny day at the beach", MaxRetries: &maxRetries, // Disable retries }) ``` ### Timeouts with Context Use Go's context for timeouts and cancellation: ```go // Timeout after 5 seconds ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "sunny day at the beach", }) if err != nil { if ctx.Err() == context.DeadlineExceeded { fmt.Println("Embedding operation timed out") } log.Fatal(err) } ``` ### Custom Headers Both `Embed` and `EmbedMany` accept an optional `Headers` parameter that you can use to add custom headers to the embedding request: ```go result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "sunny day at the beach", Headers: map[string]string{ "X-Custom-Header": "custom-value", }, }) ``` ## Response Information Both `Embed` and `EmbedMany` return response information that includes the raw provider response: ```go result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "sunny day at the beach", }) if err != nil { log.Fatal(err) } fmt.Printf("Headers: %+v\n", result.Response.Headers) fmt.Printf("Body: %+v\n", result.Response.Body) ``` ## Embedding Middleware You can enhance embedding models, e.g. to set default values, using middleware: ```go import ( "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) // Wrap model to log input and output wrappedModel := middleware.WrapEmbeddingModel(embeddingModel, []*middleware.EmbeddingModelMiddleware{ { TransformInput: func(ctx context.Context, input string, model provider.EmbeddingModel) (string, error) { // Pre-process input fmt.Printf("Embedding: %s\n", input) return input, nil }, WrapEmbed: func(ctx context.Context, doEmbed func() (*types.EmbeddingResult, error), input string, model provider.EmbeddingModel) (*types.EmbeddingResult, error) { result, err := doEmbed() // Post-process embedding if err == nil { fmt.Printf("Generated embedding with %d dimensions\n", len(result.Embedding)) } return result, err }, }, }, nil, nil) result, _ := ai.Embed(ctx, ai.EmbedOptions{ Model: wrappedModel, Input: "sunny day at the beach", }) ``` ## Practical Examples ### Semantic Search Build a simple semantic search: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type Document struct { ID string Text string Embedding []float64 } func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) embeddingModel, err := provider.EmbeddingModel("text-embedding-3-small") if err != nil { log.Fatal(err) } // Documents to search docs := []string{ "The quick brown fox jumps over the lazy dog", "Machine learning is a subset of artificial intelligence", "Go is a statically typed, compiled programming language", "Python is widely used for data science and machine learning", } // Embed all documents docsResult, _ := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: docs, }) // Create document objects documents := make([]Document, len(docs)) for i, text := range docs { documents[i] = Document{ ID: fmt.Sprintf("doc-%d", i+1), Text: text, Embedding: docsResult.Embeddings[i], } } // Search query query := "programming languages" queryResult, _ := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: query, }) // Find most similar document var bestMatch Document var bestSimilarity float64 = -1 for _, doc := range documents { similarity, err := ai.CosineSimilarity(queryResult.Embedding, doc.Embedding) if err != nil { log.Fatal(err) } if similarity > bestSimilarity { bestSimilarity = similarity bestMatch = doc } } fmt.Printf("Query: %s\n", query) fmt.Printf("Best match (%.4f): %s\n", bestSimilarity, bestMatch.Text) } ``` ### Clustering Group similar texts together: ```go func clusterTexts(ctx context.Context, texts []string, embeddingModel provider.EmbeddingModel, threshold float64) [][]int { // Embed all texts result, _ := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: texts, }) // Simple clustering based on similarity threshold var clusters [][]int visited := make([]bool, len(texts)) for i := range texts { if visited[i] { continue } // Start new cluster cluster := []int{i} visited[i] = true // Find similar texts for j := i + 1; j < len(texts); j++ { if visited[j] { continue } similarity, err := ai.CosineSimilarity( result.Embeddings[i], result.Embeddings[j], ) if err != nil { continue } if similarity >= threshold { cluster = append(cluster, j) visited[j] = true } } clusters = append(clusters, cluster) } return clusters } ``` ### RAG (Retrieval-Augmented Generation) Combine embeddings with text generation: ```go func answerWithRAG(ctx context.Context, query string, documents []Document, model provider.LanguageModel, embeddingModel provider.EmbeddingModel) (string, error) { // Embed query queryResult, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: query, }) if err != nil { return "", err } // Find top 3 most relevant documents indices, _, err := ai.RankBySimilarity(queryResult.Embedding, getEmbeddings(documents)) if err != nil { return "", err } topIndices := indices[:min(3, len(indices))] // Build context from relevant documents var context strings.Builder context.WriteString("Context:\n") for _, idx := range topIndices { context.WriteString(fmt.Sprintf("- %s\n", documents[idx].Text)) } // Generate answer using context result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("%s\n\nQuestion: %s\n\nAnswer:", context.String(), query), }) if err != nil { return "", err } return result.Text, nil } func getEmbeddings(docs []Document) [][]float64 { embeddings := make([][]float64, len(docs)) for i, doc := range docs { embeddings[i] = doc.Embedding } return embeddings } func min(a, b int) int { if a < b { return a } return b } ``` ## Embedding Providers & Models Several providers offer embedding models: | Provider | Model | Embedding Dimensions | |----------|-------|---------------------| | **OpenAI** | `text-embedding-3-large` | 3072 | | **OpenAI** | `text-embedding-3-small` | 1536 | | **OpenAI** | `text-embedding-ada-002` | 1536 | | **Google** | `gemini-embedding-001` | 3072 | | **Google** | `gemini-embedding-2-preview` | Provider-defined | | **Mistral** | `mistral-embed` | 1024 | | **Cohere** | `embed-english-v3.0` | 1024 | | **Cohere** | `embed-multilingual-v3.0` | 1024 | | **Cohere** | `embed-english-light-v3.0` | 384 | | **Cohere** | `embed-multilingual-light-v3.0` | 384 | | **Amazon Bedrock** | `amazon.titan-embed-text-v1` | 1536 | | **Amazon Bedrock** | `amazon.titan-embed-text-v2:0` | 1024 | ### Example with Different Providers ```go // OpenAI openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) openaiModel, _ := openaiProvider.EmbeddingModel("text-embedding-3-small") // Google googleProvider := google.New(google.Config{APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY")}) googleModel, _ := googleProvider.EmbeddingModel("gemini-embedding-001") // Mistral mistralProvider := mistral.New(mistral.Config{APIKey: os.Getenv("MISTRAL_API_KEY")}) mistralModel, _ := mistralProvider.EmbeddingModel("mistral-embed") // Use any model with the same API result, _ := ai.Embed(ctx, ai.EmbedOptions{ Model: openaiModel, // or googleModel, or mistralModel Input: "text to embed", }) ``` ## Performance Optimization ### Batch Processing Process large datasets efficiently: ```go func embedLargeDataset(ctx context.Context, texts []string, embeddingModel provider.EmbeddingModel, batchSize int) ([][]float64, error) { var allEmbeddings [][]float64 for i := 0; i < len(texts); i += batchSize { end := i + batchSize if end > len(texts) { end = len(texts) } batch := texts[i:end] result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: batch, MaxParallelCalls: 3, }) if err != nil { return nil, err } allEmbeddings = append(allEmbeddings, result.Embeddings...) fmt.Printf("Processed %d/%d texts\n", min(end, len(texts)), len(texts)) } return allEmbeddings, nil } ``` ### Caching Embeddings Cache embeddings to avoid recomputing: ```go type EmbeddingCache struct { cache map[string][]float64 mu sync.RWMutex } func (c *EmbeddingCache) Get(text string) ([]float64, bool) { c.mu.RLock() defer c.mu.RUnlock() embedding, ok := c.cache[text] return embedding, ok } func (c *EmbeddingCache) Set(text string, embedding []float64) { c.mu.Lock() defer c.mu.Unlock() c.cache[text] = embedding } func (c *EmbeddingCache) EmbedWithCache(ctx context.Context, text string, model provider.EmbeddingModel) ([]float64, error) { // Check cache first if embedding, ok := c.Get(text); ok { return embedding, nil } // Generate embedding result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: model, Input: text, }) if err != nil { return nil, err } // Cache result c.Set(text, result.Embedding) return result.Embedding, nil } ``` ## Best Practices 1. **Choose Right Model**: Select embedding model based on your use case (dimension size vs accuracy trade-off) 2. **Batch Processing**: Use `EmbedMany` for multiple texts to reduce API calls 3. **Cache Results**: Cache embeddings to avoid recomputing 4. **Normalize Embeddings**: Some models require normalization for accurate similarity calculations 5. **Handle Errors**: Implement retry logic for transient failures 6. **Monitor Usage**: Track token usage to manage costs 7. **Use Context**: Always pass context for cancellation and timeouts 8. **Parallel Processing**: Use `MaxParallelCalls` for large batches 9. **Storage**: Consider dimensionality reduction for storage efficiency 10. **Similarity Threshold**: Tune similarity thresholds based on your specific use case ## Next Steps - Learn about [Reranking](https://goaisdk.com/docs/ai-sdk-core/reranking.md) - Explore [RAG Patterns](https://goaisdk.com/docs/advanced/prompt-engineering.md) - See [Embedding Examples](https://github.com/digitallysavvy/go-ai/tree/main/examples) ## See Also - [Similarity Functions](https://goaisdk.com/docs/reference/ai/cosine-similarity.md) - [Providers](https://goaisdk.com/docs/providers/overview.md) - [Middleware](https://goaisdk.com/docs/ai-sdk-core/middleware.md) - [API Reference: Embed](https://goaisdk.com/docs/reference/ai/embed.md) - [API Reference: EmbedMany](https://goaisdk.com/docs/reference/ai/embed-many.md) --- # Reranking > Covers reranking documents by relevance with the Go AI SDK's Rerank function, including structured documents, settings, and reranking provider options. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/reranking Documentation index: https://goaisdk.com/llms.txt Reranking is a technique used to improve search relevance by reordering a set of documents based on their relevance to a query. Unlike embedding-based similarity search, reranking models are specifically trained to understand the relationship between queries and documents, often producing more accurate relevance scores. ## Reranking Documents The Go AI SDK provides the [`ai.Rerank()`](https://goaisdk.com/docs/reference/ai/rerank.md) function to rerank documents based on their relevance to a query. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/cohere" ) func main() { ctx := context.Background() provider := cohere.New(cohere.Config{APIKey: os.Getenv("COHERE_API_KEY")}) rerankModel, _ := provider.RerankingModel("rerank-v3.5") documents := []string{ "sunny day at the beach", "rainy afternoon in the city", "snowy night in the mountains", } topN := 2 // Return top 2 most relevant documents result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: "talk about rain", Documents: documents, TopN: &topN, }) if err != nil { log.Fatal(err) } fmt.Println("Ranked documents:") for i, rank := range result.Ranking { fmt.Printf("%d. [Score: %.4f] %s (original index: %d)\n", i+1, rank.Score, rank.Document, rank.OriginalIndex) } } // Output: // Ranked documents: // 1. [Score: 0.9000] rainy afternoon in the city (original index: 1) // 2. [Score: 0.3000] sunny day at the beach (original index: 0) ``` ## Working with Structured Documents Reranking also supports structured documents (JSON objects), making it ideal for searching through databases, emails, or other structured content: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/cohere" ) type Email struct { From string `json:"from"` Subject string `json:"subject"` Text string `json:"text"` } func main() { ctx := context.Background() provider := cohere.New(cohere.Config{APIKey: os.Getenv("COHERE_API_KEY")}) rerankModel, _ := provider.RerankingModel("rerank-v3.5") emails := []Email{ { From: "Paul Doe", Subject: "Follow-up", Text: "We are happy to give you a discount of 20% on your next order.", }, { From: "John McGill", Subject: "Missing Info", Text: "Sorry, but here is the pricing information from Oracle: $5000/month", }, } // Convert to interface{} slice for reranking documents := make([]interface{}, len(emails)) for i, email := range emails { documents[i] = email } topN := 1 result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: "Which pricing did we get from Oracle?", Documents: documents, TopN: &topN, }) if err != nil { log.Fatal(err) } // Get most relevant document (RerankedDocuments is an interface{} holding // the same slice type that was passed in as Documents) rerankedEmails := result.RerankedDocuments.([]interface{}) if len(rerankedEmails) > 0 { topEmail := rerankedEmails[0].(Email) fmt.Printf("Most relevant email:\n") fmt.Printf("From: %s\n", topEmail.From) fmt.Printf("Subject: %s\n", topEmail.Subject) fmt.Printf("Text: %s\n", topEmail.Text) } } ``` ## Understanding the Results The `Rerank` function returns a comprehensive result object: ```go result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: "talk about rain", Documents: documents, }) if err != nil { log.Fatal(err) } // result.Ranking: sorted array of ai.RerankItem structs // result.RerankedDocuments: documents sorted by relevance (convenience, same // underlying slice type as RerankOptions.Documents; assert before ranging) // result.OriginalDocuments: original documents array ``` Each item in the `Ranking` slice contains: ```go type RerankItem struct { OriginalIndex int // Position in the original documents array Score float64 // Relevance score (typically 0-1, higher is more relevant) Document interface{} // The original document } ``` ## Settings ### Top-N Results Use `TopN` to limit the number of results returned. This is useful for retrieving only the most relevant documents: ```go topN := 3 // Return only top 3 most relevant documents result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: "relevant information", Documents: []string{"doc1", "doc2", "doc3", "doc4", "doc5"}, TopN: &topN, }) ``` ### Provider Options Reranking model settings can be configured using `ProviderOptions` for provider-specific parameters: ```go result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: "talk about rain", Documents: documents, ProviderOptions: map[string]interface{}{ "cohere": map[string]interface{}{ "maxTokensPerDoc": 1000, // Limit tokens per document }, }, }) ``` ### Retries The `Rerank` function accepts an optional `MaxRetries` parameter that you can use to set the maximum number of retries for the reranking process. It defaults to `2` retries (3 attempts in total). You can set it to `0` to disable retries. ```go maxRetries := 0 result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: "talk about rain", Documents: documents, MaxRetries: &maxRetries, // Disable retries }) ``` ### Timeouts with Context Use Go's context for timeouts and cancellation: ```go // Timeout after 5 seconds ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: "talk about rain", Documents: documents, }) if err != nil { if ctx.Err() == context.DeadlineExceeded { fmt.Println("Reranking operation timed out") } log.Fatal(err) } ``` ### Custom Headers The `Rerank` function accepts an optional `Headers` parameter that you can use to add custom headers to the reranking request: ```go result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: "talk about rain", Documents: documents, Headers: map[string]string{ "X-Custom-Header": "custom-value", }, }) ``` ## Response Information The `Rerank` function returns response information that includes the raw provider response: ```go result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: "talk about rain", Documents: documents, }) if err != nil { log.Fatal(err) } fmt.Printf("Response ID: %s\n", result.Response.ID) fmt.Printf("Model: %s\n", result.Response.ModelID) fmt.Printf("Timestamp: %v\n", result.Response.Timestamp) ``` ## Practical Examples ### RAG with Reranking Combine embeddings with reranking for improved retrieval: ```go func retrieveWithReranking( ctx context.Context, query string, documents []string, embeddingModel provider.EmbeddingModel, rerankModel provider.RerankingModel, ) ([]string, error) { // Step 1: Embed query and documents queryResult, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: query, }) if err != nil { return nil, err } docsResult, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: documents, }) if err != nil { return nil, err } // Step 2: Find top candidates using embedding similarity indices, _, err := ai.RankBySimilarity(queryResult.Embedding, docsResult.Embeddings) if err != nil { return nil, err } // Get top 10 candidates candidateCount := 10 if len(indices) < candidateCount { candidateCount = len(indices) } candidates := make([]string, candidateCount) for i := 0; i < candidateCount; i++ { candidates[i] = documents[indices[i]] } // Step 3: Rerank candidates for more accurate ordering rerankTopN := 3 // Get top 3 most relevant rerankResult, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: query, Documents: toInterfaceSlice(candidates), TopN: &rerankTopN, }) if err != nil { return nil, err } // Convert back to strings rerankedDocs := rerankResult.RerankedDocuments.([]interface{}) finalResults := make([]string, len(rerankedDocs)) for i, doc := range rerankedDocs { finalResults[i] = doc.(string) } return finalResults, nil } func toInterfaceSlice(strs []string) []interface{} { result := make([]interface{}, len(strs)) for i, s := range strs { result[i] = s } return result } ``` ### Semantic Search with Reranking Build a robust search system: ```go type Document struct { ID string Title string Content string } type SearchEngine struct { documents []Document embeddings [][]float64 embeddingModel provider.EmbeddingModel rerankModel provider.RerankingModel } func (se *SearchEngine) Search(ctx context.Context, query string, limit int) ([]Document, error) { // Embed query queryResult, err := ai.Embed(ctx, ai.EmbedOptions{ Model: se.embeddingModel, Input: query, }) if err != nil { return nil, err } // Find top candidates using embeddings (fast first pass) indices, _, err := ai.RankBySimilarity(queryResult.Embedding, se.embeddings) if err != nil { return nil, err } candidateCount := min(limit*3, len(indices)) // Get 3x candidates for reranking candidates := make([]Document, candidateCount) candidateTexts := make([]interface{}, candidateCount) for i := 0; i < candidateCount; i++ { candidates[i] = se.documents[indices[i]] candidateTexts[i] = candidates[i].Content } // Rerank for accuracy (precise second pass) rerankResult, err := ai.Rerank(ctx, ai.RerankOptions{ Model: se.rerankModel, Query: query, Documents: candidateTexts, TopN: &limit, }) if err != nil { return nil, err } // Return reranked documents rerankedDocs := rerankResult.RerankedDocuments.([]interface{}) results := make([]Document, len(rerankedDocs)) for i, doc := range rerankedDocs { // Find original document content := doc.(string) for _, candidate := range candidates { if candidate.Content == content { results[i] = candidate break } } } return results, nil } func min(a, b int) int { if a < b { return a } return b } ``` ### Multi-Query Reranking Rerank with multiple related queries: ```go func rerankWithMultipleQueries( ctx context.Context, queries []string, documents []string, rerankModel provider.RerankingModel, ) ([]string, error) { // Aggregate scores from multiple queries scoreMap := make(map[int]float64) for _, query := range queries { result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: query, Documents: toInterfaceSlice(documents), }) if err != nil { return nil, err } // Accumulate scores for _, rank := range result.Ranking { scoreMap[rank.OriginalIndex] += rank.Score } } // Sort by aggregated scores type scoredDoc struct { index int score float64 } scored := make([]scoredDoc, 0, len(scoreMap)) for idx, score := range scoreMap { scored = append(scored, scoredDoc{index: idx, score: score / float64(len(queries))}) } sort.Slice(scored, func(i, j int) bool { return scored[i].score > scored[j].score }) // Return sorted documents results := make([]string, len(scored)) for i, s := range scored { results[i] = documents[s.index] } return results, nil } ``` ### Filtering Before Reranking Filter documents before reranking to improve efficiency: ```go func searchWithFiltering( ctx context.Context, query string, documents []Document, filters map[string]interface{}, rerankModel provider.RerankingModel, ) ([]Document, error) { // Filter documents first var filtered []Document for _, doc := range documents { if matchesFilters(doc, filters) { filtered = append(filtered, doc) } } if len(filtered) == 0 { return nil, nil } // Prepare for reranking docTexts := make([]interface{}, len(filtered)) for i, doc := range filtered { docTexts[i] = doc.Content } // Rerank filtered results topN := 10 result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: query, Documents: docTexts, TopN: &topN, }) if err != nil { return nil, err } // Map back to original documents rerankedDocs := result.RerankedDocuments.([]interface{}) reranked := make([]Document, len(rerankedDocs)) for i, doc := range rerankedDocs { content := doc.(string) for _, filteredDoc := range filtered { if filteredDoc.Content == content { reranked[i] = filteredDoc break } } } return reranked, nil } func matchesFilters(doc Document, filters map[string]interface{}) bool { // Implement your filtering logic here return true } ``` ## Reranking Providers & Models Several providers offer reranking models: | Provider | Model | Description | |----------|-------|-------------| | **Cohere** | `rerank-v3.5` | Latest multilingual reranking model | | **Cohere** | `rerank-english-v3.0` | English-optimized reranking | | **Cohere** | `rerank-multilingual-v3.0` | Multilingual reranking | | **Amazon Bedrock** | `amazon.rerank-v1:0` | Amazon's reranking model | | **Amazon Bedrock** | `cohere.rerank-v3-5:0` | Cohere via Bedrock | | **Together.ai** | `Salesforce/Llama-Rank-v1` | Llama-based reranking | | **Together.ai** | `mixedbread-ai/Mxbai-Rerank-Large-V2` | Mixedbread reranking | ### Example with Different Providers ```go // Cohere cohereProvider := cohere.New(cohere.Config{APIKey: os.Getenv("COHERE_API_KEY")}) cohereModel, _ := cohereProvider.RerankingModel("rerank-v3.5") // Bedrock bedrockProvider := bedrock.New(bedrock.Config{ Region: "us-east-1", }) bedrockModel, _ := bedrockProvider.RerankingModel("amazon.rerank-v1:0") // Together.ai togetherProvider := together.New(together.Config{APIKey: os.Getenv("TOGETHER_API_KEY")}) togetherModel, _ := togetherProvider.RerankingModel("Salesforce/Llama-Rank-v1") // Use any model with the same API result, _ := ai.Rerank(ctx, ai.RerankOptions{ Model: cohereModel, // or bedrockModel, or togetherModel Query: "query", Documents: documents, }) ``` ## Performance Optimization ### Batch Reranking For large document sets, batch process for efficiency: ```go func rerankLargeDataset( ctx context.Context, query string, documents []string, rerankModel provider.RerankingModel, batchSize int, ) ([]string, error) { var allResults []ai.RerankItem for i := 0; i < len(documents); i += batchSize { end := i + batchSize if end > len(documents) { end = len(documents) } batch := documents[i:end] result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Query: query, Documents: toInterfaceSlice(batch), }) if err != nil { return nil, err } // Adjust indices for global ordering for _, rank := range result.Ranking { rank.OriginalIndex += i allResults = append(allResults, rank) } fmt.Printf("Processed %d/%d documents\n", min(end, len(documents)), len(documents)) } // Sort all results by score sort.Slice(allResults, func(i, j int) bool { return allResults[i].Score > allResults[j].Score }) // Convert to string slice reranked := make([]string, len(allResults)) for i, rank := range allResults { reranked[i] = documents[rank.OriginalIndex] } return reranked, nil } ``` ## Best Practices 1. **Use with Embeddings**: Combine reranking with embedding-based retrieval for best results 2. **Limit Candidates**: Rerank only top-N candidates from initial retrieval (typically 10-100) 3. **Set TopN**: Use `TopN` to limit results and reduce latency 4. **Handle Errors**: Implement retry logic for transient failures 5. **Monitor Performance**: Track reranking latency and adjust batch sizes 6. **Cache Results**: Cache reranking results for repeated queries 7. **Use Context**: Always pass context for cancellation and timeouts 8. **Filter First**: Apply filters before reranking to reduce processing 9. **Provider Choice**: Choose provider based on language and use case 10. **Score Thresholds**: Consider score thresholds for quality filtering ## When to Use Reranking **Use Reranking When:** - You need high precision for top results - Working with multi-modal or structured data - Initial retrieval returns many candidates - Query-document relationship is complex - Building production search systems **Use Embeddings When:** - Speed is critical - Working with large document sets (millions+) - Semantic similarity is sufficient - Building initial retrieval stage **Best Practice**: Use embeddings for fast initial retrieval, then rerank top candidates for precision. ## Next Steps - Learn about [Image Generation](https://goaisdk.com/docs/ai-sdk-core/image-generation.md) - Explore [Advanced RAG Patterns](https://goaisdk.com/docs/advanced/prompt-engineering.md) - See [Embeddings](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) ## See Also - [Embeddings](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) - [Providers](https://goaisdk.com/docs/providers/overview.md) - [API Reference: Rerank](https://goaisdk.com/docs/reference/ai/rerank.md) --- # Image Generation > Shows how to generate images in Go with ai.GenerateImage, covering settings, error handling, supported image models, and advanced generation features. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/image-generation Documentation index: https://goaisdk.com/llms.txt The Go AI SDK provides the [`ai.GenerateImage()`](https://goaisdk.com/docs/reference/ai/generate-image.md) function to generate images based on a given prompt using an image model. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) imageModel, _ := provider.ImageModel("dall-e-3") result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Santa Claus driving a Cadillac", }) if err != nil { log.Fatal(err) } // Access image data fmt.Printf("Image MIME type: %s\n", result.Image.MediaType) fmt.Printf("Image size: %d bytes\n", len(result.Image.Data)) // Save image to file err = os.WriteFile("generated.png", result.Image.Data, 0644) if err != nil { log.Fatal(err) } fmt.Println("Image saved to generated.png") } ``` You can access the image data in different formats: ```go // As bytes imageBytes := result.Image.Data // As base64 string base64String := result.Image.Base64() // Get image MIME type mimeType := result.Image.MediaType ``` ## Settings ### Size and Aspect Ratio Depending on the model, you can specify either the size or the aspect ratio. #### Size The size is specified as a string in the format `{width}x{height}`. Models only support specific sizes, and the supported sizes vary by model and provider. ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Santa Claus driving a Cadillac", Size: "1024x1024", }) ``` #### Aspect Ratio The aspect ratio is specified as a string in the format `{width}:{height}`. Models only support specific aspect ratios, which vary by model and provider. ```go provider := google.New(google.Config{APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY")}) imageModel, _ := provider.ImageModel("imagen-4.0-generate-001") result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Santa Claus driving a Cadillac", Size: "1920x1080", // Automatically converts to 16:9 aspect ratio }) ``` ### Generating Multiple Images `GenerateImage` supports generating multiple images at once: ```go n := 4 // Number of images to generate result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Santa Claus driving a Cadillac", N: &n, }) if err != nil { log.Fatal(err) } // Access all images for i, image := range result.Images { filename := fmt.Sprintf("image-%d.png", i+1) err := os.WriteFile(filename, image.Data, 0644) if err != nil { log.Fatal(err) } fmt.Printf("Saved %s\n", filename) } ``` > **Note:** `GenerateImage` will automatically batch requests as needed to generate the requested number of images. Each image model has an internal limit on how many images it can generate in a single API call. The SDK manages this automatically by batching requests appropriately. By default, the SDK uses provider-documented limits (e.g., DALL-E 3 can only generate 1 image per call, while DALL-E 2 supports up to 10). You can override this behavior using `MaxImagesPerCall`: ```go n := 10 // Will make 2 calls of 5 images each result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Santa Claus driving a Cadillac", MaxImagesPerCall: 5, // Override default batch size N: &n, }) ``` ### Providing a Seed You can provide a seed to control the output of the image generation process. If supported by the model, the same seed will always produce the same image. ```go seed := 1234567890 result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Santa Claus driving a Cadillac", Seed: &seed, }) ``` ### Provider-Specific Settings Image models often have provider- or model-specific settings. You can pass such settings using the `ProviderOptions` parameter: ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Santa Claus driving a Cadillac", Size: "1024x1024", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "style": "vivid", "quality": "hd", }, }, }) ``` ### Timeouts with Context Use Go's context for timeouts and cancellation: ```go // Timeout after 60 seconds ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second) defer cancel() result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Santa Claus driving a Cadillac", }) if err != nil { if ctx.Err() == context.DeadlineExceeded { fmt.Println("Image generation timed out") } log.Fatal(err) } ``` ### Custom Headers Add custom headers to the image generation request: ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Santa Claus driving a Cadillac", Headers: map[string]string{ "X-Custom-Header": "custom-value", }, }) ``` ### Warnings If the model returns warnings, they will be available in the `Warnings` field: ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Santa Claus driving a Cadillac", }) if err != nil { log.Fatal(err) } for _, warning := range result.Warnings { fmt.Printf("Warning: %s\n", warning.Message) } ``` ### Additional Provider Metadata Some providers expose additional metadata for the result: ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Santa Claus driving a Cadillac", }) if err != nil { log.Fatal(err) } // Access provider-specific metadata if metadata, ok := result.ProviderMetadata["openai"].(map[string]interface{}); ok { if images, ok := metadata["images"].([]interface{}); ok && len(images) > 0 { if imageMetadata, ok := images[0].(map[string]interface{}); ok { if revisedPrompt, ok := imageMetadata["revisedPrompt"].(string); ok { fmt.Printf("Original prompt: %s\n", "Santa Claus driving a Cadillac") fmt.Printf("Revised prompt: %s\n", revisedPrompt) } } } } ``` ## Error Handling Handle image generation errors: ```go import providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Generate an image", }) if err != nil { // Check for specific error types var noImageErr *ai.NoImageGeneratedError var rateLimitErr *providererrors.RateLimitError if errors.As(err, &noImageErr) { fmt.Println("No image was generated") } else if errors.As(err, &rateLimitErr) { fmt.Println("Rate limit exceeded") } else { fmt.Printf("Error: %v\n", err) } return } ``` ## Practical Examples ### Saving Images in Different Formats ```go func saveImage(image []byte, mimeType string, filename string) error { ext := "png" if mimeType == "image/jpeg" { ext = "jpg" } return os.WriteFile(fmt.Sprintf("%s.%s", filename, ext), image, 0644) } result, _ := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "Beautiful sunset", }) saveImage(result.Image.Data, result.Image.MediaType, "sunset") ``` ### Generating Image Variations Generate multiple variations of the same prompt: ```go func generateVariations(ctx context.Context, model provider.ImageModel, prompt string, count int) ([]types.GeneratedFile, error) { result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: prompt, N: &count, }) if err != nil { return nil, err } return result.Images, nil } images, _ := generateVariations(ctx, imageModel, "Abstract art with vibrant colors", 4) for i, img := range images { os.WriteFile(fmt.Sprintf("variation-%d.png", i+1), img.Data, 0644) } ``` ### Batch Image Generation Generate multiple different images: ```go prompts := []string{ "A serene mountain landscape", "A bustling city street at night", "An underwater coral reef", "A cozy coffee shop interior", } for i, prompt := range prompts { result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: prompt, }) if err != nil { log.Printf("Error generating image %d: %v\n", i+1, err) continue } filename := fmt.Sprintf("batch-%d.png", i+1) os.WriteFile(filename, result.Image.Data, 0644) fmt.Printf("Generated %s: %s\n", filename, prompt) } ``` ### Image Generation with Retry Logic ```go func generateImageWithRetry(ctx context.Context, model provider.ImageModel, prompt string, maxRetries int) (*ai.GenerateImageResult, error) { var lastErr error for attempt := 0; attempt <= maxRetries; attempt++ { if attempt > 0 { fmt.Printf("Retry attempt %d/%d\n", attempt, maxRetries) time.Sleep(time.Duration(attempt) * time.Second) } result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: prompt, }) if err == nil { return result, nil } lastErr = err // Don't retry on certain errors if errors.Is(err, context.Canceled) || errors.Is(err, context.DeadlineExceeded) { break } } return nil, fmt.Errorf("failed after %d attempts: %w", maxRetries+1, lastErr) } ``` ### Google Image Generation Examples #### Google Generative AI - Imagen 4.0 ```go // High-quality image generation with Imagen 4.0 googleProvider := google.New(google.Config{ APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"), }) imagenModel, _ := googleProvider.ImageModel("imagen-4.0-generate-001") n := 1 result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imagenModel, Prompt: "A serene mountain landscape at sunset with dramatic clouds", N: &n, Size: "1024x1024", // Square 1:1 aspect ratio }) if err != nil { log.Fatal(err) } os.WriteFile("imagen_4_output.png", result.Image.Data, 0644) fmt.Printf("Generated with Imagen 4.0: %d bytes\n", len(result.Image.Data)) ``` #### Google Generative AI - Gemini Flash Image ```go // Fast image generation with Gemini geminiModel, _ := googleProvider.ImageModel("gemini-2.5-flash-image") result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: geminiModel, Prompt: "Abstract colorful art with geometric shapes and vibrant patterns", Size: "1920x1080", // Widescreen 16:9 }) if err != nil { log.Fatal(err) } os.WriteFile("gemini_output.png", result.Image.Data, 0644) ``` #### Google Vertex AI - Imagen 3.0 ```go // Enterprise-grade image generation with Vertex AI vertexProvider, err := googlevertex.New(googlevertex.Config{ Project: os.Getenv("GOOGLE_VERTEX_PROJECT"), Location: os.Getenv("GOOGLE_VERTEX_LOCATION"), AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"), }) if err != nil { log.Fatal(err) } vertexModel, _ := vertexProvider.ImageModel("imagen-3.0-generate-001") n := 1 result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: vertexModel, Prompt: "A futuristic cityscape at night with neon lights", N: &n, Size: "1920x1080", // 16:9 aspect ratio }) if err != nil { log.Fatal(err) } os.WriteFile("vertex_imagen_output.png", result.Image.Data, 0644) ``` #### Comparing Google Models ```go prompts := []string{ "A magical forest with glowing mushrooms", "Cyberpunk city street at night", "Mountain landscape with aurora borealis", } models := map[string]provider.ImageModel{ "imagen-4.0": imagenModel, "gemini-flash": geminiModel, "vertex-imagen": vertexModel, } for name, model := range models { for i, prompt := range prompts { result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: prompt, Size: "1024x1024", }) if err != nil { log.Printf("Error with %s: %v\n", name, err) continue } filename := fmt.Sprintf("%s_%d.png", name, i+1) os.WriteFile(filename, result.Image.Data, 0644) fmt.Printf("✓ Generated %s\n", filename) } } ``` ## Image Models Several providers offer image generation models: | Provider | Model | Supported Sizes/Aspect Ratios | |----------|-------|------------------------------| | **xAI** | `grok-imagine-image` | Flexible aspect ratios (1:1, 16:9, 9:16, 4:3, etc.) | | **OpenAI** | `gpt-image-1` | 1024x1024, 1536x1024, 1024x1536 | | **OpenAI** | `dall-e-3` | 1024x1024, 1792x1024, 1024x1792 | | **OpenAI** | `dall-e-2` | 256x256, 512x512, 1024x1024 | | **Amazon Bedrock** | `amazon.nova-canvas-v1:0` | 320-4096 (multiples of 16) | | **Fal** | `fal-ai/flux/dev` | 1:1, 3:4, 4:3, 9:16, 16:9, 9:21, 21:9 | | **Fal** | `fal-ai/flux-pro/v1.1-ultra` | 1:1, 3:4, 4:3, 9:16, 16:9, 9:21, 21:9 | | **Google** | `imagen-4.0-generate-001` | 1:1, 3:4, 4:3, 9:16, 16:9 | | **Google** | `gemini-2.5-flash-image` | 1:1, 3:4, 4:3, 9:16, 16:9 | | **Google Vertex** | `imagen-3.0-generate-001` | 1:1, 3:4, 4:3, 9:16, 16:9 | | **Black Forest Labs** | `flux-pro-1.1-ultra` | 3:7 to 7:3 aspect ratios | | **Stability AI** | Various via providers | Multiple sizes supported | ### Example with Different Providers ```go // XAI (Grok) xaiProvider := xai.New(xai.Config{APIKey: os.Getenv("XAI_API_KEY")}) grokModel, _ := xaiProvider.ImageModel("grok-imagine-image") // OpenAI openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) dalleModel, _ := openaiProvider.ImageModel("dall-e-3") // Google Generative AI - Imagen 4.0 googleProvider := google.New(google.Config{APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY")}) imagenModel, _ := googleProvider.ImageModel("imagen-4.0-generate-001") // Google Generative AI - Gemini Image geminiImageModel, _ := googleProvider.ImageModel("gemini-2.5-flash-image") // Google Vertex AI - Imagen 3.0 vertexProvider, _ := googlevertex.New(googlevertex.Config{ Project: os.Getenv("GOOGLE_VERTEX_PROJECT"), Location: os.Getenv("GOOGLE_VERTEX_LOCATION"), AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"), }) vertexImagenModel, _ := vertexProvider.ImageModel("imagen-3.0-generate-001") // Fal falProvider := fal.New(fal.Config{APIKey: os.Getenv("FAL_API_KEY")}) fluxModel, _ := falProvider.ImageModel("fal-ai/flux/dev") // Use any model with the same API result, _ := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imagenModel, // or grokModel, dalleModel, geminiImageModel, vertexImagenModel, or fluxModel Prompt: "A beautiful landscape", }) ``` ## Advanced Features ### Image Editing Some providers support editing existing images with natural language instructions: ```go // XAI image editing imageData, _ := os.ReadFile("source.jpg") result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: grokModel, Prompt: "Change the sky to vibrant sunset colors", Files: []provider.ImageFile{ { Type: "file", Data: imageData, MediaType: "image/jpeg", }, }, }) ``` ### Checking File and Mask Support Not every model that accepts `Files` / `Mask` actually consumes them for editing; some silently ignore them and return an unrelated fresh generation. A model can optionally advertise support through `SupportsFileInputs() *bool` / `SupportsMaskInputs() *bool`, read via `provider.ImageModelSupportsFileInputs` / `provider.ImageModelSupportsMaskInputs`. `nil` means unknown — only route an editing request to a model when the result is a non-nil `true`: ```go if supported := provider.ImageModelSupportsFileInputs(model); supported != nil && *supported { // safe to pass Files for image editing } ``` ### Image Inpainting Fill specific areas of an image with new content using masks: ```go sourceImage, _ := os.ReadFile("source.jpg") maskImage, _ := os.ReadFile("mask.png") result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: grokModel, Prompt: "Add a rainbow in the masked area", Files: []provider.ImageFile{ {Type: "file", Data: sourceImage, MediaType: "image/jpeg"}, }, Mask: &provider.ImageFile{ Type: "file", Data: maskImage, MediaType: "image/png", }, }) ``` **Mask Requirements:** - White pixels (255,255,255) indicate areas to modify - Black pixels (0,0,0) indicate areas to preserve - Mask should match source image dimensions ### Image Variations Generate multiple variations of a source image: ```go sourceImage, _ := os.ReadFile("source.jpg") n := 4 result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: grokModel, Prompt: "Generate creative variations with different styles", N: &n, Files: []provider.ImageFile{ {Type: "file", Data: sourceImage, MediaType: "image/jpeg"}, }, }) ``` Providers that can return multiple images expose every image through `result.Images`. Low-level provider results also preserve provider-returned base64 strings in `result.Base64Images`; the high-level `ai.GenerateImage` helper decodes those strings into `GeneratedFile.Data` so JSON output contains the expected base64 image data. ## Best Practices 1. **Set Timeouts**: Image generation can be slow; always use context with timeout 2. **Handle Errors**: Implement proper error handling and retry logic 3. **Validate Prompts**: Ensure prompts are appropriate and within provider guidelines 4. **Choose Right Model**: Different models excel at different styles and use cases 5. **Use Appropriate Sizes**: Request sizes supported by the model 6. **Save Efficiently**: Use appropriate image formats and compression 7. **Monitor Costs**: Track image generation usage and costs 8. **Respect Rate Limits**: Implement rate limiting for bulk generation 9. **Cache Results**: Cache generated images when appropriate 10. **Test Prompts**: Iterate on prompts to get desired results ## Next Steps - Learn about [Speech Generation](https://goaisdk.com/docs/ai-sdk-core/speech.md) - Explore [Provider Documentation](https://goaisdk.com/docs/providers/overview.md) - See [Examples](https://github.com/digitallysavvy/go-ai/tree/main/examples) ## See Also - [Providers](https://goaisdk.com/docs/providers/overview.md) - [Settings](https://goaisdk.com/docs/ai-sdk-core/settings.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [API Reference: GenerateImage](https://goaisdk.com/docs/reference/ai/generate-image.md) --- # Transcription > Shows how to transcribe audio in Go with ai.Transcribe, covering supported audio input formats, accessing results, settings, and transcription models. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/transcription Documentation index: https://goaisdk.com/llms.txt The Go AI SDK provides the [`ai.Transcribe()`](https://goaisdk.com/docs/reference/ai/transcribe.md) function to transcribe audio using a transcription model. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) transcriptionModel, _ := provider.TranscriptionModel("whisper-1") // Read audio file audioData, err := os.ReadFile("audio.mp3") if err != nil { log.Fatal(err) } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, }) if err != nil { log.Fatal(err) } fmt.Println("Transcript:", result.Text) } ``` ## Audio Input Formats `Transcribe` accepts raw bytes, base64/base64url data, or a URL that the SDK downloads before calling the provider. ### From File ```go audioData, err := os.ReadFile("audio.mp3") if err != nil { log.Fatal(err) } result, _ := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, }) ``` ### From URL ```go result, _ := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, AudioURL: "https://example.com/audio.mp3", }) ``` ### From Base64 String ```go base64Audio := "UklGRi..." // base64 encoded audio result, _ := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, AudioBase64: base64Audio, }) ``` ## Accessing Results To access the generated transcript and metadata: ```go result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, }) if err != nil { log.Fatal(err) } // Transcript text text := result.Text // e.g. "Hello, world!" // Segments with timestamps (if available) for _, segment := range result.Segments { fmt.Printf("[%.2f-%.2f] %s\n", segment.StartSecond, segment.EndSecond, segment.Text) } // Detected language (if available) if result.Language != "" { fmt.Printf("Detected language: %s\n", result.Language) } // Audio duration if result.DurationInSeconds != nil { fmt.Printf("Duration: %.2f seconds\n", *result.DurationInSeconds) } ``` ## Settings ### Provider Options Provider-specific controls such as language hints, diarization, timestamp granularity, and formatting are passed through `ProviderOptions` under the provider key: ```go result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "language": "en", }, }, }) ``` ### Timestamps Providers return timestamped transcript portions through `Segments` when the provider response includes them. Request provider-specific timestamp behavior with `ProviderOptions`: ```go result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "timestampGranularities": []string{"word"}, }, }, }) for _, segment := range result.Segments { fmt.Printf("[%.2f-%.2f] %s\n", segment.StartSecond, segment.EndSecond, segment.Text) } ``` ### MIME Type Specify the audio format explicitly: ```go result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, MimeType: "audio/mpeg", // or "audio/wav", "audio/webm", etc. }) ``` ### Provider-Specific Settings Transcription models often have provider- or model-specific settings which you can set using the `ProviderOptions` parameter: ```go result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, MimeType: "audio/wav", ProviderOptions: map[string]interface{}{ "xai": map[string]interface{}{ "language": "en", "diarize": true, "keyterm": []string{"Go-AI", "Grok"}, "fillerWords": true, }, }, }) ``` ### Timeouts with Context Use Go's context for timeouts and cancellation: ```go import "time" // Timeout after 30 seconds ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, }) if err != nil { if ctx.Err() == context.DeadlineExceeded { fmt.Println("Transcription timed out") } log.Fatal(err) } ``` ### Custom Headers Add custom headers to the transcription request: ```go result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, Headers: map[string]string{ "X-Custom-Header": "custom-value", }, }) ``` ### Warnings If the model returns warnings, they will be available in the `Warnings` field: ```go result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, }) if err != nil { log.Fatal(err) } for _, warning := range result.Warnings { fmt.Printf("Warning: %s\n", warning.Message) } ``` ## Error Handling When `ai.Transcribe()` cannot generate a valid transcript, it returns an error: ```go import "errors" result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, }) if err != nil { if ai.IsNoTranscriptGeneratedError(err) { fmt.Println("No transcript was generated") var noTranscript *ai.NoTranscriptGeneratedError if errors.As(err, &noTranscript) { fmt.Printf("Responses: %d\n", len(noTranscript.Responses)) } } else { fmt.Printf("Transcription error: %v\n", err) } return } ``` Common transcription errors: - **NoTranscriptGeneratedError**: The model returned no transcript text - **ProviderError**: Provider-specific HTTP or API failure - **InvalidArgumentError**: Invalid input such as negative `MaxRetries` or invalid base64 audio - **context.DeadlineExceeded**: Operation timed out - **context.Canceled**: Operation was canceled ## Practical Examples ### Transcribe with Automatic Language Detection ```go func transcribeAudio(ctx context.Context, model provider.TranscriptionModel, audioPath string) (*ai.TranscribeResult, error) { audioData, err := os.ReadFile(audioPath) if err != nil { return nil, fmt.Errorf("failed to read audio: %w", err) } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: model, Audio: audioData, }) if err != nil { return nil, err } fmt.Printf("Detected language: %s\n", result.Language) if result.DurationInSeconds != nil { fmt.Printf("Duration: %.2f seconds\n", *result.DurationInSeconds) } return result, nil } ``` ### Batch Transcription ```go func transcribeMultipleFiles(ctx context.Context, model provider.TranscriptionModel, audioPaths []string) ([]string, error) { transcripts := make([]string, len(audioPaths)) for i, path := range audioPaths { audioData, err := os.ReadFile(path) if err != nil { return nil, fmt.Errorf("failed to read %s: %w", path, err) } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: model, Audio: audioData, }) if err != nil { return nil, fmt.Errorf("failed to transcribe %s: %w", path, err) } transcripts[i] = result.Text fmt.Printf("Transcribed %s: %d characters\n", path, len(result.Text)) } return transcripts, nil } ``` ### Parallel Transcription with Goroutines ```go func transcribeParallel(ctx context.Context, model provider.TranscriptionModel, audioPaths []string) ([]string, error) { type result struct { index int text string err error } resultChan := make(chan result, len(audioPaths)) transcripts := make([]string, len(audioPaths)) // Start goroutines for i, path := range audioPaths { go func(index int, audioPath string) { audioData, err := os.ReadFile(audioPath) if err != nil { resultChan <- result{index: index, err: err} return } transcription, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: model, Audio: audioData, }) if err != nil { resultChan <- result{index: index, err: err} return } resultChan <- result{index: index, text: transcription.Text} }(i, path) } // Collect results for i := 0; i < len(audioPaths); i++ { res := <-resultChan if res.err != nil { return nil, fmt.Errorf("transcription failed for file %d: %w", res.index, res.err) } transcripts[res.index] = res.text } return transcripts, nil } ``` ### Transcription with Timestamps ```go func transcribeWithTimestamps(ctx context.Context, model provider.TranscriptionModel, audioPath string) error { audioData, err := os.ReadFile(audioPath) if err != nil { return err } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: model, Audio: audioData, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "timestampGranularities": []string{"word"}, }, }, }) if err != nil { return err } // Print timestamped transcript fmt.Println("Timestamped Transcript:") fmt.Println("======================") for _, segment := range result.Segments { minutes := int(segment.StartSecond) / 60 seconds := int(segment.StartSecond) % 60 fmt.Printf("[%02d:%02d] %s\n", minutes, seconds, segment.Text) } return nil } ``` ### Transcription with Retry Logic ```go func transcribeWithRetry(ctx context.Context, model provider.TranscriptionModel, audioData []byte, maxRetries int) (*ai.TranscribeResult, error) { return ai.Transcribe(ctx, ai.TranscribeOptions{ Model: model, Audio: audioData, MaxRetries: &maxRetries, }) } ``` ### Save Transcript to File ```go func transcribeAndSave(ctx context.Context, model provider.TranscriptionModel, audioPath, outputPath string) error { audioData, err := os.ReadFile(audioPath) if err != nil { return fmt.Errorf("failed to read audio: %w", err) } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: model, Audio: audioData, }) if err != nil { return fmt.Errorf("transcription failed: %w", err) } // Save to file err = os.WriteFile(outputPath, []byte(result.Text), 0644) if err != nil { return fmt.Errorf("failed to write transcript: %w", err) } fmt.Printf("Transcript saved to %s\n", outputPath) return nil } ``` ## Transcription Models Several providers offer transcription models: | Provider | Model | Notes | |----------|-------|-------| | **OpenAI** | `whisper-1` | General-purpose transcription | | **OpenAI** | `gpt-4o-transcribe` | GPT-4o with transcription | | **OpenAI** | `gpt-4o-mini-transcribe` | Faster, cost-effective | | **ElevenLabs** | `scribe_v1` | High-quality transcription | | **ElevenLabs** | `scribe_v1_experimental` | Experimental features | | **Groq** | `whisper-large-v3-turbo` | Fast transcription | | **Groq** | `whisper-large-v3` | High accuracy | | **Azure OpenAI** | `whisper-1` | Azure-hosted Whisper | | **Azure OpenAI** | `gpt-4o-transcribe` | Azure GPT-4o | | **Azure OpenAI** | `gpt-4o-mini-transcribe` | Azure GPT-4o Mini | | **Rev.ai** | `machine` | Standard machine transcription | | **Rev.ai** | `low_cost` | Budget-friendly option | | **Rev.ai** | `fusion` | Hybrid human+machine | | **Deepgram** | `base` | Base model (+ variants) | | **Deepgram** | `enhanced` | Enhanced accuracy | | **Deepgram** | `nova` | Latest model | | **Deepgram** | `nova-2` | Nova 2nd generation | | **Deepgram** | `nova-3` | Nova 3rd generation | | **Gladia** | `default` | Standard transcription | | **AssemblyAI** | `best` | Highest accuracy | | **AssemblyAI** | `nano` | Fast, lightweight | | **Fal** | `whisper` | Whisper model | | **Fal** | `wizper` | Optimized variant | ### Example with Different Providers ```go // OpenAI openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) whisperModel, _ := openaiProvider.TranscriptionModel("whisper-1") // Groq groqProvider := groq.New(groq.Config{APIKey: os.Getenv("GROQ_API_KEY")}) groqModel, _ := groqProvider.TranscriptionModel("whisper-large-v3-turbo") // AssemblyAI assemblyProvider := assemblyai.New(assemblyai.Config{APIKey: os.Getenv("ASSEMBLYAI_API_KEY")}) assemblyModel, _ := assemblyProvider.TranscriptionModel("best") // Use any model with the same API result, _ := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: whisperModel, // or groqModel, or assemblyModel Audio: audioData, }) ``` ## Supported Audio Formats Most transcription providers support common audio formats: - **MP3** (audio/mpeg, audio/mp3) - **WAV** (audio/wav) - **M4A** (audio/mp4, audio/m4a) - **WebM** (audio/webm) - **OGG** (audio/ogg) - **FLAC** (audio/flac) Maximum file sizes and durations vary by provider. Check provider documentation for specific limits. ## Best Practices 1. **Choose the Right Model**: Balance speed, accuracy, and cost based on your needs 2. **Set Timeouts**: Transcription can be slow; always use context with timeout 3. **Handle Errors**: Implement proper error handling and retry logic 4. **Provide Language Hints**: When known, language hints improve accuracy 5. **Use Appropriate Format**: Some providers work better with specific audio formats 6. **Monitor Costs**: Track transcription usage and costs 7. **Respect Rate Limits**: Implement rate limiting for bulk transcription 8. **Use Parallel Processing**: For multiple files, use goroutines 9. **Validate Audio**: Check file size and format before transcription 10. **Cache Results**: Store transcripts to avoid redundant API calls ## Performance Tips ### Optimize Audio Files ```go // Check file size before transcription func checkAudioSize(path string, maxSizeMB int) error { info, err := os.Stat(path) if err != nil { return err } sizeMB := float64(info.Size()) / (1024 * 1024) if sizeMB > float64(maxSizeMB) { return fmt.Errorf("audio file too large: %.2f MB (max: %d MB)", sizeMB, maxSizeMB) } return nil } ``` ### Rate Limiting ```go import "golang.org/x/time/rate" // Create rate limiter (e.g., 10 requests per minute) limiter := rate.NewLimiter(rate.Every(6*time.Second), 1) func transcribeWithRateLimit(ctx context.Context, model provider.TranscriptionModel, audioData []byte) (*ai.TranscribeResult, error) { // Wait for rate limiter if err := limiter.Wait(ctx); err != nil { return nil, err } return ai.Transcribe(ctx, ai.TranscribeOptions{ Model: model, Audio: audioData, }) } ``` ## Next Steps - Learn about [Speech Generation](https://goaisdk.com/docs/ai-sdk-core/speech.md) - Explore [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - See [Provider Documentation](https://goaisdk.com/docs/providers/overview.md) ## See Also - [Providers](https://goaisdk.com/docs/providers/overview.md) - [Settings](https://goaisdk.com/docs/ai-sdk-core/settings.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [API Reference: Transcribe](https://goaisdk.com/docs/reference/ai/transcribe.md) --- # Speech Generation > Shows how to generate speech from text in Go with ai.GenerateSpeech, covering audio data access, settings, supported formats, and performance tips. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/speech Documentation index: https://goaisdk.com/llms.txt The Go AI SDK provides the [`ai.GenerateSpeech()`](https://goaisdk.com/docs/reference/ai/generate-speech.md) function to generate speech from text using a speech model. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) speechModel, _ := provider.SpeechModel("tts-1") result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "alloy", }) if err != nil { log.Fatal(err) } // Save audio to file err = os.WriteFile("output.mp3", result.Audio.Data, 0644) if err != nil { log.Fatal(err) } fmt.Println("Audio saved to output.mp3") } ``` ## Accessing Audio Data To access the generated audio: ```go speed := 1.25 result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "alloy", }) // Audio bytes audioData := result.Audio.Data // MIME type (e.g., "audio/mpeg") mimeType := result.Audio.MediaType fmt.Printf("Generated %d bytes as %s\n", len(audioData), mimeType) ``` ## Settings ### Voice Selection Different models support different voices. Check your provider's documentation for available voices: ```go // OpenAI voices: alloy, echo, fable, onyx, nova, shimmer result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "nova", }) ``` ### Speech Speed Control the speed of speech generation (typically 0.25 to 4.0): ```go speed := 1.5 // 1.5x speed result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "alloy", Speed: &speed, }) ``` ### Language Setting You can specify the language for speech generation (provider support varies): ```go result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hola, mundo!", Voice: "alloy", Language: "es", // Spanish }) ``` ### Audio Format Some providers allow you to specify the output audio format: ```go result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "alloy", OutputFormat: "mp3", // or "wav", "opus", "flac", etc. }) ``` ### Provider-Specific Settings You can set model-specific settings with the `ProviderOptions` parameter: ```go result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "alloy", OutputFormat: "opus", Speed: &speed, }) ``` ### Timeouts with Context Use Go's context for timeouts and cancellation: ```go import "time" // Timeout after 30 seconds ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "alloy", }) if err != nil { if ctx.Err() == context.DeadlineExceeded { fmt.Println("Speech generation timed out") } log.Fatal(err) } ``` ### Custom Headers Add custom headers to the speech generation request: ```go result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "alloy", Headers: map[string]string{ "X-Custom-Header": "custom-value", }, }) ``` ### Warnings If the model returns warnings, they will be available in the `Warnings` field: ```go result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "alloy", }) if err != nil { log.Fatal(err) } for _, warning := range result.Warnings { fmt.Printf("Warning: %s\n", warning.Message) } ``` ## Error Handling When `ai.GenerateSpeech()` cannot generate valid audio, it returns an error: ```go import "errors" result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello, world!", Voice: "alloy", }) if err != nil { if ai.IsNoSpeechGeneratedError(err) { fmt.Println("No speech was generated") var noSpeech *ai.NoSpeechGeneratedError if errors.As(err, &noSpeech) { fmt.Printf("Responses: %d\n", len(noSpeech.Responses)) } } else { fmt.Printf("Speech generation error: %v\n", err) } return } ``` Common speech generation errors: - **NoSpeechGeneratedError**: The model returned no generated audio - **ProviderError**: Provider-specific HTTP or API failure - **InvalidArgumentError**: Invalid input such as negative `MaxRetries` - **context.DeadlineExceeded**: Operation timed out - **context.Canceled**: Operation was canceled ## Practical Examples ### Generate and Save Audio ```go func generateAndSaveAudio(ctx context.Context, model provider.SpeechModel, text, voice, outputPath string) error { result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: model, Text: text, Voice: voice, }) if err != nil { return fmt.Errorf("speech generation failed: %w", err) } err = os.WriteFile(outputPath, result.Audio.Data, 0644) if err != nil { return fmt.Errorf("failed to save audio: %w", err) } fmt.Printf("Audio saved to %s (%d bytes)\n", outputPath, len(result.Audio.Data)) return nil } ``` ### Generate Multiple Voices ```go func generateWithMultipleVoices(ctx context.Context, model provider.SpeechModel, text string, voices []string) error { for i, voice := range voices { result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: model, Text: text, Voice: voice, }) if err != nil { return fmt.Errorf("failed to generate with voice %s: %w", voice, err) } filename := fmt.Sprintf("output-%s.mp3", voice) err = os.WriteFile(filename, result.Audio.Data, 0644) if err != nil { return fmt.Errorf("failed to save %s: %w", filename, err) } fmt.Printf("%d. Generated %s\n", i+1, filename) } return nil } // Usage voices := []string{"alloy", "echo", "fable", "onyx", "nova", "shimmer"} generateWithMultipleVoices(ctx, speechModel, "Hello, world!", voices) ``` ### Text-to-Speech with Retry Logic ```go func generateSpeechWithRetry(ctx context.Context, model provider.SpeechModel, text, voice string, maxRetries int) (*ai.GenerateSpeechResult, error) { return ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: model, Text: text, Voice: voice, MaxRetries: &maxRetries, }) } ``` ### Batch Text-to-Speech ```go func batchGenerateSpeech(ctx context.Context, model provider.SpeechModel, texts []string, voice string) ([][]byte, error) { audioFiles := make([][]byte, len(texts)) for i, text := range texts { result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: model, Text: text, Voice: voice, }) if err != nil { return nil, fmt.Errorf("failed to generate audio for text %d: %w", i, err) } audioFiles[i] = result.Audio.Data fmt.Printf("Generated audio %d/%d (%d bytes)\n", i+1, len(texts), len(result.Audio.Data)) } return audioFiles, nil } ``` ### Parallel Speech Generation with Goroutines ```go func generateSpeechParallel(ctx context.Context, model provider.SpeechModel, texts []string, voice string) ([][]byte, error) { type result struct { index int audio []byte err error } resultChan := make(chan result, len(texts)) audioFiles := make([][]byte, len(texts)) // Start goroutines for i, text := range texts { go func(index int, textContent string) { speech, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: model, Text: textContent, Voice: voice, }) if err != nil { resultChan <- result{index: index, err: err} return } resultChan <- result{index: index, audio: speech.Audio.Data} }(i, text) } // Collect results for i := 0; i < len(texts); i++ { res := <-resultChan if res.err != nil { return nil, fmt.Errorf("speech generation failed for text %d: %w", res.index, res.err) } audioFiles[res.index] = res.audio } return audioFiles, nil } ``` ### Generate Speech with Different Speeds ```go func generateAtDifferentSpeeds(ctx context.Context, model provider.SpeechModel, text, voice string) error { speeds := []float64{0.5, 0.75, 1.0, 1.25, 1.5, 2.0} for _, speed := range speeds { result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: model, Text: text, Voice: voice, Speed: &speed, }) if err != nil { return fmt.Errorf("failed to generate at speed %.2fx: %w", speed, err) } filename := fmt.Sprintf("output-%.2fx.mp3", speed) err = os.WriteFile(filename, result.Audio.Data, 0644) if err != nil { return fmt.Errorf("failed to save %s: %w", filename, err) } fmt.Printf("Generated %s at %.2fx speed\n", filename, speed) } return nil } ``` ### Streaming to HTTP Response ```go import "net/http" func speechHandler(w http.ResponseWriter, r *http.Request) { text := r.URL.Query().Get("text") voice := r.URL.Query().Get("voice") if text == "" || voice == "" { http.Error(w, "Missing text or voice parameter", http.StatusBadRequest) return } provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) speechModel, _ := provider.SpeechModel("tts-1") result, err := ai.GenerateSpeech(r.Context(), ai.GenerateSpeechOptions{ Model: speechModel, Text: text, Voice: voice, }) if err != nil { http.Error(w, fmt.Sprintf("Speech generation failed: %v", err), http.StatusInternalServerError) return } // Set appropriate headers w.Header().Set("Content-Type", result.Audio.MediaType) w.Header().Set("Content-Length", fmt.Sprintf("%d", len(result.Audio.Data))) w.Header().Set("Content-Disposition", "inline; filename=\"speech.mp3\"") // Write audio data w.Write(result.Audio.Data) } ``` ### Long Text with Chunking ```go import "strings" func generateLongText(ctx context.Context, model provider.SpeechModel, longText, voice string, maxChunkSize int) ([]byte, error) { // Split text into chunks sentences := strings.Split(longText, ". ") var chunks []string var currentChunk strings.Builder for _, sentence := range sentences { if currentChunk.Len()+len(sentence) > maxChunkSize { chunks = append(chunks, currentChunk.String()) currentChunk.Reset() } currentChunk.WriteString(sentence) currentChunk.WriteString(". ") } if currentChunk.Len() > 0 { chunks = append(chunks, currentChunk.String()) } // Generate audio for each chunk var allAudio []byte for i, chunk := range chunks { fmt.Printf("Generating chunk %d/%d...\n", i+1, len(chunks)) result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: model, Text: chunk, Voice: voice, }) if err != nil { return nil, fmt.Errorf("failed to generate chunk %d: %w", i+1, err) } allAudio = append(allAudio, result.Audio.Data...) } return allAudio, nil } ``` ## Speech Models Several providers offer speech synthesis models: | Provider | Model | Voices | Languages | |----------|-------|--------|-----------| | **OpenAI** | `tts-1` | alloy, echo, fable, onyx, nova, shimmer | Multiple | | **OpenAI** | `tts-1-hd` | alloy, echo, fable, onyx, nova, shimmer | Multiple | | **OpenAI** | `gpt-4o-mini-tts` | alloy, echo, fable, onyx, nova, shimmer | Multiple | | **ElevenLabs** | `eleven_v3` | Custom voices | 30+ languages | | **ElevenLabs** | `eleven_multilingual_v2` | Custom voices | Multilingual | | **ElevenLabs** | `eleven_flash_v2_5` | Custom voices | Fast, low latency | | **ElevenLabs** | `eleven_flash_v2` | Custom voices | Fast | | **ElevenLabs** | `eleven_turbo_v2_5` | Custom voices | Ultra-fast | | **ElevenLabs** | `eleven_turbo_v2` | Custom voices | Ultra-fast | | **LMNT** | `aurora` | Various | English | | **LMNT** | `blizzard` | Various | English | | **Hume** | `default` | Expressive | English | ### Example with Different Providers ```go // OpenAI openaiProvider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) openaiModel, _ := openaiProvider.SpeechModel("tts-1-hd") // ElevenLabs elevenProvider := elevenlabs.New(elevenlabs.Config{APIKey: os.Getenv("ELEVENLABS_API_KEY")}) elevenModel, _ := elevenProvider.SpeechModel("eleven_multilingual_v2") // LMNT lmntProvider := lmnt.New(lmnt.Config{APIKey: os.Getenv("LMNT_API_KEY")}) lmntModel, _ := lmntProvider.SpeechModel("aurora") // Use any model with the same API result, _ := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: openaiModel, // or elevenModel, or lmntModel Text: "Hello, world!", Voice: "alloy", // voice name depends on provider }) ``` ## Supported Audio Formats Common audio formats supported by providers: - **MP3** (audio/mpeg) - Most common, good compression - **WAV** (audio/wav) - Uncompressed, high quality - **Opus** (audio/opus) - Efficient for speech, low latency - **AAC** (audio/aac) - Good quality, streaming-friendly - **FLAC** (audio/flac) - Lossless compression - **PCM** (audio/pcm) - Raw audio data Format availability varies by provider. Check provider documentation for specific support. ## Best Practices 1. **Choose the Right Model**: Balance quality, speed, and cost 2. **Set Timeouts**: Speech generation can be slow; always use context with timeout 3. **Handle Errors**: Implement proper error handling and retry logic 4. **Select Appropriate Voice**: Test different voices for your use case 5. **Monitor Costs**: Track character usage and generation costs 6. **Respect Rate Limits**: Implement rate limiting for bulk generation 7. **Use Parallel Processing**: For multiple texts, use goroutines 8. **Chunk Long Texts**: Split long texts into manageable chunks 9. **Cache Results**: Store generated audio when appropriate 10. **Test Quality**: Validate audio quality before production use ## Performance Tips ### Rate Limiting ```go import "golang.org/x/time/rate" // Create rate limiter (e.g., 10 requests per minute) limiter := rate.NewLimiter(rate.Every(6*time.Second), 1) func generateSpeechWithRateLimit(ctx context.Context, model provider.SpeechModel, text, voice string) (*ai.GenerateSpeechResult, error) { // Wait for rate limiter if err := limiter.Wait(ctx); err != nil { return nil, err } return ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: model, Text: text, Voice: voice, }) } ``` ### Character Count Validation ```go func validateTextLength(text string, maxChars int) error { if len(text) > maxChars { return fmt.Errorf("text too long: %d characters (max: %d)", len(text), maxChars) } return nil } // Before generating speech if err := validateTextLength(text, 4096); err != nil { log.Fatal(err) } ``` ## Speech Translation (Experimental) Speech-to-speech translation streams source-language audio in and receives translated audio (and text) out, over a WebSocket connection. It is separate from text-to-speech (`GenerateSpeech`) and speech-to-text (`Transcribe`). Google (`SpeechTranslationModel`, Live Translation over the BidiGenerateContent WebSocket) and OpenAI (`SpeechTranslationModel`, the `/realtime/translations` WebSocket) implement `provider.SpeechTranslationModel`, accessed through `Provider.SpeechTranslationModel(modelID)` or the TS-compatible alias `Provider.Translation(modelID)`: ```go prov := google.New(google.Config{APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY")}) translationModel, err := prov.SpeechTranslationModel("gemini-2.0-flash-live-001") if err != nil { log.Fatal(err) } result, err := ai.ExperimentalStreamTranslate(ctx, ai.StreamTranslateOptions{ Model: translationModel, Audio: audioStream, // provider.AudioStream of raw input chunks InputAudioFormat: provider.AudioFormat{Type: "audio/pcm", Rate: intPtr(16000)}, TargetLanguage: "es", }) if err != nil { log.Fatal(err) } translation, err := result.TranslationText() // waits for the stream to finish if err != nil { log.Fatal(err) } fmt.Println(translation) ``` This feature is experimental and its API may change. See the package docs under `pkg/providers/google` and `pkg/providers/openai` for the current `provider.AudioStream` and streamed-event shapes. ## Next Steps - Learn about [Transcription](https://goaisdk.com/docs/ai-sdk-core/transcription.md) - Explore [Settings](https://goaisdk.com/docs/ai-sdk-core/settings.md) - See [Provider Documentation](https://goaisdk.com/docs/providers/overview.md) ## See Also - [Providers](https://goaisdk.com/docs/providers/overview.md) - [Settings](https://goaisdk.com/docs/ai-sdk-core/settings.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [API Reference: GenerateSpeech](https://goaisdk.com/docs/reference/ai/generate-speech.md) --- # Video Generation > Covers text-to-video and image-to-video generation across multiple providers in the Go AI SDK, including async polling with DoStart and DoStatus. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/video-generation Documentation index: https://goaisdk.com/llms.txt The Go AI SDK supports video generation through multiple providers, enabling you to create videos from text descriptions (text-to-video) or animate static images (image-to-video). ## Quick Start ```go package main import ( "context" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/google" ) func main() { ctx := context.Background() // Create provider prov := google.New(google.Config{ APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"), }) // Get video model model, _ := prov.VideoModel("gemini-2.0-flash") // Generate video result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: ai.VideoPrompt{ Text: "A serene sunset over calm ocean waters", }, AspectRatio: "16:9", Resolution: "1920x1080", }) if err != nil { panic(err) } // Save video os.WriteFile("output.mp4", result.Video.Data, 0644) } ``` ## Text-to-Video Generation Generate videos from text descriptions: ```go result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: ai.VideoPrompt{ Text: "A cat playing piano in a jazz club", }, AspectRatio: "16:9", Resolution: "1920x1080", Duration: floatPtr(5.0), // 5 seconds }) ``` ### Prompt Tips - Be specific and descriptive - Include details about motion, lighting, and atmosphere - Mention camera movements if desired - Keep prompts clear and concise (under 200 words) **Good Prompts:** - "A serene mountain lake at sunrise with mist rising from the water" - "Close-up of raindrops falling on a window with bokeh city lights in background" - "Slow pan across a field of sunflowers swaying in the breeze" **Poor Prompts:** - "video" (too vague) - "make it cool" (no visual description) - Long paragraphs with multiple unrelated scenes ## Image-to-Video Generation Animate static images: ```go imageData, _ := os.ReadFile("input.jpg") result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: ai.VideoPrompt{ Text: "Animate with gentle movement and natural motion", Image: &ai.VideoPromptImage{ Data: imageData, MediaType: "image/jpeg", }, }, Duration: floatPtr(5.0), }) ``` ### Image-to-Video Tips - Use high-quality images (at least 1280x720) - Match aspect ratio to your image dimensions - Keep animation prompts simple - Shorter videos (3-5 seconds) generally work better ## Supported Providers ### XAI (Grok) **Model:** `grok-imagine-video` ```go prov := xai.New(xai.Config{ APIKey: os.Getenv("XAI_API_KEY"), }) model, _ := prov.VideoModel("grok-imagine-video") ``` **Features:** - Text-to-video generation - Image-to-video animation - Video editing with text instructions - 480p and 720p resolution support - Polling-based async generation **API Key:** Get your key at [x.ai](https://x.ai) **Supported Operations:** - **Text-to-Video**: Generate videos from text descriptions - **Image-to-Video**: Animate static images into videos - **Video Editing**: Modify existing videos with natural language prompts ### Google Generative AI **Model:** `gemini-2.0-flash` ```go prov := google.New(google.Config{ APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"), }) model, _ := prov.VideoModel("gemini-2.0-flash") ``` **Features:** - High-quality video generation - Text-to-video and image-to-video - Polling-based async generation - 720p, 1080p, and 4K resolution support **API Key:** Get your key at [Google AI Studio](https://makersuite.google.com/app/apikey) ### FAL **Models:** `fal-ai/luma-dream-machine`, `fal-ai/hunyuan-video` ```go prov := fal.New(fal.Config{ APIKey: os.Getenv("FAL_KEY"), }) model := fal.NewVideoModel(prov, "fal-ai/luma-dream-machine") ``` **Features:** - Fast generation (15-30 seconds) - High-quality results - Good for prototyping - Provider-specific options (motion strength, loop) **API Key:** Get your key at [fal.ai](https://fal.ai) ### Replicate **Models:** Various open-source models including Minimax, Stable Video Diffusion ```go prov := replicate.New(replicate.Config{ APIKey: os.Getenv("REPLICATE_API_TOKEN"), }) model := replicate.NewVideoModel(prov, "minimax/video-01") ``` **Features:** - Wide selection of open-source models - Good for experimentation - Flexible pricing - Model-specific parameters **API Token:** Get your token at [replicate.com](https://replicate.com) ### Google Vertex AI ```go prov, err := googlevertex.New(googlevertex.Config{ Project: os.Getenv("GOOGLE_VERTEX_PROJECT"), Location: "us-central1", }) if err != nil { log.Fatal(err) } model, _ := prov.VideoModel(googlevertex.ModelVeo30Generate001) ``` **Features:** - Veo video generation on Vertex AI - Batch generation (up to 4 videos per call, `MaxVideosPerCall`) - Text-to-video and image-to-video - Polling-based async generation, same options as Google Generative AI ## Async Video: DoStart / DoStatus xAI, fal, Google, Google Vertex, and Replicate implement the optional `provider.VideoModelStarter` / `VideoModelStatusChecker` interfaces so you can start a job and poll (or receive a webhook) instead of blocking on `GenerateVideo` for the whole generation: ```go started, err := ai.ExperimentalStartVideo(ctx, ai.StartVideoOptions{ Model: model, Prompt: ai.VideoPrompt{Text: "A cat playing piano in a jazz club"}, }) if err != nil { log.Fatal(err) } // Persist started.Operation (json.RawMessage) and poll later, from any // process: status, err := ai.ExperimentalGetVideoStatus(ctx, model, ai.GetVideoStatusOptions{ Operation: started.Operation, }) ``` fal and Replicate also support webhook-based completion via `HandleWebhookOption` instead of polling. Durable webhook suspension inside a workflow (TS's `experimental_generateVideo` Vercel Workflow DevKit integration) is not ported — see [Known Differences](https://goaisdk.com/docs/migration-guides/known-differences.md). ## API Reference ### GenerateVideo Options ```go type GenerateVideoOptions struct { // Model to use for video generation (required) Model provider.VideoModelV3 // Prompt for video generation (required) Prompt VideoPrompt // Number of videos to generate (default: 1) N int // Maximum videos per API call. If unset, uses the model default. MaxVideosPerCall *int // Aspect ratio in format "width:height" (e.g., "16:9", "9:16", "1:1") AspectRatio string // Resolution in format "widthxheight" (e.g., "1920x1080", "1280x720") Resolution string // Duration in seconds Duration *float64 // Frames per second (24, 30, 60) FPS *int // Seed for reproducible generation Seed *int // Provider-specific options ProviderOptions map[string]interface{} // Maximum retries per call (default: 2). Set to 0 to disable retries. MaxRetries *int // Additional HTTP headers Headers map[string]string // Custom bytes-only download function for URL video outputs Download URLDownloadFunction // Custom download function for URL video outputs that also returns media type DownloadWithMetadata URLDownloadWithMetadataFunction } ``` ### VideoPrompt ```go type VideoPrompt struct { // Text prompt (required for text-to-video, optional for image-to-video) Text string // Image for image-to-video generation (optional) Image *VideoPromptImage } type VideoPromptImage struct { // URL to image (or) URL string // Base64, base64url, or data URL image string DataString string // Raw image data Data []byte // Media type (e.g., "image/png", "image/jpeg") MediaType string } ``` ### GenerateVideoResult ```go type GenerateVideoResult struct { // Primary generated video (first in Videos array) Video *types.GeneratedFile // All generated videos Videos []*types.GeneratedFile // Warnings from the generation process Warnings []types.Warning // Metadata for each API call Responses []VideoModelResponseMetadata // Provider-specific metadata ProviderMetadata map[string]interface{} } type VideoModelResponseMetadata struct { Timestamp time.Time `json:"timestamp"` ModelID string `json:"modelId"` Headers map[string]string `json:"headers,omitempty"` ProviderMetadata map[string]interface{} `json:"providerMetadata,omitempty"` } ``` `types.GeneratedFile` contains the generated video bytes in `Data`, the detected `MediaType`, and TypeScript-compatible helper methods: ```go base64Data := result.Video.Base64() videoBytes := result.Video.Uint8Array() ``` ## Provider-Specific Options ### XAI ```go ProviderOptions: map[string]interface{}{ "xai": map[string]interface{}{ "pollIntervalMs": 5000, // Poll every 5 seconds "pollTimeoutMs": 600000, // 10 minute timeout "resolution": "720p", // or "480p" }, } ``` ### Google ```go ProviderOptions: map[string]interface{}{ "google": map[string]interface{}{ "pollIntervalMs": 3000, // Poll every 3 seconds "pollTimeoutMs": 600000, // 10 minute timeout "negativePrompt": "blurry, distorted", }, } ``` ### FAL ```go ProviderOptions: map[string]interface{}{ "fal": map[string]interface{}{ "loop": true, // Create looping video "motionStrength": 0.8, "negativePrompt": "static, still image", }, } ``` ### Replicate ```go ProviderOptions: map[string]interface{}{ "replicate": map[string]interface{}{ "guidance_scale": 7.5, "num_inference_steps": 50, }, } ``` ## Polling and Async Generation All video providers use async generation with polling: 1. **Submit Request**: Initial API call submits the generation job 2. **Polling**: SDK automatically polls for completion 3. **Download**: URL video outputs are downloaded into `types.GeneratedFile.Data` 4. **Return**: Complete video data returned to your code Default polling configuration (fal, Google, Google Vertex, Replicate; matches the TypeScript AI SDK's `generateVideo` defaults): - **Interval**: 5 seconds - **Timeout**: 10 minutes - **Backoff**: None (configurable per provider) You can customize polling: ```go result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: ai.VideoPrompt{Text: "A mountain landscape"}, ProviderOptions: map[string]interface{}{ "google": map[string]interface{}{ "pollIntervalMs": 5000, // Poll every 5 seconds "pollTimeoutMs": 600000, // 10 minute timeout }, }, }) ``` ## Error Handling ```go result, err := ai.GenerateVideo(ctx, opts) if err != nil { if noVideo, ok := err.(*ai.NoVideoGeneratedError); ok { fmt.Printf("No video generated; responses: %d\n", len(noVideo.Responses)) } if videoErr, ok := err.(*providererrors.VideoGenerationError); ok { fmt.Printf("Provider: %s\n", videoErr.Provider) fmt.Printf("Model: %s\n", videoErr.ModelID) fmt.Printf("Message: %s\n", videoErr.Message) } return err } // Check for warnings for _, warning := range result.Warnings { fmt.Printf("Warning (%s): %s\n", warning.Type, warning.Message) } ``` Common errors: - **Timeout**: Generation took longer than configured timeout - **Invalid Prompt**: Prompt violates content policy - **Quota Exceeded**: API quota or rate limit reached - **Network Error**: Connection issues during polling ## Best Practices ### 1. Choose the Right Provider - **XAI**: Text-to-video, image animation, video editing support - **Google**: Production use, high quality - **FAL**: Fast iteration, prototyping - **Replicate**: Experimentation, specific models ### 2. Optimize Prompts - Be specific about motion and camera movement - Include lighting and atmosphere details - Keep prompts under 200 words - Test variations to find what works best ### 3. Handle Long Generation Times ```go ctx, cancel := context.WithTimeout(context.Background(), 10*time.Minute) defer cancel() result, err := ai.GenerateVideo(ctx, opts) ``` ### 4. Save Videos Efficiently ```go if result.Video != nil { filename := fmt.Sprintf("video_%d.mp4", time.Now().Unix()) if err := os.WriteFile(filename, result.Video.Data, 0644); err != nil { return fmt.Errorf("failed to save video: %w", err) } } ``` ### 5. Monitor Costs Video generation can be expensive: - Google: ~$0.05-0.15 per video (depending on resolution/duration) - FAL: ~$0.03-0.08 per video - Replicate: Varies by model Track usage and set appropriate timeouts. ## Troubleshooting ### Videos Not Generating 1. Check API keys are set correctly 2. Verify model ID is correct 3. Check prompt doesn't violate content policy 4. Ensure sufficient API quota/credits ### Slow Generation 1. Increase timeout if hitting limits 2. Try shorter videos (3-5 seconds) 3. Use lower resolution (720p instead of 1080p) 4. Check provider status pages ### Poor Quality Results 1. Improve prompt specificity 2. Try different aspect ratios 3. Adjust resolution settings 4. Use provider-specific quality options ## Examples See complete examples in the [examples/video](https://github.com/digitallysavvy/go-ai/tree/main/examples/video) directory: - [Text-to-Video](https://github.com/digitallysavvy/go-ai/tree/main/examples/video/text-to-video) - [Image-to-Video](https://github.com/digitallysavvy/go-ai/tree/main/examples/video/image-to-video) - [Multi-Provider](https://github.com/digitallysavvy/go-ai/tree/main/examples/video/multi-provider) ## See Also - [Image Generation](https://goaisdk.com/docs/ai-sdk-core/image-generation.md) - [Provider Management](https://goaisdk.com/docs/ai-sdk-core/provider-management.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) --- # Realtime Sessions > Documents provider-neutral realtime session helpers in the Go AI SDK, exposing provider.Experimental_RealtimeModelV4 through ai.ConnectRealtime. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/realtime Documentation index: https://goaisdk.com/llms.txt The Go SDK exposes the TypeScript AI SDK experimental realtime model surface through `provider.Experimental_RealtimeModelV4` and the provider-neutral helper `ai.ConnectRealtime`. ```go model, err := openaiProvider.RealtimeModel(openai.ModelGPT4oRealtimePreview) if err != nil { log.Fatal(err) } instructions := "Answer briefly." session, err := ai.ConnectRealtime(ctx, model, ai.RealtimeSessionOptions{ SessionConfig: &provider.RealtimeSessionConfig{ Instructions: &instructions, }, }) if err != nil { log.Fatal(err) } defer session.Close() ``` The helper creates a provider client secret, opens the provider WebSocket, sends an initial session update when `SessionConfig` is supplied, and normalizes provider server events. Unknown provider events are preserved on `RealtimeServerEvent.Raw` with their raw provider event type so callers can handle new provider events before the SDK adds first-class fields. ## Providers - OpenAI realtime: `openai.Provider.RealtimeModel`, `ExperimentalRealtimeModel`, and `GetRealtimeToken`. - Google Gemini Live: `google.Provider.RealtimeModel`, `ExperimentalRealtimeModel`, and `GetRealtimeToken`. - xAI realtime: `xai.Provider.RealtimeModel`, `ExperimentalRealtimeModel`, and `GetRealtimeToken`. Gateway realtime is not exposed because the June 6 TypeScript range did not include a Gateway realtime model, factory, tests, or changeset. ## Not Applicable Browser Pieces The TypeScript SDK also ships browser-only realtime transports and React helpers. They are intentionally not implemented in Go: - `BrowserRealtimeTransport` - `BrowserRealtimeAudio` - `useRealtime` Go callers provide a server-side WebSocket dialer through `RealtimeSessionOptions.Dialer` when they need custom transport behavior. --- # Language Model Middleware > Learn how to use middleware to enhance the behavior of language models Canonical URL: https://goaisdk.com/docs/ai-sdk-core/middleware Documentation index: https://goaisdk.com/llms.txt Language model middleware is a way to enhance the behavior of language models by intercepting and modifying the calls to the language model. It can be used to add features like guardrails, RAG, caching, and logging in a language model agnostic way. Such middleware can be developed and distributed independently from the language models that they are applied to. ## Using Language Model Middleware You can use language model middleware with the `middleware.WrapLanguageModel` function. It takes a language model and one or more middleware instances and returns a new language model that incorporates the middleware. For parity-oriented API discovery, equivalent wrappers are also exposed from `pkg/ai` (for example `ai.WrapLanguageModel` and `ai.DefaultSettingsMiddleware`). ```go import ( "os" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) baseModel, _ := provider.LanguageModel("gpt-6-astra") wrappedModel := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{yourMiddleware}, nil, // optional modelID override nil, // optional providerID override ) ``` The wrapped language model can be used just like any other language model, e.g. in `ai.GenerateText` or `ai.StreamText`: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: wrappedModel, Prompt: "What cities are in the United States?", }) ``` ## Multiple Middlewares You can provide multiple middlewares to the `middleware.WrapLanguageModel` function. The middlewares will be applied in the order they are provided. ```go wrappedModel := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{ firstMiddleware, secondMiddleware, }, nil, nil, ) // Applied as: firstMiddleware(secondMiddleware(baseModel)) ``` ## Built-in Middleware The Go AI SDK comes with several built-in middlewares: ### Default Settings Middleware The `DefaultSettingsMiddleware` applies default settings to a language model. Settings provided in individual calls will override these defaults. ```go import ( "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" ) temperature := 0.5 maxTokens := 800 model := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &temperature, MaxTokens: &maxTokens, }), }, nil, nil, ) // All calls with this model will use temperature=0.5 and maxTokens=800 by default result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain photosynthesis", }) ``` ## Implementing Custom Language Model Middleware > **Note:** Implementing language model middleware is advanced functionality and requires a solid understanding of the language model specification in the provider package. You can implement any of the following functions to modify the behavior of the language model: 1. **TransformParams**: Transforms the parameters before they are passed to the language model, for both `DoGenerate` and `DoStream`. 2. **WrapGenerate**: Wraps the `DoGenerate` method of the language model. You can modify the parameters, call the language model, and modify the result. 3. **WrapStream**: Wraps the `DoStream` method of the language model. You can modify the parameters, call the language model, and modify the result. ## Middleware Structure ```go type LanguageModelMiddleware struct { // SpecificationVersion should be "v3" for the current version SpecificationVersion string // OverrideProvider allows overriding the provider name OverrideProvider func(model provider.LanguageModel) string // OverrideModelID allows overriding the model ID OverrideModelID func(model provider.LanguageModel) string // TransformParams transforms parameters before passing to the model TransformParams func( ctx context.Context, callType string, params *provider.GenerateOptions, model provider.LanguageModel, ) (*provider.GenerateOptions, error) // WrapGenerate wraps the generate operation WrapGenerate func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) // WrapStream wraps the stream operation WrapStream func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (provider.TextStream, error) } ``` ## Examples > **Note:** These examples are not meant to be used in production as-is. They demonstrate how to use middleware to enhance language model behavior. ### Logging Middleware This example shows how to log the parameters and generated text of a language model call. ```go package main import ( "context" "encoding/json" "io" "log" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func LogMiddleware() *middleware.LanguageModelMiddleware { return &middleware.LanguageModelMiddleware{ SpecificationVersion: "v3", WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { log.Println("DoGenerate called") paramsJSON, _ := json.MarshalIndent(params, "", " ") log.Printf("Params: %s\n", paramsJSON) result, err := doGenerate() if err != nil { log.Printf("DoGenerate failed: %v\n", err) return nil, err } log.Println("DoGenerate finished") log.Printf("Generated text: %s\n", result.Text) return result, nil }, WrapStream: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (provider.TextStream, error) { log.Println("DoStream called") paramsJSON, _ := json.MarshalIndent(params, "", " ") log.Printf("Params: %s\n", paramsJSON) stream, err := doStream() if err != nil { log.Printf("DoStream failed: %v\n", err) return nil, err } return &loggingStream{base: stream}, nil }, } } // loggingStream implements provider.TextStream type loggingStream struct { base provider.TextStream fullText string } func (s *loggingStream) Next() (*provider.StreamChunk, error) { chunk, err := s.base.Next() if err == io.EOF { log.Println("DoStream finished") log.Printf("Generated text: %s\n", s.fullText) return nil, io.EOF } if err != nil { return nil, err } if chunk.Type == provider.ChunkTypeText { s.fullText += chunk.Text } return chunk, nil } func (s *loggingStream) Err() error { return s.base.Err() } func (s *loggingStream) Close() error { return s.base.Close() } ``` ### Caching Middleware This example shows how to build a simple cache for the generated text of a language model call. ```go package main import ( "context" "crypto/sha256" "encoding/json" "fmt" "sync" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func CacheMiddleware() *middleware.LanguageModelMiddleware { cache := &sync.Map{} return &middleware.LanguageModelMiddleware{ SpecificationVersion: "v3", WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { // Generate cache key from parameters cacheKey := generateCacheKey(params) // Check cache if cached, ok := cache.Load(cacheKey); ok { fmt.Println("Cache hit!") return cached.(*types.GenerateResult), nil } // Call model result, err := doGenerate() if err != nil { return nil, err } // Store in cache cache.Store(cacheKey, result) return result, nil }, // Note: Implementing caching for streaming is more complex // as you need to collect the full stream before caching } } func generateCacheKey(params *provider.GenerateOptions) string { data, _ := json.Marshal(params) hash := sha256.Sum256(data) return fmt.Sprintf("%x", hash) } ``` ### Retrieval Augmented Generation (RAG) Middleware This example shows how to use RAG as middleware. ```go package main import ( "context" "fmt" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func RAGMiddleware(vectorDB VectorDatabase) *middleware.LanguageModelMiddleware { return &middleware.LanguageModelMiddleware{ SpecificationVersion: "v3", TransformParams: func( ctx context.Context, callType string, params *provider.GenerateOptions, model provider.LanguageModel, ) (*provider.GenerateOptions, error) { // Extract the last user message lastUserMessage := getLastUserMessage(params) if lastUserMessage == "" { return params, nil // No modification needed } // Find relevant sources from vector database sources, err := vectorDB.FindSimilar(ctx, lastUserMessage, 5) if err != nil { return nil, fmt.Errorf("failed to find sources: %w", err) } // Build instruction with sources instruction := "Use the following information to answer the question:\n" for _, source := range sources { instruction += fmt.Sprintf("\n%s", source.Content) } // Add instruction to the last user message modifiedParams := addToLastUserMessage(params, instruction) return modifiedParams, nil }, } } // Helper functions (implementation depends on your use case) func getLastUserMessage(params *provider.GenerateOptions) string { if len(params.Prompt.Messages) == 0 { return "" } lastMsg := params.Prompt.Messages[len(params.Prompt.Messages)-1] if lastMsg.Role != types.RoleUser { return "" } for _, part := range lastMsg.Content { if textPart, ok := part.(types.TextContent); ok { return textPart.Text } } return "" } func addToLastUserMessage(params *provider.GenerateOptions, text string) *provider.GenerateOptions { modified := *params if len(modified.Prompt.Messages) == 0 { return params } messages := make([]types.Message, len(modified.Prompt.Messages)) copy(messages, modified.Prompt.Messages) lastIdx := len(messages) - 1 lastMsg := messages[lastIdx] // Prepend text to the last user message newContent := []types.ContentPart{types.TextContent{Text: text + "\n\n"}} newContent = append(newContent, lastMsg.Content...) messages[lastIdx] = types.Message{ Role: lastMsg.Role, Content: newContent, } modified.Prompt.Messages = messages return &modified } // VectorDatabase interface (example) type VectorDatabase interface { FindSimilar(ctx context.Context, query string, limit int) ([]Document, error) } type Document struct { Content string Score float64 } ``` ### Guardrails Middleware Guardrails ensure that the generated text of a language model call is safe and appropriate. ```go package main import ( "context" "regexp" "strings" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func GuardrailMiddleware() *middleware.LanguageModelMiddleware { // Define patterns to filter sensitivePatterns := []*regexp.Regexp{ regexp.MustCompile(`\d{3}-\d{2}-\d{4}`), // SSN pattern regexp.MustCompile(`\d{16}`), // Credit card pattern regexp.MustCompile(`\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b`), // Email } return &middleware.LanguageModelMiddleware{ SpecificationVersion: "v3", WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { result, err := doGenerate() if err != nil { return nil, err } // Filter sensitive information cleanedText := result.Text for _, pattern := range sensitivePatterns { cleanedText = pattern.ReplaceAllString(cleanedText, "") } // Replace inappropriate words cleanedText = strings.ReplaceAll(cleanedText, "badword", "") result.Text = cleanedText return result, nil }, // Note: Streaming guardrails are more complex because you don't know // the full content until the stream is finished } } ``` ### Retry Middleware This middleware automatically retries failed requests with exponential backoff. ```go package main import ( "context" "fmt" "time" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func RetryMiddleware(maxRetries int) *middleware.LanguageModelMiddleware { return &middleware.LanguageModelMiddleware{ SpecificationVersion: "v3", WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { var lastErr error for attempt := 0; attempt <= maxRetries; attempt++ { if attempt > 0 { // Exponential backoff backoff := time.Duration(1< # Provider & Model Management > Explains centrally managing multiple providers and models in Go using the provider registry, custom providers with middleware, and registry utilities. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/provider-management Documentation index: https://goaisdk.com/llms.txt When you work with multiple providers and models, it is often desirable to manage them in a central place and access the models through simple string IDs. The Go AI SDK offers a **provider registry** and **middleware-based custom providers** for this purpose: - With **middleware**, you can pre-configure model settings, provide model name aliases, and enhance model behavior. - The **provider registry** lets you mix multiple providers and access them through simple string IDs. You can mix and match middleware, the provider registry, and custom provider configurations in your application. ## Provider Registry The provider registry allows you to register multiple providers and access their models through simple string identifiers in the format `provider:model`. ### Basic Setup ```go package main import ( "os" "github.com/digitallysavvy/go-ai/pkg/registry" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func init() { // Register providers registry.RegisterProvider("openai", openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), })) registry.RegisterProvider("anthropic", anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), })) } ``` ### Using the Registry #### Language Models Access language models using the `provider:model` format: ```go import ( "context" "github.com/digitallysavvy/go-ai/pkg/registry" "github.com/digitallysavvy/go-ai/pkg/ai" ) // Resolve model from registry model, err := registry.ResolveLanguageModel("openai:gpt-6-astra") if err != nil { log.Fatal(err) } // Use with AI SDK functions result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing", }) ``` #### Embedding Models ```go embeddingModel, err := registry.ResolveEmbeddingModel("openai:text-embedding-3-small") if err != nil { log.Fatal(err) } result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "sunny day at the beach", }) ``` ### Model Aliases Create short aliases for commonly used models: ```go func init() { // Register providers registry.RegisterProvider("openai", openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), })) registry.RegisterProvider("anthropic", anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), })) // Register aliases for easy access registry.RegisterAlias("opus", "anthropic:claude-opus-4-5") registry.RegisterAlias("sonnet", "anthropic:claude-sonnet-4-5") registry.RegisterAlias("haiku", "anthropic:claude-haiku-4-5") registry.RegisterAlias("gpt-6-astra", "openai:gpt-6-astra") registry.RegisterAlias("gpt-6-sol", "openai:gpt-6-sol") } // Use aliases model, err := registry.ResolveLanguageModel("opus") ``` ### Custom Registry Instances You can create custom registry instances instead of using the global registry: ```go import "github.com/digitallysavvy/go-ai/pkg/registry" // Create custom registry myRegistry := registry.NewRegistry() // Register providers myRegistry.RegisterProvider("openai", openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), })) // Use custom registry model, err := myRegistry.ResolveLanguageModel("openai:gpt-6-astra") ``` ### Registry Options `NewRegistry` accepts functional options: ```go myRegistry := registry.NewRegistry( // Use a different provider:model separator (default ":"). registry.WithSeparator("/"), // Apply middleware to every language model resolved through the registry. registry.WithLanguageModelMiddleware(loggingMiddleware), // Apply middleware to every image model resolved through the registry. registry.WithImageModelMiddleware(imageLoggingMiddleware), ) model, err := myRegistry.ResolveLanguageModel("openai/gpt-6-astra") // custom separator ``` ### NoSuchModelError A model ID that has no `provider:model` separator (or names an unregistered provider or model) returns a typed error from `pkg/provider/errors`: ```go import providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" model, err := registry.ResolveLanguageModel("gpt-6-astra") // missing "provider:" prefix if err != nil { var noSuchModel *providererrors.NoSuchModelError if errors.As(err, &noSuchModel) { log.Printf("no such model: %s (%s)", noSuchModel.ModelID, noSuchModel.ModelType) } } ``` ## Custom Providers with Middleware You can create custom provider configurations using middleware to pre-configure settings, create aliases, or limit available models. ### Example: Pre-configured Model Settings ```go import ( "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/registry" ) func setupCustomProviders() { openaiProvider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) // Create a base model baseModel, _ := openaiProvider.LanguageModel("gpt-6-astra") // Wrap with default settings temperature := 0.7 maxTokens := 2000 customModel := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &temperature, MaxTokens: &maxTokens, }), }, nil, nil, ) // Register the custom model // Note: Since we can't register a single model, we create a wrapper provider // (implementation shown below) } ``` ### Example: Model Name Aliases with Different Settings Create a configuration package that exposes different model configurations: ```go package models import ( "os" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) var ( // Fast model with lower temperature for quick, focused responses Fast provider.LanguageModel // Creative model with high temperature for diverse outputs Creative provider.LanguageModel // Precise model with low temperature for deterministic responses Precise provider.LanguageModel ) func init() { anthropicProvider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) // Fast: Claude Haiku with optimized settings haiku, _ := anthropicProvider.LanguageModel("claude-haiku-4-5") temp := 0.5 Fast = middleware.WrapLanguageModel( haiku, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &temp, }), }, nil, nil, ) // Creative: Claude Sonnet with high temperature sonnet, _ := anthropicProvider.LanguageModel("claude-sonnet-4-5") creativeTemp := 0.9 Creative = middleware.WrapLanguageModel( sonnet, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &creativeTemp, }), }, nil, nil, ) // Precise: Claude Opus with low temperature opus, _ := anthropicProvider.LanguageModel("claude-opus-4-5") preciseTemp := 0.1 Precise = middleware.WrapLanguageModel( opus, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &preciseTemp, }), }, nil, nil, ) } ``` Usage: ```go skip-compile import "myapp/models" // Use pre-configured models result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: models.Creative, Prompt: "Write a creative story", }) ``` ### Example: Limited Model Set Create a constrained provider that only exposes specific models: ```go package providers import ( "fmt" "os" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) type LimitedProvider struct { models map[string]provider.LanguageModel } func NewLimitedProvider() *LimitedProvider { openaiProvider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) anthropicProvider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) // Only expose specific models models := make(map[string]provider.LanguageModel) textMedium, _ := anthropicProvider.LanguageModel("claude-sonnet-4-5") textSmall, _ := openaiProvider.LanguageModel("gpt-6-sol") models["text-medium"] = textMedium models["text-small"] = textSmall return &LimitedProvider{models: models} } func (p *LimitedProvider) LanguageModel(modelID string) (provider.LanguageModel, error) { model, ok := p.models[modelID] if !ok { return nil, fmt.Errorf("model not available: %s (available: text-medium, text-small)", modelID) } return model, nil } func (p *LimitedProvider) EmbeddingModel(modelID string) (provider.EmbeddingModel, error) { return nil, fmt.Errorf("embedding models not available") } func (p *LimitedProvider) ImageModel(modelID string) (provider.ImageModel, error) { return nil, fmt.Errorf("image models not available") } func (p *LimitedProvider) SpeechModel(modelID string) (provider.SpeechModel, error) { return nil, fmt.Errorf("speech models not available") } func (p *LimitedProvider) TranscriptionModel(modelID string) (provider.TranscriptionModel, error) { return nil, fmt.Errorf("transcription models not available") } func (p *LimitedProvider) RerankingModel(modelID string) (provider.RerankingModel, error) { return nil, fmt.Errorf("reranking models not available") } ``` Usage: ```go limited := providers.NewLimitedProvider() registry.RegisterProvider("limited", limited) // Only "text-medium" and "text-small" are available model, _ := registry.ResolveLanguageModel("limited:text-medium") ``` ## Centralized Provider Configuration Create a central configuration file that sets up all providers and models: ```go package config import ( "os" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/providers/google" "github.com/digitallysavvy/go-ai/pkg/registry" ) func init() { setupProviders() setupAliases() setupCustomModels() } func setupProviders() { // Register standard providers registry.RegisterProvider("openai", openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), })) registry.RegisterProvider("anthropic", anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), })) registry.RegisterProvider("google", google.New(google.Config{ APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"), })) } func setupAliases() { // Short aliases for common models registry.RegisterAlias("fast", "anthropic:claude-haiku-4-5") registry.RegisterAlias("balanced", "anthropic:claude-sonnet-4-5") registry.RegisterAlias("powerful", "anthropic:claude-opus-4-5") registry.RegisterAlias("gpt-6-astra", "openai:gpt-6-astra") registry.RegisterAlias("gpt-6-sol", "openai:gpt-6-sol") registry.RegisterAlias("gemini", "google:gemini-2.0-flash") } func setupCustomModels() { // Set up models with pre-configured settings // (Implementation depends on your needs) } ``` ## Combining Registry, Middleware, and Custom Configuration Here's a comprehensive example combining all concepts: ```go package main import ( "context" "fmt" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/registry" ) func init() { // 1. Register base providers registry.RegisterProvider("openai", openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), })) registry.RegisterProvider("anthropic", anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), })) // 2. Set up aliases for common models registry.RegisterAlias("default", "anthropic:claude-sonnet-4-5") registry.RegisterAlias("fast", "anthropic:claude-haiku-4-5") registry.RegisterAlias("smart", "anthropic:claude-opus-4-5") // 3. Set up wrapped models with custom settings // (Shown in CustomModels below) } // CustomModels provides pre-configured models with specific settings type CustomModels struct { Creative provider.LanguageModel // High temperature for creative tasks Factual provider.LanguageModel // Low temperature for factual responses Concise provider.LanguageModel // Limited tokens for brief responses } func GetCustomModels() *CustomModels { anthropicProvider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) baseModel, _ := anthropicProvider.LanguageModel("claude-sonnet-4-5") // Creative model creativeTemp := 0.9 creative := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &creativeTemp, }), }, nil, nil, ) // Factual model factualTemp := 0.1 factual := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &factualTemp, }), }, nil, nil, ) // Concise model conciseTemp := 0.5 conciseTokens := 100 concise := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &conciseTemp, MaxTokens: &conciseTokens, }), }, nil, nil, ) return &CustomModels{ Creative: creative, Factual: factual, Concise: concise, } } func main() { ctx := context.Background() customModels := GetCustomModels() // Use registry with aliases model1, _ := registry.ResolveLanguageModel("default") result1, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model1, Prompt: "Explain quantum computing", }) fmt.Println("Default:", result1.Text) // Use custom pre-configured models result2, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: customModels.Creative, Prompt: "Write a creative story", }) fmt.Println("Creative:", result2.Text) // Use direct provider:model syntax model3, _ := registry.ResolveLanguageModel("openai:gpt-6-astra") result3, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model3, Prompt: "Explain photosynthesis", }) fmt.Println("GPT:", result3.Text) } ``` ## Registry Utility Functions ### List Registered Providers `ListProviders` and `ListAliases` are methods on a `*registry.Registry`, not package-level functions. Use `registry.GetGlobalRegistry()` to reach the global registry that `registry.RegisterProvider` et al. populate: ```go providers := registry.GetGlobalRegistry().ListProviders() fmt.Println("Available providers:", providers) // Output: Available providers: [openai anthropic google] ``` ### List Registered Aliases ```go aliases := registry.GetGlobalRegistry().ListAliases() for alias, target := range aliases { fmt.Printf("%s -> %s\n", alias, target) } // Output: // default -> anthropic:claude-sonnet-4-5 // fast -> anthropic:claude-haiku-4-5 // smart -> anthropic:claude-opus-4-5 ``` ### Get Provider Directly ```go provider, err := registry.GetProvider("openai") if err != nil { log.Fatal(err) } model, _ := provider.LanguageModel("gpt-6-astra") ``` ## Best Practices ### 1. Centralize Configuration Keep all provider setup in a single configuration package: ``` myapp/ ├── config/ │ └── providers.go # All provider setup here ├── models/ │ └── custom.go # Custom model configurations └── main.go ``` ### 2. Use Environment Variables Store API keys and configuration in environment variables: ```go openaiProvider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), BaseURL: os.Getenv("OPENAI_BASE_URL"), // Optional custom endpoint }) ``` ### 3. Create Semantic Aliases Use meaningful names that reflect the model's purpose: ```go // Good: Semantic aliases registry.RegisterAlias("writing-assistant", "anthropic:claude-sonnet-4-5") registry.RegisterAlias("code-reviewer", "openai:gpt-6-astra") registry.RegisterAlias("quick-qa", "anthropic:claude-haiku-4-5") // Bad: Technical aliases registry.RegisterAlias("model1", "anthropic:claude-sonnet-4-5") registry.RegisterAlias("fast-one", "anthropic:claude-haiku-4-5") ``` ### 4. Document Available Models Create documentation for your team about available models: ```go // models/README.md // // Available Models: // // default - Claude Sonnet 4.5 (balanced performance) // fast - Claude Haiku 4.5 (quick responses) // smart - Claude Opus 4.5 (complex reasoning) // writing - larger model, tuned for writing // code - larger model, tuned for code generation ``` ### 5. Handle Errors Gracefully Always check for errors when resolving models: ```go model, err := registry.ResolveLanguageModel("nonexistent:model") if err != nil { // Log error and fall back to default log.Printf("Failed to resolve model: %v, using default", err) model, _ = registry.ResolveLanguageModel("default") } ``` ### 6. Use Type-Safe Model References For critical applications, create constants for model references: ```go package models const ( DefaultModel = "anthropic:claude-sonnet-4-5" FastModel = "anthropic:claude-haiku-4-5" PowerfulModel = "anthropic:claude-opus-4-5" CodeModel = "openai:gpt-6-astra" ) ``` Then resolve a model by its constant: ```go model, _ := registry.ResolveLanguageModel(models.DefaultModel) ``` ## See Also - [Middleware](https://goaisdk.com/docs/ai-sdk-core/middleware.md) - [Settings](https://goaisdk.com/docs/ai-sdk-core/settings.md) - [Providers](https://goaisdk.com/docs/providers/overview.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) --- # Error Handling > Covers Go AI SDK error handling: regular and streaming errors, context cancellation, specific error types, and a comprehensive error-handling example. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/error-handling Documentation index: https://goaisdk.com/llms.txt The Go AI SDK provides robust error handling with specific error types for different failure scenarios. All errors follow Go's standard error handling patterns. ## Handling Regular Errors Regular errors are returned from functions and should be checked using Go's idiomatic error handling: ```go import ( "context" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a vegetarian lasagna recipe for 4 people.", }) if err != nil { // Handle error log.Printf("Error generating text: %v", err) return } fmt.Println(result.Text) } ``` See [Error Types](#error-types) below for more information on the different types of errors that may be returned. ## Handling Streaming Errors When errors occur during streaming, you should check for errors both while reading from the stream channel and after the stream completes. ### Simple Text Streaming ```go import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" ) func main() { ctx := context.Background() stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a vegetarian lasagna recipe for 4 people.", }) if err != nil { log.Fatal(err) } // Read from stream for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } // Check for errors after stream completes if err := stream.Err(); err != nil { log.Printf("Stream error: %v", err) return } } ``` ### Full Stream with Error Chunks Full streams support error chunks within the stream itself. You should handle both chunk errors and post-stream errors: ```go import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" ) func main() { ctx := context.Background() stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a vegetarian lasagna recipe for 4 people.", }) if err != nil { log.Fatal(err) } // Read from full stream for chunk := range stream.Chunks() { switch chunk.Type { case provider.ChunkTypeText: fmt.Print(chunk.Text) case provider.ChunkTypeToolCall: fmt.Printf("Tool call: %s\n", chunk.ToolCall.ToolName) case provider.ChunkTypeError: // Handle error chunk log.Printf("Error in stream: %v", chunk.Err) case provider.ChunkTypeFinish: fmt.Printf("\nFinished: %s\n", chunk.FinishReason) } } // Check for errors after stream completes if err := stream.Err(); err != nil { log.Printf("Stream error: %v", err) return } } ``` ## Handling Context Cancellation Go uses `context.Context` for cancellation and timeouts. When a context is canceled, operations return a context error: ```go import ( "context" "fmt" "time" ) func main() { // Create context with timeout ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a very long story...", }) if err != nil { log.Fatal(err) } for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } // Check if operation was canceled or timed out if err := stream.Err(); err != nil { if ctx.Err() == context.DeadlineExceeded { fmt.Println("\nOperation timed out") } else if ctx.Err() == context.Canceled { fmt.Println("\nOperation was canceled") } else { fmt.Printf("\nStream error: %v\n", err) } } } ``` ### OnFinish Callback with Context The `OnFinish` callback is only called when the operation completes normally. It is **not** called when the context is canceled: ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a story...", OnFinish: func(result *ai.StreamTextResult) { // Called only on normal completion, NOT on cancellation fmt.Printf("Completed after %d steps\n", len(result.Steps())) fmt.Printf("Total tokens used: %d\n", result.Usage().GetTotalTokens()) }, }) // Handle context cancellation separately for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } if ctx.Err() == context.Canceled { fmt.Println("Stream was canceled") // Perform cleanup operations here } ``` ## Error Types The Go AI SDK provides several specific error types for different failure scenarios. Use `errors.Is()` and `errors.As()` to check for specific error types. ### Provider Error Represents an error from an AI provider (e.g., invalid API key, model not found, rate limiting): ```go import ( "errors" "fmt" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" ) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", }) if err != nil { var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { fmt.Printf("Provider: %s\n", providerErr.Provider) fmt.Printf("Status Code: %d\n", providerErr.StatusCode) fmt.Printf("Error Code: %s\n", providerErr.ErrorCode) fmt.Printf("Message: %s\n", providerErr.Message) return } } ``` ### Rate Limit Error Represents a rate limit error from a provider: ```go import ( "errors" "time" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" ) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", }) if err != nil { var rateLimitErr *providererrors.RateLimitError if errors.As(err, &rateLimitErr) { fmt.Printf("Rate limit exceeded for %s\n", rateLimitErr.Provider) if rateLimitErr.RetryAfterSeconds != nil { fmt.Printf("Retry after %d seconds\n", *rateLimitErr.RetryAfterSeconds) time.Sleep(time.Duration(*rateLimitErr.RetryAfterSeconds) * time.Second) // Retry the request... } return } } ``` ### Validation Error Represents a validation error (e.g., invalid parameters, schema validation failure): ```go import ( "errors" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" ) result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: invalidSchema, Prompt: "Generate object", }) if err != nil { var validationErr *providererrors.ValidationError if errors.As(err, &validationErr) { fmt.Printf("Validation failed on field: %s\n", validationErr.Field) fmt.Printf("Error: %s\n", validationErr.Message) return } } ``` ### Tool Execution Error Represents an error during tool execution: ```go import ( "errors" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" ) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Use the weather tool", Tools: []types.Tool{weatherTool}, }) if err != nil { var toolErr *providererrors.ToolExecutionError if errors.As(err, &toolErr) { fmt.Printf("Tool %s failed (call ID: %s)\n", toolErr.ToolName, toolErr.ToolCallID) fmt.Printf("Error: %s\n", toolErr.Message) return } } ``` ### Stream Error Represents an error during streaming: ```go import ( "errors" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" ) stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate text", }) if err != nil { log.Fatal(err) } for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } if err := stream.Err(); err != nil { var streamErr *providererrors.StreamError if errors.As(err, &streamErr) { fmt.Printf("Stream error: %s\n", streamErr.Message) return } } ``` ### Standard Errors The SDK also provides standard error variables for common cases: ```go import ( "errors" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" ) // Check for specific standard errors if errors.Is(err, providererrors.ErrInvalidInput) { fmt.Println("Invalid input parameters") } if errors.Is(err, providererrors.ErrModelNotFound) { fmt.Println("Model not found") } if errors.Is(err, providererrors.ErrProviderNotFound) { fmt.Println("Provider not found") } if errors.Is(err, providererrors.ErrToolNotFound) { fmt.Println("Tool not found") } if errors.Is(err, providererrors.ErrValidationFailed) { fmt.Println("Validation failed") } if errors.Is(err, providererrors.ErrUnsupportedFeature) { fmt.Println("Feature not supported by provider") } ``` ## Comprehensive Error Handling Example Here's a complete example showing robust error handling: ```go package main import ( "context" "errors" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func generateWithRetry(ctx context.Context, model provider.LanguageModel, prompt string, maxRetries int) (*ai.GenerateTextResult, error) { var lastErr error for attempt := 0; attempt <= maxRetries; attempt++ { if attempt > 0 { fmt.Printf("Retry attempt %d/%d\n", attempt, maxRetries) // Exponential backoff backoff := time.Duration(1<= 500 && providerErr.StatusCode < 600 { fmt.Printf("Provider error %d, retrying...\n", providerErr.StatusCode) continue } // 4xx errors are not retryable return nil, fmt.Errorf("non-retryable provider error: %w", err) } // Don't retry on context errors if errors.Is(err, context.Canceled) || errors.Is(err, context.DeadlineExceeded) { return nil, err } // Don't retry on validation errors if providererrors.IsValidationError(err) { return nil, fmt.Errorf("validation error: %w", err) } } return nil, fmt.Errorf("failed after %d attempts: %w", maxRetries+1, lastErr) } func main() { ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, err := provider.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } result, err := generateWithRetry(ctx, model, "Explain quantum computing", 3) if err != nil { log.Fatalf("Failed to generate text: %v", err) } fmt.Println(result.Text) } ``` ## Best Practices ### 1. Always Check Errors Never ignore errors - always check and handle them appropriately: ```go // Good result, err := ai.GenerateText(ctx, options) if err != nil { return fmt.Errorf("failed to generate text: %w", err) } // Bad result, _ := ai.GenerateText(ctx, options) ``` ### 2. Use Error Wrapping Wrap errors to provide context: ```go result, err := ai.GenerateText(ctx, options) if err != nil { return fmt.Errorf("failed to process user request: %w", err) } ``` ### 3. Check Specific Error Types Use `errors.As()` and `errors.Is()` to check for specific errors: ```go var rateLimitErr *providererrors.RateLimitError if errors.As(err, &rateLimitErr) { // Handle rate limit specifically } ``` ### 4. Handle Context Cancellation Always respect context cancellation: ```go for chunk := range stream.Chunks() { select { case <-ctx.Done(): return ctx.Err() default: fmt.Print(chunk.Text) } } ``` ### 5. Use Defer for Cleanup Use defer to ensure cleanup happens even on error: ```go ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() // Always called, even on error result, err := ai.GenerateText(ctx, options) // ... ``` ### 6. Log Errors Appropriately Log errors with sufficient context for debugging: ```go if err != nil { log.Printf("Failed to generate text for user %s: %v", userID, err) return err } ``` ### 7. Don't Retry Non-Retryable Errors Some errors shouldn't be retried (validation errors, 4xx errors): ```go if providererrors.IsValidationError(err) { return err // Don't retry } var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) && providerErr.StatusCode >= 400 && providerErr.StatusCode < 500 { return err // Don't retry 4xx errors } ``` ### 8. Implement Exponential Backoff When retrying, use exponential backoff to avoid overwhelming the service: ```go for attempt := 0; attempt <= maxRetries; attempt++ { if attempt > 0 { backoff := time.Duration(1< # Testing > Describes strategies for testing Go code that calls non-deterministic, slow, and costly language models, with examples and integration-testing tips. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/testing Documentation index: https://goaisdk.com/llms.txt Testing language models can be challenging, because they are non-deterministic and calling them is slow and expensive. The Go AI SDK is designed with testability in mind. All models are interfaces, which makes it easy to create mock implementations for testing. This allows you to test your code in a repeatable and deterministic way without actually calling a language model provider. ## Testing Strategies ### 1. Use the SDK mocks The `pkg/testutil` package ships ready-made mocks for every model type: `MockLanguageModel`, `MockEmbeddingModel`, `MockImageModel`, `MockSpeechModel`, `MockTranscriptionModel`, `MockRerankingModel` and `MockProvider`. Each mock has a `Do...Func` field you can set to control the response, and records the calls it receives. ```go package myapp import ( "context" "testing" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/testutil" ) func TestMockLanguageModel(t *testing.T) { inputTokens, outputTokens, totalTokens := int64(10), int64(20), int64(30) model := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { return &types.GenerateResult{ Text: "Mock response", FinishReason: types.FinishReasonStop, Usage: types.Usage{ InputTokens: &inputTokens, OutputTokens: &outputTokens, TotalTokens: &totalTokens, }, }, nil }, } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Hello", }) if err != nil { t.Fatal(err) } if result.Text != "Mock response" { t.Errorf("unexpected text: %q", result.Text) } if len(model.GenerateCalls) != 1 { t.Errorf("expected 1 call, got %d", len(model.GenerateCalls)) } } ``` With no `DoGenerateFunc` set, `MockLanguageModel` returns a default `"mock response"` result. With no `DoStreamFunc`, it streams `"mock "` and `"response"`. ### 2. Mock text stream Use `testutil.NewMockTextStream` to build a stream from a fixed list of chunks: ```go package myapp import ( "context" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/testutil" ) func newStreamingModel(text string) *testutil.MockLanguageModel { return &testutil.MockLanguageModel{ DoStreamFunc: func(ctx context.Context, opts *provider.GenerateOptions) (provider.TextStream, error) { return testutil.NewMockTextStream([]provider.StreamChunk{ {Type: provider.ChunkTypeText, Text: text}, {Type: provider.ChunkTypeFinish, FinishReason: types.FinishReasonStop}, }), nil }, } } ``` Use `testutil.NewMockTextStreamWithError` to simulate a stream that fails. ## Testing Examples ### Testing GenerateText ```go package myapp import ( "context" "testing" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/testutil" ) func TestGenerateText(t *testing.T) { ctx := context.Background() tests := []struct { name string prompt string mockResponse string expectedResult string expectError bool }{ { name: "successful generation", prompt: "Hello, test!", mockResponse: "Hello, world!", expectedResult: "Hello, world!", expectError: false, }, { name: "empty response", prompt: "Test", mockResponse: "", expectedResult: "", expectError: false, }, } for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { // Create mock model mockModel := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { return &types.GenerateResult{ Text: tt.mockResponse, FinishReason: types.FinishReasonStop, }, nil }, } // Call function under test result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: mockModel, Prompt: tt.prompt, }) // Assertions if tt.expectError { if err == nil { t.Errorf("expected error, got nil") } return } if err != nil { t.Fatalf("unexpected error: %v", err) } if result.Text != tt.expectedResult { t.Errorf("expected %q, got %q", tt.expectedResult, result.Text) } }) } } ``` ### Testing StreamText ```go package myapp import ( "context" "strings" "testing" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/testutil" ) func TestStreamText(t *testing.T) { ctx := context.Background() tests := []struct { name string prompt string mockChunks []string expectedResult string }{ { name: "streaming response", prompt: "Hello, test!", mockChunks: []string{"Hello", ", ", "world", "!"}, expectedResult: "Hello, world!", }, } for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { // Create mock model mockModel := &testutil.MockLanguageModel{ DoStreamFunc: func(ctx context.Context, opts *provider.GenerateOptions) (provider.TextStream, error) { return newMockStreamWithChunks(tt.mockChunks), nil }, } // Call function under test stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: mockModel, Prompt: tt.prompt, }) if err != nil { t.Fatalf("unexpected error: %v", err) } // Collect streamed text var builder strings.Builder for chunk := range stream.Chunks() { builder.WriteString(chunk.Text) } if err := stream.Err(); err != nil { t.Fatalf("stream error: %v", err) } result := builder.String() if result != tt.expectedResult { t.Errorf("expected %q, got %q", tt.expectedResult, result) } }) } } // newMockStreamWithChunks creates a mock stream with multiple chunks func newMockStreamWithChunks(chunks []string) provider.TextStream { streamChunks := make([]provider.StreamChunk, 0, len(chunks)+1) for _, chunk := range chunks { streamChunks = append(streamChunks, provider.StreamChunk{ Type: provider.ChunkTypeText, Text: chunk, }) } streamChunks = append(streamChunks, provider.StreamChunk{ Type: provider.ChunkTypeFinish, FinishReason: types.FinishReasonStop, }) return testutil.NewMockTextStream(streamChunks) } ``` ### Testing GenerateObject ```go package myapp import ( "context" "encoding/json" "testing" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/schema" "github.com/digitallysavvy/go-ai/pkg/testutil" ) func TestGenerateObject(t *testing.T) { ctx := context.Background() objSchema := map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "age": map[string]interface{}{"type": "number"}, }, "required": []string{"name", "age"}, } tests := []struct { name string mockResponse string expectedName string expectedAge float64 expectError bool }{ { name: "valid object", mockResponse: `{"name":"John","age":30}`, expectedName: "John", expectedAge: 30, expectError: false, }, } for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { mockModel := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { return &types.GenerateResult{ Text: tt.mockResponse, FinishReason: types.FinishReasonStop, }, nil }, } result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: mockModel, Schema: schema.NewSimpleJSONSchema(objSchema), Prompt: "Generate a person", }) if tt.expectError { if err == nil { t.Errorf("expected error, got nil") } return } if err != nil { t.Fatalf("unexpected error: %v", err) } // Parse the object (result.Object is already unmarshaled into a // generic interface{}; re-marshal it to decode into a typed struct) var person struct { Name string `json:"name"` Age float64 `json:"age"` } objBytes, err := json.Marshal(result.Object) if err != nil { t.Fatalf("failed to marshal object: %v", err) } if err := json.Unmarshal(objBytes, &person); err != nil { t.Fatalf("failed to unmarshal object: %v", err) } if person.Name != tt.expectedName { t.Errorf("expected name %q, got %q", tt.expectedName, person.Name) } if person.Age != tt.expectedAge { t.Errorf("expected age %f, got %f", tt.expectedAge, person.Age) } }) } } ``` ### Testing Tool Calling ```go package myapp import ( "context" "testing" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/testutil" ) func TestToolCalling(t *testing.T) { ctx := context.Background() weatherTool := types.Tool{ Name: "weather", Description: "Get the weather", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{"type": "string"}, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{ "temperature": 72, "condition": "sunny", }, nil }, } mockModel := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { // Simulate tool call return &types.GenerateResult{ ToolCalls: []types.ToolCall{ { ID: "call-1", ToolName: "weather", Arguments: map[string]interface{}{ "location": "San Francisco", }, }, }, FinishReason: types.FinishReasonToolCalls, }, nil }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: mockModel, Prompt: "What's the weather?", Tools: []types.Tool{weatherTool}, }) if err != nil { t.Fatalf("unexpected error: %v", err) } if len(result.ToolResults) == 0 { t.Errorf("expected tool results, got none") } } ``` ### Testing with Embedding Models ```go package myapp import ( "context" "testing" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/testutil" ) func TestEmbed(t *testing.T) { ctx := context.Background() mockModel := &testutil.MockEmbeddingModel{} result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: mockModel, Input: "test text", }) if err != nil { t.Fatalf("unexpected error: %v", err) } if len(result.Embedding) != 5 { t.Errorf("expected 5-dimensional embedding, got %d", len(result.Embedding)) } } ``` ## Testing Best Practices ### 1. Use Table-Driven Tests Go's table-driven tests are ideal for testing multiple scenarios: ```go func TestMultipleScenarios(t *testing.T) { tests := []struct { name string input string expected string }{ {"scenario 1", "input1", "output1"}, {"scenario 2", "input2", "output2"}, {"scenario 3", "input3", "output3"}, } for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { // Test logic here }) } } ``` ### 2. Test Error Conditions Always test error paths: ```go func TestErrorHandling(t *testing.T) { ctx := context.Background() mockModel := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { return nil, errors.New("mock error") }, } _, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: mockModel, Prompt: "test", }) if err == nil { t.Error("expected error, got nil") } } ``` ### 3. Test Context Cancellation Test that your code respects context cancellation: ```go func TestContextCancellation(t *testing.T) { ctx, cancel := context.WithCancel(context.Background()) cancel() // Cancel immediately mockModel := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { // A real model call would fail on a canceled context; the mock // must check it explicitly since it never makes an HTTP request. if err := ctx.Err(); err != nil { return nil, err } return &types.GenerateResult{Text: "unreachable"}, nil }, } _, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: mockModel, Prompt: "test", }) if err == nil { t.Error("expected context cancellation error") } } ``` ### 4. Use Test Helpers Create reusable test helpers: ```go func newTestModel(response string) *testutil.MockLanguageModel { return &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { return &types.GenerateResult{ Text: response, FinishReason: types.FinishReasonStop, }, nil }, } } func TestWithHelper(t *testing.T) { model := newTestModel("test response") // Use model in test... } ``` ### 5. Test Callbacks Test that callbacks are called correctly: ```go func TestOnFinishCallback(t *testing.T) { ctx := context.Background() callbackCalled := false mockModel := &testutil.MockLanguageModel{} _, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: mockModel, Prompt: "test", OnFinish: func(ctx context.Context, result *ai.GenerateTextResult, userContext interface{}) { callbackCalled = true }, }) if err != nil { t.Fatalf("unexpected error: %v", err) } if !callbackCalled { t.Error("OnFinish callback was not called") } } ``` ### 6. Use Subtests for Complex Scenarios Break down complex tests into subtests: ```go func TestComplexScenario(t *testing.T) { t.Run("setup", func(t *testing.T) { // Setup test }) t.Run("execution", func(t *testing.T) { // Execute test }) t.Run("validation", func(t *testing.T) { // Validate results }) } ``` ### 7. Mock External Dependencies When testing code that uses AI SDK, mock the model interface: ```go type MyService struct { model provider.LanguageModel } func (s *MyService) ProcessText(ctx context.Context, text string) (string, error) { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: s.model, Prompt: text, }) if err != nil { return "", err } return result.Text, nil } func TestMyService(t *testing.T) { mockModel := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { return &types.GenerateResult{ Text: "processed: " + opts.Prompt.Messages[0].Content[0].(types.TextContent).Text, }, nil }, } service := &MyService{model: mockModel} result, err := service.ProcessText(context.Background(), "test") if err != nil { t.Fatalf("unexpected error: %v", err) } if result != "processed: test" { t.Errorf("unexpected result: %s", result) } } ``` ## Integration Testing For integration tests with real providers, use build tags to separate them from unit tests: ```go //go:build integration // +build integration package myapp import ( "context" "os" "testing" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func TestRealProvider(t *testing.T) { if os.Getenv("OPENAI_API_KEY") == "" { t.Skip("OPENAI_API_KEY not set") } ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Say hello", }) if err != nil { t.Fatalf("unexpected error: %v", err) } if result.Text == "" { t.Error("expected non-empty response") } } ``` Run integration tests with: ```bash go test -tags=integration ./... ``` ## See Also - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - [Middleware](https://goaisdk.com/docs/ai-sdk-core/middleware.md) --- # Telemetry > Shows how to enable OpenTelemetry tracing in the Go AI SDK Core, covering telemetry settings, setup in Go, collected data, and semantic conventions. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/telemetry Documentation index: https://goaisdk.com/llms.txt The Go AI SDK uses [OpenTelemetry](https://opentelemetry.io/) to collect telemetry data. OpenTelemetry is an open-source observability framework designed to provide standardized instrumentation for collecting telemetry data. Check out the AI SDK Observability Integrations to see providers that offer monitoring and tracing for AI SDK applications. ## Enabling Telemetry Telemetry is active by default when at least one telemetry integration is registered. Use `Telemetry` for per-call options; `ExperimentalTelemetry` remains as a deprecated alias for migration. ```go import ( "context" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/telemetry" ) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a short story about a cat.", Telemetry: &telemetry.Options{}, }) ``` When telemetry is enabled, you can control whether to record input values and output values. By default, both are enabled. You can disable them for privacy, data transfer, or performance reasons: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a short story about a cat.", Telemetry: &telemetry.Options{ IsEnabled: telemetry.Bool(true), RecordInputs: false, // Don't record sensitive inputs RecordOutputs: true, // Record outputs }, }) ``` ## Telemetry Settings ### FunctionID Provide a `FunctionID` to identify the function that the telemetry data is for: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a short story about a cat.", Telemetry: &telemetry.Options{ IsEnabled: telemetry.Bool(true), FunctionID: "story-generator", }, }) ``` ### Context Inclusion Runtime and tool context is excluded from telemetry by default. Opt in to specific top-level keys: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a short story about a cat.", RuntimeContext: map[string]interface{}{"request_id": "req-123", "secret": "token"}, ToolsContext: map[string]interface{}{ "weather": map[string]interface{}{"city": "Paris", "api_key": "secret"}, }, Telemetry: &telemetry.Options{ IncludeRuntimeContext: map[string]bool{"request_id": true}, IncludeToolsContext: map[string]map[string]bool{ "weather": map[string]bool{"city": true}, }, }, }) ``` ### Custom Tracer Provide a custom OpenTelemetry tracer instead of using the global tracer. The tracer belongs to the integration, not to `telemetry.Options`: construct the integration with it and register it once at startup. ```go import ( "go.opentelemetry.io/otel/sdk/trace" "github.com/digitallysavvy/go-ai/pkg/telemetry" ) // Create custom tracer provider tp := trace.NewTracerProvider() tracer := tp.Tracer("my-custom-tracer") // GenAI semantic-convention spans. Use telemetry.NewLegacyOpenTelemetry // with telemetry.LegacyOpenTelemetryOptions{Tracer: tracer} for the // ai.* span names instead. telemetry.RegisterTelemetryIntegration( telemetry.NewOpenTelemetry(telemetry.OpenTelemetryOptions{Tracer: tracer}), ) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a short story about a cat.", Telemetry: &telemetry.Options{ IsEnabled: telemetry.Bool(true), }, }) ``` ### Custom Span Attributes Use `EnrichSpan` to attach observability-specific attributes when OTel spans are created. SDK-managed attributes take precedence when a custom key overlaps with a standard AI SDK or GenAI semantic-convention key. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a deployment summary.", RuntimeContext: map[string]interface{}{ "user_id": "user-123", }, Telemetry: &telemetry.Options{ EnrichSpan: func(ctx context.Context, span telemetry.EnrichSpanOptions) map[string]interface{} { attrs := map[string]interface{}{ "app.span_type": string(span.SpanType), } if span.RuntimeContext["user_id"] != nil { attrs["app.user_id"] = span.RuntimeContext["user_id"] } return attrs }, }, }) ``` ### Event Names Telemetry integrations should implement `OnEnd`, `OnEmbedEnd`, and `OnRerankEnd`. The older `OnFinish`, `OnEmbedFinish`, and `OnRerankFinish` names are kept as deprecated compatibility fallbacks. Streaming `OnChunk` telemetry is no longer emitted; application-level `StreamTextOptions.OnChunk` callbacks still receive stream chunks. `OnEmbedEnd` and `OnRerankEnd` now also fire for a failed model-call attempt (a retry that errored), not only on success. `EmbeddingModelCallEndEvent` and `RerankingModelCallEndEvent` each gained an `Error error` field: when set, the built-in integrations record error status on the span instead of setting usage/embeddings/ranking attributes. Custom integrations that implement `OnEmbedEnd` / `OnRerankEnd` should check `event.Error` before assuming a successful call. Step lifecycle telemetry uses `OnStepEnd` as the canonical name. `OnStepFinish` is still emitted as a deprecated compatibility event for migration, and `OnStepEnd` takes precedence when both callback names are configured. The built-in `telemetry.OpenTelemetry` and `telemetry.LegacyOpenTelemetry` integrations set the OTel span status to `codes.Error` on the language-model-call, step, and root spans whenever the corresponding event's `FinishReason` is `"error"` (for example a provider stream that ends with an error chunk), even when no separate `OnError`/`OnStepError` event fires. Spans finishing with any other reason keep the default `Unset` status. ### Abort and Performance Events Telemetry integrations can implement `OnAbort(context.Context, telemetry.TelemetryAbortEvent)`. The SDK emits this event when generation is canceled or aborted and closes the active OpenTelemetry span so traces do not remain open. Model calls are made inside the telemetry context returned by `OnStart` and `OnStepStart`. If an integration stores span context or request-scoped values in the returned context, provider calls and tool execution receive that context. Language model call performance uses TypeScript-compatible names: ```go type LanguageModelCallPerformance struct { TimeToFirstOutputMs *int64 TimeBetweenOutputChunksMs *types.OutputChunkTimingStats } ``` File outputs and tool calls count as output for `TimeToFirstOutputMs`. Runtime and tool context remain excluded from telemetry by default; when included through `IncludeRuntimeContext` or `IncludeToolsContext`, nested objects are emitted as separate attributes instead of a single collapsed object. ## Fluent API for Settings Use the fluent API to build telemetry settings: ```go telemetrySettings := telemetry.DefaultSettings(). WithEnabled(true). WithFunctionID("story-generator"). WithRecordInputs(false) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate text", Telemetry: telemetrySettings, }) ``` ## Setting Up OpenTelemetry in Go Before using telemetry, you need to configure OpenTelemetry in your application: ### Basic Setup ```go skip-compile package main import ( "context" "log" "go.opentelemetry.io/otel" "go.opentelemetry.io/otel/exporters/stdout/stdouttrace" "go.opentelemetry.io/otel/sdk/trace" ) func initTelemetry() func() { // Create stdout exporter exporter, err := stdouttrace.New(stdouttrace.WithPrettyPrint()) if err != nil { log.Fatalf("failed to create exporter: %v", err) } // Create tracer provider tp := trace.NewTracerProvider( trace.WithBatcher(exporter), ) // Set global tracer provider otel.SetTracerProvider(tp) // Return cleanup function return func() { if err := tp.Shutdown(context.Background()); err != nil { log.Printf("Error shutting down tracer provider: %v", err) } } } func main() { // Initialize telemetry cleanup := initTelemetry() defer cleanup() // Your application code... } ``` ### Jaeger Setup ```go skip-compile import ( "go.opentelemetry.io/otel" "go.opentelemetry.io/otel/exporters/jaeger" "go.opentelemetry.io/otel/sdk/trace" ) func initJaeger() func() { // Create Jaeger exporter exporter, err := jaeger.New(jaeger.WithCollectorEndpoint( jaeger.WithEndpoint("http://localhost:14268/api/traces"), )) if err != nil { log.Fatalf("failed to create Jaeger exporter: %v", err) } // Create tracer provider tp := trace.NewTracerProvider( trace.WithBatcher(exporter), trace.WithSampler(trace.AlwaysSample()), ) otel.SetTracerProvider(tp) return func() { if err := tp.Shutdown(context.Background()); err != nil { log.Printf("Error shutting down tracer provider: %v", err) } } } ``` ### OTLP (OpenTelemetry Protocol) Setup ```go import ( "go.opentelemetry.io/otel" "go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp" "go.opentelemetry.io/otel/sdk/trace" ) func initOTLP() func() { // Create OTLP exporter exporter, err := otlptracehttp.New(context.Background(), otlptracehttp.WithEndpoint("localhost:4318"), otlptracehttp.WithInsecure(), ) if err != nil { log.Fatalf("failed to create OTLP exporter: %v", err) } // Create tracer provider tp := trace.NewTracerProvider( trace.WithBatcher(exporter), ) otel.SetTracerProvider(tp) return func() { if err := tp.Shutdown(context.Background()); err != nil { log.Printf("Error shutting down tracer provider: %v", err) } } } ``` ## Collected Data The Go AI SDK records spans and attributes following OpenTelemetry conventions. ### GenerateText Records the following spans: **`ai.generateText` span:** - Full length of the generateText call - Contains one or more `ai.generateText.doGenerate` spans - Attributes: - `operation.name`: `ai.generateText` + function ID - `ai.operationId`: `"ai.generateText"` - `ai.model.id`: Model identifier - `ai.model.provider`: Provider name - `ai.prompt`: The prompt used - `ai.response.text`: Generated text (if RecordOutputs is true) - `ai.response.finishReason`: Finish reason - `ai.usage.promptTokens`: Input tokens used - `ai.usage.completionTokens`: Output tokens used - `ai.usage.totalTokens`: Total tokens used - `ai.telemetry.functionId`: Function ID from settings - Custom metadata attributes **`ai.generateText.doGenerate` span:** - Individual provider call - Attributes: - `operation.name`: `ai.generateText.doGenerate` + function ID - `ai.operationId`: `"ai.generateText.doGenerate"` - `ai.prompt.messages`: Messages sent to provider - `ai.response.text`: Generated text - `ai.response.finishReason`: Finish reason - `ai.response.model`: Actual model used by provider - `gen_ai.system`: Provider name - `gen_ai.request.model`: Requested model - `gen_ai.request.temperature`: Temperature setting - `gen_ai.request.max_tokens`: Max tokens setting - `gen_ai.usage.input_tokens`: Input tokens - `gen_ai.usage.output_tokens`: Output tokens ### StreamText Records the following spans: **`ai.streamText` span:** - Full length of the streamText call - Contains `ai.streamText.doStream` span - Same attributes as `ai.generateText`, plus: - `ai.response.msToFirstChunk`: Time to first chunk (milliseconds) - `ai.response.msToFinish`: Time to completion (milliseconds) - `ai.response.effectiveOutputTokensPerSecond`: Output tokens per request second - `ai.response.effectiveTotalTokensPerSecond`: Input plus output tokens per request second - `ai.response.outputTokensPerSecond`: Streaming output tokens per output-stream second, when available - `ai.response.inputTokensPerSecond`: Streaming input tokens per time-to-first-output-token second, when available **`ai.streamText.doStream` span:** - Individual provider stream call - Emits `ai.stream.firstChunk` event when first chunk received - Same attributes as `ai.generateText.doGenerate`, plus streaming metrics ### GenerateObject Records the following spans: **`ai.generateObject` span:** - Full length of the generateObject call - Attributes: - Same as `ai.generateText`, plus: - `ai.schema`: Stringified JSON schema - `ai.schema.name`: Schema name - `ai.schema.description`: Schema description - `ai.response.object`: Generated object (stringified JSON) - `ai.settings.output`: Output type (`object`, `array`, `no-schema`, etc.) **`ai.generateObject.doGenerate` span:** - Individual provider call - Attributes similar to `ai.generateText.doGenerate` ### StreamObject Records the following spans: **`ai.streamObject` span:** - Full length of the streamObject call - Same attributes as `ai.generateObject`, plus streaming metrics **`ai.streamObject.doStream` span:** - Individual provider stream call - Emits `ai.stream.firstChunk` event ### Embed Records the following spans: **`ai.embed` span:** - Full length of the embed call - Attributes: - `operation.name`: `ai.embed` + function ID - `ai.operationId`: `"ai.embed"` - `ai.model.id`: Model identifier - `ai.model.provider`: Provider name - `ai.value`: Input value (if RecordInputs is true) - `ai.embedding`: Stringified embedding vector (if RecordOutputs is true) - `ai.usage.tokens`: Tokens used **`ai.embed.doEmbed` span:** - Provider embed call - Similar attributes ### EmbedMany Records the following spans: **`ai.embedMany` span:** - Full length of the embedMany call - Attributes: - `operation.name`: `ai.embedMany` + function ID - `ai.operationId`: `"ai.embedMany"` - `ai.values`: Input values array - `ai.embeddings`: Array of stringified embeddings - `ai.usage.tokens`: Total tokens used **`ai.embedMany.doEmbed` span:** - Individual provider call - May be multiple spans if batching occurs ### Speech and Transcription `GenerateSpeech`, `Transcribe`, and `ExperimentalStreamTranscribe` emit a single root span covering the whole call (including retries for the non-streaming calls) — there is no nested per-attempt span, unlike `ai.embed`/`ai.generateText`. All three accept a `Telemetry`/`ExperimentalTelemetry` option, dispatch `OnStart`/`OnEnd`/`OnError` to every registered integration, and publish the same diagnostics-channel events (`onStart`, `onEnd`, `onError`) as every other operation. **`ai.generateSpeech` span:** - Attributes: - `operation.name`: `ai.generateSpeech` + function ID - `ai.operationId`: `"ai.generateSpeech"` - `gen_ai.output.type`: `"speech"` (GenAI integration only) - `ai.request.text`: Input text (if RecordInputs is true) - `ai.response.audio.size` / `ai.response.audio.mediaType` / `ai.response.audio.format`: Generated audio metadata (if RecordOutputs is true) - `ai.usage.` / `gen_ai.usage.`: Provider-reported usage, flattened to numeric attributes (e.g. `ai.usage.characters` for Deepgram TTS; see "Provider usage attributes" below) - `ai.response.usage`: Raw provider usage object, JSON-encoded - `ai.response.providerMetadata`: Provider-specific metadata (if enabled) - Ends with `codes.Error` status on failure, including a provider response with no audio. **`ai.transcribe` / `ai.streamTranscribe` spans:** - Attributes: - `operation.name`: `ai.transcribe`/`ai.streamTranscribe` + function ID - `ai.operationId`: `"ai.transcribe"` / `"ai.streamTranscribe"` - `gen_ai.request.stream`: `true` for `ai.streamTranscribe`, `false` for `ai.transcribe` (GenAI integration only) - `ai.request.audio.size` / `ai.request.audio.mediaType`: Input audio metadata (if RecordInputs is true). For `ai.streamTranscribe`, the byte count accumulates as audio chunks are read and is only set once the stream reaches its final part. - `ai.response.text`: Transcript text (if RecordOutputs is true) - `ai.usage.` / `gen_ai.usage.`: Provider-reported usage, flattened to numeric attributes - `ai.response.usage`: Raw provider usage object, JSON-encoded - `ai.response.providerMetadata`: Provider-specific metadata (if enabled) - Ends with `codes.Error` status on failure, including a provider response with no transcript. **Provider usage attributes:** Speech/transcription providers report usage in their own native shape (e.g. Deepgram TTS: `{"characters": N}`; Deepgram transcription: `{"seconds": N}`; Google: `{"promptTokenCount": N, ...}`). The telemetry layer flattens this generically: common field spellings (`inputTokens`, `promptTokens`, `promptTokenCount`, `totalInputTokens` → `input_tokens`; `outputTokens`, `completionTokens`, `completionTokenCount`, `candidatesTokenCount` → `output_tokens`; `totalTokens`, `totalTokenCount` → `total_tokens`) are aliased to a shared attribute name, and everything else is converted from camelCase to `snake_case` (e.g. `characters` stays `characters`, `promptAudioSeconds` becomes `prompt_audio_seconds`). Nested usage objects are flattened with dotted paths. ### Tool Calls When tools are executed, `ai.toolCall` spans are recorded: - `operation.name`: `"ai.toolCall"` - `ai.operationId`: `"ai.toolCall"` - `ai.toolCall.name`: Tool name - `ai.toolCall.id`: Tool call ID - `ai.toolCall.args`: Tool input arguments - `ai.toolCall.result`: Tool output (if successful and serializable) ## Example: Complete Telemetry Setup Here's a complete example showing telemetry setup and usage: ```go skip-compile package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/telemetry" "go.opentelemetry.io/otel" "go.opentelemetry.io/otel/exporters/stdout/stdouttrace" "go.opentelemetry.io/otel/sdk/trace" ) func initTelemetry() func() { exporter, err := stdouttrace.New(stdouttrace.WithPrettyPrint()) if err != nil { log.Fatalf("failed to create exporter: %v", err) } tp := trace.NewTracerProvider( trace.WithBatcher(exporter), trace.WithSampler(trace.AlwaysSample()), ) otel.SetTracerProvider(tp) return func() { if err := tp.Shutdown(context.Background()); err != nil { log.Printf("Error shutting down tracer provider: %v", err) } } } func main() { // Initialize telemetry cleanup := initTelemetry() defer cleanup() ctx := context.Background() // Set up provider and model provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Configure telemetry settings telemetrySettings := &telemetry.Options{ IsEnabled: telemetry.Bool(true), RecordInputs: true, RecordOutputs: true, FunctionID: "story-generator", } // Generate text with telemetry result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a short story about a cat.", Telemetry: telemetrySettings, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Best Practices ### 1. Disable Sensitive Data Recording Disable input/output recording for sensitive data: ```go telemetrySettings := &telemetry.Options{ IsEnabled: telemetry.Bool(true), RecordInputs: false, // Don't record sensitive prompts RecordOutputs: false, // Don't record sensitive responses FunctionID: "sensitive-operation", } ``` ### 2. Use Meaningful Function IDs Use descriptive function IDs to group related operations: ```go // Good FunctionID: "user-chat-response" FunctionID: "document-summarization" FunctionID: "code-generation" // Bad FunctionID: "func1" FunctionID: "test" ``` ### 3. Include Safe Context Explicitly Include only non-sensitive context keys that help with debugging and analysis: ```go Telemetry: &telemetry.Options{ IncludeRuntimeContext: map[string]bool{ "request_id": true, "feature_flag": true, }, } ``` ### 4. Use Sampling for High-Volume Applications For high-traffic applications, use sampling to reduce overhead: ```go import "go.opentelemetry.io/otel/sdk/trace" tp := trace.NewTracerProvider( trace.WithBatcher(exporter), trace.WithSampler(trace.TraceIDRatioBased(0.1)), // Sample 10% of traces ) ``` ### 5. Configure Appropriate Exporters Choose exporters based on your observability stack: - **Stdout**: Development and debugging - **Jaeger**: Self-hosted tracing - **OTLP**: Cloud providers (Datadog, New Relic, Honeycomb, etc.) - **Zipkin**: Existing Zipkin infrastructure ### 6. Monitor Performance Impact Telemetry has a small performance cost. Monitor and adjust: ```go // Disable telemetry for performance-critical paths criticalOp, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, Telemetry: &telemetry.Options{ IsEnabled: telemetry.Bool(false), }, }) // Enable telemetry for monitored paths monitoredOp, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, Telemetry: telemetrySettings, }) ``` ### 7. Use Environment-Based Configuration Configure telemetry based on environment: ```go func getTelemetrySettings(env string) *telemetry.Options { if env == "production" { return &telemetry.Options{ IsEnabled: telemetry.Bool(true), RecordInputs: false, // Privacy in production RecordOutputs: false, FunctionID: "prod-operation", } } // Full telemetry in development return &telemetry.Options{ IsEnabled: telemetry.Bool(true), RecordInputs: true, RecordOutputs: true, FunctionID: "dev-operation", } } ``` ## OpenTelemetry Semantic Conventions The Go AI SDK follows [OpenTelemetry Semantic Conventions for GenAI operations](https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-spans/), including: - `gen_ai.system`: Provider identifier - `gen_ai.request.model`: Requested model - `gen_ai.request.temperature`: Temperature setting - `gen_ai.request.max_tokens`: Maximum tokens - `gen_ai.response.finish_reasons`: Finish reasons - `gen_ai.usage.input_tokens`: Input token count - `gen_ai.usage.output_tokens`: Output token count ## See Also - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [Testing](https://goaisdk.com/docs/ai-sdk-core/testing.md) - [Middleware](https://goaisdk.com/docs/ai-sdk-core/middleware.md) - [OpenTelemetry Go Documentation](https://opentelemetry.io/docs/instrumentation/go/) --- # Development Tools & Debugging > Covers debugging, profiling, and testing Go AI SDK applications, including debug mode, request/response inspection, token usage, and performance profiling. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/devtools Documentation index: https://goaisdk.com/llms.txt The Go AI SDK provides several tools and techniques for debugging, observability, and testing your AI applications. This guide covers logging, request inspection, performance profiling, and testing utilities. ## Debug Mode ### Enabling Debug Logging There is no built-in `DebugMode` flag. Every provider's `Config` accepts an `HTTPClient` override, so you can log requests and responses by supplying an `http.Client` with a logging `http.RoundTripper` (see [Intercepting HTTP Requests](#intercepting-http-requests) below for the `LoggingTransport` implementation): ```go package main import ( "context" "log" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // simpleLoggingTransport logs the method and URL of every request. // See LoggingTransport under "Intercepting HTTP Requests" below for a // version that also logs headers and bodies. type simpleLoggingTransport struct { Transport http.RoundTripper } func (t *simpleLoggingTransport) RoundTrip(req *http.Request) (*http.Response, error) { log.Printf("Request: %s %s", req.Method, req.URL) resp, err := t.Transport.RoundTrip(req) if err == nil { log.Printf("Response Status: %s", resp.Status) } return resp, err } func main() { ctx := context.Background() // Log all HTTP requests and responses via a custom HTTPClient provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), HTTPClient: &http.Client{ Transport: &simpleLoggingTransport{Transport: http.DefaultTransport}, }, }) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello, world!", }) if err != nil { log.Fatal(err) } log.Printf("Result: %s", result.Text) } ``` ### Provider-Specific Debug Options The same `HTTPClient` override works for every provider: ```go // Anthropic anthropicProvider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), HTTPClient: httpClient, }) // Google googleProvider := google.New(google.Config{ APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"), HTTPClient: httpClient, }) // Bedrock does not expose an HTTPClient override; use Headers for // provider-visible debug markers, or inspect traffic at the network layer. bedrockProvider := bedrock.New(bedrock.Config{ Region: "us-east-1", }) ``` ## Request & Response Inspection ### Logging Request Details Inspect full request configuration before sending: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") temperature := 0.7 maxTokens := 500 options := ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing", Temperature: &temperature, MaxTokens: &maxTokens, System: "You are a physics professor.", } // Log request configuration fmt.Println("=== Request Configuration ===") fmt.Printf("Model: %T\n", options.Model) fmt.Printf("Temperature: %.2f\n", *options.Temperature) fmt.Printf("MaxTokens: %d\n", *options.MaxTokens) fmt.Printf("System: %s\n", options.System) fmt.Printf("Prompt: %s\n", options.Prompt) fmt.Println() result, err := ai.GenerateText(ctx, options) if err != nil { log.Fatal(err) } // Log response details fmt.Println("=== Response Details ===") fmt.Printf("Text: %s\n", result.Text) fmt.Printf("Finish Reason: %s\n", result.FinishReason) fmt.Printf("Model: %s\n", result.Response.ModelID) fmt.Println() // Log usage statistics fmt.Println("=== Usage Statistics ===") fmt.Printf("Prompt Tokens: %d\n", result.Usage.GetInputTokens()) fmt.Printf("Completion Tokens: %d\n", result.Usage.GetOutputTokens()) fmt.Printf("Total Tokens: %d\n", result.Usage.GetTotalTokens()) } ``` ### Intercepting HTTP Requests Use a custom HTTP client to intercept and log requests: ```go package main import ( "bytes" "context" "io" "log" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // LoggingTransport wraps http.RoundTripper to log requests type LoggingTransport struct { Transport http.RoundTripper } func (t *LoggingTransport) RoundTrip(req *http.Request) (*http.Response, error) { // Log request log.Printf("Request: %s %s", req.Method, req.URL) log.Printf("Headers: %v", req.Header) // Read and log body if req.Body != nil { body, _ := io.ReadAll(req.Body) log.Printf("Body: %s", string(body)) req.Body = io.NopCloser(bytes.NewBuffer(body)) } // Execute request resp, err := t.Transport.RoundTrip(req) if err != nil { return nil, err } // Log response log.Printf("Response Status: %s", resp.Status) log.Printf("Response Headers: %v", resp.Header) return resp, nil } func main() { ctx := context.Background() // Create custom HTTP client with logging httpClient := &http.Client{ Transport: &LoggingTransport{ Transport: http.DefaultTransport, }, } provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), HTTPClient: httpClient, }) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { log.Fatal(err) } log.Printf("Result: %s", result.Text) } ``` ## Token Usage Monitoring ### Tracking Token Consumption Monitor token usage to optimize costs: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type TokenTracker struct { TotalPromptTokens int64 TotalCompletionTokens int64 TotalCost float64 } func (t *TokenTracker) Track(result *ai.GenerateTextResult) { t.TotalPromptTokens += result.Usage.GetInputTokens() t.TotalCompletionTokens += result.Usage.GetOutputTokens() // Example pricing promptCost := float64(result.Usage.GetInputTokens()) * 0.03 / 1000 completionCost := float64(result.Usage.GetOutputTokens()) * 0.06 / 1000 t.TotalCost += promptCost + completionCost } func (t *TokenTracker) Report() { fmt.Println("=== Token Usage Report ===") fmt.Printf("Total Prompt Tokens: %d\n", t.TotalPromptTokens) fmt.Printf("Total Completion Tokens: %d\n", t.TotalCompletionTokens) fmt.Printf("Total Tokens: %d\n", t.TotalPromptTokens+t.TotalCompletionTokens) fmt.Printf("Estimated Cost: $%.4f\n", t.TotalCost) } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") tracker := &TokenTracker{} // Make multiple requests prompts := []string{ "What is Go?", "Explain concurrency in Go", "What are goroutines?", } for _, prompt := range prompts { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { log.Fatal(err) } tracker.Track(result) fmt.Printf("Prompt: %s\n", prompt) fmt.Printf("Response: %s\n", result.Text) fmt.Printf("Tokens: %d\n\n", result.Usage.GetTotalTokens()) } tracker.Report() } ``` ### Setting Token Budgets Enforce token limits to prevent runaway costs: ```go package main import ( "context" "errors" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type BudgetedAI struct { Provider *openai.Provider TokenBudget int64 TokensUsed int64 } func (b *BudgetedAI) GenerateText(ctx context.Context, opts ai.GenerateTextOptions) (*ai.GenerateTextResult, error) { if b.TokensUsed >= b.TokenBudget { return nil, errors.New("token budget exceeded") } result, err := ai.GenerateText(ctx, opts) if err != nil { return nil, err } b.TokensUsed += result.Usage.GetTotalTokens() fmt.Printf("Tokens used: %d/%d (%.1f%%)\n", b.TokensUsed, b.TokenBudget, float64(b.TokensUsed)/float64(b.TokenBudget)*100) return result, nil } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) budgetedAI := &BudgetedAI{ Provider: provider, TokenBudget: 1000, // Maximum 1000 tokens } model, _ := provider.LanguageModel("gpt-6-astra") result, err := budgetedAI.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a short story", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Performance Profiling ### Using Go's pprof Profile CPU and memory usage: ```go package main import ( "context" "log" "os" "runtime/pprof" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { // Start CPU profiling cpuFile, err := os.Create("cpu.prof") if err != nil { log.Fatal(err) } defer cpuFile.Close() pprof.StartCPUProfile(cpuFile) defer pprof.StopCPUProfile() // Your AI operations ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") for i := 0; i < 10; i++ { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { log.Fatal(err) } log.Printf("Iteration %d: %s", i, result.Text) } // Write heap profile heapFile, err := os.Create("heap.prof") if err != nil { log.Fatal(err) } defer heapFile.Close() pprof.WriteHeapProfile(heapFile) } ``` Analyze profiles: ```bash # View CPU profile go tool pprof cpu.prof # View memory profile go tool pprof heap.prof # Generate visualization go tool pprof -http=:8080 cpu.prof ``` ### Measuring Request Latency Track operation duration: ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type LatencyTracker struct { Latencies []time.Duration } func (t *LatencyTracker) Track(duration time.Duration) { t.Latencies = append(t.Latencies, duration) } func (t *LatencyTracker) Report() { if len(t.Latencies) == 0 { return } var total time.Duration min := t.Latencies[0] max := t.Latencies[0] for _, d := range t.Latencies { total += d if d < min { min = d } if d > max { max = d } } avg := total / time.Duration(len(t.Latencies)) fmt.Println("=== Latency Report ===") fmt.Printf("Requests: %d\n", len(t.Latencies)) fmt.Printf("Min: %v\n", min) fmt.Printf("Max: %v\n", max) fmt.Printf("Avg: %v\n", avg) } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") tracker := &LatencyTracker{} for i := 0; i < 5; i++ { start := time.Now() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) duration := time.Since(start) tracker.Track(duration) if err != nil { log.Fatal(err) } fmt.Printf("Request %d: %v - %s\n", i+1, duration, result.Text) } tracker.Report() } ``` ## Testing Utilities ### Mocking AI Responses Create mock providers for testing: ```go package main import ( "context" "testing" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) // MockLanguageModel implements provider.LanguageModel type MockLanguageModel struct { Response string Error error } func int64Ptr(v int64) *int64 { return &v } func (m *MockLanguageModel) SpecificationVersion() string { return "v3" } func (m *MockLanguageModel) Provider() string { return "mock" } func (m *MockLanguageModel) ModelID() string { return "mock-model" } func (m *MockLanguageModel) SupportsTools() bool { return false } func (m *MockLanguageModel) SupportsStructuredOutput() bool { return false } func (m *MockLanguageModel) SupportsImageInput() bool { return false } func (m *MockLanguageModel) DoGenerate(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { if m.Error != nil { return nil, m.Error } return &types.GenerateResult{ Text: m.Response, FinishReason: types.FinishReasonStop, Usage: types.Usage{ InputTokens: int64Ptr(10), OutputTokens: int64Ptr(20), TotalTokens: int64Ptr(30), }, }, nil } func (m *MockLanguageModel) DoStream(ctx context.Context, opts *provider.GenerateOptions) (provider.TextStream, error) { return nil, nil // Implement if needed } // Example test func TestGenerateText(t *testing.T) { ctx := context.Background() mock := &MockLanguageModel{ Response: "This is a mocked response", } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: mock, Prompt: "Test prompt", }) if err != nil { t.Fatalf("Expected no error, got %v", err) } if result.Text != "This is a mocked response" { t.Errorf("Expected 'This is a mocked response', got '%s'", result.Text) } if result.Usage.GetTotalTokens() != 30 { t.Errorf("Expected 30 total tokens, got %d", result.Usage.GetTotalTokens()) } } ``` ### Integration Testing Test with real providers in controlled environments: ```go package main import ( "context" "os" "testing" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func TestOpenAIIntegration(t *testing.T) { if testing.Short() { t.Skip("Skipping integration test") } apiKey := os.Getenv("OPENAI_API_KEY") if apiKey == "" { t.Skip("OPENAI_API_KEY not set") } ctx := context.Background() provider := openai.New(openai.Config{ APIKey: apiKey, }) model, err := provider.LanguageModel("gpt-5.4-nano") if err != nil { t.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Say 'test'", }) if err != nil { t.Fatalf("GenerateText failed: %v", err) } if result.Text == "" { t.Error("Expected non-empty response") } if result.Usage.GetTotalTokens() == 0 { t.Error("Expected non-zero token usage") } } ``` Run integration tests: ```bash # Run all tests including integration tests go test ./... # Skip integration tests (faster) go test -short ./... # Run with verbose output go test -v ./... ``` ### Snapshot Testing Test output consistency: ```go package main import ( "context" "encoding/json" "os" "path/filepath" "testing" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) // MockLanguageModel implements provider.LanguageModel (same definition as // shown earlier under "Mocking AI Responses"). type MockLanguageModel struct { Response string Error error } func (m *MockLanguageModel) SpecificationVersion() string { return "v3" } func (m *MockLanguageModel) Provider() string { return "mock" } func (m *MockLanguageModel) ModelID() string { return "mock-model" } func (m *MockLanguageModel) SupportsTools() bool { return false } func (m *MockLanguageModel) SupportsStructuredOutput() bool { return false } func (m *MockLanguageModel) SupportsImageInput() bool { return false } func (m *MockLanguageModel) DoGenerate(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { if m.Error != nil { return nil, m.Error } return &types.GenerateResult{ Text: m.Response, FinishReason: types.FinishReasonStop, }, nil } func (m *MockLanguageModel) DoStream(ctx context.Context, opts *provider.GenerateOptions) (provider.TextStream, error) { return nil, nil } func TestGenerateTextSnapshot(t *testing.T) { ctx := context.Background() mock := &MockLanguageModel{ Response: "Consistent response for snapshot", } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: mock, Prompt: "Test prompt", }) if err != nil { t.Fatal(err) } // Create snapshot snapshotPath := filepath.Join("testdata", "snapshots", "generate_text.json") data, _ := json.MarshalIndent(result, "", " ") // Check if snapshot exists if _, err := os.Stat(snapshotPath); os.IsNotExist(err) { // Create snapshot os.MkdirAll(filepath.Dir(snapshotPath), 0755) os.WriteFile(snapshotPath, data, 0644) t.Log("Created new snapshot") return } // Compare with snapshot expected, _ := os.ReadFile(snapshotPath) if string(data) != string(expected) { t.Errorf("Snapshot mismatch.\nExpected:\n%s\n\nGot:\n%s", expected, data) } } ``` ## Observability ### Structured Logging Use structured logging for better observability: ```go package main import ( "context" "log/slog" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { // Set up structured logger logger := slog.New(slog.NewJSONHandler(os.Stdout, &slog.HandlerOptions{ Level: slog.LevelDebug, })) ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") logger.Info("Starting AI request", "model", "gpt-6-astra", "prompt", "Hello!", ) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { logger.Error("AI request failed", "error", err, ) return } logger.Info("AI request completed", "model", result.Response.ModelID, "finish_reason", result.FinishReason, "prompt_tokens", result.Usage.GetInputTokens(), "completion_tokens", result.Usage.GetOutputTokens(), "total_tokens", result.Usage.GetTotalTokens(), ) } ``` ### Tracing with OpenTelemetry Add distributed tracing: ```go package main import ( "context" "log" "os" "go.opentelemetry.io/otel" "go.opentelemetry.io/otel/attribute" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func generateWithTracing(ctx context.Context, model provider.LanguageModel, prompt string) (*ai.GenerateTextResult, error) { tracer := otel.Tracer("go-ai-app") ctx, span := tracer.Start(ctx, "ai.GenerateText") defer span.End() span.SetAttributes( attribute.String("ai.model", "gpt-6-astra"), attribute.String("ai.prompt", prompt), ) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { span.RecordError(err) return nil, err } span.SetAttributes( attribute.Int64("ai.tokens.prompt", result.Usage.GetInputTokens()), attribute.Int64("ai.tokens.completion", result.Usage.GetOutputTokens()), attribute.Int64("ai.tokens.total", result.Usage.GetTotalTokens()), attribute.String("ai.finish_reason", string(result.FinishReason)), ) return result, nil } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") result, err := generateWithTracing(ctx, model, "Hello!") if err != nil { log.Fatal(err) } log.Printf("Result: %s", result.Text) } ``` ## Debugging Common Issues ### Rate Limiting Debug rate limit errors: ```go if err != nil { if strings.Contains(err.Error(), "rate_limit") { log.Printf("Rate limit hit: %v", err) log.Printf("Retry after delay...") time.Sleep(60 * time.Second) // Retry logic } } ``` ### Context Cancellation Debug timeout issues: ```go ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { if ctx.Err() == context.DeadlineExceeded { log.Println("Request timed out after 30 seconds") } else { log.Printf("Request failed: %v", err) } } ``` ### Tool Execution Errors Debug tool calling issues: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Tools: tools, Prompt: "Use the weather tool", OnStepFinish: func(ctx context.Context, step types.StepResult, userContext interface{}) { log.Printf("Step %d:", step.StepNumber) for _, call := range step.ToolCalls { log.Printf(" Tool: %s", call.ToolName) log.Printf(" Args: %+v", call.Arguments) } for _, res := range step.ToolResults { if res.Error != nil { log.Printf(" Error: %v", res.Error) } else { log.Printf(" Result: %+v", res.Result) } } }, }) ``` ## Best Practices ### 1. Log HTTP traffic in development ```go // ✅ Good: Debug mode in development cfg := openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")} if os.Getenv("ENV") == "development" { cfg.HTTPClient = &http.Client{ Transport: &LoggingTransport{Transport: http.DefaultTransport}, } } provider := openai.New(cfg) ``` ### 2. Monitor Token Usage ```go // ✅ Good: Track and log token usage if result.Usage.GetTotalTokens() > 10000 { log.Printf("WARNING: High token usage: %d", result.Usage.GetTotalTokens()) } ``` ### 3. Use Structured Logging ```go // ✅ Good: Structured logging logger.Info("AI request", "model", "gpt-6-astra", "tokens", result.Usage.GetTotalTokens(), "duration", duration, ) ``` ### 4. Test with Mocks ```go // ✅ Good: Use mocks for unit tests mock := &MockLanguageModel{Response: "test"} result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: mock, }) ``` ### 5. Profile Production Issues ```go // ✅ Good: Enable profiling in production import _ "net/http/pprof" go func() { log.Println(http.ListenAndServe("localhost:6060", nil)) }() ``` ## Tools Summary | Tool | Purpose | Use Case | |------|---------|----------| | Debug Mode | Log HTTP requests/responses | Development debugging | | Token Tracking | Monitor API costs | Cost optimization | | pprof | CPU/memory profiling | Performance tuning | | Latency Tracking | Measure request duration | Performance monitoring | | Mock Providers | Unit testing | Test automation | | Structured Logging | Observability | Production monitoring | | OpenTelemetry | Distributed tracing | Microservices debugging | ## See Also - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [Prompt Engineering](https://goaisdk.com/docs/ai-sdk-core/prompt-engineering.md) - [Advanced: Rate Limiting](https://goaisdk.com/docs/advanced/rate-limiting.md) - [Advanced: Performance Optimization](https://goaisdk.com/docs/advanced/model-as-router.md) --- # Event callbacks > Subscribe to lifecycle events in GenerateText, StreamText, Embed, EmbedMany, and Rerank calls. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/event-callbacks Documentation index: https://goaisdk.com/llms.txt The Go AI SDK provides per-call event callbacks that you can pass to `GenerateText`, `StreamText`, `Embed`, `EmbedMany`, and `Rerank` to observe lifecycle events. This is useful for building observability tools, logging systems, analytics, and debugging utilities. ## Basic usage Pass callbacks directly to `GenerateText` or `StreamText`: ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() p := openai.New(openai.Config{APIKey: "your-api-key"}) model, err := p.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What is the weather in San Francisco?", OnStart: func(ctx context.Context, event ai.OnStartEvent) { fmt.Println("Generation started:", event.ModelID) }, OnFinishEvent: func(ctx context.Context, event ai.OnFinishEvent) { fmt.Println("Total tokens:", event.TotalUsage.GetTotalTokens()) }, }) if err != nil { log.Fatalf("Generation failed: %v", err) } fmt.Println(result.Text) } ``` ## Available callbacks ### GenerateText / StreamText | Callback | Event type | Description | |----------|-----------|-------------| | `OnStart` | `OnStartEvent` | Called when generation begins, before any LLM calls. | | `OnStepStart` | `OnStepStartEvent` | Called when a step (LLM call) begins, before the provider is called. | | `OnToolExecutionStart` | `OnToolCallStartEvent` | Called when a tool's Execute function is about to run. | | `OnToolExecutionEnd` | `OnToolCallFinishEvent` | Called when a tool's Execute function completes or errors. | | `OnStepFinishEvent` | `OnStepFinishEvent` | Called when a step (LLM call) completes. | | `OnFinishEvent` | `OnFinishEvent` | Called when the entire generation completes (all steps finished). | All six callbacks receive `(ctx context.Context, event )` and are panic-safe. They fire in addition to the legacy `OnStepFinish` and `OnFinish` callbacks. `OnToolCallStart` and `OnToolCallFinish` remain as deprecated aliases for `OnToolExecutionStart` and `OnToolExecutionEnd`. ### Embed / EmbedMany | Callback | Event type | Description | |----------|-----------|-------------| | `ExperimentalOnStart` | `EmbedOnStartEvent` | Called before the embedding model is invoked. | | `ExperimentalOnFinish` | `EmbedOnFinishEvent` | Called after the embedding model returns. | Embed callbacks receive `(event )` without a context parameter. ### Rerank | Callback | Event type | Description | |----------|-----------|-------------| | `ExperimentalOnStart` | `RerankOnStartEvent` | Called before the reranking model is invoked. | | `ExperimentalOnFinish` | `RerankOnFinishEvent` | Called after the reranking model returns. | Rerank callbacks receive `(event )` without a context parameter. ## Event reference ### GenerateText / StreamText #### OnStartEvent Called when the generation begins, before any LLM calls are made. | Field | Type | Description | |-------|------|-------------| | `ModelProvider` | `string` | The provider name (e.g., `"openai"`). | | `ModelID` | `string` | The model identifier (e.g., `"gpt-6-astra"`). | | `System` | `string` | The system message provided to the model. | | `Prompt` | `string` | The prompt string if using the Prompt option. | | `Messages` | `[]types.Message` | The messages array if using the Messages option. | | `Tools` | `[]types.Tool` | The tools available for this generation. | | `Temperature` | `*float64` | Sampling temperature. | | `MaxTokens` | `*int` | Maximum number of tokens to generate. | | `TopP` | `*float64` | Top-p (nucleus) sampling parameter. | | `TopK` | `*int` | Top-k sampling parameter. | | `FrequencyPenalty` | `*float64` | Frequency penalty. | | `PresencePenalty` | `*float64` | Presence penalty. | | `StopSequences` | `[]string` | Sequences that stop generation. | | `Seed` | `*int` | Random seed for reproducible generation. | | `ExperimentalContext` | `interface{}` | User-defined context flowing through the lifecycle. | #### OnStepStartEvent Called before each step (LLM call) begins. Useful for tracking multi-step generations with tool loops. | Field | Type | Description | |-------|------|-------------| | `StepNumber` | `int` | 0-indexed step number. | | `ModelProvider` | `string` | The provider name. | | `ModelID` | `string` | The model identifier. | | `System` | `string` | The system message for this step. | | `Messages` | `[]types.Message` | The messages sent to the model for this step. | | `Tools` | `[]types.Tool` | The tools available for this step. | | `PreviousSteps` | `[]types.StepResult` | Results from previous steps (empty for the first step). | | `ExperimentalContext` | `interface{}` | User-defined context object. | #### OnToolCallStartEvent Called before a tool's `Execute` function runs. | Field | Type | Description | |-------|------|-------------| | `ToolCallID` | `string` | Unique identifier for this tool call. | | `ToolName` | `string` | Name of the tool being called. | | `Args` | `map[string]any` | Input arguments passed to the tool. | | `StepNumber` | `int` | 0-indexed step where this tool call occurs. | | `ModelProvider` | `string` | The provider name. | | `ModelID` | `string` | The model identifier. | | `Messages` | `[]types.Message` | The conversation messages at tool execution time. | | `ExperimentalContext` | `interface{}` | User-defined context object. | #### OnToolCallFinishEvent Called after a tool's `Execute` function completes or errors. Exactly one of `Result` or `Error` will be non-nil. | Field | Type | Description | |-------|------|-------------| | `ToolCallID` | `string` | Unique identifier for this tool call. | | `ToolName` | `string` | Name of the tool that was called. | | `Args` | `map[string]any` | Input arguments passed to the tool. | | `Result` | `any` | The tool's return value (nil on error). | | `Error` | `error` | The error from tool execution (nil on success). | | `DurationMs` | `int64` | Execution time of the tool call in milliseconds. | | `StepNumber` | `int` | 0-indexed step where this tool call occurred. | | `ModelProvider` | `string` | The provider name. | | `ModelID` | `string` | The model identifier. | | `Messages` | `[]types.Message` | The conversation messages at tool execution time. | | `ExperimentalContext` | `interface{}` | User-defined context object. | #### OnStepFinishEvent Called after each step (LLM call) completes. | Field | Type | Description | |-------|------|-------------| | `StepNumber` | `int` | 0-indexed step number. | | `ModelProvider` | `string` | The provider name. | | `ModelID` | `string` | The model identifier. | | `Text` | `string` | The generated text from this step. | | `ToolCalls` | `[]types.ToolCall` | Tool calls made during this step. | | `ToolResults` | `[]types.ToolResult` | Results of tool calls from this step. | | `FinishReason` | `types.FinishReason` | Why the generation finished (`"stop"`, `"length"`, `"tool-calls"`, etc.). | | `Usage` | `types.Usage` | Token usage for this step. | | `Warnings` | `[]types.Warning` | Warnings from the provider. | | `ExperimentalContext` | `interface{}` | User-defined context object. | #### OnFinishEvent Called when the entire generation completes (all steps finished). Includes aggregated data. | Field | Type | Description | |-------|------|-------------| | `Text` | `string` | The full generated text. | | `ToolCalls` | `[]types.ToolCall` | Tool calls aggregated across all steps. | | `ToolResults` | `[]types.ToolResult` | Tool results aggregated across all steps. | | `FinishReason` | `types.FinishReason` | Why the final step finished. | | `Steps` | `[]types.StepResult` | Results from all steps in the generation. | | `TotalUsage` | `types.Usage` | Aggregated token usage across all steps. | | `Warnings` | `[]types.Warning` | Warnings aggregated across all steps. | | `ExperimentalContext` | `interface{}` | Final state of the user-defined context. | ### Embed / EmbedMany #### EmbedOnStartEvent Called when the embedding operation begins, before the embedding model is called. Both `Embed` and `EmbedMany` share the same event type; the `OperationID` field distinguishes them (`"ai.embed"` vs `"ai.embedMany"`). | Field | Type | Description | |-------|------|-------------| | `CallID` | `string` | Unique identifier for this embed call. | | `OperationID` | `string` | Operation type: `"ai.embed"` or `"ai.embedMany"`. | | `Provider` | `string` | The embedding provider name. | | `ModelID` | `string` | The embedding model identifier. | | `Values` | `[]string` | The input text(s) being embedded. | | `MaxRetries` | `int` | Maximum retries for failed requests. | | `Headers` | `map[string]string` | Additional HTTP headers sent with the request. | | `ProviderOptions` | `map[string]interface{}` | Provider-specific options. | | `FunctionID` | `string` | Telemetry function identifier. | | `Metadata` | `map[string]any` | Additional telemetry metadata. | #### EmbedOnFinishEvent Called when the embedding operation completes. For `Embed`, `Embeddings` contains a single vector. For `EmbedMany`, it contains one vector per input value. | Field | Type | Description | |-------|------|-------------| | `CallID` | `string` | Matches the `CallID` from the corresponding start event. | | `OperationID` | `string` | Operation type: `"ai.embed"` or `"ai.embedMany"`. | | `Provider` | `string` | The embedding provider name. | | `ModelID` | `string` | The embedding model identifier. | | `Value` | `[]string` | The input text(s) that were embedded. | | `Embeddings` | `[][]float64` | The resulting embedding vectors. | | `Usage` | `types.EmbeddingUsage` | Token usage for the embedding operation. | | `Warnings` | `[]types.Warning` | Warnings from the provider. | | `ProviderMetadata` | `json.RawMessage` | Provider-specific metadata. | | `Responses` | `[]types.EmbeddingResponse` | HTTP response metadata (headers, body). Length 1 for `Embed`; one per batch chunk for `EmbedMany`. | | `FunctionID` | `string` | Telemetry function identifier. | | `Metadata` | `map[string]any` | Additional telemetry metadata. | ### Rerank #### RerankOnStartEvent Called when the reranking operation begins, before the reranking model is called. | Field | Type | Description | |-------|------|-------------| | `CallID` | `string` | Unique identifier for this rerank call. | | `OperationID` | `string` | Always `"ai.rerank"`. | | `Provider` | `string` | The reranking provider name. | | `ModelID` | `string` | The reranking model identifier. | | `Query` | `string` | The query to rerank documents against. | | `Documents` | `interface{}` | The documents being reranked (`[]string` or `[]map[string]interface{}`). | | `TopN` | `*int` | Number of top documents to return (nil means all). | | `MaxRetries` | `int` | Maximum retries for failed requests. | | `Headers` | `map[string]string` | Additional HTTP headers sent with the request. | | `ProviderOptions` | `map[string]interface{}` | Provider-specific options. | | `FunctionID` | `string` | Telemetry function identifier. | #### RerankOnFinishEvent Called when the reranking operation completes, after the reranking model returns. | Field | Type | Description | |-------|------|-------------| | `CallID` | `string` | Matches the `CallID` from the corresponding start event. | | `OperationID` | `string` | Always `"ai.rerank"`. | | `Provider` | `string` | The reranking provider name. | | `ModelID` | `string` | The reranking model identifier. | | `Documents` | `interface{}` | The documents that were reranked. | | `Query` | `string` | The query that documents were reranked against. | | `Ranking` | `[]ai.RerankItem` | Reranked results sorted by relevance score (descending). Each item has `OriginalIndex`, `Score`, and `Document`. | | `Warnings` | `[]types.Warning` | Warnings from the provider. | | `Response` | `types.RerankResponse` | Response metadata (id, timestamp, modelId, headers). | | `ProviderMetadata` | `json.RawMessage` | Provider-specific metadata. | | `FunctionID` | `string` | Telemetry function identifier. | ## Use cases ### Logging and debugging Track the full lifecycle of a generation with timestamps: ```go package main import ( "context" "fmt" "log" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() p := openai.New(openai.Config{APIKey: "your-api-key"}) model, err := p.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", OnStart: func(ctx context.Context, event ai.OnStartEvent) { fmt.Printf("[%s] Generation started: model=%s provider=%s\n", time.Now().Format(time.RFC3339), event.ModelID, event.ModelProvider) }, OnStepFinishEvent: func(ctx context.Context, event ai.OnStepFinishEvent) { fmt.Printf("[%s] Step %d finished: reason=%s tokens=%d\n", time.Now().Format(time.RFC3339), event.StepNumber, event.FinishReason, event.Usage.GetTotalTokens()) }, OnFinishEvent: func(ctx context.Context, event ai.OnFinishEvent) { fmt.Printf("[%s] Generation complete: steps=%d totalTokens=%d\n", time.Now().Format(time.RFC3339), len(event.Steps), event.TotalUsage.GetTotalTokens()) }, }) if err != nil { log.Fatalf("Generation failed: %v", err) } fmt.Println(result.Text) } ``` ### Tool execution monitoring Track tool call durations and detect failures: ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() p := openai.New(openai.Config{APIKey: "your-api-key"}) model, err := p.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What is the weather?", // Tools: []types.Tool{getWeather}, OnToolExecutionStart: func(ctx context.Context, event ai.OnToolCallStartEvent) { fmt.Printf("Tool %q starting...\n", event.ToolName) }, OnToolExecutionEnd: func(ctx context.Context, event ai.OnToolCallFinishEvent) { if event.Error == nil { fmt.Printf("Tool %q completed in %dms\n", event.ToolName, event.DurationMs) } else { fmt.Printf("Tool %q failed: %v\n", event.ToolName, event.Error) } }, }) if err != nil { log.Fatalf("Generation failed: %v", err) } fmt.Println(result.Text) } ``` ### Embedding observability Monitor embedding operations with token usage tracking: ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() p := openai.New(openai.Config{APIKey: "your-api-key"}) model, err := p.EmbeddingModel("text-embedding-3-small") if err != nil { log.Fatal(err) } result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: model, Inputs: []string{"sunny day at the beach", "rainy afternoon in the city"}, ExperimentalOnStart: func(event ai.EmbedOnStartEvent) { fmt.Printf("Embedding started (%s): model=%s values=%d\n", event.OperationID, event.ModelID, len(event.Values)) }, ExperimentalOnFinish: func(event ai.EmbedOnFinishEvent) { fmt.Printf("Embedding complete (%s): tokens=%d vectors=%d\n", event.OperationID, event.Usage.TotalTokens, len(event.Embeddings)) }, }) if err != nil { log.Fatalf("Embedding failed: %v", err) } fmt.Printf("Generated %d embeddings\n", len(result.Embeddings)) } ``` ## Error handling Errors that occur inside callbacks are caught internally and **do not break** the generation, embedding, or reranking flow. The `Notify` helper wraps each callback invocation in a deferred recover, so panics are silently absorbed. This ensures that monitoring code cannot disrupt your application: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", OnStart: func(ctx context.Context, event ai.OnStartEvent) { panic("this panic is recovered internally") // Generation continues normally }, }) // err is nil — the callback panic did not propagate ``` If you need to capture callback errors for debugging, handle them within the callback itself (e.g., log them or send them to an error tracking service). ## Related documentation - [Generating text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - [Embeddings](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) - [Reranking](https://goaisdk.com/docs/ai-sdk-core/reranking.md) - [Tools and tool calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - [Telemetry](https://goaisdk.com/docs/ai-sdk-core/telemetry.md) --- # AI SDK Core > Learn about AI SDK Core and how to work with Large Language Models in Go Canonical URL: https://goaisdk.com/docs/ai-sdk-core/ Documentation index: https://goaisdk.com/llms.txt The core API provides a comprehensive set of functions for working with Large Language Models, embeddings, image generation, speech, and more. All functions are designed with Go's concurrency patterns and best practices in mind. ## Core Functions ### Text Generation - **[Overview](https://goaisdk.com/docs/ai-sdk-core/overview.md)** - Learn about AI SDK Core and how to work with Large Language Models (LLMs). - **[Generating Text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md)** - Learn how to generate text using `GenerateText` and `StreamText`. - **[Generating Structured Data](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md)** - Learn how to generate structured data with type-safe JSON output. - **[Tool Calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md)** - Learn how to do tool calling with AI SDK Core. ### Configuration - **[Settings](https://goaisdk.com/docs/ai-sdk-core/settings.md)** - Learn how to set up settings for language model generations. - **[Upload File and Upload Skill](https://goaisdk.com/docs/ai-sdk-core/upload-file-and-skill.md)** - Learn how to upload files and multi-file skills for provider references. ### Embeddings and Ranking - **[Embeddings](https://goaisdk.com/docs/ai-sdk-core/embeddings.md)** - Learn how to use embeddings with AI SDK Core. - **[Reranking](https://goaisdk.com/docs/ai-sdk-core/reranking.md)** - Learn how to rerank documents using reranking models. ### Media Generation - **[Image Generation](https://goaisdk.com/docs/ai-sdk-core/image-generation.md)** - Learn how to generate images with AI SDK Core. - **[Transcription](https://goaisdk.com/docs/ai-sdk-core/transcription.md)** - Learn how to transcribe audio with AI SDK Core. - **[Speech](https://goaisdk.com/docs/ai-sdk-core/speech.md)** - Learn how to generate speech with AI SDK Core. - **[Realtime](https://goaisdk.com/docs/ai-sdk-core/realtime.md)** - Learn how to build realtime voice conversations with Go. ### Advanced Features - **[Middleware](https://goaisdk.com/docs/ai-sdk-core/middleware.md)** - Learn how to use middleware to wrap AI provider calls. - **[Provider Management](https://goaisdk.com/docs/ai-sdk-core/provider-management.md)** - Learn how to work with multiple providers and implement fallbacks. - **[Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md)** - Learn how to handle errors with AI SDK Core. - **[Testing](https://goaisdk.com/docs/ai-sdk-core/testing.md)** - Learn how to test your AI-powered Go applications. - **[Telemetry](https://goaisdk.com/docs/ai-sdk-core/telemetry.md)** - Learn how to use telemetry to monitor your AI applications. ## Next Steps - Explore [Advanced Topics](https://goaisdk.com/docs/advanced.md) for production patterns - Learn about [Building Agents](https://goaisdk.com/docs/agents.md) for autonomous workflows - Review [Foundations](https://goaisdk.com/docs/foundations.md) for core concepts --- # MCP Tool Serialization > Documents GetSerializableTools() in the Go AI SDK for retrieving MCP tool definitions in a storable, transmittable format, with pagination support. Canonical URL: https://goaisdk.com/docs/ai-sdk-core/mcp-serialization Documentation index: https://goaisdk.com/llms.txt The Model Context Protocol (MCP) allows servers to expose tools that can be used by AI models. The Go AI SDK provides methods to retrieve tool definitions in a format that can be stored, transmitted, or cached. ## GetSerializableTools() The `GetSerializableTools()` method returns tool definitions from an MCP server in a JSON-serializable format with full pagination support. ### Method Signature ```go func (c *MCPClient) GetSerializableTools(ctx context.Context) (*ListToolsResult, error) ``` ### Return Type ```go type ListToolsResult struct { Tools []MCPTool `json:"tools"` NextCursor string `json:"nextCursor,omitempty"` } type MCPTool struct { Name string `json:"name"` Title string `json:"title,omitempty"` Description string `json:"description,omitempty"` InputSchema map[string]interface{} `json:"inputSchema"` OutputSchema map[string]interface{} `json:"outputSchema,omitempty"` Annotations map[string]interface{} `json:"annotations,omitempty"` Meta map[string]interface{} `json:"_meta,omitempty"` } ``` `MCPTool` preserves the current MCP tool JSON shape, including `title`, `outputSchema`, `annotations`, and `_meta`. MCP Apps metadata is stored in `_meta.ui` or the legacy flat `_meta["ui/resourceUri"]` field and can be normalized with `mcp.GetMCPAppToolMeta`. ## Basic Usage ```go package main import ( "context" "encoding/json" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/mcp" ) func main() { // Create MCP client transport := mcp.NewStdioTransport(mcp.StdioTransportConfig{ Command: "npx", Args: []string{"-y", "@modelcontextprotocol/server-everything"}, }) client := mcp.NewMCPClient(transport, mcp.MCPClientConfig{ ClientName: "my-app", ClientVersion: "1.0.0", }) // Connect to server ctx := context.Background() if err := client.Connect(ctx); err != nil { log.Fatal(err) } defer client.Close() // Get serializable tools tools, err := client.GetSerializableTools(ctx) if err != nil { log.Fatal(err) } // Tools can now be serialized and stored data, _ := json.MarshalIndent(tools, "", " ") fmt.Printf("Serialized tools:\n%s\n", string(data)) } ``` ## Storing Tool Definitions Tool definitions can be cached to avoid repeated server queries: ```go import ( "encoding/json" "os" ) // Cache tools to file func cacheTools(tools *mcp.ListToolsResult, filename string) error { data, err := json.Marshal(tools) if err != nil { return err } return os.WriteFile(filename, data, 0644) } // Load tools from cache func loadCachedTools(filename string) (*mcp.ListToolsResult, error) { data, err := os.ReadFile(filename) if err != nil { return nil, err } var tools mcp.ListToolsResult if err := json.Unmarshal(data, &tools); err != nil { return nil, err } return &tools, nil } // Usage tools, err := client.GetSerializableTools(ctx) if err != nil { log.Fatal(err) } // Cache for later use if err := cacheTools(tools, "mcp-tools-cache.json"); err != nil { log.Printf("Warning: Failed to cache tools: %v", err) } ``` ## Transmitting Tool Definitions Tool definitions can be sent over the network: ```go import ( "bytes" "encoding/json" "net/http" ) // Send tools to remote service func sendToolsToService(tools *mcp.ListToolsResult, url string) error { data, err := json.Marshal(tools) if err != nil { return err } resp, err := http.Post(url, "application/json", bytes.NewReader(data)) if err != nil { return err } defer resp.Body.Close() return nil } // Usage tools, err := client.GetSerializableTools(ctx) if err != nil { log.Fatal(err) } err = sendToolsToService(tools, "https://api.example.com/mcp-tools") if err != nil { log.Printf("Failed to send tools: %v", err) } ``` ## Pagination Support When MCP servers have many tools, results may be paginated: ```go func getAllTools(client *mcp.MCPClient, ctx context.Context) ([]mcp.MCPTool, error) { var allTools []mcp.MCPTool // Get first page result, err := client.GetSerializableTools(ctx) if err != nil { return nil, err } allTools = append(allTools, result.Tools...) // Handle pagination if present. if result.NextCursor != "" { log.Printf("More tools available (cursor: %s)", result.NextCursor) } return allTools, nil } ``` ## Comparison: ListTools vs GetSerializableTools The SDK provides two methods for retrieving tools: | Method | Return Type | Use Case | |--------|-------------|----------| | `ListTools()` | `[]MCPTool` | Convenience method for immediate use | | `GetSerializableTools()` | `*ListToolsResult` | Full result with pagination support for storage/transmission | ```go // Convenience method - just the tools tools, err := client.ListTools(ctx) // Returns: []MCPTool // Serialization method - complete result result, err := client.GetSerializableTools(ctx) // Returns: *ListToolsResult with Tools + NextCursor ``` ## Use Cases ### 1. Tool Registry Service Build a centralized registry of available tools: ```go type ToolRegistry struct { tools map[string]*mcp.ListToolsResult } func (r *ToolRegistry) RegisterServer(serverName string, tools *mcp.ListToolsResult) { r.tools[serverName] = tools } func (r *ToolRegistry) GetAllTools() []*mcp.ListToolsResult { var all []*mcp.ListToolsResult for _, tools := range r.tools { all = append(all, tools) } return all } ``` ### 2. Offline Tool Discovery Cache tools for offline browsing: ```go func discoverAndCache(servers []string) error { for _, server := range servers { client := connectToServer(server) tools, err := client.GetSerializableTools(ctx) if err != nil { log.Printf("Failed to get tools from %s: %v", server, err) continue } filename := fmt.Sprintf("tools-%s.json", server) if err := cacheTools(tools, filename); err != nil { log.Printf("Failed to cache tools for %s: %v", server, err) } } return nil } ``` ### 3. Tool Versioning Track tool changes over time: ```go type ToolSnapshot struct { Timestamp time.Time `json:"timestamp"` Tools *mcp.ListToolsResult `json:"tools"` } func snapshotTools(client *mcp.MCPClient) (*ToolSnapshot, error) { tools, err := client.GetSerializableTools(context.Background()) if err != nil { return nil, err } return &ToolSnapshot{ Timestamp: time.Now(), Tools: tools, }, nil } ``` ## Best Practices 1. **Cache tool definitions**: Avoid repeated queries by caching results 2. **Handle pagination**: Check for `NextCursor` and implement pagination when needed 3. **Validate after deserialization**: Ensure tools loaded from cache are still valid 4. **Version your cache**: Include timestamps or version numbers in cached data 5. **Implement TTL**: Refresh cached tools periodically to catch updates ## Error Handling ```go tools, err := client.GetSerializableTools(ctx) if err != nil { var mcpErr *mcp.MCPClientError switch { case errors.Is(err, context.DeadlineExceeded): log.Fatal("Timeout getting tools") case errors.As(err, &mcpErr): log.Fatalf("MCP error %d: %s", mcpErr.Code, mcpErr.Message) default: log.Fatalf("Failed to get tools: %v", err) } } ``` ## See Also - [MCP Integration Guide](https://goaisdk.com/docs/ai-sdk-core/mcp-tools.md) - [MCP Client API](https://goaisdk.com/docs/ai-sdk-core/mcp-tools.md) - [Tool Execution](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) --- # Coding Agents > How to build AI agents that generate, review, and iteratively refine code using the Go AI SDK. Canonical URL: https://goaisdk.com/docs/advanced/coding-agents Documentation index: https://goaisdk.com/llms.txt > **Looking to run Claude Code or Codex?** This page builds a coding agent from your own tools: a model that writes code and a loop you control. To drive an existing coding agent such as Claude Code or Codex in a sandbox, see [Coding agents with the harness](https://goaisdk.com/docs/build-a-chat-app/coding-agents-harness.md). Coding agents use LLMs to generate, review, test, and refine code through multi-step workflows. Unlike simple code generation prompts, coding agents maintain context across iterations, execute code, and correct errors automatically. ## Core Pattern: Generate → Execute → Refine The fundamental coding agent loop: 1. **Generate** — The model writes code based on a specification 2. **Execute** — Run the code in a sandboxed environment 3. **Observe** — Capture output, errors, and test results 4. **Refine** — Feed results back to the model for corrections 5. **Repeat** until tests pass or the step limit is reached ```go package main import ( "context" "fmt" "log" "os" "os/exec" "strings" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() p := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) model, _ := p.LanguageModel(anthropic.ClaudeSonnet4_5) // Build a coding agent with shell execution tools codingAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: `You are an expert Go programmer. Your job is to write correct, idiomatic Go code. Workflow: 1. Write the requested code 2. Use the write_file tool to save it 3. Use the run_go tool to test it 4. Fix any errors and repeat until the code passes 5. When the code passes, report the final solution`, Tools: []types.Tool{ writeFileTool(), runGoTool(), }, MaxSteps: 15, }) result, err := codingAgent.Execute(ctx, "Write a Go function `ReverseWords(s string) string` that reverses the order of words in a sentence. Include a test that verifies it.") if err != nil { log.Fatalf("Agent failed: %v", err) } fmt.Println("=== Final Result ===") fmt.Println(result.Text) fmt.Printf("\nCompleted in %d steps\n", len(result.Steps)) } // writeFileTool writes content to a file in a temp directory. func writeFileTool() types.Tool { return types.Tool{ Name: "write_file", Description: "Write content to a file. Use this to save generated Go code.", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "path": map[string]interface{}{"type": "string", "description": "File path (relative to /tmp/coding-agent/)"}, "content": map[string]interface{}{"type": "string", "description": "File content"}, }, "required": []string{"path", "content"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { path := "/tmp/coding-agent/" + input["path"].(string) content := input["content"].(string) os.MkdirAll("/tmp/coding-agent", 0755) if err := os.WriteFile(path, []byte(content), 0644); err != nil { return nil, fmt.Errorf("failed to write file: %w", err) } return map[string]interface{}{"status": "ok", "path": path}, nil }, } } // runGoTool runs `go test` or `go run` in the temp directory. func runGoTool() types.Tool { return types.Tool{ Name: "run_go", Description: "Run `go test ./...` or `go run ` in the working directory. Returns stdout, stderr, and exit code.", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "command": map[string]interface{}{ "type": "string", "description": "The go command to run: 'test', 'run ', or 'build'", }, }, "required": []string{"command"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { command := input["command"].(string) parts := strings.Fields("go " + command) cmd := exec.CommandContext(ctx, parts[0], parts[1:]...) cmd.Dir = "/tmp/coding-agent" out, err := cmd.CombinedOutput() exitCode := 0 if err != nil { if exitErr, ok := err.(*exec.ExitError); ok { exitCode = exitErr.ExitCode() } } return map[string]interface{}{ "output": string(out), "exit_code": exitCode, "success": exitCode == 0, }, nil }, } } ``` ## Code Review Agent A review agent examines existing code and suggests improvements without needing to execute anything: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type ReviewResult struct { Issues []CodeIssue Suggestions []string OverallScore int // 1-10 } type CodeIssue struct { Severity string // "critical", "warning", "info" Line int Description string Fix string } func reviewCode(ctx context.Context, code string) (*ReviewResult, error) { p := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := p.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: `You are an expert Go code reviewer. Analyze code for: - Security vulnerabilities (SQL injection, command injection, etc.) - Memory leaks and resource management issues - Race conditions and concurrency bugs - Performance problems - Style and idiomatic Go violations Return a structured JSON review.`, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: fmt.Sprintf("Review this Go code:\n```go\n%s\n```", code)}, }, }, }, Output: ai.ObjectOutput[ReviewResult](ai.ObjectOutputOptions{ Schema: ai.SchemaFor[ReviewResult](), Name: "code_review", Description: "Structured code review with issues and suggestions", }), }) if err != nil { return nil, err } review, ok := result.Output.(ReviewResult) if !ok { return nil, fmt.Errorf("unexpected output type") } return &review, nil } func main() { ctx := context.Background() code := ` func processUserInput(db *sql.DB, username string) error { query := "SELECT * FROM users WHERE name = '" + username + "'" rows, err := db.Query(query) if err != nil { return err } defer rows.Close() return nil }` review, err := reviewCode(ctx, code) if err != nil { log.Fatal(err) } fmt.Printf("Overall Score: %d/10\n\n", review.OverallScore) for _, issue := range review.Issues { fmt.Printf("[%s] Line %d: %s\n Fix: %s\n\n", issue.Severity, issue.Line, issue.Description, issue.Fix) } } ``` ## Multi-Step Code Refinement Some tasks require multiple rounds of generation and feedback. Use `GenerateText` directly for fine-grained control: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) // CodeRefinementSession manages a multi-step code refinement workflow. type CodeRefinementSession struct { model provider.LanguageModel history []types.Message maxRounds int } func NewCodeRefinementSession(model provider.LanguageModel, maxRounds int) *CodeRefinementSession { return &CodeRefinementSession{model: model, maxRounds: maxRounds} } // Generate produces or refines code based on the current conversation state. func (s *CodeRefinementSession) Generate(ctx context.Context, input string) (string, error) { s.history = append(s.history, types.Message{ Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: input}}, }) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: s.model, System: "You are an expert Go programmer. Write production-quality, well-documented Go code.", Messages: s.history, }) if err != nil { return "", err } s.history = append(s.history, types.Message{ Role: types.RoleAssistant, Content: []types.ContentPart{types.TextContent{Text: result.Text}}, }) return result.Text, nil } func main() { ctx := context.Background() p := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) model, _ := p.LanguageModel(anthropic.ClaudeSonnet4_5) session := NewCodeRefinementSession(model, 3) // Round 1: Initial implementation code, err := session.Generate(ctx, "Implement a concurrent worker pool in Go that processes jobs from a channel.") if err != nil { log.Fatal(err) } fmt.Println("Round 1 - Initial implementation:") fmt.Println(code) // Round 2: Add error handling code, err = session.Generate(ctx, "Add proper error handling so failed jobs don't silently fail. Workers should send errors to an error channel.") if err != nil { log.Fatal(err) } fmt.Println("\nRound 2 - With error handling:") fmt.Println(code) // Round 3: Add graceful shutdown code, err = session.Generate(ctx, "Add graceful shutdown so the pool waits for all in-progress jobs to complete before exiting.") if err != nil { log.Fatal(err) } fmt.Println("\nRound 3 - With graceful shutdown:") fmt.Println(code) } ``` ## Streaming Code Generation For interactive code generation (showing code as it's written), use `StreamText`: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() p := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := p.LanguageModel("gpt-6-astra") result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, System: "You are an expert Go programmer. Write clean, idiomatic code with inline comments.", Prompt: "Implement a simple HTTP rate limiter middleware for a Go web server using the token bucket algorithm.", }) if err != nil { log.Fatal(err) } defer result.Close() fmt.Println("Generating rate limiter implementation...") fmt.Println() // Print code tokens as they arrive for chunk := range result.Chunks() { fmt.Print(chunk.Text) } fmt.Println() } ``` ## Best Practices ### 1. Sandbox Execution Always run generated code in an isolated environment: ```go // Use Docker for strong isolation cmd := exec.CommandContext(ctx, "docker", "run", "--rm", "--memory=256m", // Memory limit "--cpus=0.5", // CPU limit "--network=none", // No network access "--read-only", // Read-only filesystem "-v", "/tmp/code:/workspace:rw", "golang:1.22", "sh", "-c", "cd /workspace && go test ./...", ) ``` ### 2. Timeout Control Always set timeouts on execution: ```go ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() result, err := codingAgent.Execute(ctx, specification) if err != nil { if ctx.Err() == context.DeadlineExceeded { log.Println("Code generation timed out") } } ``` ### 3. Clear Specifications The quality of generated code directly reflects the quality of the specification: ```go // Good — specific, testable requirements specification := ` Implement a Go function with this signature: func ParseDuration(s string) (time.Duration, error) Requirements: - Accept formats: "1h30m", "90m", "1.5h", "5400s" - Return an error for invalid formats - Include table-driven tests covering valid inputs, edge cases, and invalid inputs ` // Bad — vague specification := "Write a duration parser" ``` ### 4. Iterative Feedback Include specific error context when feeding back failures: ```go feedbackPrompt := fmt.Sprintf(`The code failed with: Test output: %s Please fix the failing tests. Focus on: 1. The exact error messages above 2. Don't change the function signature 3. Don't modify the test file`, testOutput) ``` ### 5. Context Window Management For long coding sessions, manage token usage: ```go // Use summarization for long refinement sessions if len(session.history) > 20 { summary, _ := session.Generate(ctx, "Summarize the key requirements and design decisions from our discussion so far in 2-3 paragraphs.") // Reset history with just the summary session.history = []types.Message{ {Role: types.RoleSystem, Content: []types.ContentPart{ types.TextContent{Text: "Previous context: " + summary}, }}, } } ``` ## See Also - [Building Agents](https://goaisdk.com/docs/agents/building-agents.md) — agent configuration and patterns - [Loop Control](https://goaisdk.com/docs/agents/loop-control.md) — StopWhen and MaxSteps - [Memory Management](https://goaisdk.com/docs/advanced/memory-management.md) — managing conversation history - [Tools and Tool Calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) — building execution tools - [Generating Structured Data](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) — typed code review results --- # Advanced Guides > In-depth guides for advanced AI application patterns using the Go AI SDK. Canonical URL: https://goaisdk.com/docs/advanced-guides Documentation index: https://goaisdk.com/llms.txt This section covers advanced application patterns for building production-grade AI systems with the Go AI SDK. ## Topics - **[Memory Management](https://goaisdk.com/docs/advanced/memory-management.md)** - Strategies for managing conversation history: sliding window, summarization-based compaction, Anthropic automatic compaction, and external memory stores. - **[Coding Agents](https://goaisdk.com/docs/advanced/coding-agents.md)** - Build AI agents that generate, execute, and iteratively refine code. Covers the Generate → Execute → Refine loop, code review with structured output, streaming code generation, and sandbox best practices. ## See Also - **[Advanced SDK Patterns](https://goaisdk.com/docs/advanced.md)** - Streaming patterns, caching, rate limiting, model-as-router, and more. - **[Building Agents](https://goaisdk.com/docs/agents.md)** - Agent configuration, tool calling, and loop control. - **[Providers](https://goaisdk.com/docs/providers.md)** - Provider-specific features and configuration. --- # Memory Management > Patterns for managing conversation history in long-running AI applications. Canonical URL: https://goaisdk.com/docs/advanced/memory-management Documentation index: https://goaisdk.com/llms.txt Long-running agents and multi-turn applications need strategies to keep conversation history within a model's context window. This guide covers three patterns: in-memory history, summarization-based compaction, and external memory stores. ## The Problem: Context Window Limits Every language model has a fixed context window — the maximum number of tokens it can process in a single request. Conversations that exceed this limit fail or truncate silently. Check your provider's documentation for the window of the model you use. For reference: | Model | Context Window | |-------|---------------| | claude-sonnet-4-5 | 200K tokens | | gemini-2.0-flash | 1M tokens | Even with large windows, longer conversations become expensive. Sending 100K tokens per request adds up quickly in production. ## Pattern 1: In-Memory History The simplest approach — keep all messages in a Go slice and pass it on each call. Suitable for short sessions or when context window size isn't a concern. ```go package main import ( "bufio" "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) // MessageHistory holds the in-memory conversation history. type MessageHistory struct { messages []types.Message } // Add appends a message to the history. func (h *MessageHistory) Add(role types.MessageRole, text string) { h.messages = append(h.messages, types.Message{ Role: role, Content: []types.ContentPart{ types.TextContent{Text: text}, }, }) } // Messages returns all messages. func (h *MessageHistory) Messages() []types.Message { return h.messages } func main() { ctx := context.Background() p := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) model, _ := p.LanguageModel(anthropic.ClaudeSonnet4_5) history := &MessageHistory{} scanner := bufio.NewScanner(os.Stdin) fmt.Println("Chat (Ctrl+C to quit):") for { fmt.Print("> ") if !scanner.Scan() { break } userInput := scanner.Text() history.Add(types.RoleUser, userInput) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a helpful assistant.", Messages: history.Messages(), }) if err != nil { log.Printf("Error: %v", err) continue } history.Add(types.RoleAssistant, result.Text) fmt.Printf("Assistant: %s\n\n", result.Text) } } ``` **Limitation:** History grows unbounded. Once it exceeds the model's context window, requests fail. ## Pattern 2: Sliding Window Keep only the N most recent message pairs. Simple to implement, but loses older context entirely. ```go // SlidingWindowHistory keeps only the most recent maxTurns exchanges. type SlidingWindowHistory struct { messages []types.Message maxTurns int // each "turn" = 1 user message + 1 assistant message } func NewSlidingWindowHistory(maxTurns int) *SlidingWindowHistory { return &SlidingWindowHistory{maxTurns: maxTurns} } func (h *SlidingWindowHistory) Add(role types.MessageRole, text string) { h.messages = append(h.messages, types.Message{ Role: role, Content: []types.ContentPart{types.TextContent{Text: text}}, }) h.trim() } func (h *SlidingWindowHistory) trim() { // Each turn has 2 messages (user + assistant); keep at most maxTurns*2 maxMessages := h.maxTurns * 2 if len(h.messages) > maxMessages { h.messages = h.messages[len(h.messages)-maxMessages:] } } func (h *SlidingWindowHistory) Messages() []types.Message { return h.messages } ``` ## Pattern 3: Summarization-Based Compaction When the conversation gets long, ask the model to summarize older turns and replace them with the summary. This preserves semantic meaning while reducing token count. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) const ( // Compact when the history exceeds this many messages. compactThreshold = 20 // After compaction, keep this many recent messages verbatim. keepRecentMessages = 6 ) // CompactingHistory maintains conversation history with automatic compaction. type CompactingHistory struct { messages []types.Message model provider.LanguageModel system string } func NewCompactingHistory(model provider.LanguageModel, system string) *CompactingHistory { return &CompactingHistory{model: model, system: system} } func (h *CompactingHistory) Add(role types.MessageRole, text string) { h.messages = append(h.messages, types.Message{ Role: role, Content: []types.ContentPart{types.TextContent{Text: text}}, }) } // MaybeCompact compacts older messages if the history is too long. // It preserves the keepRecentMessages most recent messages verbatim, // and replaces older messages with a model-generated summary. func (h *CompactingHistory) MaybeCompact(ctx context.Context) error { if len(h.messages) < compactThreshold { return nil } // Split: older messages to summarize, recent ones to keep splitAt := len(h.messages) - keepRecentMessages toSummarize := h.messages[:splitAt] toKeep := h.messages[splitAt:] // Build a summary prompt from the older messages var summaryPrompt string summaryPrompt = "Summarize the following conversation history concisely, preserving key facts, decisions, and context:\n\n" for _, msg := range toSummarize { for _, part := range msg.Content { if tc, ok := part.(types.TextContent); ok { summaryPrompt += fmt.Sprintf("[%s]: %s\n", msg.Role, tc.Text) } } } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: h.model, Prompt: summaryPrompt, }) if err != nil { return fmt.Errorf("compaction failed: %w", err) } // Replace older history with summary + recent messages summaryMessage := types.Message{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{ Text: fmt.Sprintf("[Previous conversation summary]: %s", result.Text), }, }, } h.messages = append([]types.Message{summaryMessage}, toKeep...) return nil } func (h *CompactingHistory) Messages() []types.Message { return h.messages } func main() { ctx := context.Background() p := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) model, _ := p.LanguageModel(anthropic.ClaudeSonnet4_5) system := "You are a helpful assistant with a long memory." history := NewCompactingHistory(model, system) // Simulate a long conversation exchanges := []string{ "What is Go's approach to concurrency?", "Can you explain goroutines in more detail?", "How do channels work in Go?", "What is a select statement?", "How does Go handle errors compared to exceptions?", // ... many more turns would trigger compaction } for _, userMsg := range exchanges { history.Add(types.RoleUser, userMsg) // Compact if needed before generating if err := history.MaybeCompact(ctx); err != nil { log.Printf("Warning: compaction failed: %v", err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: system, Messages: history.Messages(), }) if err != nil { log.Fatalf("Generation error: %v", err) } history.Add(types.RoleAssistant, result.Text) fmt.Printf("Q: %s\nA: %s\n\n", userMsg, result.Text) } } ``` ## Pattern 4: Automatic Compaction with Anthropic Context Management For Anthropic models, the API supports automatic context management — the model handles compaction server-side without a separate summarization call. This is more efficient than client-side summarization. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() p := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) // Create a model with automatic compaction enabled. // The API compacts the conversation server-side when input tokens exceed 50K. model, err := p.LanguageModelWithOptions( anthropic.ClaudeSonnet4_5, &anthropic.ModelOptions{ ContextManagement: &anthropic.ContextManagement{ Edits: []anthropic.ContextManagementEdit{ anthropic.NewCompactEdit(). WithTrigger(50000). WithInstructions("Preserve key decisions, facts, and user preferences. Discard verbose explanations."), }, }, }, ) if err != nil { log.Fatal(err) } // Use the model normally — compaction happens automatically messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "Let's work through a complex project together."}}, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a helpful project assistant.", Messages: messages, }) if err != nil { log.Fatal(err) } // Check if compaction was applied if result.ContextManagement != nil { cmr, ok := result.ContextManagement.(*anthropic.ContextManagementResponse) if ok { for _, edit := range cmr.AppliedEdits { if _, isCompact := edit.(*anthropic.AppliedCompactEdit); isCompact { fmt.Println("Context automatically compacted by the API") } } } } fmt.Println(result.Text) } ``` See the [Anthropic Provider](https://goaisdk.com/docs/providers/anthropic.md) guide for more details on `clear_tool_uses` and `clear_thinking` context management strategies. ## Pattern 5: External Memory Store For persistent memory across sessions, store conversation history in a database or key-value store. Retrieve relevant history on demand rather than loading everything. ```go package main import ( "context" "encoding/json" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // MemoryStore is an interface for persistent conversation storage. type MemoryStore interface { Save(sessionID string, messages []types.Message) error Load(sessionID string) ([]types.Message, error) // LoadRecent returns the most recent n messages for a session. LoadRecent(sessionID string, n int) ([]types.Message, error) } // FileMemoryStore is a simple file-based memory store (for illustration). // In production, use a database like PostgreSQL or Redis. type FileMemoryStore struct { dir string } func NewFileMemoryStore(dir string) *FileMemoryStore { os.MkdirAll(dir, 0755) return &FileMemoryStore{dir: dir} } func (s *FileMemoryStore) Save(sessionID string, messages []types.Message) error { path := fmt.Sprintf("%s/%s.json", s.dir, sessionID) data, err := json.MarshalIndent(messages, "", " ") if err != nil { return err } return os.WriteFile(path, data, 0644) } func (s *FileMemoryStore) Load(sessionID string) ([]types.Message, error) { path := fmt.Sprintf("%s/%s.json", s.dir, sessionID) data, err := os.ReadFile(path) if os.IsNotExist(err) { return nil, nil // New session } if err != nil { return nil, err } var messages []types.Message return messages, json.Unmarshal(data, &messages) } func (s *FileMemoryStore) LoadRecent(sessionID string, n int) ([]types.Message, error) { all, err := s.Load(sessionID) if err != nil || len(all) == 0 { return all, err } if len(all) <= n { return all, nil } return all[len(all)-n:], nil } func main() { ctx := context.Background() p := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := p.LanguageModel("gpt-6-astra") store := NewFileMemoryStore("/tmp/chat-sessions") sessionID := fmt.Sprintf("session-%d", time.Now().Unix()) // Load recent history (last 10 messages) history, err := store.LoadRecent(sessionID, 10) if err != nil { log.Fatalf("Failed to load history: %v", err) } userInput := "Continue our discussion about Go best practices." history = append(history, types.Message{ Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: userInput}}, }) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a Go programming expert.", Messages: history, }) if err != nil { log.Fatalf("Generation error: %v", err) } history = append(history, types.Message{ Role: types.RoleAssistant, Content: []types.ContentPart{types.TextContent{Text: result.Text}}, }) // Persist the updated history if err := store.Save(sessionID, history); err != nil { log.Printf("Warning: failed to save history: %v", err) } fmt.Println(result.Text) } ``` ## Choosing the Right Pattern | Pattern | Best For | Trade-off | |---------|----------|-----------| | In-memory history | Short sessions, prototyping | Fails on long conversations | | Sliding window | Simple apps, bounded memory | Loses older context | | Summarization compaction | Medium-length sessions | Extra API call for summarization | | Automatic compaction (Anthropic) | Anthropic-only apps | Provider-specific; most efficient | | External store | Multi-session, persistent apps | Requires database infrastructure | For most production applications, combine patterns: - Use a **sliding window** or **summarization** for in-flight messages - Use an **external store** for cross-session persistence - Use **automatic compaction** (Anthropic) for long single-session conversations ## See Also - [Anthropic Provider](https://goaisdk.com/docs/providers/anthropic.md) — server-side context management - [Building Agents](https://goaisdk.com/docs/agents/building-agents.md) — agent memory patterns - [Coding Agents](https://goaisdk.com/docs/advanced/coding-agents.md) — memory in coding agent workflows --- # Provider System Overview > Complete guide to the Go AI SDK provider architecture and capabilities Canonical URL: https://goaisdk.com/docs/providers/overview Documentation index: https://goaisdk.com/llms.txt The Go AI SDK provides a unified interface for working with 50 AI providers (`pkg/providers/*`, one shared internal package excluded), from OpenAI and Anthropic to open-source platforms and specialized services. Five of them — LMNT, Ollama, Stability AI, Vercel, and You.com (search tools) — are Go-only and have no upstream TypeScript SDK equivalent. See the [full provider index](https://goaisdk.com/docs/providers.md) for the complete list. ## Architecture ### Unified Interface All providers implement a consistent interface, allowing you to switch providers without changing your code: ```go // Works with any provider var p provider.Provider = openai.New(openai.Config{APIKey: key}) // OR p = anthropic.New(anthropic.Config{APIKey: key}) // OR p = google.New(google.Config{APIKey: key}) // Same API for all providers model, err := p.LanguageModel("model-id") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) ``` ### Provider Interface Every provider implements the `Provider` interface: ```go type Provider interface { Name() string LanguageModel(modelID string) (LanguageModel, error) EmbeddingModel(modelID string) (EmbeddingModel, error) ImageModel(modelID string) (ImageModel, error) SpeechModel(modelID string) (SpeechModel, error) TranscriptionModel(modelID string) (TranscriptionModel, error) RerankingModel(modelID string) (RerankingModel, error) } ``` ## Provider Types ### 1. Language Model Providers Providers offering text generation and chat capabilities: **Tier 1 (Frontier Models)** - OpenAI - GPT-6, GPT-5, o3, o4-mini - Anthropic - Claude Opus 4.5, Sonnet 4.5 - Google - Gemini 2.0 Flash, Pro **Tier 2 (High Performance)** - Mistral AI - Large 2, Pixtral - Cohere - Command R+ - xAI - Grok models - DeepSeek - DeepSeek-V3 **Tier 3 (Fast Inference)** - Groq - Ultra-fast LPU inference - Cerebras - Wafer-scale inference - Fireworks AI - Optimized serving **Tier 4 (Open Source Hosting)** - Together AI - Replicate - HuggingFace - DeepInfra **Tier 5 (Local Deployment)** - Ollama - Run models locally - Baseten - Self-hosted deployment **Enterprise Platforms** - Azure OpenAI - AWS Bedrock - Google Vertex AI ### 2. Embedding Providers Specialized in vector embeddings: - OpenAI - text-embedding-3-large - Cohere - embed-english-v3.0 - Google - gemini-embedding-001 - Mistral - mistral-embed - Voyage AI (via OpenAI-compatible API) ### 3. Image Generation Providers Text-to-image and image-to-image: - Stability AI - Stable Diffusion XL, SD3 - Black Forest Labs - FLUX.1 Pro, Dev, Schnell - FAL - Fast image/video generation - Replicate - Various image models - Together AI - Stable Diffusion hosting - Topaz Labs - Wonder 3.5 image enhancement (upscale, denoise, sharpen, restore) ### 3a. Video Generation Providers Text-to-video and image-to-video: - KlingAI - Kling v2.6 text-to-video and image-to-video - Alibaba - Wan video generation - ByteDance - Seedance text-to-video and image-to-video (Volcengine Ark) - Topaz Labs - Proteus and Starlight Precise 2.6 video enhancement (upscale, denoise, sharpen, restore an existing video) ### 4. Speech Synthesis Providers Text-to-speech capabilities: - ElevenLabs - High-quality voice synthesis - OpenAI - TTS HD models - Google - Text-to-Speech API ### 5. Transcription Providers Speech-to-text capabilities: - Deepgram - Real-time transcription - AssemblyAI - Audio intelligence - OpenAI - Whisper models - Google - Speech-to-Text API ## Configuration Patterns ### Basic Configuration ```go provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) ``` ### Advanced Configuration ```go provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), BaseURL: "https://custom-proxy.com", HTTPClient: &http.Client{ Timeout: time.Second * 60, }, Headers: map[string]string{ "X-Custom-Header": "value", }, }) ``` ### Regional Configuration ```go // Azure OpenAI with regional endpoint azureProvider, err := azure.New(azure.Config{ ResourceName: os.Getenv("AZURE_RESOURCE_NAME"), APIKey: os.Getenv("AZURE_API_KEY"), APIVersion: "v1", }) if err != nil { log.Fatal(err) } // AWS Bedrock with region bedrockProvider := bedrock.New(bedrock.Config{ Region: "us-west-2", AWSAccessKeyID: os.Getenv("AWS_ACCESS_KEY_ID"), AWSSecretAccessKey: os.Getenv("AWS_SECRET_ACCESS_KEY"), }) ``` ### OpenAI-Compatible Providers Many providers offer OpenAI-compatible APIs: ```go // Together AI provider := openai.New(openai.Config{ APIKey: os.Getenv("TOGETHER_API_KEY"), BaseURL: "https://api.together.xyz/v1", }) // Groq provider := openai.New(openai.Config{ APIKey: os.Getenv("GROQ_API_KEY"), BaseURL: "https://api.groq.com/openai/v1", }) // Ollama (local) provider := openai.New(openai.Config{ APIKey: "ollama", // Not used but required BaseURL: "http://localhost:11434/v1", }) ``` ## Model Selection ### By Use Case **General Purpose Chat** ```go // Best: GPT-6, Claude Opus 5.5, Gemini 3.5 Flash model, _ := provider.LanguageModel("gpt-6-astra") ``` **Code Generation** ```go // Best: GPT-6, Claude Sonnet 5.5, DeepSeek V4 model, _ := provider.LanguageModel("claude-sonnet-4.5") ``` **Reasoning Tasks** ```go // Best: o1, o3, Claude Opus 4.5, DeepSeek-R1 model, _ := provider.LanguageModel("o1") ``` **Fast Responses** ```go // Best: Groq, Cerebras, Gemini 2.0 Flash model, _ := provider.LanguageModel("llama-3.3-70b-versatile") ``` **Cost-Effective** ```go // Best: Gemini Flash, Claude Haiku, GPT-5.4 mini model, _ := provider.LanguageModel("gemini-2.0-flash") ``` **Long Context** ```go // Best: Gemini Pro (1M+), Claude Opus (200K) model, _ := provider.LanguageModel("gemini-2.5-pro") ``` ### By Capability **Tool Calling** - OpenAI GPT-6 - Anthropic Claude 4.5 - Google Gemini 2.0 - Mistral Large 2 - Cohere Command R+ **Structured Output** - OpenAI (native JSON mode) - Anthropic (guided generation) - Google Gemini - Mistral **Vision** - OpenAI GPT-6 - Anthropic Claude 4.5 - Google Gemini 2.0 Flash - Mistral Pixtral **Streaming** - All major providers support streaming - Groq and Cerebras offer fastest streaming ## Provider Capabilities Matrix ### Language Models | Provider | Streaming | Tools | Vision | JSON Mode | Reasoning | |----------|-----------|-------|--------|-----------|-----------| | OpenAI | Yes | Yes | Yes | Yes | Yes (o1/o3) | | Anthropic | Yes | Yes | Yes | Yes | Yes | | Google | Yes | Yes | Yes | Yes | Yes | | Mistral | Yes | Yes | Yes | Yes | No | | Cohere | Yes | Yes | No | Yes | No | | Groq | Yes | Yes | Yes | Yes | No | | DeepSeek | Yes | Yes | No | Yes | Yes (R1) | | xAI | Yes | Yes | Yes | Yes | No | ### Image Generation | Provider | Models | Speed | Quality | Upscaling | |----------|--------|-------|---------|-----------| | Stability AI | SDXL, SD3 | Fast | High | Yes | | Black Forest Labs | FLUX.1 | Medium | Excellent | Yes | | FAL | Multiple | Very Fast | High | Yes | | Replicate | Various | Variable | Variable | Yes | | Topaz Labs | Wonder 3.5 | Medium | High | Yes (enhancement only, no generation) | ### Video Generation | Provider | Models | Text-to-Video | Image-to-Video | Async Polling | |----------|--------|---------------|----------------|---------------| | KlingAI | Kling v2.6 | Yes | Yes | Yes | | Alibaba | Wan | Yes | Yes | Yes | | ByteDance | Seedance 1.x | Yes | Yes | Yes | | Topaz Labs | Proteus, Starlight Precise 2.6 | No | No | Yes | ### Speech & Audio | Provider | TTS | STT | Real-time | Languages | |----------|-----|-----|-----------|-----------| | ElevenLabs | Yes | No | Yes | 30+ | | Deepgram | No | Yes | Yes | 40+ | | AssemblyAI | No | Yes | No | 100+ | | OpenAI | Yes | Yes | No | 50+ | ## Error Handling ### Common Error Types ```go import "github.com/digitallysavvy/go-ai/pkg/ai" result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { switch providerErr.StatusCode { case 401: log.Fatal("Invalid API key") case 404: log.Fatal("Model not available") case 429: // Implement backoff time.Sleep(time.Second * 5) default: log.Printf("Provider error: %s", providerErr.Message) } } else { log.Printf("Unexpected error: %v", err) } } ``` ### Provider-Specific Errors ```go var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { log.Printf("%s error code: %s", providerErr.Provider, providerErr.ErrorCode) } ``` ### Retry Logic ```go func generateWithRetry(ctx context.Context, model provider.LanguageModel, prompt string, maxRetries int) (*ai.GenerateTextResult, error) { var result *ai.GenerateTextResult var err error for i := 0; i < maxRetries; i++ { result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err == nil { return result, nil } var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) && providerErr.StatusCode == 429 { backoff := time.Duration(math.Pow(2, float64(i))) * time.Second time.Sleep(backoff) continue } return nil, err // Non-retryable error } return nil, fmt.Errorf("max retries exceeded: %w", err) } ``` ## Multi-Provider Strategies ### Provider Fallback ```go type MultiProvider struct { providers []provider.Provider } func (m *MultiProvider) GenerateText(prompt string) (*ai.GenerateTextResult, error) { var lastErr error for _, provider := range m.providers { model, err := provider.LanguageModel("default-model") if err != nil { lastErr = err continue } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) if err == nil { return result, nil } lastErr = err log.Printf("Provider failed: %v, trying next", err) } return nil, fmt.Errorf("all providers failed: %w", lastErr) } ``` ### Load Balancing ```go type LoadBalancer struct { providers []provider.Provider current int mu sync.Mutex } func (lb *LoadBalancer) NextProvider() provider.Provider { lb.mu.Lock() defer lb.mu.Unlock() provider := lb.providers[lb.current] lb.current = (lb.current + 1) % len(lb.providers) return provider } ``` ### Cost Optimization ```go type CostOptimizer struct { cheapProvider provider.Provider premiumProvider provider.Provider } func (co *CostOptimizer) GenerateText(prompt string, priority string) (*ai.GenerateTextResult, error) { var provider provider.Provider if priority == "high" { provider = co.premiumProvider } else { provider = co.cheapProvider } model, err := provider.LanguageModel("default-model") if err != nil { return nil, err } return ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) } ``` ## Performance Optimization ### Connection Pooling ```go provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), HTTPClient: &http.Client{ Transport: &http.Transport{ MaxIdleConns: 100, MaxIdleConnsPerHost: 10, IdleConnTimeout: 90 * time.Second, }, Timeout: 60 * time.Second, }, }) ``` ### Request Batching ```go // For embedding providers type BatchEmbedder struct { model provider.EmbeddingModel batchSize int } func (be *BatchEmbedder) EmbedBatch(texts []string) ([][]float64, error) { var allEmbeddings [][]float64 for i := 0; i < len(texts); i += be.batchSize { end := min(i+be.batchSize, len(texts)) batch := texts[i:end] result, err := ai.EmbedMany(context.Background(), ai.EmbedManyOptions{ Model: be.model, Inputs: batch, }) if err != nil { return nil, err } allEmbeddings = append(allEmbeddings, result.Embeddings...) } return allEmbeddings, nil } ``` ### Caching ```go type CachedProvider struct { provider provider.Provider cache map[string]*ai.GenerateTextResult mu sync.RWMutex } func (cp *CachedProvider) GenerateText(ctx context.Context, model provider.LanguageModel, prompt string) (*ai.GenerateTextResult, error) { cacheKey := hash(prompt) cp.mu.RLock() if result, ok := cp.cache[cacheKey]; ok { cp.mu.RUnlock() return result, nil } cp.mu.RUnlock() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return nil, err } cp.mu.Lock() cp.cache[cacheKey] = result cp.mu.Unlock() return result, nil } ``` ## Monitoring & Observability ### Token Usage Tracking ```go type UsageTracker struct { totalTokens map[string]int64 mu sync.Mutex } func (ut *UsageTracker) TrackGeneration(provider string, result *ai.GenerateTextResult) { ut.mu.Lock() defer ut.mu.Unlock() ut.totalTokens[provider] += result.Usage.GetTotalTokens() } func (ut *UsageTracker) GetUsage(provider string) int64 { ut.mu.Lock() defer ut.mu.Unlock() return ut.totalTokens[provider] } ``` ### Request Logging ```go type LoggingProvider struct { provider provider.Provider logger *log.Logger } func (lp *LoggingProvider) GenerateText(ctx context.Context, model provider.LanguageModel, prompt string) (*ai.GenerateTextResult, error) { start := time.Now() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) duration := time.Since(start) lp.logger.Printf( "provider=%s duration=%v tokens=%d error=%v", "openai", duration, result.Usage.GetTotalTokens(), err, ) return result, err } ``` ## Best Practices ### 1. Configuration Management ```go // Use a config struct type AIConfig struct { OpenAIKey string AnthropicKey string DefaultModel string MaxRetries int Timeout time.Duration } func NewAIClient(cfg AIConfig) *AIClient { return &AIClient{ openai: openai.New(openai.Config{APIKey: cfg.OpenAIKey}), anthropic: anthropic.New(anthropic.Config{APIKey: cfg.AnthropicKey}), config: cfg, } } ``` ### 2. Error Handling Always handle provider-specific errors gracefully with fallbacks and retries. ### 3. Resource Management ```go // Clean up resources defer func() { if stream != nil { stream.Close() } }() ``` ### 4. Security - Never log API keys - Use environment variables for credentials - Implement rate limiting on client side - Validate user inputs before sending to providers ### 5. Testing ```go // Use mock providers for testing type MockProvider struct { responses []*ai.GenerateTextResult index int } func (m *MockProvider) LanguageModel(id string) (provider.LanguageModel, error) { return &MockLanguageModel{provider: m}, nil } ``` ## See Also - [Provider Index](https://goaisdk.com/docs/providers.md) - Complete provider list - [OpenAI Provider](https://goaisdk.com/docs/providers/openai.md) - OpenAI setup guide - [Anthropic Provider](https://goaisdk.com/docs/providers/anthropic.md) - Anthropic setup guide - [API Reference](https://goaisdk.com/docs/reference/ai/generate-text.md) - Complete API documentation --- # OpenAI Provider > Setup and usage guide for OpenAI models (GPT-6, GPT-5, o-series) with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/openai Documentation index: https://goaisdk.com/llms.txt OpenAI provides industry-leading language models including GPT-6, GPT-5, and reasoning models like o3 and o4-mini. Known for high-quality responses, extensive capabilities, and comprehensive API features. ## Setup ### Installation The OpenAI provider is included in the Go AI SDK: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) ``` ### Configuration ```go provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, err := provider.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } ``` `LanguageModel` matches the TypeScript OpenAI provider default and returns a Responses API model. Use `ChatModel` when you need the Chat Completions endpoint or `CompletionModel` for legacy instruct-style completions: ```go chatModel, err := provider.ChatModel("gpt-6-astra") completionModel, err := provider.CompletionModel("gpt-3.5-turbo-instruct") ``` ### Get API Key 1. Sign up at [platform.openai.com](https://platform.openai.com) 2. Navigate to API Keys section 3. Create new secret key 4. Set environment variable: ```bash export OPENAI_API_KEY=sk-... ``` ## Available Models Use the constants in `pkg/providers/openai/model_ids.go` instead of raw strings. For context windows and pricing, see [OpenAI's model documentation](https://platform.openai.com/docs/models). ### Language models | Model ID | Go constant | Notes | |----------|-------------|-------| | `gpt-6.1-sol` | `openai.ModelGPT61Sol` | GPT-6.1 series | | `gpt-6-astra` | `openai.ModelGPT6Astra` | GPT-6 series | | `gpt-6-luna` | `openai.ModelGPT6Luna` | GPT-6 series | | `gpt-6-sol` | `openai.ModelGPT6Sol` | GPT-6 series | | `gpt-5.6`, `gpt-5.6-luna`, `gpt-5.6-sol`, `gpt-5.6-terra` | `openai.ModelGPT56`, `ModelGPT56Luna`, `ModelGPT56Sol`, `ModelGPT56Terra` | GPT-5.6 series | | `gpt-5.5` | `openai.ModelGPT55` | Also pinned as `gpt-5.5-2026-04-23` | | `gpt-5.4`, `gpt-5.4-mini`, `gpt-5.4-nano`, `gpt-5.4-pro` | `openai.ModelGPT54`, `ModelGPT54Mini`, `ModelGPT54Nano`, `ModelGPT54Pro` | GPT-5.4 series | | `gpt-5.3-codex` | `openai.ModelGPT53Codex` | Coding | | `gpt-5.2`, `gpt-5.2-pro`, `gpt-5.2-codex` | `openai.ModelGPT52`, `ModelGPT52Pro`, `ModelGPT52Codex` | GPT-5.2 series | | `gpt-5.1`, `gpt-5.1-codex`, `gpt-5.1-codex-mini`, `gpt-5.1-codex-max` | `openai.ModelGPT51`, `ModelGPT51Codex`, `ModelGPT51CodexMini`, `ModelGPT51CodexMax` | GPT-5.1 series | | `gpt-5`, `gpt-5-mini`, `gpt-5-nano`, `gpt-5-pro` | `openai.ModelGPT5`, `ModelGPT5Mini`, `ModelGPT5Nano`, `ModelGPT5Pro` | GPT-5 series | | `o3`, `o4-mini` | `openai.ModelO3`, `openai.ModelO4Mini` | Reasoning models | Earlier models such as the GPT-4 and GPT-4o families, `o1`, `o3-mini` and GPT-3.5 still have constants. Prefer the models above for new code. ### Embedding models | Model ID | Go constant | |----------|-------------| | `text-embedding-3-large` | `openai.ModelTextEmbedding3Large` | | `text-embedding-3-small` | `openai.ModelTextEmbedding3Small` | | `text-embedding-ada-002` | `openai.ModelTextEmbeddingAda002` | ### Image generation models | Model ID | Go constant | |----------|-------------| | `gpt-image-2.5-flare` | `openai.ModelGPTImage25Flare` | | `gpt-image-2.5-sunburst` | `openai.ModelGPTImage25Sunburst` | | `gpt-image-2` | `openai.ModelGPTImage2` | | `gpt-image-1.5` | `openai.ModelGPTImage15` | | `gpt-image-1` | `openai.ModelGPTImage1` | | `gpt-image-1-mini` | `openai.ModelGPTImage1Mini` | | `dall-e-3`, `dall-e-2` | `openai.ModelDallE3`, `openai.ModelDallE2` | ### Speech and transcription models | Model ID | Type | Go constant | |----------|------|-------------| | `gpt-4o-mini-tts` | TTS | `openai.ModelGPT4oMiniTTS` | | `tts-1-hd` | TTS | `openai.ModelTTS1HD` | | `tts-1` | TTS | `openai.ModelTTS1` | | `gpt-4o-transcribe` | STT | `openai.ModelGPT4oTranscribe` | | `gpt-4o-mini-transcribe` | STT | `openai.ModelGPT4oMiniTranscribe` | | `whisper-1` | STT | `openai.ModelWhisper1` | ## Provider-Specific Features ### Transport and Headers OpenAI provider configuration supports the same transport flexibility as the TypeScript SDK: custom base URL, custom HTTP client, custom provider name, organization/project headers, and additional headers. ```go provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), BaseURL: "https://proxy.example.com/v1", Organization: "org_...", Project: "proj_...", Headers: map[string]string{ "X-Trace-ID": "trace-123", }, HTTPClient: http.DefaultClient, }) ``` Provider-level headers are forwarded to language, Responses, embeddings, image, speech, and transcription requests. Request-level headers override provider-level headers when both are present, matching TypeScript `headers` merge behavior. ### Structured Output (JSON Mode) OpenAI supports native JSON mode for reliable structured output: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/schema" ) personSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]string{"type": "string"}, "age": map[string]string{"type": "number"}, "email": map[string]string{"type": "string"}, }, "required": []string{"name", "age"}, }) result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: personSchema, Prompt: "Extract person info: John Doe, 30 years old, john@example.com", }) if err != nil { log.Fatal(err) } fmt.Printf("Structured data: %+v\n", result.Object) ``` JSON schemas are normalized for OpenAI's structured-outputs dialect before being sent: `propertyNames` and lookaround `pattern`s are stripped (with a `compatibility` warning), and a singleton-reference `allOf` wrapper (as recursive schemas commonly produce) is rewritten to a direct `$ref`, or inlined at the schema root, since OpenAI rejects `allOf` and requires an object at the root. ### Vision Capabilities Current GPT models support image understanding: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What's in this image?"}, types.FileContent{ URL: "https://example.com/image.jpg", MediaType: "image/jpeg", }, }, }, }, }) ``` For OpenAI Responses image/file parts, set the provider option `imageDetail` to forward the API `detail` field: ```go types.FileContent{ URL: "https://example.com/chart.png", MediaType: "image/png", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "imageDetail": "high", }, }, } ``` ### Function Calling Define tools that the model can call: ```go weatherTool := types.Tool{ Name: "get_weather", Description: "Get current weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]string{ "type": "string", "description": "City name", }, "unit": map[string]interface{}{ "type": "string", "enum": []string{"celsius", "fahrenheit"}, }, }, "required": []string{"location"}, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in San Francisco?", Tools: []types.Tool{weatherTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if result.ToolCalls != nil { for _, call := range result.ToolCalls { fmt.Printf("Tool: %s, Args: %v\n", call.ToolName, call.Arguments) } } ``` Responses API tool calls preserve OpenAI provider metadata, including `function_call.namespace`. Tool calls with no input serialize `{}` rather than `null`, and assistant messages serialize `content: null` only when tool calls are present and no text content is available. Non-tool assistant turns are not sent as null-content messages. Function tools can be grouped into OpenAI Responses namespaces by setting `providerOptions.openai.namespace` on the tool definition. Go serializes the namespace wrapper exactly as the TypeScript provider does and rejects conflicting namespace descriptions. When a reasoning effort is set for Responses models and no summary is provided, Go defaults `reasoning.summary` to `"detailed"`, matching the OpenAI provider default. For Responses API calls, use `ProviderOptions["openai"]["allowedTools"]` to restrict callable function tools while keeping the full `tools` array in the request. This preserves prompt caching behavior across different allowlists and overrides request-level `ToolChoice`, matching the TypeScript SDK. ```go result, err := model.DoGenerate(ctx, &provider.GenerateOptions{ Prompt: prompt, Tools: tools, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "allowedTools": map[string]interface{}{ "toolNames": []string{"weather"}, "mode": "auto", // optional: "auto" or "required" }, }, }, }) ``` ### Image Generation Options Use `OpenAIImageModelOptions` under the `"openai"` provider key for typed image options. The struct maps camelCase Go JSON names to the OpenAI wire fields used by the TypeScript SDK. ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "A studio product photo of a matte ceramic cup", ProviderOptions: map[string]interface{}{ "openai": openai.OpenAIImageModelOptions{ Size: "1024x1024", Quality: "high", OutputFormat: "png", OutputCompression: ptr(80), Background: "transparent", Moderation: "auto", }, }, }) ``` ### Reasoning Models (o1/o3) o1 and o3 models use extended thinking for complex problems: ```go // o1 uses internal reasoning tokens (not visible in response) model, err := provider.LanguageModel("o1") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Solve this complex math problem: Find the derivative of f(x) = x^3 * sin(x)", }) fmt.Println(result.Text) // Output includes detailed step-by-step reasoning ``` Note: o1 models do not support: - System messages - Streaming - Temperature settings - Function calling (use o4-mini for some tool support) ### Streaming Responses Stream text generation for real-time output: ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long story about AI", }) if err != nil { log.Fatal(err) } defer stream.Close() for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } if stream.Err() != nil { log.Fatal(stream.Err()) } ``` ### Response Format Control output format: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate a JSON object", ResponseFormat: &provider.ResponseFormat{ Type: "json_object", }, }) ``` ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, err := provider.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing in simple terms", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) fmt.Printf("Tokens used: %d\n", result.Usage.GetTotalTokens()) } ``` ### Chat Conversation ```go messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "How do I reverse a string in Go?"}, }, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a helpful coding assistant", Messages: messages, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) // Continue conversation messages = append(messages, types.Message{ Role: types.RoleAssistant, Content: []types.ContentPart{ types.TextContent{Text: result.Text}, }, }, types.Message{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Can you make it more efficient?"}, }, }, ) result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a helpful coding assistant", Messages: messages, }) ``` ### Image Generation with DALL-E ```go imageModel, err := provider.ImageModel("dall-e-3") if err != nil { log.Fatal(err) } result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "A futuristic city with flying cars", Size: "1024x1024", Quality: "hd", Style: "vivid", }) if err != nil { log.Fatal(err) } fmt.Printf("Image bytes: %d\n", len(result.Images[0].Data)) ``` `Quality` accepts `"standard"`/`"hd"` (DALL-E) or `"low"`/`"medium"`/`"high"`/`"xhigh"`/`"max"`/`"auto"` (GPT Image models) — `"xhigh"` and `"max"` are new this cycle. ### Text Embeddings ```go embeddingModel, err := provider.EmbeddingModel("text-embedding-3-large") if err != nil { log.Fatal(err) } texts := []string{ "The quick brown fox", "jumps over the lazy dog", } result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: texts, }) if err != nil { log.Fatal(err) } for i, embedding := range result.Embeddings { fmt.Printf("Text %d: %d dimensions\n", i, len(embedding)) } ``` ### Speech Synthesis ```go ttsModel, err := provider.SpeechModel("tts-1-hd") if err != nil { log.Fatal(err) } result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: ttsModel, Text: "Hello, welcome to Go AI SDK", Voice: "alloy", OutputFormat: "mp3", }) if err != nil { log.Fatal(err) } // Save audio file os.WriteFile("output.mp3", result.Audio.Data, 0644) ``` ### Transcription with Whisper ```go transcriptionModel, err := provider.TranscriptionModel("whisper-1") if err != nil { log.Fatal(err) } audioData, err := os.ReadFile("audio.mp3") if err != nil { log.Fatal(err) } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, MimeType: "audio/mpeg", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "language": "en", }, }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` ## Advanced Configuration ### Custom HTTP Client ```go import "net/http" provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), HTTPClient: &http.Client{ Timeout: time.Second * 120, Transport: &http.Transport{ MaxIdleConns: 10, IdleConnTimeout: 90 * time.Second, DisableCompression: false, }, }, }) ``` ### Custom Base URL (for proxies) ```go provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), BaseURL: "https://your-proxy.com/v1", }) ``` ### Organization and Project IDs ```go provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), Organization: "org-xxxxx", Project: "proj-xxxxx", }) ``` ## Error Handling ### Common Errors ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { switch providerErr.ErrorCode { case "invalid_api_key": log.Fatal("Invalid API key") case "model_not_found": log.Fatal("Model not found") case "context_length_exceeded": log.Fatal("Prompt too long") case "rate_limit_exceeded": log.Println("Rate limited, retrying...") time.Sleep(time.Second * 5) case "insufficient_quota": log.Fatal("Insufficient quota") default: log.Printf("OpenAI error: %s - %s", providerErr.ErrorCode, providerErr.Message) } } log.Fatal(err) } ``` ### Rate Limit Handling ```go import "github.com/digitallysavvy/go-ai/pkg/ai" func generateWithBackoff(ctx context.Context, model provider.LanguageModel, prompt string) (*ai.GenerateTextResult, error) { maxRetries := 3 baseDelay := time.Second for i := 0; i < maxRetries; i++ { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err == nil { return result, nil } var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) && providerErr.StatusCode == 429 { delay := baseDelay * time.Duration(math.Pow(2, float64(i))) log.Printf("Rate limited, waiting %v", delay) time.Sleep(delay) continue } return nil, err } return nil, fmt.Errorf("max retries exceeded") } ``` ## Best Practices 1. **Model Selection** - Use `gpt-6-astra` for general tasks with vision support - Use `gpt-5.4-mini` for cost-effective, fast responses - Use `o1` for complex reasoning and math problems - Use `gpt-5.4-nano` for simple, high-volume tasks 2. **Cost Optimization** - Monitor token usage with `result.Usage` - Use shorter system messages - Implement response caching for repeated queries - Use `max_tokens` to limit response length 3. **Performance** - Use streaming for long responses - Implement connection pooling for high-volume apps - Batch embedding requests (up to 2048 inputs) - Use `gpt-5.4-mini` for latency-sensitive applications 4. **Error Handling** - Always implement retry logic for rate limits - Handle context length errors gracefully - Log API errors for debugging - Implement fallback models 5. **Security** - Never expose API keys in client code - Use environment variables for credentials - Implement rate limiting on your API - Validate user inputs before sending ## Rate limits and pricing Rate limits depend on your usage tier. See [OpenAI's rate limits](https://platform.openai.com/docs/guides/rate-limits) and [pricing](https://openai.com/api/pricing/) pages for current values. ### Cost estimation Pricing changes often, so load rates from your own configuration instead of hard-coding them: ```go func estimateCost(promptTokens, completionTokens int, inputPerToken, outputPerToken float64) float64 { return float64(promptTokens)*inputPerToken + float64(completionTokens)*outputPerToken } ``` ## Realtime `Provider.RealtimeModel(modelID)` / `ExperimentalRealtimeModel` route known Live model IDs (currently `"gpt-live-1"`) to the OpenAI Live WebSocket and everything else (e.g. `"gpt-realtime"`) to the classic Realtime API: ```go import providerapi "github.com/digitallysavvy/go-ai/pkg/provider" model, err := provider.RealtimeModel("gpt-realtime") if err != nil { log.Fatal(err) } secret, err := provider.GetRealtimeToken(ctx, providerapi.RealtimeFactoryGetTokenOptions{ Model: "gpt-realtime", }) if err != nil { log.Fatal(err) } wsConfigProvider, ok := model.(providerapi.RealtimeWebSocketConfigProvider) if !ok { log.Fatal("model does not support client-side WebSocket config") } ws := wsConfigProvider.GetWebSocketConfig(secret.Token, secret.URL) ``` > **Fixed this cycle:** the WebSocket handshake's `Authorization` bearer > token is now extracted case-insensitively with any whitespace after the > scheme (`/^bearer\s+(.+)$/i`, TS parity). Previously `BEARER ` or a > tab after `Bearer` were not recognized and the connection silently > dropped the API key subprotocol. This also affects > `gpt-realtime-whisper` realtime transcription. ## Speech Translation (experimental) `Provider.SpeechTranslationModel(modelID)` (alias `Provider.Translation`) streams speech-to-speech translation over `/realtime/translations` (default model `"gpt-realtime-translate"`), for use with `ai.ExperimentalStreamTranslate`: ```go translationModel, err := provider.SpeechTranslationModel("gpt-realtime-translate") if err != nil { log.Fatal(err) } result, err := ai.ExperimentalStreamTranslate(ctx, ai.StreamTranslateOptions{ Model: translationModel, Audio: audioStream, TargetLanguage: "es", }) ``` ## Workflow Serialization OpenAI embedding, image, speech, and transcription models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel` (language models could already be serialized). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; realtime and speech-translation models are not yet serializable. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Core Concepts: Language Models](https://goaisdk.com/docs/foundations/providers-and-models.md) - [OpenAI API Documentation](https://platform.openai.com/docs/api-reference) - [Azure OpenAI Provider](https://goaisdk.com/docs/providers/azure.md) - For enterprise deployment ## May 2026 parity updates ### Model IDs The OpenAI provider includes GPT-5.5 chat model constants and `gpt-image-2` image generation support. | Capability | Go constant | Model ID | | --- | --- | --- | | Chat | `openai.ModelGPT55` | `gpt-5.5` | | Chat | `openai.ModelGPT552026_04_23` | `gpt-5.5-2026-04-23` | | Image | `openai.ModelGPTImage2` | `gpt-image-2` | ```go p := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), Headers: map[string]string{ "OpenAI-Beta": "example", }, }) model, err := p.LanguageModel(openai.ModelGPT55) ``` ### Responses API metadata `LanguageModel` and `ResponsesModel` both use OpenAI Responses API behavior. The provider preserves response IDs, headers, provider metadata, encrypted reasoning, function-call namespace values, and generated files in the standard result fields. ```go model, err := p.ResponsesModel(openai.ModelGPT55) ``` Continue an OpenAI Conversation by passing the conversation ID through provider options. This is distinct from `previousResponseId`; if both are set, the request forwards both fields and returns an unsupported warning matching the TypeScript SDK. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: messages, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "conversation": "conv_123", }, }, }) ``` ### Image detail and file content Image detail provider options are forwarded on image parts, including tool-result image content. Prefer `types.FileContent` with tagged `types.FileData` for new multimodal code. `types.ImageContent` remains available for compatibility. For OpenAI Responses compatibility, string file data with the default `file-` prefix is sent as a `file_id` reference instead of inline base64 data, matching the TypeScript provider. To disable or customize this deprecated compatibility path, set `openai.Config.FileIDPrefixes`. ```go messages := []types.Message{{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Describe the image."}, types.FileContent{ FileData: types.FileData{ Type: types.FileDataTypeURL, URL: "https://example.com/chart.png", MediaType: "image/png", }, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{"imageDetail": "high"}, }, }, }, }} ``` ### Assistant null content behavior Assistant messages with tool calls serialize `content: null` only when the target OpenAI-compatible API requires it for tool-call assistant turns. Plain assistant messages keep their text content. ## October 2026 parity updates ### Model IDs GPT-6.1 Sol is available as a Chat Completions and Responses model ID. | Capability | Go constant | Model ID | | --- | --- | --- | | Chat / Responses | `openai.ModelGPT61Sol` | `gpt-6.1-sol` | ```go model, err := p.LanguageModel(openai.ModelGPT61Sol) ``` ### `reasoningSummary` warning on Chat Completions models `providerOptions.openai.reasoningSummary` is a Responses API option. Passing it to a Chat Completions model (`p.LanguageModel("o3")`) now returns an `unsupported` warning instead of being silently dropped: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, // a Chat Completions model, e.g. p.LanguageModel("o3") Prompt: "Explain quantum tunnelling.", ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{"reasoningSummary": "detailed"}, }, }) // result.Warnings contains {Type: "unsupported", Feature: "reasoningSummary", ...} ``` Use `p.ResponsesModel(...)` instead to enable reasoning summaries. --- # Anthropic Provider > Setup and usage guide for Anthropic Claude models (Sonnet, Opus, Haiku) with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/anthropic Documentation index: https://goaisdk.com/llms.txt Anthropic provides the Claude family of models, known for their helpfulness, honesty, and harmlessness. Claude excels at analysis, coding, creative writing, and extended conversations with industry-leading context windows. ## Setup ### Installation The Anthropic provider is included in the Go AI SDK: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) ``` ### Configuration ```go provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, err := provider.LanguageModel("claude-sonnet-4-5") if err != nil { log.Fatal(err) } ``` ### Get API Key 1. Sign up at [console.anthropic.com](https://console.anthropic.com) 2. Navigate to API Keys section 3. Create new API key 4. Set environment variable: ```bash export ANTHROPIC_API_KEY=sk-ant-... ``` ## Claude Platform on AWS Use `pkg/providers/anthropicaws` when calling Claude Platform through AWS. The provider reuses the Anthropic language model, Files API, Skills API, and Anthropic tool factories, but signs requests for AWS or sends the Anthropic AWS API key headers. ### API Key Authentication ```go provider, err := anthropicaws.New(anthropicaws.Config{ Region: os.Getenv("AWS_REGION"), WorkspaceID: os.Getenv("ANTHROPIC_AWS_WORKSPACE_ID"), APIKey: os.Getenv("ANTHROPIC_AWS_API_KEY"), }) if err != nil { log.Fatal(err) } model, err := provider.LanguageModel(anthropic.ClaudeOpus4_8) ``` ### SigV4 Authentication When `APIKey` and `ANTHROPIC_AWS_API_KEY` are empty, the provider signs POST requests with SigV4. Credentials may come from `Config`, environment variables, or `CredentialProvider`. ```go provider, err := anthropicaws.New(anthropicaws.Config{ Region: os.Getenv("AWS_REGION"), WorkspaceID: os.Getenv("ANTHROPIC_AWS_WORKSPACE_ID"), AccessKeyID: os.Getenv("AWS_ACCESS_KEY_ID"), SecretAccessKey: os.Getenv("AWS_SECRET_ACCESS_KEY"), SessionToken: os.Getenv("AWS_SESSION_TOKEN"), }) ``` The default base URL is `https://aws-external-anthropic.{region}.api.aws/v1`. `WorkspaceID` is required and is sent as `anthropic-workspace-id`. ## Available Models Use the constants in `pkg/providers/anthropic/model_ids.go` instead of raw strings. The list below is the current set. For context windows and pricing, see [Anthropic's model overview](https://docs.anthropic.com/en/docs/about-claude/models/overview). ### Current models | Model ID | Go constant | Notes | |----------|-------------|-------| | `claude-opus-5-5` | `anthropic.ClaudeOpus5_5` | Flagship, always-adaptive thinking | | `claude-opus-5` | `anthropic.ClaudeOpus5` | Supports `fallbacks: 'default'` | | `claude-sonnet-5-5` | `anthropic.ClaudeSonnet5_5` | Between-tools thinking; rejects disabled thinking and forced tool use | | `claude-sonnet-5` | `anthropic.ClaudeSonnet5` | Current-generation Sonnet | | `claude-fable-5-1` | `anthropic.ClaudeFable5_1` | Always-adaptive thinking, system-message `clearAt` and effort | | `claude-fable-5` | `anthropic.ClaudeFable5` | Server-side fallbacks | | `claude-opus-4-8` | `anthropic.ClaudeOpus4_8` | Native structured output | | `claude-opus-4-7` | `anthropic.ClaudeOpus4_7` | `xhigh` effort | | `claude-opus-4-6` | `anthropic.ClaudeOpus4_6` | Adaptive thinking, fast mode | | `claude-sonnet-4-6` | `anthropic.ClaudeSonnet4_6` | Balanced performance and capability | | `claude-haiku-4-5` | `anthropic.ClaudeHaiku4_5` | Fast and cost-effective | ### Earlier generations These IDs still have constants. Prefer the current models above for new code. | Model ID | Go constant | |----------|-------------| | `claude-opus-4-5` | `anthropic.ClaudeOpus4_5` | | `claude-sonnet-4-5` | `anthropic.ClaudeSonnet4_5` | | `claude-opus-4-1` | `anthropic.ClaudeOpus4_1` | | `claude-opus-4-0` | `anthropic.ClaudeOpus4_0` | | `claude-sonnet-4-0` | `anthropic.ClaudeSonnet4_0` | Dated variants such as `claude-sonnet-4-5-20250929` and `claude-haiku-4-5-20251001` also have constants. The Claude 3 and 3.5 models are deprecated by Anthropic. ## Provider-Specific Features ### Inference Region Set `InferenceGeo` on Anthropic model options to forward the TypeScript SDK `inferenceGeo` provider option as `inference_geo`. Supported values are `"us"` and `"global"`. `anthropic.ClaudeFable5` is available for Claude Fable 5. `ModelOptions.Fallbacks` forwards Anthropic's server-side fallback configuration and enables the `server-side-fallback-2026-06-01` beta header. Fallback response blocks are consumed like the TypeScript provider and are not returned as content; fallback activity is visible through the raw usage `iterations` metadata. ```go model := anthropic.NewLanguageModel(provider, anthropic.ClaudeOpus4_7, &anthropic.ModelOptions{ InferenceGeo: "us", }) ``` Anthropic JSON schema serialization sanitizes unsupported validation keywords before sending requests, matching the TypeScript provider behavior. This keeps structured output/tool schemas compatible with Anthropic without requiring callers to manually strip incompatible fields. ### Extended Context (200K tokens) Claude models support up to 200K token context windows: ```go // Process entire codebases or documents largeDocument := readLargeFile("document.txt") // up to 200K tokens result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Analyze this document and provide insights:\n\n%s", largeDocument), }) ``` ### Vision Capabilities Claude 4.5 and 3.5 models support image understanding: ```go // Read image imageData, err := os.ReadFile("chart.png") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Analyze this chart"}, types.FileContent{ Data: imageData, MediaType: "image/png", }, }, }, }, }) ``` ### Tool Use (Function Calling) Claude excels at tool use with precise parameter extraction: ```go searchTool := types.Tool{ Name: "search_database", Description: "Search the customer database", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{ "type": "string", "description": "Search query", }, "filters": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "status": map[string]interface{}{ "type": "string", "enum": []string{"active", "inactive"}, }, "created_after": map[string]string{ "type": "string", }, }, }, }, "required": []string{"query"}, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Find all active customers who signed up in 2024", Tools: []types.Tool{searchTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) for _, call := range result.ToolCalls { fmt.Printf("Tool: %s\nArgs: %v\n", call.ToolName, call.Arguments) } ``` ### Prompt Caching Reduce costs by caching frequently used prompt segments: ```go // Cache system prompt and context result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are an expert Go programmer...", Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Review this code: " + largeCodebase}, }, }, { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "How can I improve error handling?"}, }, }, }, ProviderOptions: map[string]interface{}{ "anthropic": anthropic.ModelOptions{AutomaticCaching: true}, }, }) // Subsequent requests with same cached content cost significantly less ``` Cache writes cost more than normal input tokens and cache reads cost much less. See [Anthropic's pricing page](https://www.anthropic.com/pricing) for current multipliers. Cached content is valid for 5 minutes by default. ### Thinking (Extended Reasoning) Claude 4.5 supports extended thinking for complex problems: ```go thinkingModel, err := provider.LanguageModelWithOptions("claude-sonnet-4-5", &anthropic.ModelOptions{ Thinking: &anthropic.ThinkingConfig{Type: anthropic.ThinkingTypeAdaptive}, }) if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: thinkingModel, Prompt: "Solve this complex optimization problem...", }) // Response includes thinking process as ordered reasoning content parts for _, part := range result.Content { if reasoning, ok := part.(types.ReasoningContent); ok { fmt.Println("Claude's reasoning:", reasoning.Text) } } fmt.Println("Final answer:", result.Text) ``` When streaming (`ai.StreamText`) extended thinking together with tool calls across multiple steps, each thinking block's cryptographic signature streams as its own `signature_delta` event, after the block's text deltas. The provider carries it through as the reasoning content part's `Signature`, so a later step's request replays the thinking block with its signature intact (required for Anthropic to accept the replayed block) instead of dropping it. #### Between-Tools Thinking (Claude Sonnet 5.5) `anthropic.ClaudeSonnet5_5` rejects `thinking.type` `"disabled"` and budget-based thinking, and adds `anthropic.ThinkingTypeBetweenTools`, its lowest thinking setting: no upfront thinking, but short progress notes between tool calls stream back as summarized thinking blocks. It's only accepted at `"low"`, `"medium"`, and `"high"` effort — `"xhigh"`/`"max"` is lowered to `"high"` with a warning. `Reasoning: types.ReasoningNone` and `ThinkingConfig{Type: anthropic.ThinkingTypeDisabled}` both map to `between_tools` automatically for this model, with a warning in the latter case. ```go model, err := provider.LanguageModelWithOptions(anthropic.ClaudeSonnet5_5, &anthropic.ModelOptions{ Thinking: &anthropic.ThinkingConfig{Type: anthropic.ThinkingTypeBetweenTools}, Effort: anthropic.EffortMedium, }) ``` ### Context Management (Beta) Automatically manage conversation history to prevent context window overflows. Claude can intelligently clear old tool calls, thinking blocks, or compact entire conversations: ```go import "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" // Create model with context management model, err := provider.LanguageModelWithOptions("claude-sonnet-4-5", &anthropic.ModelOptions{ ContextManagement: &anthropic.ContextManagement{ Edits: []anthropic.ContextManagementEdit{ // Clear old tool calls after 10,000 input tokens anthropic.NewClearToolUsesEdit(). WithInputTokensTrigger(10000). WithKeepToolUses(3), // Clear thinking blocks, keeping last 2 turns anthropic.NewClearThinkingEdit(). WithKeepRecentTurns(2), // Compact conversation at 50,000 tokens anthropic.NewCompactEdit(). WithTrigger(50000), }, }, }) // Use in long conversations - context automatically managed result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: longConversation, }) // Check if context management was applied if result.ContextManagement != nil { cmr := result.ContextManagement.(*anthropic.ContextManagementResponse) for _, edit := range cmr.AppliedEdits { switch e := edit.(type) { case *anthropic.AppliedClearToolUsesEdit: fmt.Printf("Cleared %d tool uses, freed %d tokens\n", e.ClearedToolUses, e.ClearedInputTokens) case *anthropic.AppliedClearThinkingEdit: fmt.Printf("Cleared %d thinking turns, freed %d tokens\n", e.ClearedThinkingTurns, e.ClearedInputTokens) case *anthropic.AppliedCompactEdit: fmt.Println("Conversation compacted") } } } ``` **Available Edit Types:** 1. **ClearToolUses** - Removes old tool call details - Configure trigger thresholds (token count or tool use count) - Keep recent tool uses - Exclude specific tools - Control minimum tokens to clear 2. **ClearThinking** - Removes extended thinking blocks - Keep all thinking or specify number of recent turns - Preserves final conclusions 3. **Compact** - Summarizes entire conversation - Most aggressive space-saving - Optional pause after compaction - Custom compaction instructions **Beta Headers Required:** - `context-management-2025-06-27` for clear_tool_uses and clear_thinking - `compact-2026-01-12` for compact edits (Headers are automatically added by the SDK) ### Streaming Stream responses for real-time output: ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a detailed essay about AI safety"}) if err != nil { log.Fatal(err) } defer stream.Close() for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } if stream.Err() != nil { log.Fatal(stream.Err()) } ``` ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, err := provider.LanguageModel("claude-sonnet-4-5") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain the Go concurrency model with examples", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) fmt.Printf("\nTokens - Input: %d, Output: %d\n", result.Usage.GetInputTokens(), result.Usage.GetOutputTokens()) } ``` ### Multi-Turn Conversation ```go messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Review this function:\n\n" + code}, }, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a helpful code review assistant", Messages: messages, }) if err != nil { log.Fatal(err) } fmt.Println("Review:", result.Text) // Continue conversation messages = append(messages, types.Message{ Role: types.RoleAssistant, Content: []types.ContentPart{ types.TextContent{Text: result.Text}, }, }, types.Message{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Can you suggest specific improvements?"}, }, }, ) result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a helpful code review assistant", Messages: messages, }) ``` ### Document Analysis with Vision ```go // Analyze a PDF page rendered as image imageData, err := os.ReadFile("invoice.png") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Extract all line items from this invoice with prices"}, types.FileContent{ Data: imageData, MediaType: "image/png", }, }, }, }, ResponseFormat: &provider.ResponseFormat{Type: "json_object"}, }) fmt.Println("Extracted data:", result.Text) ``` ### Code Generation with Tool Use ```go fileReadTool := types.Tool{ Name: "read_file", Description: "Read contents of a file", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "path": map[string]string{ "type": "string", "description": "File path", }, }, "required": []string{"path"}, }, } fileWriteTool := types.Tool{ Name: "write_file", Description: "Write content to a file", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "path": map[string]string{"type": "string"}, "content": map[string]string{"type": "string"}, }, "required": []string{"path", "content"}, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Read main.go, add error handling, and save the improved version", Tools: []types.Tool{fileReadTool, fileWriteTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) // Process tool calls for _, call := range result.ToolCalls { switch call.ToolName { case "read_file": // Execute file read case "write_file": // Execute file write } } ``` ### Large Document Processing with Caching ```go // Load large codebase codebase := loadCodebase("./src") // ~100K tokens messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Here is the codebase:\n\n" + codebase}, // Cached }, }, } // First query - caches codebase result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are an expert code reviewer", Messages: append(messages, types.Message{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "List all exported functions"}, }, }), ProviderOptions: map[string]interface{}{ "anthropic": anthropic.ModelOptions{AutomaticCaching: true}, }, }) // Second query - uses cached codebase (90% cost reduction) result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are an expert code reviewer", Messages: append(messages, types.Message{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Find potential security issues"}, }, }), ProviderOptions: map[string]interface{}{ "anthropic": anthropic.ModelOptions{AutomaticCaching: true}, }, }) ``` ## Advanced Configuration ### Custom HTTP Client ```go provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), HTTPClient: &http.Client{ Timeout: time.Minute * 5, Transport: &http.Transport{ MaxIdleConns: 100, MaxIdleConnsPerHost: 10, IdleConnTimeout: 90 * time.Second, }, }, }) ``` ### Custom Base URL > **Breaking change:** the default base URL now includes `/v1` (TS parity). > A bare `https://api.anthropic.com` is normalized automatically, but a > custom `BaseURL` (proxy, test server) must include `/v1` itself. ```go provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), BaseURL: "https://your-proxy.com/v1", }) ``` ### Beta Features Most beta headers (prompt caching, extended thinking, context management, and so on) are added automatically based on the `ModelOptions` you set. To send an additional `anthropic-beta` flag that isn't covered by a dedicated option, set `AnthropicBeta` on `ModelOptions`: ```go provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, err := provider.LanguageModelWithOptions("claude-sonnet-4-5", &anthropic.ModelOptions{ AnthropicBeta: []string{"some-other-beta-flag"}, }) ``` ## Error Handling ### Common Errors ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { switch providerErr.ErrorCode { case "invalid_request_error": log.Printf("Invalid request: %s", providerErr.Message) case "authentication_error": log.Fatal("Invalid API key") case "permission_error": log.Fatal("Permission denied") case "not_found_error": log.Fatal("Model not found") case "rate_limit_error": log.Printf("Rate limited: %s", providerErr.Message) time.Sleep(time.Second * 5) case "overloaded_error": log.Println("Service overloaded, retrying...") time.Sleep(time.Second * 10) default: log.Printf("Anthropic error: %s", providerErr.Message) } } return nil, err } ``` ### Retry with Exponential Backoff ```go func generateWithRetry(ctx context.Context, model provider.LanguageModel, prompt string) (*ai.GenerateTextResult, error) { maxRetries := 5 baseDelay := time.Second for attempt := 0; attempt < maxRetries; attempt++ { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err == nil { return result, nil } var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { if providerErr.ErrorCode == "rate_limit_error" || providerErr.ErrorCode == "overloaded_error" { delay := baseDelay * time.Duration(math.Pow(2, float64(attempt))) log.Printf("Attempt %d failed, waiting %v", attempt+1, delay) time.Sleep(delay) continue } } return nil, err } return nil, fmt.Errorf("max retries exceeded") } ``` ## Best Practices 1. **Model Selection** - Use `claude-sonnet-4-5` for best balance of quality and speed - Use `claude-opus-4-5` for most complex tasks requiring deep analysis - Use `claude-haiku-4-5` for high-volume, latency-sensitive applications 2. **Prompt Engineering** - Be specific and detailed in instructions - Use XML tags for clear structure: ``, ``, `` - Provide examples when possible - Place most important information at the start or end 3. **Cost Optimization** - Enable prompt caching for repeated context (large documents, system prompts) - Use Haiku for simple tasks - Monitor token usage with `result.Usage` - Cache responses on your side for identical queries 4. **Context Management** - Take advantage of 200K context window - Include full context rather than summarizing - Use structured formats (JSON, XML) for large data 5. **Tool Use** - Provide clear, detailed tool descriptions - Use precise parameter schemas - Handle tool results in follow-up messages - Chain multiple tool calls for complex workflows ## Rate Limits & Pricing ### Rate limits Rate limits depend on your account tier and change over time. See [Anthropic's rate limits page](https://docs.anthropic.com/en/api/rate-limits) for current values. ### Token Counting ```go // Estimate tokens (approximate: 1 token ≈ 4 characters) func estimateTokens(text string) int { return len(text) / 4 } // Calculate cost func calculateCost(usage types.Usage, model string) float64 { // Example rates only. Check Anthropic's pricing page for current prices. rates := map[string][2]float64{ "claude-opus-4-5": {15.00 / 1_000_000, 75.00 / 1_000_000}, "claude-sonnet-4-5": {3.00 / 1_000_000, 15.00 / 1_000_000}, "claude-haiku-4-5": {0.80 / 1_000_000, 4.00 / 1_000_000}, } rate := rates[model] inputCost := float64(usage.GetInputTokens()) * rate[0] outputCost := float64(usage.GetOutputTokens()) * rate[1] return inputCost + outputCost } ``` ### Batch API Anthropic implements the batch API (Message Batches) used by `ai.ExperimentalStartBatch` / `ExperimentalGetBatchStatus` / `ExperimentalGetBatchResults` — process large volumes at a discount with async turnaround instead of `ai.GenerateText` per request: ```go started, err := ai.ExperimentalStartBatch(ctx, ai.StartBatchOptions{ Provider: provider, // an *anthropic.Provider (implements provider.BatchProvider) Requests: []ai.BatchRequest{ {Text: &ai.BatchTextRequest{ ID: "req-1", Model: "claude-sonnet-4-6", Prompt: "Summarize the Go concurrency model.", }}, }, }) if err != nil { log.Fatal(err) } status, err := ai.ExperimentalGetBatchStatus(ctx, ai.GetBatchStatusOptions{ Provider: provider, Batch: started.BatchReference, }) ``` ## Forwarding Code-Execution Containers Across Steps `anthropic.ForwardContainerIDFromLastStep(steps)` builds `ProviderOptions` that reuse the code-execution container from the most recent step that had one, so a multi-step agent loop doesn't pay to spin up a fresh sandbox on every step. Use it from `PrepareStep`: ```go // tools includes anthropicTools.CodeExecution20260120() — see // docs/providers/anthropic-advanced-features.md#code-execution-tool-2026-01-20 // for the full code execution tool setup. result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Run some Python to analyze this data, then summarize.", Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, PrepareStep: func(ctx context.Context, step ai.PrepareStepOptions) ai.PrepareStepOptions { if opts := anthropic.ForwardContainerIDFromLastStep(step.Steps); opts != nil { step.ProviderOptions = opts } return step }, }) ``` ### Pruned Programmatic Tool History If history is pruned between steps (e.g. to control context size) and that removes a code-execution source call (`code_execution_20250825`/`code_execution_20260120`) while a dependent tool call still references it via `caller`, the provider now omits that dangling `caller` metadata — rather than sending it as-is and having Anthropic reject the request with "source tool ... not found" — and returns one warning per affected tool call. The call, its result, ordering, and cache controls are otherwise unchanged; a caller that still resolves to an emitted `server_tool_use` block, or that is in an active continuation not yet followed by a new user message, is left untouched. ## Finish Reasons `result.FinishReason` and `result.RawFinishReason` (and their `StreamChunk` equivalents) report Anthropic's raw `stop_reason` alongside the normalized value. Since v0.5.0, three additional raw reasons map to non-`"other"` finish reasons, matching TS: | Anthropic `stop_reason` | `FinishReason` | |--------------------------|-----------------| | `end_turn`, `stop_sequence` | `stop` | | `pause_turn` | `stop` | | `refusal` | `content-filter` | | `model_context_window_exceeded` | `length` | | `max_tokens` | `length` | | `tool_use` | `tool-calls` | `RawFinishReason` always carries Anthropic's original string (e.g. `"pause_turn"`), so you can distinguish these cases from a plain `end_turn` even though they share the same normalized `FinishReason`. ## Workflow Serialization Anthropic language models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. Anthropic has no embedding, image, speech, transcription, or video models, so there is nothing else to serialize. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Core Concepts: Language Models](https://goaisdk.com/docs/foundations/providers-and-models.md) - [Anthropic API Documentation](https://docs.anthropic.com/claude/reference) - [Prompt Engineering Guide](https://docs.anthropic.com/claude/docs/prompt-engineering) ## May 2026 parity updates ### Claude Opus 4.7 Use `anthropic.ClaudeOpus4_7` for Claude Opus 4.7. The provider maps top-level `Reasoning` to Anthropic thinking settings and supports x-high effort on the newest Opus models. ```go p := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}) model, err := p.LanguageModel(anthropic.ClaudeOpus4_7) ``` ### Inference geography and task budget `ModelOptions` includes `InferenceGeo` and `TaskBudget`, matching the TypeScript provider options. ```go remaining := 12000 model, err := p.LanguageModelWithOptions(anthropic.ClaudeOpus4_7, &anthropic.ModelOptions{ InferenceGeo: "us", TaskBudget: &anthropic.TaskBudget{ Type: "tokens", Total: 32000, Remaining: &remaining, }, }) ``` ### Schema and beta header behavior The provider sanitizes unsupported JSON Schema validation properties before sending tools or structured output requests to Anthropic. The obsolete fine-grained tool streaming beta header is no longer injected. --- # Google Provider > Setup and usage guide for Google's Gemini models in the Go AI SDK, covering available models, multimodal features, video generation, and configuration. Canonical URL: https://goaisdk.com/docs/providers/google Documentation index: https://goaisdk.com/llms.txt Google provides the Gemini family of models through Google AI Studio. Gemini models offer exceptional multimodal capabilities, massive context windows (up to 1M+ tokens), and competitive pricing. ## Setup ### Installation The Google provider is included in the Go AI SDK: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" providerapi "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/google" ) ``` ### Configuration ```go googleProvider := google.New(google.Config{ APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"), }) model, err := googleProvider.LanguageModel("gemini-2.0-flash") if err != nil { log.Fatal(err) } ``` ### Get API Key 1. Visit [makersuite.google.com/app/apikey](https://makersuite.google.com/app/apikey) 2. Create new API key 3. Set environment variable: ```bash export GOOGLE_GENERATIVE_AI_API_KEY=AI... ``` ## Available Models Use the constants in `pkg/providers/google/model_ids.go` instead of raw strings. For context windows and pricing, see [Google's model documentation](https://ai.google.dev/gemini-api/docs/models). ### Language models | Model ID | Go constant | Best for | |----------|-------------|----------| | `gemini-3.8-flash` | `google.ModelGemini38Flash` | Fast, current-generation Flash | | `gemini-3.7-flash` | `google.ModelGemini37Flash` | Fast, current-generation Flash | | `gemini-3.6-flash` | `google.ModelGemini36Flash` | Fast, current-generation Flash | | `gemini-3.5-flash` | `google.ModelGemini35Flash` | Fast multimodal tasks | | `gemini-3.5-flash-lite` | `google.ModelGemini35FlashLite` | High-volume, low-cost tasks | | `gemini-3.1-pro-preview` | `google.ModelGemini31ProPreview` | Advanced reasoning and generation | | `gemini-3.1-pro-preview-customtools` | `google.ModelGemini31ProPreviewCustom` | Advanced generation with custom tools | | `gemini-3.1-flash-lite-preview` | `google.ModelGemini31FlashLitePreview` | Low-cost preview model | | `gemini-3-pro-preview` | `google.ModelGemini3ProPreview` | Advanced preview generation | | `gemini-3-flash-preview` | `google.ModelGemini3FlashPreview` | Fast preview generation | | `gemini-2.5-pro` | `google.ModelGemini25Pro` | Complex reasoning and multimodal tasks | | `gemini-2.5-flash` | `google.ModelGemini25Flash` | Fast multimodal tasks | | `gemini-2.5-flash-lite` | `google.ModelGemini25FlashLite` | Cost-effective high-volume tasks | | `gemini-pro-latest`, `gemini-flash-latest`, `gemini-flash-lite-latest` | `google.ModelGeminiProLatest`, `ModelGeminiFlashLatest`, `ModelGeminiFlashLiteLatest` | Latest aliases | Earlier Gemini 2.0, 1.5 and Gemma models still have constants. Prefer the models above for new code. ### Image Generation Models > **Breaking change:** non-Gemini Imagen models (`imagen-4.0-*`, etc.) were > removed from the Google provider, matching the TypeScript SDK. Calling > `provider.ImageModel(id)` with a model ID that doesn't start with > `gemini-` now fails with: "Google image models other than Gemini are no > longer supported. Use a model ID that starts with `gemini-`." Use Google > Vertex AI if you still need Imagen. | Model ID | Go constant | Best for | |----------|-------------|----------| | `gemini-3.1-flash-image-preview` | `google.ModelGemini31FlashImagePreview` | Fast Gemini generation | | `gemini-3-pro-image-preview` | `google.ModelGemini3ProImagePreview` | Advanced generation | | `gemini-2.5-flash-image` | `google.ModelGemini25FlashImage` | Fast Gemini generation | Gemini image models ignore `N` (they generate one image per call and are auto-batched by `ai.GenerateImage` to reach the requested count) instead of erroring on `N > 1`. Gemini image models can use Google Search grounding during image generation through provider options: ```go result, err := imageModel.DoGenerate(ctx, &provider.ImageGenerateOptions{ Prompt: "create a current-events infographic", ProviderOptions: map[string]interface{}{ "google": map[string]interface{}{ "googleSearch": map[string]interface{}{}, }, }, }) ``` `googleSearch` is only supported for Gemini image models — the only image models this provider exposes now that Imagen has been removed. ### Embedding Models | Model ID | Go constant | Dimensions | Best for | |----------|-------------|------------|----------| | `gemini-embedding-001` | `google.EmbeddingModelGeminiEmbedding001` | 3072 | Multimodal semantic search | | `gemini-embedding-2` | `google.EmbeddingModelGeminiEmbedding2` | 3072 | Stable Gemini embedding workloads | | `gemini-embedding-2-preview` | `google.EmbeddingModelGeminiEmbedding2Preview` | Provider-defined | Preview embedding workloads | ### Gemini TTS Speech Models Google Gemini TTS is exposed through `Provider.SpeechModel` and `Provider.Speech`. | Model ID | Go constant | |----------|-------------| | gemini-2.5-flash-preview-tts | `google.ModelGemini25FlashTTS` | | gemini-2.5-pro-preview-tts | `google.ModelGemini25ProTTS` | | gemini-3.1-flash-tts-preview | `google.ModelGemini31FlashTTSPreview` | ```go speechModel, err := googleProvider.SpeechModel(google.ModelGemini25FlashTTS) if err != nil { log.Fatal(err) } result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Welcome to Go-AI.", Voice: "Kore", }) if err != nil { log.Fatal(err) } fmt.Println(result.Audio.MediaType) ``` Gemini TTS defaults to the `Kore` voice when no voice is supplied. Use `ProviderOptions["google"]` with `google.GoogleSpeechModelOptions` to pass Gemini-specific speech options such as multi-speaker voice configuration. ### Realtime Models Gemini Live is exposed through `Provider.RealtimeModel`, `ExperimentalRealtimeModel`, and `GetRealtimeToken`. Model IDs are strings, matching the TypeScript `GoogleRealtimeModelId` surface. ```go model, err := googleProvider.RealtimeModel("gemini-3.1-flash-live-preview") if err != nil { log.Fatal(err) } secretCreator, ok := model.(providerapi.RealtimeClientSecretCreator) if !ok { log.Fatal("model does not support client secret creation") } token, err := secretCreator.DoCreateClientSecret(ctx, providerapi.ClientSecretOptions{}) if err != nil { log.Fatal(err) } wsConfigProvider, ok := model.(providerapi.RealtimeWebSocketConfigProvider) if !ok { log.Fatal("model does not support client-side WebSocket config") } ws := wsConfigProvider.GetWebSocketConfig(token.Token, token.URL) fmt.Println(ws.URL) ``` Use `ai.ConnectRealtime` for a provider-neutral Go session helper. Browser-only TypeScript pieces such as `BrowserRealtimeTransport`, `BrowserRealtimeAudio`, and `useRealtime` are not part of the Go runtime. Gemini Live Translate configuration can be passed through `ProviderOptions["google"]["translationConfig"]`. The provider merges it into the realtime session `generationConfig.translationConfig`, matching the TypeScript Live Translate request shape. Background-reasoning Live models (model IDs matching `gemini-.-live...thinking`, e.g. `gemini-3.8-live-extended-thinking`) require a `thinkingLevel` or `thinkingBudget` in the session setup. When neither is set — via the typed `ProviderOptions["google"]["thinkingConfig"]` or a raw `ProviderOptions["generationConfig"]["thinkingConfig"]` — the provider adds `thinkingLevel: "low"` while keeping any other fields (e.g. `includeThoughts`) and without disturbing an explicit `thinkingBudget: 0`: ```go setup := map[string]interface{}{ "providerOptions": map[string]interface{}{ "google": map[string]interface{}{ "thinkingConfig": map[string]interface{}{"includeThoughts": true}, }, }, } // -> generationConfig.thinkingConfig == {"includeThoughts": true, "thinkingLevel": "low"} ``` Input transcription events (`input-transcription-completed`) accumulate consecutive fragments for one user utterance under a stable synthetic item ID, and allocate a new ID once Google signals the utterance is finished or the turn completes; this keeps multi-turn and barge-in (interrupted) conversations from overwriting or misgrouping earlier user messages. ### Speech Translation (experimental) `Provider.SpeechTranslationModel(modelID)` (alias `Provider.Translation`) returns a speech-to-speech translation model over the same Live API `BidiGenerateContent` WebSocket, for use with `ai.ExperimentalStreamTranslate`: ```go translationModel, err := googleProvider.SpeechTranslationModel("gemini-3.5-live-translate-preview") if err != nil { log.Fatal(err) } result, err := ai.ExperimentalStreamTranslate(ctx, ai.StreamTranslateOptions{ Model: translationModel, Audio: audioStream, TargetLanguage: "es", }) if err != nil { log.Fatal(err) } stream, err := result.FullStream() if err != nil { log.Fatal(err) } defer stream.Close() for { part, err := stream.Next() if err == io.EOF { break } if err != nil { log.Fatal(err) } if part.Type == "audio" { // part.AudioData carries translated PCM audio. } } ``` ## Provider-Specific Features ### Service Tier For the Gemini API, pass `serviceTier` in Google provider options: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "hello", ProviderOptions: map[string]interface{}{ "google": map[string]interface{}{ "serviceTier": "priority", }, }, }) ``` ### Massive Context Windows Gemini supports up to 2M tokens of context: ```go // Process entire codebases or books hugeDocument := loadDocument("book.txt") // up to 2M tokens result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Summarize this book:\n\n%s", hugeDocument), }) ``` ### Multimodal Capabilities Gemini excels at processing text, images, video, and audio together: ```go // Image understanding result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Describe this image in detail"}, types.FileContent{ URL: "https://example.com/photo.jpg", MediaType: "image/jpeg", }, }, }, }, }) // Multiple images result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Compare these images"}, types.FileContent{ URL: "https://example.com/before.jpg", MediaType: "image/jpeg", }, types.FileContent{ URL: "https://example.com/after.jpg", MediaType: "image/jpeg", }, }, }, }, }) // Video understanding result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What happens in this video?"}, types.FileContent{ URL: "https://example.com/video.mp4", MediaType: "video/mp4", }, }, }, }, }) ``` ### Function Calling Gemini supports sophisticated function calling: ```go weatherTool := types.Tool{ Name: "get_weather", Description: "Get current weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "City and country", }, "unit": map[string]interface{}{ "type": "string", "enum": []string{"celsius", "fahrenheit"}, }, }, "required": []string{"location"}, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in Tokyo?", Tools: []types.Tool{weatherTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) for _, call := range result.ToolCalls { fmt.Printf("Function: %s\nArgs: %v\n", call.ToolName, call.Arguments) } ``` ### Interactions API The Google provider exposes the Gemini Interactions API through `Provider.Interactions(modelID)` and agent presets through `Provider.InteractionsAgent(agent)`. Interactions models implement the standard `provider.LanguageModel` interface, so they can be used with `DoGenerate`, `DoStream`, and higher-level AI SDK helpers. ```go store := false interactionsModel, err := googleProvider.Interactions(google.InteractionsModelGemini25Flash) if err != nil { log.Fatal(err) } result, err := interactionsModel.DoGenerate(ctx, &provider.GenerateOptions{ Prompt: types.Prompt{Text: "Give a concise answer."}, ProviderOptions: map[string]interface{}{ "google": google.GoogleInteractionsProviderOptions{ Store: &store, ResponseModalities: []string{"text"}, ThinkingLevel: "low", }, }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` Use `PreviousInteractionID` with `Store` for stateful follow-up calls. Agent models run as background interactions and are polled until a terminal status; `PollingTimeoutMs` controls that wait. Agent model IDs include `deep-research-pro-preview-12-2025`, `deep-research-preview-04-2026`, `deep-research-max-preview-04-2026`, and `antigravity-preview-05-2026`. ```go agentModel, err := googleProvider.InteractionsAgent(google.InteractionsAgentDeepResearchProPreview) if err != nil { log.Fatal(err) } result, err := agentModel.DoGenerate(ctx, &provider.GenerateOptions{ Prompt: types.Prompt{Text: "Research the main tradeoffs in retrieval augmented generation."}, ProviderOptions: map[string]interface{}{ "google": google.GoogleInteractionsProviderOptions{ PollingTimeoutMs: 120000, }, }, }) ``` Interactions responses preserve ordered text, reasoning, files, sources, and tool-call content. Google metadata is available under `ProviderMetadata["google"]`, including the `interactionId` when returned by the API. ### Grounding with Google Search Connect Gemini to real-time web information: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What are the latest developments in quantum computing?", Tools: []types.Tool{google.GoogleSearchTool()}, }) // Response includes citations for _, source := range result.Sources { fmt.Printf("Source: %s\n", source.URL) } ``` ### JSON Mode Structured output generation: ```go import "github.com/digitallysavvy/go-ai/pkg/schema" bookSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "title": map[string]string{"type": "string"}, "author": map[string]string{"type": "string"}, "year": map[string]string{"type": "integer"}, "genre": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, }, }, "required": []string{"title", "author", "year"}, }) result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: bookSchema, Prompt: "Extract book metadata: '1984 by George Orwell, published 1949'", }) ``` ### Streaming Real-time response streaming: ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long story", }) if err != nil { log.Fatal(err) } defer stream.Close() for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } ``` ### Thinking Mode Gemini 2.0 Flash Thinking uses extended reasoning: ```go model, err := googleProvider.LanguageModel(google.ModelGemini20FlashThinkingExp) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Solve this logic puzzle: Three friends...", }) // Reasoning content is preserved in ordered content parts when returned. for _, part := range result.Content { if reasoning, ok := part.(types.ReasoningContent); ok { fmt.Println("Reasoning:", reasoning.Text) } } fmt.Println("Answer:", result.Text) ``` Gemini 3 thinking and tool-call streams preserve `thoughtSignature` in provider metadata so multi-step calls can safely round-trip model-generated reasoning and function calls. Streaming no-argument function calls are emitted with `{}` input instead of being dropped. Usage includes per-modality token details when the API returns them. Text, image, audio, and video token counts are preserved in the result usage details rather than collapsed into only total input/output tokens. Use the `"google"` provider key for AI Studio options. Vertex options use the `"googleVertex"` key (see the [Google Vertex AI provider](https://goaisdk.com/docs/providers/google-vertex.md#provider-specific-options) for its lookup order) — the hyphenated `"google-vertex"` string is only the Vertex model's `Provider()` identity, not a `ProviderOptions` key. ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/google" ) func main() { ctx := context.Background() googleProvider := google.New(google.Config{ APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"), }) model, err := googleProvider.LanguageModel("gemini-2.0-flash") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain machine learning in simple terms", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) fmt.Printf("Tokens: %d\n", result.Usage.GetTotalTokens()) } ``` ### Image Analysis ```go imageData, err := os.ReadFile("chart.png") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Analyze this chart and extract key insights"}, types.FileContent{ Data: imageData, MediaType: "image/png", }, }, }, }, }) fmt.Println(result.Text) ``` ### Video Analysis ```go // Upload video file videoData, err := os.ReadFile("demo.mp4") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Describe what happens in this video"}, types.FileContent{ Data: videoData, MediaType: "video/mp4", }, }, }, }, }) fmt.Println(result.Text) ``` ### Document Processing with Large Context ```go // Load multiple documents docs := []string{ loadFile("doc1.txt"), loadFile("doc2.txt"), loadFile("doc3.txt"), } combined := strings.Join(docs, "\n\n---\n\n") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Analyze these documents and identify common themes:\n\n%s", combined), }) ``` ### Grounded Search Responses ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What are the current COVID-19 vaccination rates in major countries?", Tools: []types.Tool{google.GoogleSearchTool()}, }) fmt.Println("Answer:", result.Text) fmt.Println("\nSources:") for _, source := range result.Sources { fmt.Printf("- %s: %s\n", source.Title, source.URL) } ``` ### Text Embeddings ```go embeddingModel, err := googleProvider.EmbeddingModel(google.EmbeddingModelGeminiEmbedding001) if err != nil { log.Fatal(err) } dimensions := 768 result, err := embeddingModel.DoEmbed(ctx, "Go is a programming language", &provider.EmbedModelOptions{ ProviderOptions: map[string]interface{}{ "google": google.GoogleEmbeddingProviderOptions{ TaskType: "RETRIEVAL_DOCUMENT", OutputDimensionality: &dimensions, }, }, }) if err != nil { log.Fatal(err) } fmt.Printf("Dimensions: %d\n", len(result.Embedding)) ``` Google embeddings also support multimodal content and provider-hosted files through `fileData`: ```go result, err := embeddingModel.DoEmbed(ctx, "Summarize this document for retrieval.", &provider.EmbedModelOptions{ ProviderOptions: map[string]interface{}{ "google": google.GoogleEmbeddingProviderOptions{ TaskType: "RETRIEVAL_DOCUMENT", Content: [][]google.EmbeddingPart{ { google.FileDataEmbeddingPart{ MimeType: "application/pdf", FileURI: "files/abc123", }, }, }, }, }, }) if err != nil { log.Fatal(err) } ``` ### Multi-Turn Conversation ```go messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What is the capital of France?"}, }, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: messages, }) messages = append(messages, types.Message{ Role: types.RoleAssistant, Content: []types.ContentPart{ types.TextContent{Text: result.Text}, }, }, types.Message{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What's the population?"}, }, }, ) result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: messages, }) ``` ### Image Generation Google's text-to-image generation is now **Gemini-only** — Imagen model IDs (`imagen-4.0-*`) were removed from this provider; use Google Vertex AI for Imagen. **Gemini Image Models** offer fast generation: ```go // Gemini 2.5 Flash - Fast image generation geminiImageModel, err := googleProvider.ImageModel("gemini-2.5-flash-image") if err != nil { log.Fatal(err) } result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: geminiImageModel, Prompt: "Abstract colorful art with geometric shapes and vibrant patterns", Size: "1920x1080", // 16:9 aspect ratio }) if err != nil { log.Fatal(err) } fmt.Printf("Image generated: %d bytes\n", len(result.Images[0].Data)) os.WriteFile("gemini_output.png", result.Images[0].Data, 0644) ``` **Supported Aspect Ratios:** - `1:1` - Square (1024x1024, 512x512) - `4:3` - Standard (1024x768) - `3:4` - Portrait (768x1024) - `16:9` - Widescreen (1920x1080, 1792x1024) - `9:16` - Vertical (1080x1920, 1024x1792) **Model Comparison:** | Model | Speed | Quality | Use Case | |-------|-------|---------|----------| | gemini-2.5-flash-image | Very Fast | Good | Quick iterations | | gemini-3-pro-image-preview | Fast | High | Advanced generation | **Note:** For Imagen models and image editing, use Google Vertex AI. ## Video Generation `provider.VideoModel(modelID)` submits to Veo's `:predictLongRunning` endpoint and polls the long-running operation until it completes (`MaxVideosPerCall()` returns 4 — Veo accepts up to 4 videos per call, and `ai.GenerateVideo` fans out across multiple calls automatically when you ask for more): ```go videoModel, err := googleProvider.VideoModel("veo-3.1-generate-preview") if err != nil { log.Fatal(err) } result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{Text: "A drone shot flying over a coastline at sunrise"}, }) if err != nil { log.Fatal(err) } for _, video := range result.Videos { fmt.Println(video.URL) } ``` `VideoModel` also implements `provider.VideoModelStarter` / `VideoModelStatusChecker`, so it works with the async `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus` calls. ## Advanced Configuration ### Custom HTTP Client ```go googleProvider := google.New(google.Config{ APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"), HTTPClient: &http.Client{ Timeout: time.Minute * 5, Transport: &http.Transport{ MaxIdleConns: 100, IdleConnTimeout: 90 * time.Second, }, }, }) ``` ### Safety Settings ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, ProviderOptions: map[string]interface{}{ "google": map[string]interface{}{ "safetySettings": []map[string]string{ {"category": "HARM_CATEGORY_HATE_SPEECH", "threshold": "BLOCK_LOW_AND_ABOVE"}, {"category": "HARM_CATEGORY_DANGEROUS_CONTENT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"}, {"category": "HARM_CATEGORY_SEXUALLY_EXPLICIT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"}, {"category": "HARM_CATEGORY_HARASSMENT", "threshold": "BLOCK_MEDIUM_AND_ABOVE"}, }, }, }, }) ``` ### Generation Config ```go temperature := 0.9 topP := 0.95 topK := 40 maxTokens := 2048 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, Temperature: &temperature, TopP: &topP, TopK: &topK, MaxTokens: &maxTokens, }) ``` ## Error Handling ### Common Errors ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { switch providerErr.StatusCode { case 400: log.Printf("Invalid request: %s", providerErr.Message) case 401: log.Fatal("Invalid API key") case 403: log.Fatal("Permission denied or quota exceeded") case 404: log.Fatal("Model not found") case 429: log.Println("Rate limited, retrying...") time.Sleep(time.Second * 5) case 500, 503: log.Println("Service error, retrying...") time.Sleep(time.Second * 10) default: log.Printf("Google AI error: %d - %s", providerErr.StatusCode, providerErr.Message) } } return nil, err } ``` ### Content Filtering ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { if providerErr.StatusCode == 400 && strings.Contains(providerErr.Message, "SAFETY") { log.Println("Content blocked by safety filters") // Try with adjusted safety settings } } } ``` ## Best Practices 1. **Model Selection** - Use `gemini-2.0-flash` for fast, multimodal tasks - Use `gemini-2.5-pro` for complex reasoning with huge context - Use `gemini-2.5-flash-lite` for cost-effective high-volume tasks 2. **Context Management** - Take advantage of massive context windows - Include full documents rather than chunking - Use structured prompts for better results 3. **Multimodal Input** - Compress images appropriately (max 20MB) - Use clear, specific instructions for vision tasks - Combine multiple modalities for richer context 4. **Cost Optimization** - Monitor context length (pricing changes after 128K) - Use flash-8b for simple tasks - Cache repeated queries on your side 5. **Safety** - Adjust safety settings based on use case - Handle content filtering gracefully - Log filtered content for review ## Rate Limits & Pricing ### Rate limits Rate limits depend on your model and tier. See [Google's rate limits](https://ai.google.dev/gemini-api/docs/rate-limits) and [pricing](https://ai.google.dev/gemini-api/docs/pricing) pages for current values. Pricing can also depend on prompt length. ## Workflow Serialization Google embedding, image, and (unary) transcription models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel` (language models could already be serialized). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; speech, speech-translation, video, and realtime models are not yet serializable. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Google Vertex AI Provider](https://goaisdk.com/docs/providers/google-vertex.md) - Enterprise deployment - [Google AI Studio Documentation](https://ai.google.dev/docs) - [Gemini API Reference](https://ai.google.dev/api/rest) ## May 2026 parity updates ### Model refresh and token usage The Google provider tracks the refreshed Gemini model IDs in `pkg/providers/google/model_ids.go`. Usage details include per-modality token fields through `types.Usage.InputDetails.TextTokens` and `types.Usage.InputDetails.ImageTokens` when Google returns modality-specific counts. ### Thought signatures and no-arg tool calls The provider preserves Google `thoughtSignature` metadata on text and tool-call parts so multi-turn requests can forward model reasoning state. Streaming tool calls with no arguments are emitted with an empty argument object, matching the TypeScript SDK. ### Provider options Use the standard `ProviderOptions` map for Google-specific options such as thinking configuration. The current Go provider accepts the same nested `thinkingConfig` request shape that the TypeScript provider sends to Gemini. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Use Google search if useful.", ProviderOptions: map[string]interface{}{ "google": map[string]interface{}{ "thinkingConfig": map[string]interface{}{ "thinkingBudget": 512, "includeThoughts": true, }, }, }, }) ``` --- # Azure OpenAI Provider > Setup guide for Azure OpenAI Service in the Go AI SDK, covering deployment names as model IDs, Entra ID authentication, and regional/private endpoints. Canonical URL: https://goaisdk.com/docs/providers/azure Documentation index: https://goaisdk.com/llms.txt Azure OpenAI Service provides OpenAI-compatible models through Azure resources and deployments. In the Go AI SDK, the deployment name is the model ID you pass to the provider. ## Setup ```go import ( "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/azure" ) ``` Set credentials: ```bash export AZURE_API_KEY=your-key export AZURE_RESOURCE_NAME=your-resource-name ``` Create a Responses model: ```go provider, err := azure.New(azure.Config{ APIKey: os.Getenv("AZURE_API_KEY"), ResourceName: os.Getenv("AZURE_RESOURCE_NAME"), }) if err != nil { log.Fatal(err) } model, err := provider.LanguageModel("my-gpt-deployment") if err != nil { log.Fatal(err) } ``` `LanguageModel` matches the TypeScript Azure provider default and uses the Responses API at `/openai/v1/responses?api-version=v1`. Use `ChatModel` when you need the Chat Completions endpoint or `CompletionModel` when you need the legacy Completions endpoint. ## Microsoft Entra ID Authentication Use `ADTokenProvider` when Azure OpenAI should authenticate with Microsoft Entra ID instead of an API key. The token provider is called with the request context and sets `Authorization: Bearer ` across Responses, Chat Completions, Completions, embeddings, image, speech, and transcription request surfaces. ```go provider, err := azure.New(azure.Config{ ResourceName: os.Getenv("AZURE_RESOURCE_NAME"), ADTokenProvider: func(ctx context.Context) (string, error) { return acquireToken(ctx) }, }) if err != nil { log.Fatal(err) } ``` Passing both `APIKey` and `ADTokenProvider` returns an invalid-argument error, matching the TypeScript `createAzure` validation. ## Responses Provider Tools Azure exposes the OpenAI Responses provider tools as Azure helpers: - `azure.CodeInterpreter` - `azure.FileSearch` - `azure.ImageGeneration` - `azure.WebSearch` - `azure.WebSearchPreview` Use them with Responses-compatible Azure deployments. ```go searchContext := "low" result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Summarize how to upgrade AI SDK from 5 to 6.", Tools: []types.Tool{ azure.WebSearch(openaitool.WebSearchConfig{ SearchContextSize: searchContext, Filters: &openaitool.WebSearchFilters{ AllowedDomains: []string{"ai-sdk.dev"}, }, }), }, }) if err != nil { log.Fatal(err) } defer result.Close() ``` This matches the TypeScript `azure.tools` behavior by serializing provider-executed tools as OpenAI Responses hosted tools. ## Regional And Private Endpoints For the standard Azure URL format, set `ResourceName`. For custom regional, proxy, or private-link endpoints, set `BaseURL` to the Azure OpenAI `/openai` prefix: `ResourceName` is interpolated directly into the request host (`https://{resourceName}.openai.azure.com/...`), so `New` rejects a value that isn't a single DNS label (letters, digits, and hyphens, 1-63 characters, no leading/trailing hyphen) with an `InvalidArgumentError` before building the client — a value like `user@internal:8080/#` would otherwise rewrite the destination host. This only applies when `BaseURL` is unset; a custom `BaseURL` is used as-is and `ResourceName` is ignored. ```go provider, err := azure.New(azure.Config{ BaseURL: "https://your-resource-eastus.openai.azure.com/openai", APIKey: os.Getenv("AZURE_API_KEY"), }) if err != nil { log.Fatal(err) } ``` The default URL shape matches TypeScript: `{baseURL}/v1{path}?api-version=v1`. This only applies to a recognized Azure OpenAI / AI Foundry / Cognitive Services host (`*.openai.azure.com`, `*.services.ai.azure.com`, `*.cognitiveservices.azure.com`); a `BaseURL` host outside those patterns is treated as a custom gateway and used exactly as given, with no `/v1` suffix or `api-version` query param appended — verify your proxy's URL matches one of the recognized patterns if you expect Azure's own query-string behavior. Set `UseDeploymentBasedURLs` only for Azure deployments that still require `{baseURL}/deployments/{deployment}{path}`: ```go provider, err := azure.New(azure.Config{ ResourceName: os.Getenv("AZURE_RESOURCE_NAME"), APIKey: os.Getenv("AZURE_API_KEY"), UseDeploymentBasedURLs: true, }) if err != nil { log.Fatal(err) } ``` ## Other Model Types Azure also supports completions, embeddings, image generation, speech, and transcription through deployment names: ```go completionModel, err := provider.CompletionModel("gpt-35-turbo-instruct") embeddingModel, err := provider.EmbeddingModel("text-embedding-3-small") imageModel, err := provider.ImageModel("dall-e-3") speechModel, err := provider.SpeechModel("tts-1") transcriptionModel, err := provider.TranscriptionModel("whisper-1") ``` Azure DeepSeek deployments are available through `provider.DeepSeek("deployment-name")`. The model uses Azure authentication, the Azure `api-version` query parameter, and the `"azure"` provider option namespace while preserving DeepSeek reasoning-effort request behavior. ## MAI-Transcribe And MAI-Voice (Azure Speech) `provider.TranscriptionModel` and `provider.SpeechModel` route Microsoft AI's MAI-Transcribe and MAI-Voice model families through the Azure Speech API instead of Azure OpenAI, since they use a different protocol and endpoint. Routing is automatic by model ID, and can be overridden per call with `providerOptions.azure.api` (`"openai"`, `"speech"`, or `"mai"`): | Model ID | Default API | | --- | --- | | `mai-transcribe-2`, `mai-transcribe-1.5` | Speech (file transcription) | | `mai-transcribe-2-streaming` | MAI realtime (streaming transcription only) | | `mai-voice-2`, `mai-voice-2-flash`, `mai-voice-2.1`, `mai-voice-2.1-flash` | Speech (SSML text-to-speech) | | any other deployment name | OpenAI | ```go azureProvider, err := azure.New(azure.Config{ ResourceName: os.Getenv("AZURE_RESOURCE_NAME"), APIKey: os.Getenv("AZURE_API_KEY"), }) if err != nil { log.Fatal(err) } transcriptionModel, err := azureProvider.TranscriptionModel("mai-transcribe-2") if err != nil { log.Fatal(err) } result, err := transcriptionModel.DoTranscribe(ctx, &provider.TranscriptionOptions{ Audio: audioBytes, MimeType: "audio/wav", ProviderOptions: map[string]interface{}{ "azure": map[string]interface{}{ "timestamps": "segment", // "word" | "segment" | "none"; default "segment" "locales": []interface{}{"en"}, }, }, }) ``` `providerOptions.azure` is validated strictly for both models: an unrecognized key, or an option a given MAI model doesn't support (e.g. `diarization` on MAI-Transcribe-1.5), returns an `InvalidArgumentError` instead of being silently dropped; an option that only applies to a different backend (e.g. `style` routed to OpenAI TTS) is reported as an `unsupported` warning on the result instead. MAI-Voice speech picks a default voice from `Language` when `Voice` is unset, and accepts `providerOptions.azure.style` / `styleDegree` for Azure's `mstts:express-as` speaking styles: ```go speechModel, err := provider.SpeechModel("mai-voice-2") result, err := speechModel.DoGenerate(ctx, &provider.SpeechGenerateOptions{ Text: "Hello from Azure.", Language: "es", // selects es-MX-Valeria when Voice is unset ProviderOptions: map[string]interface{}{ "azure": map[string]interface{}{"style": "excited", "styleDegree": 1.5}, }, }) ``` MAI-Transcribe-2-Streaming only supports `ai.ExperimentalStreamTranscribe` (streaming transcription), over a WebSocket to the MAI realtime API; it requires 16-bit mono PCM audio at 16000 or 24000 Hz. `gpt-realtime-whisper` deployments (and dated snapshots, e.g. `gpt-realtime-whisper-2026-01-01`) also support `ai.ExperimentalStreamTranscribe`, streaming over the OpenAI-compatible realtime WebSocket at your Azure resource's `/realtime?intent=transcription` endpoint instead of the REST transcription endpoint; `DoTranscribe` (non-streaming) is unsupported for these deployments, matching OpenAI's own `gpt-realtime-whisper` restriction. Two optional `azure.Config` settings control the endpoints these models use, both defaulting from `ResourceName` and validated the same way as `ResourceName` itself (a single DNS label) before being used: - `SpeechBaseURL` — the Azure Speech endpoint prefix (default `https://{ResourceName}.cognitiveservices.azure.com`), used for both MAI-Transcribe file transcription and MAI-Voice speech. - `MaiBaseURL` — the MAI realtime endpoint prefix (default `https://{ResourceName}.services.ai.azure.com/mai/v1`), used for MAI-Transcribe-2-Streaming. ```go provider, err := azure.New(azure.Config{ ResourceName: os.Getenv("AZURE_RESOURCE_NAME"), APIKey: os.Getenv("AZURE_API_KEY"), SpeechBaseURL: "https://eastus.api.cognitive.microsoft.com", }) ``` ## Custom Headers Use `Headers` for request metadata that should be sent with every provider call: ```go provider, err := azure.New(azure.Config{ ResourceName: os.Getenv("AZURE_RESOURCE_NAME"), APIKey: os.Getenv("AZURE_API_KEY"), Headers: map[string]string{ "X-Request-ID": "unique-request-id", }, }) if err != nil { log.Fatal(err) } ``` Per-call headers can also be supplied through generation options and override provider-level headers. --- # AWS Bedrock Provider > Setup and usage guide for AWS Bedrock multi-model platform with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/bedrock Documentation index: https://goaisdk.com/llms.txt AWS Bedrock provides access to multiple foundation models from different providers through a single API. Includes models from Anthropic, Meta, Cohere, AI21, Amazon, and more with enterprise-grade security and compliance. Cohere embedding models are detected both as direct model IDs such as `cohere.embed-v4` and cross-region inference profile IDs such as `us.cohere.embed-v4:0`. The request body uses the Cohere embedding shape in both cases. ## Setup ### Installation The AWS Bedrock provider is included in the Go AI SDK: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/bedrock" ) ``` ### Configuration ```go provider := bedrock.New(bedrock.Config{ Region: "us-east-1", AWSAccessKeyID: os.Getenv("AWS_ACCESS_KEY_ID"), AWSSecretAccessKey: os.Getenv("AWS_SECRET_ACCESS_KEY"), }) model, err := provider.LanguageModel("anthropic.claude-sonnet-5-5") if err != nil { log.Fatal(err) } ``` ### Get Started 1. Sign up for AWS account at [aws.amazon.com](https://aws.amazon.com) 2. Request model access in [Bedrock Console](https://console.aws.amazon.com/bedrock) 3. Create IAM user with Bedrock permissions 4. Set environment variables: ```bash export AWS_ACCESS_KEY_ID=AKIA... export AWS_SECRET_ACCESS_KEY=... export AWS_REGION=us-east-1 ``` ## Available Models Use the constants in `pkg/providers/bedrock/model_ids.go` instead of raw strings. Availability and pricing vary by AWS Region. See the [Amazon Bedrock model catalog](https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html) and [pricing page](https://aws.amazon.com/bedrock/pricing/). ### Anthropic Claude models | Model ID | Go constant | |----------|-------------| | `anthropic.claude-opus-5-5` | `bedrock.ModelAnthropicClaudeOpus5_5` | | `anthropic.claude-opus-5` | `bedrock.ModelAnthropicClaudeOpus5` | | `anthropic.claude-sonnet-5-5` | `bedrock.ModelAnthropicClaudeSonnet5_5` | | `anthropic.claude-sonnet-5` | `bedrock.ModelAnthropicClaudeSonnet5` | | `anthropic.claude-fable-5-1` | `bedrock.ModelAnthropicClaudeFable5_1` | | `anthropic.claude-fable-5` | `bedrock.ModelAnthropicClaudeFable5` | | `anthropic.claude-opus-4-8` | `bedrock.ModelAnthropicClaudeOpus4_8` | | `anthropic.claude-opus-4-7` | `bedrock.ModelAnthropicClaudeOpus4_7` | | `anthropic.claude-sonnet-4-6-v1` | `bedrock.ModelAnthropicClaudeSonnet4_6V1` | | `anthropic.claude-haiku-4-5-20251001-v1:0` | `bedrock.ModelAnthropicClaudeHaiku45_V1` | Cross-region inference profiles use a prefix, for example `us.anthropic.claude-sonnet-5-5` (`bedrock.ModelUSAnthropicClaudeSonnet5_5`), `eu.anthropic.claude-opus-4-8` and `global.anthropic.claude-fable-5-1`. Earlier Claude generations also have constants. ### Amazon models | Model ID | Go constant | |----------|-------------| | `us.amazon.nova-2-lite-v1:0` | `bedrock.ModelAmazonNova2LiteV1` | | `us.amazon.nova-premier-v1:0` | `bedrock.ModelAmazonNovaPremierV1` | | `us.amazon.nova-pro-v1:0` | `bedrock.ModelAmazonNovaProV1` | | `us.amazon.nova-lite-v1:0` | `bedrock.ModelAmazonNovaLiteV1` | | `us.amazon.nova-micro-v1:0` | `bedrock.ModelAmazonNovaMicroV1` | | `amazon.titan-text-express-v1` | `bedrock.ModelAmazonTitanTextExpressV1` | ### Meta Llama models | Model ID | Go constant | |----------|-------------| | `us.meta.llama4-maverick-17b-instruct-v1:0` | `bedrock.ModelUSMetaLlama4Maverick17BInstructV1` | | `us.meta.llama4-scout-17b-instruct-v1:0` | `bedrock.ModelUSMetaLlama4Scout17BInstructV1` | | `us.meta.llama3-3-70b-instruct-v1:0` | `bedrock.ModelUSMetaLlama3370BInstructV1` | | `meta.llama3-1-405b-instruct-v1:0` | `bedrock.ModelMetaLlama31405BInstructV1` | | `meta.llama3-2-90b-instruct-v1:0` | `bedrock.ModelMetaLlama3290BInstructV1` | ### Other language models | Model ID | Go constant | |----------|-------------| | `cohere.command-r-plus-v1:0` | `bedrock.ModelCohereCommandRPlusV1` | | `cohere.command-r-v1:0` | `bedrock.ModelCohereCommandRV1` | | `mistral.mistral-large-2402-v1:0` | `bedrock.ModelMistralLarge2402V1` | | `us.mistral.pixtral-large-2502-v1:0` | `bedrock.ModelUSMistralPixtralLarge2502V1` | | `openai.gpt-oss-120b-1:0` | `bedrock.ModelOpenAIGPTOSS120B` | | `us.deepseek.r1-v1:0` | `bedrock.ModelUSDeepSeekR1V1` | ### Embedding models | Model ID | Dimensions | Best for | |----------|-----------|----------| | `amazon.titan-embed-text-v2:0` | 1024 | General embeddings | | `cohere.embed-english-v3` | 1024 | English text | | `cohere.embed-multilingual-v3` | 1024 | Multilingual | ### Image generation | Model ID | Best for | |----------|----------| | `amazon.titan-image-generator-v2:0` | General images | | `stability.stable-diffusion-xl-v1` | Artistic | ## Provider-Specific Features ### Model Access Control Request access to models in the [Bedrock Console](https://console.aws.amazon.com/bedrock/home#/modelaccess) before calling them; the SDK does not expose an API to list or manage model access. ### Cross-Region Inference Cross-region inference profiles are plain model IDs with a region prefix (e.g. `us.`, `eu.`, or no prefix for global profiles) — pass the profile ID directly to `LanguageModel`, there is no separate config field: ```go provider := bedrock.New(bedrock.Config{ Region: "us-east-1", AWSAccessKeyID: os.Getenv("AWS_ACCESS_KEY_ID"), AWSSecretAccessKey: os.Getenv("AWS_SECRET_ACCESS_KEY"), }) model, err := provider.LanguageModel("us.anthropic.claude-sonnet-5-5") ``` ### Guardrails Apply AWS Bedrock Guardrails for content filtering: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, ProviderOptions: map[string]interface{}{ "amazonBedrock": map[string]interface{}{ "guardrailIdentifier": "your-guardrail-id", "guardrailVersion": "1", }, }, }) ``` ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/bedrock" ) func main() { ctx := context.Background() provider := bedrock.New(bedrock.Config{ Region: "us-east-1", AWSAccessKeyID: os.Getenv("AWS_ACCESS_KEY_ID"), AWSSecretAccessKey: os.Getenv("AWS_SECRET_ACCESS_KEY"), }) model, err := provider.LanguageModel("anthropic.claude-sonnet-5-5") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain AWS Bedrock benefits", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Multi-Model Comparison ```go models := []string{ "anthropic.claude-sonnet-5-5", "us.meta.llama3-3-70b-instruct-v1:0", "cohere.command-r-plus-v1:0", } prompt := "Write a haiku about clouds" for _, modelID := range models { model, err := provider.LanguageModel(modelID) if err != nil { log.Printf("Failed to load %s: %v", modelID, err) continue } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) if err != nil { log.Printf("Failed to generate with %s: %v", modelID, err) continue } fmt.Printf("\n=== %s ===\n%s\n", modelID, result.Text) } ``` ### Image Generation with Titan ```go imageModel, err := provider.ImageModel("amazon.titan-image-generator-v2:0") if err != nil { log.Fatal(err) } result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "A serene mountain landscape at sunset", Size: "1024x1024", }) if err != nil { log.Fatal(err) } // Save image os.WriteFile("output.png", result.Images[0].Data, 0644) ``` ### Embeddings with Cohere ```go embeddingModel, err := provider.EmbeddingModel("cohere.embed-english-v3") if err != nil { log.Fatal(err) } texts := []string{ "AWS Bedrock provides access to foundation models", "Amazon offers cloud computing services", "Machine learning powers AI applications", } result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: texts, }) if err != nil { log.Fatal(err) } fmt.Printf("Generated %d embeddings\n", len(result.Embeddings)) ``` ### Streaming with Claude ```go model, err := provider.LanguageModel("anthropic.claude-sonnet-5-5") if err != nil { log.Fatal(err) } stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a detailed essay about serverless computing"}) if err != nil { log.Fatal(err) } defer stream.Close() for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } ``` ## Advanced Configuration ### IAM Role Authentication ```go // Use IAM role credentials (for EC2, ECS, Lambda) provider := bedrock.New(bedrock.Config{ Region: "us-east-1", // Credentials automatically loaded from IAM role }) ``` ### Session Token (STS) ```go provider := bedrock.New(bedrock.Config{ Region: "us-east-1", AWSAccessKeyID: os.Getenv("AWS_ACCESS_KEY_ID"), AWSSecretAccessKey: os.Getenv("AWS_SECRET_ACCESS_KEY"), SessionToken: os.Getenv("AWS_SESSION_TOKEN"), }) ``` When explicit credentials are provided, the provider uses the explicit `SessionToken` value only. It does not merge in `AWS_SESSION_TOKEN` from the environment unless credentials are being resolved from the environment, matching the TypeScript credential precedence. ### Custom Endpoint ```go provider := bedrock.New(bedrock.Config{ Region: "us-east-1", BaseURL: "https://bedrock-runtime.us-east-1.amazonaws.com", }) ``` ### Reasoning and Structured Output Bedrock drops reasoning blocks that belong to a foreign provider before request signing, so unsigned provider-specific reasoning is not forwarded accidentally. When top-level reasoning is set, custom reasoning configuration properties are merged into the Bedrock request instead of replacing the reasoning config. Anthropic-on-Bedrock supports native structured output when reasoning is not enabled. The SDK preserves empty text blocks that accompany reasoning content, matching the TypeScript content-shape behavior required by Claude reasoning models. ```go model, _ := provider.LanguageModel("anthropic.claude-opus-4-7") // invoiceSchema is a schema.Schema, e.g. schema.NewSimpleJSONSchema(map[string]interface{}{...}). result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Prompt: "Extract the invoice total.", Schema: invoiceSchema, }) ``` ## Error Handling ### Common Errors Bedrock surfaces errors as `*providererrors.ProviderError`, like every other provider. AWS's exception type is folded into `Message` as `"{type}: {message}"` (there is no separate error-code field), so match on the message prefix: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) if err != nil { var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { switch { case strings.HasPrefix(providerErr.Message, "AccessDeniedException"): log.Fatal("Access denied - check IAM permissions") case strings.HasPrefix(providerErr.Message, "ResourceNotFoundException"): log.Fatal("Model not found - request access in Bedrock console") case strings.HasPrefix(providerErr.Message, "ThrottlingException"): log.Println("Throttled, retrying...") time.Sleep(time.Second * 5) case strings.HasPrefix(providerErr.Message, "ValidationException"): log.Printf("Invalid request: %s", providerErr.Message) case strings.HasPrefix(providerErr.Message, "ModelNotReadyException"): log.Println("Model still loading, retry later") case strings.HasPrefix(providerErr.Message, "ServiceQuotaExceededException"): log.Fatal("Quota exceeded - request limit increase") default: log.Printf("Bedrock error (%d): %s", providerErr.StatusCode, providerErr.Message) } } } ``` ### IAM Permission Requirements Minimum IAM policy for Bedrock: ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream" ], "Resource": "arn:aws:bedrock:*::foundation-model/*" } ] } ``` ## Best Practices 1. **Model Selection** - Use Claude for complex reasoning and analysis - Use Llama for cost-effective open-source alternative - Use Titan for AWS-native integration - Use Cohere for RAG and search applications 2. **Cost Management** - Compare prices across providers - Use smaller models for simple tasks - Monitor usage with AWS Cost Explorer - Set up billing alerts 3. **Security** - Use IAM roles instead of access keys - Apply least privilege IAM policies - Enable CloudTrail logging - Use VPC endpoints for private access 4. **Performance** - Choose region closest to users - Use cross-region inference for high availability - Implement caching for repeated queries - Use streaming for long responses 5. **Compliance** - Review model provider compliance certifications - Use Guardrails for content filtering - Enable logging for audit trails - Configure data retention policies ## Rate Limits & Pricing ### Rate Limits Varies by model and region: | Model Family | Default Tokens/Min | Burst | |--------------|-------------------|-------| | Claude | 100K - 400K | 200K | | Llama | 200K - 800K | 400K | | Titan | 400K - 1.6M | 800K | Request quota increases through AWS Support. ### Pricing Pricing varies by model and AWS Region. See the [Amazon Bedrock pricing page](https://aws.amazon.com/bedrock/pricing/). ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Anthropic Provider](https://goaisdk.com/docs/providers/anthropic.md) - Direct Anthropic access - [AWS Bedrock Documentation](https://docs.aws.amazon.com/bedrock/) - [Bedrock Model Access](https://console.aws.amazon.com/bedrock/home#/modelaccess) ## May 2026 parity updates ### Credential precedence Explicit credentials in `bedrock.Config` take precedence over environment and shared credential file values. When `AWSAccessKeyID` and `AWSSecretAccessKey` are set, `AWS_SESSION_TOKEN` is not used unless `SessionToken` is also explicitly configured. ### Reasoning and structured output `bedrock.ModelOptions.ReasoningConfig` mirrors the TypeScript `amazonBedrock.reasoningConfig` option. Partial provider configs merge with top-level `Reasoning`. Foreign-provider reasoning blocks are dropped rather than sent unsigned. ```go budget := 4096 model, err := p.LanguageModelWithOptions("anthropic.claude-opus-4-7", &bedrock.ModelOptions{ ReasoningConfig: &bedrock.ReasoningConfig{ Type: "enabled", BudgetTokens: &budget, }, ServiceTier: "priority", }) ``` ### Transport behavior The TypeScript lazy fetch resolution change has no direct Go runtime equivalent. The Go provider uses its AWS signing request path and explicit configuration fields; no application migration is required for JavaScript `fetch` patching behavior. ### requestMetadata Pass `providerOptions.amazonBedrock.requestMetadata` (a `map[string]string`) to attach key-value pairs to a Converse/ConverseStream request for invocation-log cost attribution: ```go _, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "hello", ProviderOptions: map[string]interface{}{ "amazonBedrock": map[string]interface{}{ "requestMetadata": map[string]interface{}{"team": "search", "environment": "prod"}, }, }, }) ``` It is forwarded as a top-level `requestMetadata` field on the request body, separate from `additionalModelRequestFields`. ### Nova 2 Lite high reasoning `amazon.nova-2-lite-v1:0` rejects a request that combines `maxOutputTokens` with `high`/`xhigh`/`max` reasoning effort. The provider now drops `inferenceConfig.maxTokens` and returns an `unsupported` warning for `maxOutputTokens` instead of sending a request Bedrock would reject; `medium` (and lower) reasoning keeps `maxOutputTokens` unchanged. ### Region validation `Region` (and Bedrock Mantle's `Region`) is interpolated directly into the request host (`https://bedrock-runtime.{region}.amazonaws.com`, `https://bedrock-mantle.{region}.api.aws`). A value that isn't a single DNS label (letters, digits, and hyphens) is rejected with an `InvalidArgumentError` before any request is sent — this only applies when no explicit `BaseURL` or endpoint environment variable is configured, since those make `Region` unused. --- # Cohere Provider > Setup and usage guide for Cohere Command and Embed models with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/cohere Documentation index: https://goaisdk.com/llms.txt Cohere specializes in enterprise-grade language models optimized for production use cases including RAG, search, classification, and embeddings. Known for reliable performance and strong multilingual support. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/cohere" ) ``` ### Configuration ```go provider := cohere.New(cohere.Config{ APIKey: os.Getenv("COHERE_API_KEY"), }) model, err := provider.LanguageModel("command-r-plus") if err != nil { log.Fatal(err) } ``` ### Get API Key 1. Sign up at [cohere.com](https://cohere.com) 2. Get API key from dashboard 3. Set environment variable: ```bash export COHERE_API_KEY=... ``` ## Available Models ### Language Models | Model ID | Context | Best For | | ---------- | --------- | ---------- | | command-r-plus | 128K | RAG, search, agents | | command-r | 128K | Cost-effective tasks | | command | 4K | Legacy | | command-light | 4K | Fast responses | ### Embedding Models | Model ID | Dimensions | Best For | | ---------- | ----------- | ---------- | | embed-english-v3.0 | 1024 | English semantic search | | embed-multilingual-v3.0 | 1024 | Multilingual search | | embed-english-light-v3.0 | 384 | Fast embeddings | ## Provider-Specific Features ### Vision Input (Image Parts) Cohere chat models support image input in user messages. ```go prompt := types.Prompt{ Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Describe this image"}, types.FileContent{ FileData: types.FileData{ Type: types.FileDataTypeURL, URL: "https://example.com/cat.png", MediaType: "image/png", }, }, }, }, }, } result, err := model.DoGenerate(ctx, &provider.GenerateOptions{Prompt: prompt}) ``` ### RAG Optimization Cohere models are optimized for retrieval-augmented generation: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.FileContent{ Data: []byte("Go was created at Google in 2007"), MediaType: "text/plain", }, types.TextContent{Text: "Who created Go?"}, }, }, }, }) ``` ### Tool Use Tool calling, `TopP`/`TopK`/penalties/`Seed`/`StopSequences`, and JSON `response_format` are all forwarded to Cohere's v2 chat API: ```go searchTool := types.Tool{ Name: "search_products", Description: "Search product database", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]string{"type": "string"}, }, "required": []string{"query"}, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Find laptop under $1000", Tools: []types.Tool{searchTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) for _, call := range result.ToolCalls { fmt.Printf("Tool: %s Args: %v\n", call.ToolName, call.Arguments) } ``` `providerOptions.cohere.thinking` takes precedence over the standardized `Reasoning` call option when both are set. Unsupported provider-defined tools produce warnings instead of being silently dropped. ### Structured Output ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Extract the person's name and age from: John is 30.", ResponseFormat: &provider.ResponseFormat{ Type: "json", Schema: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]string{"type": "string"}, "age": map[string]string{"type": "integer"}, }, "required": []string{"name", "age"}, }, }, }) ``` ### Citations `DoGenerate` (non-streaming) surfaces Cohere's inline citations as ordered `types.SourceContent` parts in `result.Content` (`SourceType: "document"`), with citation span, text, and source detail under `ProviderMetadata["cohere"]`. Streaming does not accumulate citations — Cohere's `citation-start` / `citation-end` stream events carry no payload to reconstruct them from. ```go for _, part := range result.Content { if source, ok := part.(types.SourceContent); ok { fmt.Printf("Cited: %s\n", source.Title) } } ``` `STOP_SEQUENCE` maps to `types.FinishReasonStop`. ### Reranking Improve search results with reranking: ```go reranker, err := provider.RerankingModel("rerank-english-v3.0") if err != nil { log.Fatal(err) } topN := 2 result, err := reranker.DoRerank(ctx, &provider.RerankOptions{ Query: "machine learning", Documents: []string{ "ML is a subset of AI", "Weather forecast for today", "Neural networks learn patterns", }, TopN: &topN, }) if err != nil { log.Fatal(err) } for i, doc := range result.Ranking { fmt.Printf("%d. Score: %.2f - Index: %d\n", i+1, doc.RelevanceScore, doc.Index) } ``` ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/cohere" ) func main() { ctx := context.Background() provider := cohere.New(cohere.Config{ APIKey: os.Getenv("COHERE_API_KEY"), }) model, err := provider.LanguageModel("command-r-plus") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain vector databases"}) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Embeddings for Semantic Search ```go embeddingModel, err := provider.EmbeddingModel("embed-english-v3.0") if err != nil { log.Fatal(err) } // Embed documents docs := []string{ "Go is a statically typed language", "Python is dynamically typed", "JavaScript runs in browsers", } docResult, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: docs, }) // Embed query queryResult, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "Tell me about Go programming", }) // Find most similar bestMatch := findMostSimilar(queryResult.Embedding, docResult.Embeddings) ``` ## Best Practices 1. **Model Selection** - Use Command R+ for RAG and multi-step reasoning - Use Command R for cost-effective production workloads - Use embed-v3.0 for semantic search 2. **RAG Implementation** - Provide relevant documents in context - Use citations for transparency - Implement reranking for better results 3. **Embeddings** - Use input_type parameter (search_document, search_query, classification) - Batch embed operations - Cache embeddings for repeated documents ## Rate Limits & Pricing ### Rate Limits | Tier | RPM | Tokens/Min | |------|-----|------------| | Trial | 100 | 40K | | Production | 10K | 10M | ## Workflow Serialization Cohere embedding models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel` (language models could already be serialized). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; reranking models are not yet serializable. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Cohere API Documentation](https://docs.cohere.com) ## May 2026 parity updates ### Image input Cohere chat models support image input. Use `types.FileContent` for new code, or `types.ImageContent` for compatibility. Provider option `imageDetail` is forwarded when present under the `cohere` key. ```go messages := []types.Message{{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What is in this image?"}, types.FileContent{ FileData: types.FileData{ Type: types.FileDataTypeURL, URL: "https://example.com/image.png", MediaType: "image/png", }, ProviderOptions: map[string]interface{}{ "cohere": map[string]interface{}{"imageDetail": "auto"}, }, }, }, }} ``` --- # Mistral AI Provider > Setup and usage guide for Mistral AI models in the Go AI SDK, covering available models, provider-specific features, rate limits, and pricing. Canonical URL: https://goaisdk.com/docs/providers/mistral Documentation index: https://goaisdk.com/llms.txt Mistral AI provides powerful open-source and proprietary language models with strong reasoning capabilities, function calling, and competitive pricing. Known for efficient architectures and European data sovereignty. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/mistral" ) ``` ### Configuration ```go provider := mistral.New(mistral.Config{ APIKey: os.Getenv("MISTRAL_API_KEY"), }) model, err := provider.LanguageModel("mistral-large-latest") if err != nil { log.Fatal(err) } ``` ### Get API Key 1. Sign up at [console.mistral.ai](https://console.mistral.ai) 2. Create API key 3. Set environment variable: ```bash export MISTRAL_API_KEY=... ``` ## Available Models ### Language Models Use the constants in `pkg/providers/mistral/model_ids.go` instead of raw strings. See [Mistral's model documentation](https://docs.mistral.ai/getting-started/models/) for context windows and pricing. | Model ID | Go constant | Best for | |----------|-------------|----------| | `mistral-large-latest` | `mistral.ModelMistralLargeLatest` | Complex reasoning, code | | `mistral-medium-latest` | `mistral.ModelMistralMediumLatest` | Balanced quality and latency | | `mistral-medium-3.5` | `mistral.ModelMistralMedium35` | Balanced quality and latency | | `mistral-small-latest` | `mistral.ModelMistralSmallLatest` | Fast, cost-effective | | `magistral-medium-latest` | `mistral.ModelMagistralMediumLatest` | Reasoning | | `ministral-14b-latest` | `mistral.ModelMinistral14BLatest` | Small, efficient | | `codestral-latest` | `mistral.ModelCodestralLatest` | Code generation | | `voxtral-small-latest` | `mistral.ModelVoxtralSmallLatest` | Speech and audio | ### Embedding Models | Model ID | Dimensions | Best For | | ---------- | ----------- | ---------- | | mistral-embed | 1024 | Semantic search | `mistral.EmbeddingModel` batches at most 32 inputs per call (`MaxEmbeddingsPerCall() == 32`) and does not support parallel calls (`SupportsParallelCalls() == false`), matching the Mistral API's own limits — `ai.EmbedMany` automatically splits larger input sets into sequential batched requests. ### Speech and Transcription (Voxtral) ```go speechModel, err := provider.SpeechModel("voxtral-mini-tts-latest") // default when empty if err != nil { log.Fatal(err) } result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello from Mistral.", }) ``` ```go transcriptionModel, err := provider.TranscriptionModel("voxtral-mini-latest") // default when empty if err != nil { log.Fatal(err) } transcript, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioBytes, }) ``` Other Voxtral model IDs include `mistral.ModelVoxtralSmall2507` / `voxtral-small-2507` and `mistral.ModelVoxtralSmallLatest`. ## Provider-Specific Features ### Cached Token Usage Mistral usage responses can include cache fields such as: - `num_cached_tokens` - `cache_read_input_tokens` - `cache_creation_input_tokens` These are mapped into `Usage.InputDetails.CacheReadTokens`, `Usage.InputDetails.CacheWriteTokens`, and `Usage.Raw`. ### Function Calling Mistral supports parallel function calling: ```go tools := []types.Tool{ { Name: "get_weather", Description: "Get weather for location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]string{"type": "string"}, }, "required": []string{"location"}, }, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in Paris and London?", Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) // Processes both calls in parallel ``` ### JSON Mode Structured output generation: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Extract: John, 30, engineer", ResponseFormat: &provider.ResponseFormat{Type: "json_object"}, }) ``` If you supply a JSON schema but disable structured outputs (`providerOptions.mistral.structuredOutputs: false`), Mistral still receives a plain `{"type": "json_object"}` response format, but the schema is preserved by injecting it into the system message so the model still sees the intended shape: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Extract: John, 30, engineer", ResponseFormat: &provider.ResponseFormat{ Type: "json", Schema: mySchema, }, ProviderOptions: map[string]interface{}{ "mistral": map[string]interface{}{"structuredOutputs": false}, }, }) ``` ### Code Generation Codestral is optimized for code: ```go codeModel, err := provider.LanguageModel("codestral-latest") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: codeModel, Prompt: "Write a binary search function in Go", }) ``` ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/mistral" ) func main() { ctx := context.Background() provider := mistral.New(mistral.Config{ APIKey: os.Getenv("MISTRAL_API_KEY"), }) model, err := provider.LanguageModel("mistral-large-latest") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain Mistral AI advantages", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Best Practices 1. **Model Selection** - Use mistral-large for complex reasoning - Use mistral-small for fast, cost-effective tasks - Use a vision-capable model for image input - Use codestral for code generation 2. **Cost Optimization** - Use smaller models for simple tasks - Implement caching - Monitor token usage 3. **Function Calling** - Define clear tool descriptions - Handle parallel calls efficiently ## Rate Limits & Pricing ### Rate Limits | Tier | RPM | Tokens/Min | |------|-----|------------| | Free | 10 | 10K | | Premium | 1000 | 1M | ## Workflow Serialization Mistral embedding, speech, and transcription models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel` (language models could already be serialized). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Mistral AI Documentation](https://docs.mistral.ai) ## May 2026 parity updates ### Mistral Medium 3 and 3.5 Use `mistral.ModelMistralMedium3` for `mistral-medium-3` and `mistral.ModelMistralMedium35` for `mistral-medium-3.5`. ```go p := mistral.New(mistral.Config{APIKey: os.Getenv("MISTRAL_API_KEY")}) model, err := p.LanguageModel(mistral.ModelMistralMedium35) ``` ### Cached token usage Mistral cached token counts map into `types.Usage.InputDetails.CacheReadTokens` when returned by the provider. --- # Groq Provider > Setup and usage guide for Groq ultra-fast LPU inference with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/groq Documentation index: https://goaisdk.com/llms.txt Groq provides the fastest AI inference in the world using custom Language Processing Units (LPUs). Offers speeds up to 10x faster than traditional GPUs, perfect for latency-sensitive applications. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/groq" ) ``` ### Configuration ```go provider := groq.New(groq.Config{ APIKey: os.Getenv("GROQ_API_KEY"), }) model, err := provider.LanguageModel("llama-3.3-70b-versatile") if err != nil { log.Fatal(err) } ``` Groq has its own provider package (`pkg/providers/groq`); it does not support embeddings, image generation, or speech synthesis — `EmbeddingModel`, `ImageModel`, and `SpeechModel` all return errors. `LanguageModel` defaults to `"mixtral-8x7b-32768"` when the model ID is empty. ### Get API Key 1. Sign up at [console.groq.com](https://console.groq.com) 2. Create API key 3. Set environment variable: ```bash export GROQ_API_KEY=gsk_... ``` ## Available Models ### Language Models | Model ID | Context | Tokens/Sec | Best For | |----------|---------|-----------|----------| | llama-3.3-70b-versatile | 128K | 300+ | Best balance | | llama-3.2-90b-vision-preview | 128K | 250+ | Multimodal | | llama-3.1-70b-versatile | 128K | 300+ | General purpose | | mixtral-8x7b-32768 | 32K | 500+ | Very fast | | gemma2-9b-it | 8K | 800+ | Ultra-fast | All models FREE in public beta with rate limits. ## Provider-Specific Features ### Ultra-Fast Streaming Groq excels at real-time streaming: ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a story"}) if err != nil { log.Fatal(err) } defer stream.Close() // Tokens arrive at 300-800 tokens/second for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } ``` ### Tool Use Function calling with high-speed inference: ```go tools := []types.Tool{ { Name: "calculate", Description: "Perform calculation", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "expression": map[string]string{"type": "string"}, }, }, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What is 234 * 567?", Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) ``` ### JSON Mode ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Extract structured data from: John, age 30", ResponseFormat: &provider.ResponseFormat{Type: "json_object"}, }) ``` ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/groq" ) func main() { ctx := context.Background() provider := groq.New(groq.Config{ APIKey: os.Getenv("GROQ_API_KEY"), }) model, err := provider.LanguageModel("llama-3.3-70b-versatile") if err != nil { log.Fatal(err) } start := time.Now() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain Groq LPU technology"}) if err != nil { log.Fatal(err) } elapsed := time.Since(start) fmt.Println(result.Text) fmt.Printf("\nCompleted in %v (%.0f tokens/sec)\n", elapsed, float64(result.Usage.GetOutputTokens())/elapsed.Seconds()) } ``` ### Real-Time Chat Application ```go func streamChat(ctx context.Context, model provider.LanguageModel, message string) { stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: message, }) if err != nil { log.Fatal(err) } defer stream.Close() // Ultra-fast streaming for real-time UX for chunk := range stream.Chunks() { fmt.Print(chunk.Text) // No perceptible delay between tokens } fmt.Println() } ``` ## Best Practices 1. **Model Selection** - Use llama-3.3-70b for best quality - Use mixtral-8x7b for speed - Use gemma2-9b for ultra-fast responses 2. **Leverage Speed** - Use streaming for real-time UX - Build interactive applications - Enable instant feedback loops 3. **Cost Management** - Free during beta period - Monitor rate limits - Plan for future pricing ## Rate Limits ### Free Tier | Model | RPM | RPD | Tokens/Min | |-------|-----|-----|------------| | Llama 3.3 70B | 30 | 14,400 | 6,000 | | Mixtral 8x7B | 30 | 14,400 | 5,000 | | Gemma 9B | 30 | 14,400 | 15,000 | ## Error Handling ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) if err != nil { if strings.Contains(err.Error(), "rate_limit") { log.Println("Rate limited - wait before retry") time.Sleep(time.Second * 2) } } ``` ## Transcription `provider.TranscriptionModel(modelID)` (default `"whisper-large-v3-turbo"`) wraps Groq's Whisper endpoint: ```go transcriptionModel, err := provider.TranscriptionModel("whisper-large-v3-turbo") if err != nil { log.Fatal(err) } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioBytes, }) ``` ## Assistant Content With Tool Calls Groq sends the assistant message's text content verbatim on tool-call turns, including an empty string, rather than `null` — matching the TypeScript SDK. See [OpenAI-compatible providers: Assistant Content Serialization](https://goaisdk.com/docs/providers/openai-compatible.md#assistant-content-serialization) for how this compares across providers. ## Workflow Serialization Groq transcription models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel` (language models could already be serialized). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Groq Documentation](https://console.groq.com/docs) --- # xAI Provider > Setup and usage guide for xAI's Grok models in the Go AI SDK, covering the Responses API, provider options, image/video generation, and speech support. Canonical URL: https://goaisdk.com/docs/providers/xai Documentation index: https://goaisdk.com/llms.txt xAI provides Grok models with real-time knowledge through X (Twitter) integration, image/video generation, and strong reasoning capabilities. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/xai" ) ``` ### Configuration ```go provider := xai.New(xai.Config{ APIKey: os.Getenv("XAI_API_KEY"), }) // Responses API (the only language model API xAI supports) model, err := provider.LanguageModel("grok-3") ``` ### Get API key 1. Sign up at [x.ai](https://x.ai) 2. Request API access 3. Set environment variable: ```bash export XAI_API_KEY=xai-... ``` ## Available models ### Language models | Model ID | Description | |----------|-------------| | `grok-3` | Latest Grok 3 model | | `grok-3-latest` | Alias for latest Grok 3 | | `grok-3-mini` | Smaller, faster Grok 3 variant | | `grok-4` | Latest Grok 4 model | | `grok-4-latest` | Alias for latest Grok 4 | | `grok-4.3` | Current Grok 4.3 model | | `grok-4.20-multi-agent` | Multi-agent orchestration | | `grok-4.20-reasoning` | Reasoning-optimized | | `grok-4.20-non-reasoning` | Fast non-reasoning | ### Image models | Model ID | Description | |----------|-------------| | `grok-imagine-image` | Image generation | | `grok-imagine-image-pro` | Higher quality image generation | ### Video models | Model ID | Description | |----------|-------------| | `grok-imagine-video` | Video generation | ### Audio models xAI speech and transcription use provider factory methods rather than named model IDs. The Go provider method keeps the shared provider interface shape, but the model ID parameter is ignored to match the TypeScript provider default. ```go speechModel, _ := provider.SpeechModel("") transcriptionModel, _ := provider.TranscriptionModel("") ``` ## Responses API (only) `LanguageModel()` returns a Responses API model. Starting in v0.4.0 this was the default (with the legacy Chat Completions API available as an opt-in escape hatch); as of v0.5.0 the Chat Completions API (`ChatCompletionsLanguageModel()`) has been removed entirely, matching the upstream TypeScript SDK. See the migration guide if you were still using `ChatCompletionsLanguageModel()`. ```go model, _ := provider.LanguageModel("grok-3") result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's happening in AI today?", }) ``` ## Provider options ### Provider-executed web search Use `xai.WebSearch` with the Responses API. `EnableImageSearch` maps to xAI's `enable_image_search` request field. ```go enableImageSearch := true result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Show me images of SpaceX Starship on the launch pad.", Tools: []types.Tool{ xai.WebSearch(xai.WebSearchConfig{ EnableImageSearch: &enableImageSearch, }), }, }) ``` The legacy chat-completions `SearchParameters` option was removed along with `ChatCompletionsLanguageModel()`. Use the provider-executed `WebSearch` or `XSearch` tools on the Responses API instead. ### Reasoning summary Control reasoning output in Responses API: ```go result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing", ProviderOptions: map[string]interface{}{ "xai": map[string]interface{}{ "reasoningSummary": "detailed", // "auto", "concise", or "detailed" }, }, }) ``` ### File URL inputs The xAI Responses API path accepts public file URLs for non-image content, including `application/pdf` and `text/*` media types. The SDK serializes these as `input_file` with `file_url`, matching the TypeScript provider. Inline non-image bytes remain unsupported by xAI; use a public URL or provider file reference for PDFs and text files. ```go result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Summarize this report."}, types.FileContent{ MediaType: "application/pdf", FileData: types.FileData{ Type: types.FileDataTypeURL, URL: "https://example.com/report.pdf", }, }, }, }}, }) ``` ### Cost metadata xAI Responses usage metadata is exposed under `ProviderMetadata["xai"]`. `costInUsdTicks` is forwarded directly, and token costs are available in the nested `cost` map using TypeScript-compatible camelCase keys. ```go meta := result.ProviderMetadata["xai"].(map[string]interface{}) fmt.Println(meta["costInUsdTicks"]) if cost, ok := meta["cost"].(map[string]interface{}); ok { fmt.Println(cost["inputTokensCost"]) fmt.Println(cost["outputTokensCost"]) } ``` ### Logprobs ```go result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello", ProviderOptions: map[string]interface{}{ "xai": map[string]interface{}{ "logprobs": true, "topLogprobs": 5, }, }, }) ``` ## Image generation ```go imageModel, _ := provider.ImageModel("grok-imagine-image") result, _ := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "A futuristic cityscape at sunset", ProviderOptions: map[string]interface{}{ "xai": map[string]interface{}{ "quality": "high", // "low", "medium", "high" }, }, }) ``` Supports multiple input images for editing, b64_json output format, and `CostInUsdTicks` in provider metadata. ## Video generation ```go videoModel, _ := provider.VideoModel("grok-imagine-video") result, _ := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{Text: "A cat playing piano"}, }) ``` Returns `ModerationError` if content is filtered. `CostInUsdTicks` available in provider metadata. ### Async video `xai.VideoModel` implements `provider.VideoModelStarter` / `VideoModelStatusChecker`, so it also works with the fire-and-forget `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus` calls instead of the polling `ai.GenerateVideo`: ```go started, _ := ai.ExperimentalStartVideo(ctx, ai.StartVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{Text: "A cat playing piano"}, }) // Persist started.Operation, then later (even from another process): status, _ := ai.ExperimentalGetVideoStatus(ctx, videoModel, ai.GetVideoStatusOptions{ Operation: started.Operation, }) if status.Status == provider.VideoOperationStatusCompleted { fmt.Println(status.Videos) } ``` ## Speech and transcription `SpeechModel("")` maps to xAI's `/v1/tts` endpoint. It defaults to voice `"eve"`, language `"auto"`, and MP3 output, matching the TypeScript provider. Use provider options under the `"xai"` key for xAI-specific audio controls. ```go speechModel, _ := provider.SpeechModel("") result, _ := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello from Go-AI.", Voice: "ara", OutputFormat: "wav", ProviderOptions: map[string]interface{}{ "xai": map[string]interface{}{ "sampleRate": 44100, "optimizeStreamingLatency": 1, "textNormalization": true, }, }, }) ``` `TranscriptionModel("")` maps to `/v1/stt`. Multipart request fields match the TypeScript provider, including repeated `keyterm` fields and the audio `file` part last. ```go transcriptionModel, _ := provider.TranscriptionModel("") transcript, _ := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioBytes, MimeType: "audio/wav", ProviderOptions: map[string]interface{}{ "xai": map[string]interface{}{ "language": "en", "diarize": true, "keyterm": []string{"Go-AI", "Grok"}, "fillerWords": true, }, }, }) fmt.Println(transcript.Text) ``` ## Best practices 1. **Use the Responses API** — it is the default and recommended path 2. **Leverage ReasoningSummary** — get insight into model reasoning without full chain-of-thought overhead 3. **Check CostInUsdTicks** — available in image and video provider metadata for cost tracking ## See also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [xAI Documentation](https://docs.x.ai) ## May 2026 parity updates ### Non-image file URLs The xAI Responses provider accepts non-image file URLs by serializing URL file parts as `input_file` with `file_url`. ```go messages := []types.Message{{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Summarize this report."}, types.FileContent{FileData: types.FileData{ Type: types.FileDataTypeURL, URL: "https://example.com/report.pdf", MediaType: "application/pdf", }}, }, }} ``` ### Cost and ZDR metadata xAI response, image, and video calls surface `costInUsdTicks` under provider metadata when the API returns it. Encrypted reasoning is preserved for zero-data-retention follow-up turns. ### File uploads Use `ai.UploadFile` with `xai.New(...)`. The returned `ProviderReference` is `map[string]string{"xai": fileID}`. ## Workflow serialization xAI's image, speech, and transcription models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel` (language models could already be serialized). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; video models are not yet serializable. --- # DeepSeek Provider > Setup and usage guide for DeepSeek's reasoning models in the Go AI SDK, covering available models, provider-specific features, and May 2026 parity updates. Canonical URL: https://goaisdk.com/docs/providers/deepseek Documentation index: https://goaisdk.com/llms.txt DeepSeek provides advanced reasoning models with exceptional performance at competitive prices. Known for strong math, coding, and logical reasoning capabilities. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/deepseek" ) ``` ### Configuration ```go provider := deepseek.New(deepseek.Config{ APIKey: os.Getenv("DEEPSEEK_API_KEY"), }) model, err := provider.LanguageModel("deepseek-v4-flash") if err != nil { log.Fatal(err) } ``` ### Get API Key 1. Sign up at [platform.deepseek.com](https://platform.deepseek.com) 2. Create API key 3. Set environment variable: ```bash export DEEPSEEK_API_KEY=sk-... ``` ## Available Models ### Language Models Use the constants in `pkg/providers/deepseek/model_ids.go` instead of raw strings. See [DeepSeek's pricing and model list](https://api-docs.deepseek.com/quick_start/pricing) for context windows and prices. | Model ID | Go constant | Notes | |----------|-------------|-------| | `deepseek-flash` | `deepseek.ModelDeepSeekFlash` | Alias for the latest Flash model | | `deepseek-v4-flash` | `deepseek.ModelDeepSeekV4Flash` | DeepSeek V4 Flash | | `deepseek-v4-pro` | `deepseek.ModelDeepSeekV4Pro` | DeepSeek V4 Pro | | `deepseek-v4-flash-vision-exp` | `deepseek.ModelDeepSeekV4FlashVisionExp` | Experimental, image input | V4 models enable thinking mode by default. `deepseek-chat` and `deepseek-reasoner` still work but are deprecated aliases for the V3 and R1 generations. ## Provider-Specific Features ### Reasoning Mode Extended thinking for complex problems: ```go model, err := provider.LanguageModel("deepseek-v4-pro") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Solve this logic puzzle: Three friends...", }) // Response includes reasoning process if result.ReasoningText != "" { fmt.Println("Reasoning:", result.ReasoningText) } ``` ### Code Generation Optimized for programming tasks: ```go codeModel, err := provider.LanguageModel("deepseek-v4-flash") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: codeModel, Prompt: "Implement a thread-safe LRU cache in Go", }) ``` ### Vision (DeepSeek V4) `deepseek-v4*` models return `SupportsImageInput() == true`; image file parts are sent to the API instead of being rejected with an error, as on earlier DeepSeek models: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, // a deepseek-v4* model Messages: []types.Message{{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Describe this image"}, types.FileContent{MediaType: "image/png", URL: "https://example.com/image.png"}, }, }}, }) ``` ### Files API `provider.Files()` returns a `provider.FilesAPI`, or use the higher-level `ai.UploadFile` helper, to upload files that are later referenced from chat messages by `file_id`: ```go uploaded, err := ai.UploadFile(ctx, ai.UploadFileOptions{ API: provider, Data: fileBytes, MediaType: "application/pdf", }) if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Summarize this document."}, types.FileContent{FileData: types.FileData{ Type: types.FileDataTypeReference, Reference: uploaded.ProviderReference, }}, }, }}, }) ``` DeepSeek has no batch API and no additional Files methods beyond upload — this matches the upstream TypeScript `@ai-sdk/deepseek` package, which has neither. ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/deepseek" ) func main() { ctx := context.Background() provider := deepseek.New(deepseek.Config{ APIKey: os.Getenv("DEEPSEEK_API_KEY"), }) model, err := provider.LanguageModel("deepseek-v4-flash") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain gradient descent", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Complex Reasoning ```go reasoner, err := provider.LanguageModel("deepseek-v4-pro") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: reasoner, Prompt: "A farmer has 100 feet of fence. What's the maximum area he can enclose?", }) fmt.Println("Reasoning:", result.ReasoningText) fmt.Println("Answer:", result.Text) ``` ## Best Practices 1. **Model Selection** - Use deepseek-v4-flash for general tasks - Use deepseek-v4-pro for math/logic problems - Use deepseek-v4-pro for programming and long agent runs 2. **Cost Optimization** - Typically much cheaper than other frontier models - Great for high-volume applications - Excellent price-performance ratio 3. **V4 Reasoning Round-Trips** - For `deepseek-v4*` multi-turn chats, include prior assistant `ReasoningContent` so `reasoning_content` is preserved across turns. ## Rate Limits & Pricing ### Rate Limits | Model | RPM | Tokens/Min | |-------|-----|------------| | DeepSeek Chat | 60 | 200K | | DeepSeek Reasoner | 30 | 100K | ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [DeepSeek Documentation](https://platform.deepseek.com/docs) ## May 2026 parity updates ### DeepSeek v4 reasoning The DeepSeek provider preserves `reasoning_content` for DeepSeek v4-style multi-turn requests. Reasoning output maps to `types.ReasoningContent` and is forwarded when the result is reused as message history. ```go p := deepseek.New(deepseek.Config{APIKey: os.Getenv("DEEPSEEK_API_KEY")}) model, err := p.LanguageModel("deepseek-v4-pro") ``` --- # Perplexity Provider > Setup guide for Perplexity's Agent API in the Go AI SDK, covering model selection, image input, embeddings, and migrating from the Chat Completions API. Canonical URL: https://goaisdk.com/docs/providers/perplexity Documentation index: https://goaisdk.com/llms.txt Perplexity's **Agent API** (`/v1/agent`) combines LLM reasoning with real-time web search, URL fetching, and other native tools, returning cited answers. > **Breaking change:** as of v0.5.0, the Go SDK's Perplexity provider > targets the Agent API instead of the pre-migration Sonar Chat Completions > API (`/chat/completions`). This mirrors an equivalent breaking migration in > the upstream TypeScript AI SDK. See [Migrating from the Chat Completions > API](#migrating-from-the-chat-completions-api) below. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/perplexity" ) ``` ### Configuration ```go provider := perplexity.New(perplexity.Config{ APIKey: os.Getenv("PERPLEXITY_API_KEY"), }) model, err := provider.LanguageModel("low") if err != nil { log.Fatal(err) } ``` The provider sends `X-Pplx-Integration: vercel-ai-sdk` on every request so Perplexity can attribute traffic to the SDK; pass your own value in `Config.Headers` to override it: ```go provider := perplexity.New(perplexity.Config{ APIKey: os.Getenv("PERPLEXITY_API_KEY"), Headers: map[string]string{"X-Pplx-Integration": "my-app"}, }) ``` ### Get API Key 1. Sign up at [perplexity.ai](https://www.perplexity.ai) 2. Get an API key from settings 3. Set the environment variable: ```bash export PERPLEXITY_API_KEY=pplx-... ``` ## Model Selection The Agent API accepts either a **routing preset** or a **direct model ID**. ### Presets Presets (`fast`, `low`, `medium`, `high`, `xhigh`) let Perplexity choose the underlying model for you, trading off latency/cost against quality: ```go model, _ := provider.LanguageModel("low") ``` A preset is sent as `{"preset": "low"}` in the request body. ### Direct model IDs Any other string (e.g. an upstream model like `"openai/gpt-5.1"`, or a legacy Sonar model like `"sonar-pro"`) is sent as `{"model": ""}`: ```go model, _ := provider.LanguageModel("sonar-pro") ``` Legacy Sonar model IDs (`sonar`, `sonar-pro`, `sonar-reasoning`, `sonar-reasoning-pro`, `sonar-deep-research` — exposed as `perplexity.ModelSonar`, `perplexity.ModelSonarPro`, etc.) remain valid model IDs on the Agent API; they are **not** automatically mapped to a preset. ## Provider-Specific Features ### Web Search and Citations Search-augmented answers return citations as `source` content parts: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What are the latest developments in quantum computing?", }) for _, part := range result.Content { if source, ok := part.(types.SourceContent); ok { fmt.Printf("%s: %s\n", source.Title, source.URL) } } ``` ### Native Agent API Tools Enable Perplexity's built-in tools (`web_search`, `fetch_url`, `people_search`, `finance_search`, `sandbox`, `mcp`, `connector`) via `providerOptions.perplexity.tools`: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's Apple's current stock price?", ProviderOptions: map[string]interface{}{ "perplexity": map[string]interface{}{ "tools": []interface{}{ map[string]interface{}{"type": "finance_search"}, }, }, }, }) ``` ### AI SDK Function Tools Regular AI SDK function tools are also supported and are sent alongside any native tools: ```go weatherTool := types.Tool{ Name: "weather", Description: "Get the current weather for a city", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{"city": map[string]interface{}{"type": "string"}}, "required": []interface{}{"city"}, }, } ``` ### Reasoning Effort Use the standardized `Reasoning` call option, or set `providerOptions.perplexity.reasoning.effort` directly (`"minimal"`, `"low"`, `"medium"`, `"high"`, or `"xhigh"`) — the provider option takes precedence: ```go reasoning := types.ReasoningHigh result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, Reasoning: &reasoning, }) ``` ### Structured Output `response_format` with a JSON schema maps to the Agent API's `response_format.json_schema`: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, ResponseFormat: &provider.ResponseFormat{ Type: "json", Name: "answer", Schema: mySchema, }, }) ``` ### Other Agent API Options `providerOptions.perplexity` also accepts: `instructions`, `models` (a fallback model list), `max_steps`, `max_tool_calls`, `previous_response_id` (continue an earlier Agent API response), `store`, `language_preference`, and `skills`. Unrecognized keys are forwarded verbatim so new Agent API options work without an SDK update. ### Provider Metadata `result.ProviderMetadata["perplexity"]` carries usage/cost details as a `perplexity.PerplexityMetadata` value: ```go if meta, ok := result.ProviderMetadata["perplexity"].(perplexity.PerplexityMetadata); ok { fmt.Printf("search queries: %v\n", meta.Usage.NumSearchQueries) if meta.Cost != nil { fmt.Printf("total cost: %v %v\n", *meta.Cost.TotalCost, *meta.Cost.Currency) } for name, tc := range meta.ToolCalls { fmt.Printf("%s invoked %v times\n", name, *tc.Invocation) } } ``` ## Image Input The Agent API accepts image input only — PDFs and other document types supported by the pre-migration Sonar Chat Completions API are rejected: ```go messages := []types.Message{ {Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Describe this image"}, types.FileContent{MediaType: "image/png", URL: "https://example.com/image.png"}, }}, } ``` ## Embeddings Embeddings are unaffected by the Agent API migration — see `perplexity.PerplexityEmbeddingModelID` (`pplx-embed-v1-0.6b`, `pplx-embed-v1-4b`). ## Migrating from the Chat Completions API If you were using the pre-migration Sonar Chat Completions provider, the following changed: - **Endpoint**: `/chat/completions` → `/v1/agent`. - **`SupportsTools()`**, **`SupportsStructuredOutput()`**, and **`SupportsImageInput()`** now all return `true` (previously `false`). - **`providerOptions.perplexity`**: every field from the old Sonar options (`search_recency_filter`, `search_domain_filter`, `web_search_options`, `return_images`, `return_related_questions`, `disable_search`, `reasoning_effort`, `media_response`, `stream_mode`, etc.) is **removed and not aliased**. Use the new Agent API option shape (`instructions`, `tools`, `models`, `max_steps`, `max_tool_calls`, `previous_response_id`, `store`, `language_preference`, `reasoning.effort`, `skills`). - **`PerplexityMetadata`** (`result.ProviderMetadata["perplexity"]`) shape changed: - `Images` is now always `nil` (previously populated from `response.images`; the Agent API does not expose image results through this SDK version). - `Cost` gained `Currency`, `CacheCreationCost`, `CacheReadCost`, and `ToolCallsCost` fields. - A new `ToolCalls map[string]PerplexityToolCallUsage` field surfaces per-tool invocation counts. - `Usage.CitationTokens` is now always `nil` (the Agent API does not report a citation-token count); `Usage.NumSearchQueries` is now derived from `usage.tool_calls_details` instead of a flat `num_search_queries` field. - **Usage semantics**: reasoning tokens are now a **subset** of `outputTokens.total` (standard convention), not additional to it as in the pre-migration Sonar usage shape. - **PDF/document input** is no longer supported; only images. - **Sources**: `result.Content` now carries `types.SourceContent` parts (with `sourceType: "url"`) built from search results, fetched URLs, and message annotations, instead of a flat `Citations []string` list. - Tool calls and native tool results now appear as ordered `ToolCallContent` / `SourceContent` parts in `result.Content`, matching every other tool-calling provider in this SDK. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Perplexity Agent API Documentation](https://docs.perplexity.ai) --- # Together AI Provider > Setup and usage guide for Together AI open-source model hosting with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/together Documentation index: https://goaisdk.com/llms.txt Together AI provides fast inference for open-source models including Llama, Mixtral, Qwen, and more. Offers competitive pricing and high-quality model hosting. ## Setup ### Installation Together has its own dedicated package — it is not built on top of the generic `openai` package: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/together" ) ``` ### Configuration ```go provider := together.New(together.Config{ APIKey: os.Getenv("TOGETHER_API_KEY"), }) model, err := provider.LanguageModel("meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo") if err != nil { log.Fatal(err) } ``` `together.Config` also accepts `BaseURL` (default `https://api.together.xyz`) and `Headers`. ### Get API Key ```bash export TOGETHER_API_KEY=... ``` `TOGETHER_AI_API_KEY` is accepted as a deprecated fallback and logs a deprecation warning; prefer `TOGETHER_API_KEY`. ## Available Models ### Language Models | Model ID | Context | Best For | | ---------- | --------- | ---------- | | meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo | 131K | General purpose | | meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo | 131K | Fast, cheap | | mistralai/Mixtral-8x7B-Instruct-v0.1 | 32K | Balanced | | Qwen/Qwen2.5-72B-Instruct-Turbo | 32K | Multilingual | ### Image Generation | Model ID | Quality | Best For | | ---------- | --------- | ---------- | | stabilityai/stable-diffusion-xl-base-1.0 | High | Images | | black-forest-labs/FLUX.1-schnell | High | Fast generation | ### Embeddings ```go embeddingModel, err := provider.EmbeddingModel("togethercomputer/m2-bert-80M-8k-retrieval") ``` ### Reranking `provider.RerankingModel(modelID)` calls `POST /rerank`: ```go reranker, err := provider.RerankingModel("Salesforce/Llama-Rank-v1") if err != nil { log.Fatal(err) } topN := 3 result, err := reranker.DoRerank(ctx, &provider.RerankOptions{ Query: "What is the capital of France?", Documents: []string{"Paris is the capital of France.", "Berlin is the capital of Germany."}, TopN: &topN, ProviderOptions: map[string]interface{}{ "togetherai": map[string]interface{}{ // RankFields selects which keys of a JSON-object document to // rank on; defaults to every supplied key. "rankFields": []string{"text"}, }, }, }) ``` ## Provider Options Together-specific chat options go under the `"togetherai"` key (the TS provider options name); the Go package's own name, `"together"`, is also accepted as a fallback: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain open-source AI", ProviderOptions: map[string]interface{}{ "togetherai": map[string]interface{}{ "topK": 40, }, }, }) ``` Together does not support speech synthesis or transcription — those model factories return an error. ## Workflow Serialization Together image models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel` (language models could already be serialized). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; embedding and reranking models are not yet serializable. ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/together" ) func main() { ctx := context.Background() provider := together.New(together.Config{ APIKey: os.Getenv("TOGETHER_API_KEY"), }) model, err := provider.LanguageModel("meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain open-source AI"}) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Best Practices 1. **Model Selection** - Use Llama 3.1 70B for best quality - Use Llama 3.1 8B for cost efficiency - Use Mixtral for balanced performance 2. **Cost Optimization** - Open-source models are cost-effective - Monitor usage and optimize prompts ## Rate Limits Varies by plan - check dashboard for limits. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Together AI Documentation](https://docs.together.ai) --- # Fireworks AI Provider > Setup and usage guide for Fireworks AI's fast inference models in the Go AI SDK, covering available models, async models, and production best practices. Canonical URL: https://goaisdk.com/docs/providers/fireworks Documentation index: https://goaisdk.com/llms.txt Fireworks AI provides blazing-fast inference for open-source models with optimized serving infrastructure. Excellent for production deployments requiring speed and reliability. Streaming language-model requests include `stream_options.include_usage: true`, so final stream usage is available when Fireworks returns it. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/fireworks" ) ``` ### Configuration ```go prov := fireworks.New(fireworks.Config{ APIKey: os.Getenv("FIREWORKS_API_KEY"), }) model, err := prov.LanguageModel("accounts/fireworks/models/llama-v3p1-70b-instruct") ``` ### Get API Key ```bash export FIREWORKS_API_KEY=fw_... ``` ## Available Models ### Language Models | Model ID | Context | Best For | | ---------- | --------- | ---------- | | accounts/fireworks/models/llama-v3p1-70b-instruct | 128K | High quality | | accounts/fireworks/models/llama-v3p1-8b-instruct | 128K | Fast | | accounts/fireworks/models/mixtral-8x7b-instruct | 32K | Balanced | | accounts/fireworks/models/kimi-k2p5 | 128K | Extended reasoning | ### Image Models | Model ID | Type | Notes | |----------|------|-------| | accounts/fireworks/models/flux-kontext-dev | Text-to-image, image editing | Async polling | | accounts/fireworks/models/flux-kontext-pro | Text-to-image, image editing | Async polling | | accounts/fireworks/models/flux-kontext-max | Text-to-image, image editing | Async polling | | accounts/fireworks/models/stable-diffusion-xl-1024-v1-0 | Text-to-image | Sync | | accounts/fireworks/models/playground-v2-5-1024px-aesthetic | Text-to-image | Sync, supports size | ## Examples ### Language Model ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/fireworks" ) func main() { ctx := context.Background() prov := fireworks.New(fireworks.Config{ APIKey: os.Getenv("FIREWORKS_API_KEY"), }) model, err := prov.LanguageModel("accounts/fireworks/models/llama-v3p1-70b-instruct") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain Fireworks AI"}) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Async Image Generation (flux-kontext-\*) `flux-kontext-*` models use a two-phase async flow: submit a request, then poll for the result. The Go AI SDK handles this automatically. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/fireworks" ) func main() { prov := fireworks.New(fireworks.Config{ APIKey: os.Getenv("FIREWORKS_API_KEY"), }) model, err := prov.ImageModel("accounts/fireworks/models/flux-kontext-dev") if err != nil { log.Fatal(err) } // The SDK submits asynchronously and polls until ready (up to 2 minutes) result, err := model.DoGenerate(context.Background(), &provider.ImageGenerateOptions{ Prompt: "A serene mountain lake at golden hour, photorealistic", AspectRatio: "16:9", }) if err != nil { log.Fatal(err) } fmt.Println("Image URL:", result.URL) } ``` ### Image Editing with flux-kontext `flux-kontext-pro` and `flux-kontext-max` support image editing by providing an input image: ```go // Edit using a URL result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{ Prompt: "Change the sky to a dramatic sunset", Files: []provider.ImageFile{ {Type: "url", URL: "https://example.com/landscape.png"}, }, }) // Edit using file data imageBytes, _ := os.ReadFile("photo.jpg") result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{ Prompt: "Add a red barn in the background", Files: []provider.ImageFile{ {Type: "file", Data: imageBytes, MediaType: "image/jpeg"}, }, }) ``` ## Async Models — How It Works Models with IDs containing `flux-kontext` use a two-phase generation API: 1. **Submit** — POST to `/v1/workflows/{modelID}` with the generation parameters. Returns a `request_id`. 2. **Poll** — POST to `/v1/workflows/{modelID}/get_result` every 500ms (configurable via `fireworks.Config.ImagePollIntervalMs`) until status is `"Ready"`. 3. **Result** — When `"Ready"`, the response contains a URL in `result.sample`. The Go AI SDK handles this flow transparently via `DoGenerate`. Polling respects context cancellation and has a 2-minute timeout by default (configurable via `fireworks.Config.ImagePollTimeoutMs`). **Supported poll statuses:** - `Pending` / `Running` → continue polling - `Ready` → success, return image URL - `Error` / `Failed` → return descriptive error **Non-flux-kontext models** (e.g., `stable-diffusion-xl-1024-v1-0`) use a standard synchronous image generation endpoint and are unaffected. ## Best Practices 1. Use flux-kontext models for high-quality image generation and editing 2. For time-sensitive workflows, consider setting a shorter deadline via `context.WithTimeout` 3. The SDK automatically retries polling — do not add your own retry loop around `DoGenerate` ## See Also - [Fireworks AI Documentation](https://docs.fireworks.ai) - [flux-kontext Model Guide](https://docs.fireworks.ai/models/flux-kontext) --- # Replicate Provider > Setup and usage guide for Replicate's model marketplace in the Go AI SDK, covering available models, workflow serialization, and example API calls. Canonical URL: https://goaisdk.com/docs/providers/replicate Documentation index: https://goaisdk.com/llms.txt Replicate provides a marketplace to run any open-source model through a simple API. Perfect for experimentation and accessing diverse models including image generation, video, and specialized models. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/replicate" ) ``` ### Configuration ```go provider := replicate.New(replicate.Config{ APIKey: os.Getenv("REPLICATE_API_KEY"), }) model, err := provider.LanguageModel("meta/meta-llama-3-70b-instruct") ``` ### Get API Key ```bash export REPLICATE_API_KEY=r8_... ``` ## Available Models ### Language Models | Model ID | Context | Best For | | ---------- | --------- | ---------- | | meta/meta-llama-3-70b-instruct | 8K | General purpose | | mistralai/mixtral-8x7b-instruct-v0.1 | 32K | Balanced | ### Image Generation | Model ID | Quality | Best For | | ---------- | --------- | ---------- | | stability-ai/sdxl | High | Images | | black-forest-labs/flux-pro | Excellent | Professional | | black-forest-labs/flux-schnell | Good | Fast | ### Video Generation | Model ID | Quality | Best For | | ---------- | --------- | ---------- | | stability-ai/stable-video-diffusion | Medium | Video from image | `provider.VideoModel(modelID)` accepts both forms Replicate uses: - **Unversioned** — `"owner/model"` (no `:version`). Requests go to `POST /models/{owner}/{model}/predictions`, which always runs the model's latest version. This is now supported; previously only versioned IDs worked. - **Versioned** — `"owner/model:version-hash"`. Requests go to `POST /predictions` with `version` in the body, pinning an exact version. ```go videoModel, err := provider.VideoModel("stability-ai/stable-video-diffusion") if err != nil { log.Fatal(err) } result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{Text: "A boat sailing across a calm lake"}, }) ``` `VideoModel` also implements `provider.VideoModelStarter` / `VideoModelStatusChecker` for the async `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus` calls. ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/replicate" ) func main() { ctx := context.Background() provider := replicate.New(replicate.Config{ APIKey: os.Getenv("REPLICATE_API_KEY"), }) model, err := provider.LanguageModel("meta/meta-llama-3-70b-instruct") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain Replicate"}) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Image Generation ```go imageModel, err := provider.ImageModel("black-forest-labs/flux-pro") if err != nil { log.Fatal(err) } result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "A cyberpunk city at night", Size: "1024x1024", }) fmt.Printf("Image bytes: %d\n", len(result.Images[0].Data)) ``` ## Best Practices 1. **Model Discovery** - Browse Replicate marketplace for models - Try different models easily - Great for prototyping 2. **Cost Management** - Pay-per-run pricing - Monitor usage carefully - Cache results when possible 3. **Use Cases** - Experimentation and testing - Specialized model access - Image/video generation ## Workflow Serialization Replicate image models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; video models are not yet serializable. ## See Also - [Replicate Documentation](https://replicate.com/docs) - [Model Marketplace](https://replicate.com/explore) --- # HuggingFace Provider > Setup and usage guide for HuggingFace Router Responses API with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/huggingface Documentation index: https://goaisdk.com/llms.txt HuggingFace's Router exposes an OpenAI Responses-API-compatible endpoint (`https://router.huggingface.co/v1/responses`) that proxies to open-source models hosted across HuggingFace's Inference Providers. This is the only API the Go SDK's HuggingFace provider talks to — it mirrors `@ai-sdk/huggingface`. > The HuggingFace Router Responses API does not offer text embeddings or > image generation. `EmbeddingModel`/`ImageModel` return an error; use the > HuggingFace Inference API directly for those. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/huggingface" ) ``` ### Configuration ```go provider := huggingface.New(huggingface.Config{ APIKey: os.Getenv("HUGGINGFACE_API_KEY"), // falls back to this env var if empty }) model, err := provider.LanguageModel("deepseek-ai/DeepSeek-V3-0324") ``` ### Get API Key 1. Sign up at [huggingface.co](https://huggingface.co) 2. Create a token in settings 3. Set the environment variable: ```bash export HUGGINGFACE_API_KEY=hf_... ``` ## Example ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/huggingface" ) func main() { provider := huggingface.New(huggingface.Config{ APIKey: os.Getenv("HUGGINGFACE_API_KEY"), }) model, err := provider.LanguageModel("deepseek-ai/DeepSeek-V3-0324") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Explain transformers", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` Streaming (`ai.StreamText`) uses the Responses API's real Server-Sent Events stream — text, reasoning, tool-call, and source deltas are emitted as they arrive rather than being chunked after the fact. ## Tool calls, reasoning, and sources The Responses API returns tool calls (`function_call`), MCP tool calls/lists (`mcp_call`, `mcp_list_tools`), reasoning blocks, and URL citation sources alongside text. The Go SDK surfaces all of these through the standard `ai.GenerateTextResult`/`ai.StreamTextResult` content and tool-call APIs. ## Provider Options ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "...", ProviderOptions: map[string]interface{}{ "huggingface": map[string]interface{}{ "metadata": map[string]interface{}{"key": "value"}, "instructions": "Be concise", "strictJsonSchema": true, "reasoningEffort": "high", }, }, }) ``` ## Available Models The Responses API accepts any model ID listed at [router.huggingface.co/v1/models](https://router.huggingface.co/v1/models), including: | Model ID | Best For | |----------|----------| | meta-llama/Llama-3.3-70B-Instruct | General purpose | | deepseek-ai/DeepSeek-V3-0324 | General purpose | | deepseek-ai/DeepSeek-R1 | Reasoning | | Qwen/Qwen3-Coder-480B-A35B-Instruct | Coding | | moonshotai/Kimi-K2-Instruct | General purpose | ## Workflow Serialization HuggingFace language models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [HuggingFace Router Documentation](https://huggingface.co/docs/inference-providers) - [Model Hub](https://huggingface.co/models) --- # Ollama Provider > Setup and usage guide for Ollama local model deployment with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/ollama Documentation index: https://goaisdk.com/llms.txt Ollama enables running large language models locally on your machine. Perfect for development, privacy-sensitive applications, and offline usage. ## Setup ### Installation 1. Install Ollama from [ollama.com](https://ollama.com) 2. Pull a model: ```bash ollama pull llama3.1 ``` ### Configuration Ollama has its own provider package; no API key is needed: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/ollama" ) provider := ollama.New(ollama.Config{ // BaseURL defaults to http://localhost:11434 }) model, err := provider.LanguageModel("llama3.1") ``` Ollama does not support image generation, speech synthesis, transcription, or reranking — `ImageModel`, `SpeechModel`, `TranscriptionModel`, and `RerankingModel` all return errors. ## Available Models ### Popular Models | Model | Parameters | Size | Best For | |-------|-----------|------|----------| | llama3.1 | 70B/8B | 40GB/4.7GB | General purpose | | llama3.2 | 3B/1B | 2GB/1.3GB | Fast, efficient | | mistral | 7B | 4.1GB | Balanced | | phi3 | 3.8B | 2.3GB | Small, capable | | qwen2.5 | 72B/7B | 43GB/4.7GB | Multilingual | | codellama | 34B/13B/7B | 19GB/7.4GB/3.8GB | Code generation | ### Specialized Models | Model | Size | Best For | |-------|------|----------| | nomic-embed-text | 137M | Embeddings | | llava | 7B | Vision | | stable-code | 3B | Code completion | ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/ollama" ) func main() { ctx := context.Background() provider := ollama.New(ollama.Config{}) model, err := provider.LanguageModel("llama3.1") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain Ollama"}) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Streaming ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a story"}) if err != nil { log.Fatal(err) } defer stream.Close() for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } ``` ### Vision with LLaVA ```go model, err := provider.LanguageModel("llava") imageData, _ := os.ReadFile("image.jpg") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Describe this image"}, types.FileContent{ Data: imageData, MediaType: "image/jpeg", }, }, }, }, }) ``` ### Embeddings ```go embeddingModel, err := provider.EmbeddingModel("nomic-embed-text") texts := []string{ "Ollama runs models locally", "Local inference is private", } result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: texts, }) ``` ## Best Practices 1. **Hardware Requirements** - 8GB RAM minimum for 7B models - 16GB RAM for 13B models - 32GB+ RAM for 70B models - GPU recommended for better performance 2. **Model Selection** - Start with smaller models (7B) - Use quantized versions for less RAM - Test different models for your use case 3. **Performance** - Use GPU acceleration when available - Adjust context window based on needs - Monitor system resources 4. **Development** - Perfect for development and testing - No API costs - Complete privacy - Works offline ## Managing Models ### Pull Models ```bash # Pull latest version ollama pull llama3.1 # Pull specific version ollama pull llama3.1:70b # Pull with tag ollama pull llama3.1:8b-instruct-q4_0 ``` ### List Models ```bash ollama list ``` ### Remove Models ```bash ollama rm llama3.1 ``` ### Create Custom Model Create Modelfile: ``` FROM llama3.1 PARAMETER temperature 0.8 SYSTEM You are a helpful coding assistant ``` Build model: ```bash ollama create my-coding-assistant -f Modelfile ``` ## Configuration Options ### GPU Configuration ```go // Ollama automatically uses GPU if available // Configure via OLLAMA_NUM_GPU environment variable ``` ### Context Window ```go maxTokens := 4096 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, MaxTokens: &maxTokens, // Adjust based on model }) ``` ### Remote Ollama ```go provider := ollama.New(ollama.Config{ BaseURL: "http://remote-server:11434", }) ``` ## Error Handling ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) if err != nil { if strings.Contains(err.Error(), "connection refused") { log.Fatal("Ollama server not running - run: ollama serve") } if strings.Contains(err.Error(), "model not found") { log.Fatal("Model not found - run: ollama pull llama3.1") } log.Fatal(err) } ``` ## Performance Tips 1. **Use Quantized Models** - q4_0: 4-bit quantization (smallest) - q5_0: 5-bit quantization (balanced) - q8_0: 8-bit quantization (high quality) 2. **Optimize Context** - Reduce context window for faster inference - Clear conversation history periodically 3. **GPU Acceleration** - Ensure CUDA/ROCm installed - Monitor GPU memory usage - Adjust num_gpu parameter ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Ollama Documentation](https://ollama.com/docs) - [Ollama Model Library](https://ollama.com/library) --- # Google Vertex AI Provider > Setup and usage guide for Google Vertex AI enterprise platform with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/google-vertex Documentation index: https://goaisdk.com/llms.txt Google Vertex AI provides enterprise-grade AI platform with Gemini models, custom model training, and MLOps capabilities. Offers enhanced security, compliance, and integration with Google Cloud. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/googlevertex" ) ``` ### Configuration ```go provider, err := googlevertex.New(googlevertex.Config{ Project: os.Getenv("GOOGLE_VERTEX_PROJECT"), Location: "us-central1", AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"), }) if err != nil { log.Fatal(err) } model, err := provider.LanguageModel("gemini-2.0-flash") ``` ### Authentication 1. Install Google Cloud SDK 2. Authenticate: ```bash gcloud auth application-default login ``` 3. Set project: ```bash export GOOGLE_CLOUD_PROJECT=your-project-id ``` The Vertex provider reuses auth per provider instance. You can provide an explicit access token or token generator to override default Google credential discovery, which is useful for tests, service-account impersonation, and custom credential brokers. ## Available Models Use the constants in `pkg/providers/googlevertex/models.go` instead of raw strings. For context windows and pricing, see [Google Cloud's Vertex AI model documentation](https://cloud.google.com/vertex-ai/generative-ai/docs/models). ### Gemini models | Model ID | Go constant | Best for | |----------|-------------|----------| | `gemini-3.8-flash` | `googlevertex.ModelGemini38Flash` | Fast, current-generation Flash | | `gemini-3.7-flash` | `googlevertex.ModelGemini37Flash` | Fast, current-generation Flash | | `gemini-3.6-flash` | `googlevertex.ModelGemini36Flash` | Fast, current-generation Flash | | `gemini-3.5-flash` | `googlevertex.ModelGemini35Flash` | Fast multimodal tasks | | `gemini-3.5-flash-lite` | `googlevertex.ModelGemini35FlashLite` | High-volume, low-cost tasks | | `gemini-3.1-pro-preview` | `googlevertex.ModelGemini31ProPreview` | Advanced reasoning and generation | | `gemini-3.1-flash-lite-preview` | `googlevertex.ModelGemini31FlashLitePreview` | Low-cost preview model | | `gemini-3-pro-preview` | `googlevertex.ModelGemini3ProPreview` | Advanced preview generation | | `gemini-3-flash-preview` | `googlevertex.ModelGemini3FlashPreview` | Fast preview generation | | `gemini-2.5-pro` | `googlevertex.ModelGemini25Pro` | Complex reasoning and multimodal tasks | | `gemini-2.5-flash` | `googlevertex.ModelGemini25Flash` | Fast multimodal tasks | | `gemini-2.5-flash-lite` | `googlevertex.ModelGemini25FlashLite` | Cost-effective high-volume tasks | Earlier Gemini 2.0, 1.5 and 1.0 models still have constants. Prefer the models above for new code. Grok models are not reachable through `Provider.LanguageModel` — see [xAI (Grok) Sub-Provider](#xai-grok-sub-provider) and [Vertex MaaS and Grok](#vertex-maas-and-grok) below for the MaaS model IDs (e.g. `xai/grok-4.20-reasoning`). ### Embedding models | Model ID | Go constant | Best for | |----------|-------------|----------| | `gemini-embedding-2` | `googlevertex.EmbeddingModelGeminiEmbedding2` | Stable Gemini embeddings | | `gemini-embedding-2-preview` | `googlevertex.EmbeddingModelGeminiEmbedding2Preview` | Preview embeddings | | `text-multilingual-embedding-002` | `googlevertex.EmbeddingModelTextMultilingualEmbedding002` | Multilingual semantic search | | `textembedding-gecko` | `googlevertex.EmbeddingModelTextEmbeddingGecko` | Semantic search | ### Gemini TTS Speech Models Vertex reuses the Gemini speech model implementation and exposes it through `Provider.SpeechModel` and `Provider.Speech`. Vertex Chirp transcription is exposed through `Provider.TranscriptionModel`. Chirp requests use the Vertex Speech-to-Text v2 `recognize` endpoint and support language codes, word time offsets, punctuation, and billed-duration metadata. | Model ID | Go constant | |----------|-------------| | `gemini-2.5-flash-tts` | `googlevertex.SpeechModelGemini25FlashTTS` | | `gemini-2.5-pro-tts` | `googlevertex.SpeechModelGemini25ProTTS` | | `gemini-2.5-flash-lite-preview-tts` | `googlevertex.SpeechModelGemini25FlashLitePreviewTTS` | | `gemini-3.1-flash-tts-preview` | `googlevertex.SpeechModelGemini31FlashTTSPreview` | ```go provider, err := googlevertex.New(googlevertex.Config{ Project: os.Getenv("GOOGLE_VERTEX_PROJECT"), Location: "us-central1", AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"), }) if err != nil { log.Fatal(err) } speechModel, err := provider.SpeechModel(googlevertex.SpeechModelGemini25FlashTTS) if err != nil { log.Fatal(err) } result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Vertex Gemini can synthesize speech.", Voice: "Kore", }) ``` Provider metadata for Vertex speech is keyed as `google`, matching the TypeScript SDK's reuse of `GoogleSpeechModel`. ### Image Generation Models > **Breaking change:** non-Gemini Imagen models (`imagen-3.0-*`, etc.) were > removed from the Vertex provider too, matching the TypeScript SDK. Calling > `provider.ImageModel(id)` with a model ID that doesn't start with > `gemini-` now fails with: "Google image models other than Gemini are no > longer supported. Use a model ID that starts with `gemini-`." | Model ID | Go constant | Best for | |----------|-------------|----------| | `gemini-3.1-flash-image-preview` | `googlevertex.ModelGemini31FlashImagePreview` | Fast Gemini generation | | `gemini-3-pro-image-preview` | `googlevertex.ModelGemini3ProImagePreview` | Advanced generation | | `gemini-2.5-flash-image` | `googlevertex.ModelGemini25FlashImage` | Fast Gemini generation | ### Gemini Transcription `Provider.TranscriptionModel(modelID)` dispatches on the model ID: any ID beginning with `gemini` uses a Gemini transcription model instead of Chirp. - Unary (`gemini-3.5-transcribe`): calls `generateContent` once and returns the full transcript, usable with `ai.Transcribe`. - Live (any `-live` suffix, e.g. `gemini-3.5-transcribe-live`): streams over the Vertex Live WebSocket and only supports `ai.ExperimentalStreamTranscribe` (`DoStream`); it returns an error if you call the unary path. - Any other model ID still routes to Chirp (Speech-to-Text v2 `recognize`). ```go transcriptionModel, err := provider.TranscriptionModel("gemini-3.5-transcribe") if err != nil { log.Fatal(err) } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioBytes, }) ``` ```go liveModel, err := provider.TranscriptionModel("gemini-3.5-transcribe-live") if err != nil { log.Fatal(err) } stream, err := ai.ExperimentalStreamTranscribe(ctx, ai.StreamTranscribeOptions{ Model: liveModel, Audio: audioStream, }) ``` > Vertex transcription (Chirp and Gemini) does not support Express Mode API > keys — use standard Google Cloud credentials. ## Provider-Specific Features ### EU And US Multi-Region Routing When `Location` is `eu` or `us`, standard Gemini, MaaS, and Anthropic-on-Vertex requests use the regional REP hosts: - `aiplatform.eu.rep.googleapis.com` - `aiplatform.us.rep.googleapis.com` Explicit `BaseURL` values still take precedence. `Location` (and the Cloud Speech-to-Text `region` transcription option) is interpolated directly into the request host. A value that isn't a single DNS label (letters, digits, and hyphens) is rejected with an `InvalidArgumentError` before any request is sent; this only applies when no explicit `BaseURL` is configured, since that makes `Location` unused. ### PayGo Request Type Vertex AI uses request headers for PayGo tier selection. Set `sharedRequestType` and optionally `requestType` in Vertex provider options: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "hello", ProviderOptions: map[string]interface{}{ "vertex": map[string]interface{}{ "sharedRequestType": "priority", "requestType": "shared", }, }, }) ``` `serviceTier` is a Gemini API body option and is ignored for Vertex requests with a warning. ### Enterprise Security Configure IAM, VPC Service Controls, CMEK, and audit logging in Google Cloud. The SDK sends requests through the Vertex endpoint selected by `Project`, `Location`, and optional `BaseURL`. ### Model Garden Access hundreds of models from Model Garden: ```go // Deploy from Model Garden model, err := provider.LanguageModel("publishers/google/models/gemini-2.0-flash") ``` ### Batch Prediction And Custom Training Use Google Cloud Vertex AI client libraries for batch prediction and custom training jobs. The Go AI SDK focuses on the AI SDK provider surfaces for generation, embedding, image, speech, MaaS, and Anthropic-on-Vertex requests. ### Image Generation Vertex AI's text-to-image generation is now **Gemini-only** — Imagen model IDs (`imagen-3.0-*`) were removed from this provider, matching the TypeScript SDK. ```go // Create Vertex AI provider provider, err := googlevertex.New(googlevertex.Config{ Project: os.Getenv("GOOGLE_VERTEX_PROJECT"), Location: os.Getenv("GOOGLE_VERTEX_LOCATION"), AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"), }) if err != nil { log.Fatal(err) } ``` **Gemini Image Generation:** ```go // Gemini 2.5 Flash Image - Fast generation geminiModel, err := provider.ImageModel("gemini-2.5-flash-image") if err != nil { log.Fatal(err) } result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: geminiModel, Prompt: "A magical forest with glowing mushrooms and ethereal lighting", Size: "1024x1024", // Square 1:1 aspect ratio }) if err != nil { log.Fatal(err) } os.WriteFile("vertex_gemini_image.png", result.Image.Data, 0644) ``` **Supported Aspect Ratios:** - `1:1` - Square (1024x1024) - `4:3` - Standard (1024x768) - `3:4` - Portrait (768x1024) - `16:9` - Widescreen (1920x1080) - `9:16` - Vertical (1080x1920) **Model Comparison:** | Model | Speed | Quality | Best For | |-------|-------|---------|----------| | gemini-2.5-flash-image | Very Fast | Good | Quick generation | | gemini-3-pro-image-preview | Fast | High | Advanced generation | **Enterprise Features for Image Generation:** - VPC network support for secure generation - Customer-managed encryption keys (CMEK) - Audit logging for compliance - Regional data residency - Batch image generation for large-scale operations ## Video Generation `provider.VideoModel(modelID)` submits to Veo's `:predictLongRunning` endpoint (`MaxVideosPerCall()` returns 4) and works with both the polling `ai.GenerateVideo` and the async `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus` calls: ```go videoModel, err := provider.VideoModel("veo-3.1-generate-preview") if err != nil { log.Fatal(err) } result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{Text: "An aerial shot of a mountain range at dawn"}, }) if err != nil { log.Fatal(err) } for _, video := range result.Videos { fmt.Println(video.URL) } ``` ## xAI (Grok) Sub-Provider `pkg/providers/googlevertex/xai` is a standalone package (mirroring TS's separate `@ai-sdk/google-vertex/xai`) that serves Grok models through Vertex's Model-as-a-Service OpenAI-compatible endpoint. It is not reachable from `googlevertex.Provider` — construct it directly: ```go import vertexxai "github.com/digitallysavvy/go-ai/pkg/providers/googlevertex/xai" xaiProvider, err := vertexxai.New(vertexxai.Config{ Project: os.Getenv("GOOGLE_VERTEX_PROJECT"), Location: "global", // MaaS Grok endpoint default }) if err != nil { log.Fatal(err) } model, err := xaiProvider.LanguageModel(vertexxai.ModelGrok420Reasoning) if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain the Vertex MaaS Grok endpoint.", }) ``` Auth resolves the same way as `googlevertex.Config`: `AccessToken`, `AuthToken`, or `TokenSource`, falling back to Application Default Credentials when none are set. This is a separate code path from the generic `googlevertex.NewMaaS` wrapper below — either works for Grok, but `googlevertex/xai` is the TS-parity implementation for new code. ## Provider-Specific Options Pass Vertex AI-specific options through `ProviderOptions` under the `"googleVertex"` key (the model's `Provider()` method itself still returns `"google-vertex"`, which is only the model identity/error-wrapping name — not the options map key). ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Analyze this document for compliance issues.", ProviderOptions: map[string]interface{}{ "googleVertex": map[string]interface{}{ // Vertex-specific safety settings "safetySettings": []map[string]interface{}{ { "category": "HARM_CATEGORY_DANGEROUS_CONTENT", "threshold": "BLOCK_ONLY_HIGH", }, }, }, }, }) ``` Vertex-specific options are read from the `googleVertex` key (falling back to `vertex`, then `google`), matching the TypeScript SDK's own `providerOptions.googleVertex` convention. The legacy `"vertex"` key is still accepted, and a plain `"google"` key works as the last fallback so options shared between the AI Studio and Vertex providers don't need duplicating. Note that the hyphenated `"google-vertex"` string is only the model's `Provider()` identity/error-wrapping name — it was never a `ProviderOptions` key. | Provider | `ProviderOptions` key(s), in lookup order | |----------|----------------------| | Google AI Studio (`providers/google`) | `"google"` | | Google Vertex AI (`providers/googlevertex`) | `"googleVertex"` → `"vertex"` → `"google"` | Vertex streaming can request function-call argument deltas with `streamFunctionCallArguments`; the SDK only sends this field for streaming Vertex requests. It is omitted for non-streaming `DoGenerate` calls, matching the TypeScript provider. ```go ProviderOptions: map[string]interface{}{ "googleVertex": map[string]interface{}{ "streamFunctionCallArguments": true, }, } ``` ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/googlevertex" ) func main() { ctx := context.Background() provider, err := googlevertex.New(googlevertex.Config{ Project: os.Getenv("GOOGLE_VERTEX_PROJECT"), Location: "us-central1", AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"), }) if err != nil { log.Fatal(err) } model, err := provider.LanguageModel("gemini-2.0-flash") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain Vertex AI"}) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Multi-Region Deployment ```go // Deploy across regions for high availability usProvider, err := googlevertex.New(googlevertex.Config{ Project: os.Getenv("GOOGLE_VERTEX_PROJECT"), Location: "us", AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"), }) if err != nil { log.Fatal(err) } euProvider, err := googlevertex.New(googlevertex.Config{ Project: os.Getenv("GOOGLE_VERTEX_PROJECT"), Location: "eu", AccessToken: os.Getenv("GOOGLE_VERTEX_ACCESS_TOKEN"), }) if err != nil { log.Fatal(err) } // Try primary region first result, err := generateWithFallback(usProvider, euProvider, prompt) ``` ## Best Practices 1. **Enterprise Features** - Use VPC Service Controls - Enable audit logging - Configure IAM policies - Use customer-managed encryption keys 2. **Cost Management** - Monitor usage in Cloud Console - Set up budget alerts - Use committed use discounts - Optimize batch operations 3. **Performance** - Choose regions close to users - Use batch prediction for large datasets - Cache predictions when appropriate 4. **Compliance** - Leverage Google Cloud compliance certifications - Configure data residency requirements - Enable audit logs for governance ## Rate Limits Varies by model and quota: | Model | Default QPM | Max QPM | |-------|------------|---------| | Gemini 2.0 Flash | 60 | 1000+ | | Gemini 2.5 Pro | 60 | 1000+ | Request quota increases through Cloud Console. ## Workflow Serialization Vertex embedding, image, and transcription models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel` (language models could already be serialized). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; speech, video, and the xai sub-provider are not yet serializable. ## See Also - [Google Provider](https://goaisdk.com/docs/providers/google.md) - Google AI Studio - [Vertex AI Documentation](https://cloud.google.com/vertex-ai/docs) - [Model Garden](https://console.cloud.google.com/vertex-ai/model-garden) ## May 2026 parity updates ### Auth reuse Vertex providers reuse Google authentication per provider instance. Configure `googlevertex.New` once and reuse it for language, embedding, and image models instead of creating a new provider for every call. ### Vertex MaaS and Grok Use `googlevertex.NewMaaS` for Vertex Model-as-a-Service models. Grok MaaS constants include `MaaSModelGrok420Reasoning`, `MaaSModelGrok420NonReasoning`, `MaaSModelGrok41FastReasoning`, and `MaaSModelGrok41FastNonReasoning`. ```go maas := googlevertex.NewMaaS(googlevertex.MaaSConfig{ Project: os.Getenv("GOOGLE_VERTEX_PROJECT"), Location: "us-central1", }) model, err := maas.LanguageModel(string(googlevertex.MaaSModelGrok420Reasoning)) ``` ### Auth token overrides Vertex Anthropic auth token generation can be overridden through MaaS configuration when a custom token source is required. --- # Stability AI Provider > Setup and usage guide for Stability AI image generation with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/stability Documentation index: https://goaisdk.com/llms.txt The current Go provider exposes Stability text-to-image generation through the shared `ai.GenerateImage` API. ## Setup ### Installation ```go import ( "context" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/stability" ) ``` ### Configuration ```go provider := stability.New(stability.Config{ APIKey: os.Getenv("STABILITY_API_KEY"), }) model, err := provider.ImageModel("stable-diffusion-xl-1024-v1-0") ``` ### Get API Key ```bash export STABILITY_API_KEY=sk-... ``` ## Available Models `provider.ImageModel(modelID)` calls the legacy Stability v1 REST API — `POST /v1/generation/{modelID}/text-to-image` — so `modelID` must be a v1 "engine" ID. Stability's newer SD3/Stable Image v2beta models (`sd3-large`, `sd3-medium`, `core`, `ultra`, etc.) are **not** reachable through this provider, since they live under a different endpoint (`/v2beta/stable-image/...`) that this package does not implement. ### Image Generation | Model ID | Output | Best For | |----------|--------|----------| | stable-diffusion-xl-1024-v1-0 | PNG bytes | High quality text-to-image generation | | stable-diffusion-v1-6 | PNG bytes | Fast text-to-image generation | ## Provider-Specific Features ### Text-to-Image ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: "A serene landscape", Size: "1024x1024", }) ``` ## Examples ### Basic Image Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/stability" ) func main() { provider := stability.New(stability.Config{ APIKey: os.Getenv("STABILITY_API_KEY"), }) model, err := provider.ImageModel("stable-diffusion-xl-1024-v1-0") if err != nil { log.Fatal(err) } result, err := ai.GenerateImage(context.Background(), ai.GenerateImageOptions{ Model: model, Prompt: "A futuristic cityscape at sunset, photorealistic", Size: "1024x1024", }) if err != nil { log.Fatal(err) } // Save image imageData := result.Images[0].Data err = os.WriteFile("output.png", imageData, 0644) if err != nil { log.Fatal(err) } fmt.Println("Image saved to output.png") } ``` ### Batch Generation ```go prompts := []string{ "A red sports car", "A blue sports car", "A green sports car", } for i, prompt := range prompts { result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: prompt, }) if err != nil { log.Printf("Failed to generate %d: %v", i, err) continue } filename := fmt.Sprintf("car_%d.png", i) os.WriteFile(filename, result.Images[0].Data, 0644) } ``` ### Multiple Images ```go n := 4 result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: "A set of colorful product concept images", N: &n, Size: "1024x1024", }) if err != nil { log.Fatal(err) } for i, image := range result.Images { filename := fmt.Sprintf("concept_%d.png", i) os.WriteFile(filename, image.Data, 0644) } ``` ## Best Practices 1. **Prompt Engineering** - Be specific and detailed - Include style keywords - Mention quality (4k, detailed, etc) 2. **Cost Optimization** - Start with one image before increasing `N` - Batch similar requests - Cache generated images 3. **Performance** - Use appropriate model for task - Parallelize generations ## Rate Limits & Pricing ### Rate Limits Varies by plan - typical limits: - 150 images/min (standard) - 500 images/min (premium) ### Cost Management ```go func estimateImageCost(model string, count int) float64 { prices := map[string]float64{ "stable-diffusion-xl-1024-v1-0": 0.04, "stable-diffusion-v1-6": 0.02, } return prices[model] * float64(count) } ``` ## See Also - [API Reference: GenerateImage](https://goaisdk.com/docs/reference/ai/generate-image.md) - [Stability AI Documentation](https://platform.stability.ai/docs) - [Black Forest Labs Provider](https://goaisdk.com/docs/providers/bfl.md) - Alternative image generation --- # Black Forest Labs Provider > Setup and usage guide for Black Forest Labs FLUX image models with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/bfl Documentation index: https://goaisdk.com/llms.txt Black Forest Labs (founded by Stable Diffusion creators) provides FLUX models - the next generation of image synthesis with exceptional quality and prompt adherence. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/bfl" ) ``` ### Configuration ```go provider := bfl.New(bfl.Config{ APIKey: os.Getenv("BFL_API_KEY"), }) model, err := provider.ImageModel("flux-pro-1.1") ``` ### Get API Key ```bash export BFL_API_KEY=... ``` ## Available Models ### FLUX Models | Model ID | Quality | Speed | Best For | | ---------- | --------- | ------- | ---------- | | flux-pro-1.1 | Excellent | Medium | Professional work | | flux-pro | Excellent | Medium | High quality | | flux-dev | Very Good | Fast | Development | | flux-schnell | Good | Very Fast | Rapid iteration | ## Provider-Specific Features ### Exceptional Quality FLUX models produce photorealistic and artistically sophisticated images: ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: "A photorealistic portrait of an elderly person with expressive eyes", Size: "1024x1024", }) ``` ### Perfect Text Rendering FLUX excels at text in images: ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: "A storefront sign that says 'Coffee & Code' in elegant typography", }) ``` ### Advanced Prompt Understanding Complex, nuanced prompts work well: ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: "A cyberpunk street scene at night, neon reflections on wet pavement, " + "shallow depth of field, cinematic composition, blade runner aesthetic", }) ``` ## Examples ### Basic Image Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/bfl" ) func main() { provider := bfl.New(bfl.Config{ APIKey: os.Getenv("BFL_API_KEY"), }) model, err := provider.ImageModel("flux-pro-1.1") if err != nil { log.Fatal(err) } result, err := ai.GenerateImage(context.Background(), ai.GenerateImageOptions{ Model: model, Prompt: "A majestic dragon perched on a mountain peak at sunrise", Size: "1024x1024", }) if err != nil { log.Fatal(err) } os.WriteFile("dragon.png", result.Images[0].Data, 0644) fmt.Println("Image saved to dragon.png") } ``` ### Professional Photography ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: "Professional product photography: luxury watch on marble surface, " + "soft studio lighting, shallow depth of field, 85mm lens, f/1.4", Size: "1024x1024", }) ``` ### Fill Models `flux-pro-1.0-fill` uses Black Forest Labs' fill request format. Pass the source image through `Files`; the provider maps it to the `image` request field required by the upstream API. ```go model, _ := bflProvider.ImageModel("flux-pro-1.0-fill") result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{ Prompt: "replace the background with a studio wall", Files: []provider.ImageFile{{ Data: imageBytes, MediaType: "image/png", }}, }) ``` ### Artistic Styles ```go styles := []string{ "oil painting, impressionist style", "watercolor, loose brushstrokes", "digital art, concept art style", "pencil sketch, detailed cross-hatching", } basePrompt := "A serene forest path" for i, style := range styles { prompt := fmt.Sprintf("%s, %s", basePrompt, style) result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: prompt, }) if err != nil { continue } filename := fmt.Sprintf("forest_%d.png", i) os.WriteFile(filename, result.Images[0].Data, 0644) } ``` ## Video Generation (FLUX 3) `provider.VideoModel(modelID)` (e.g. `"flux-video"`) generates a single video per call (`MaxVideosPerCall() == 1`) — FLUX 3 video does not accept a custom frame rate or seed: ```go videoModel, err := provider.VideoModel("flux-video") if err != nil { log.Fatal(err) } result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{Text: "A time-lapse of clouds over a mountain range"}, }) if err != nil { log.Fatal(err) } for _, video := range result.Videos { fmt.Println(video.URL) } ``` `VideoModel` also implements `provider.VideoModelStarter` / `VideoModelStatusChecker`, so it works with `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus`. ## Workflow Serialization BFL image and video models can cross a workflow boundary with `provider.SerializeImageModel` / `DeserializeImageModel` and `provider.SerializeVideoModel` / `DeserializeVideoModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## Best Practices 1. **Prompt Engineering** - Be detailed and specific - Include style, mood, lighting - Mention technical details (lens, lighting, etc) - FLUX understands complex descriptions 2. **Model Selection** - Use flux-pro-1.1 for final production - Use flux-dev for development/testing - Use flux-schnell for rapid prototyping 3. **Quality Optimization** - FLUX needs fewer iterations than SD - Prompts can be more natural/conversational - Text in images works reliably 4. **Cost Management** - Use schnell for experimentation - Switch to pro for finals - Batch similar requests ## Rate Limits & Pricing ### Rate Limits Varies by plan - check dashboard. ### Pricing FLUX pricing varies by model. See the [BFL pricing page](https://bfl.ai/pricing) for current values. ## See Also Provider options are parsed once when the request body is built. Polling and result retrieval reuse the generated request state, matching the TypeScript provider behavior for single-parse option semantics. - [API Reference: GenerateImage](https://goaisdk.com/docs/reference/ai/generate-image.md) - [BFL Documentation](https://docs.bfl.ai) - [Stability AI Provider](https://goaisdk.com/docs/providers/stability.md) - Alternative --- # FAL Provider > Setup and usage guide for FAL fast image and video generation with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/fal Documentation index: https://goaisdk.com/llms.txt FAL provides ultra-fast inference for image and video generation models. Known for speed optimization and support for latest models including FLUX, Stable Diffusion, and video generation. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/fal" ) ``` ### Configuration ```go provider := fal.New(fal.Config{ APIKey: os.Getenv("FAL_API_KEY"), Headers: map[string]string{"X-Custom": "value"}, }) model, err := provider.ImageModel("fal-ai/flux-pro") ``` `fal.Config.APIKey` falls back to the `FAL_API_KEY` environment variable, then to `FAL_KEY`, if left empty. `fal.Config.Headers` are sent on every request across image, video, speech, and transcription models. > **Breaking change:** the default base URL is now `https://fal.run` (was > `https://fal.run/fal-ai`, which doubled the path for fully-qualified model > IDs and broke default-config requests). Always pass fully-qualified model > IDs like `fal-ai/fast-sdxl`, not `fast-sdxl`. The image and video default > models are `fal-ai/fast-sdxl` and `fal-ai/luma-ray`. ### Get API Key ```bash export FAL_API_KEY=... ``` ## Available Models ### Image Generation | Model ID | Quality | Speed | Best For | | ---------- | --------- | ------- | ---------- | | fal-ai/flux-pro | Excellent | Fast | High quality | | fal-ai/flux-schnell | Good | Very Fast | Speed | | fal-ai/stable-diffusion-xl | High | Fast | General | ### Video Generation | Model ID | Quality | Speed | Best For | | ---------- | --------- | ------- | ---------- | | fal-ai/stable-video-diffusion | Medium | Medium | Video from image | | fal-ai/runway-gen3 | High | Slow | High quality video | ## Provider-Specific Features ### Ultra-Fast Inference FAL optimizes for speed: ```go start := time.Now() result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: "A sunset over mountains", }) elapsed := time.Since(start) fmt.Printf("Generated in %v\n", elapsed) // ~1-2 seconds ``` ### Video Generation Create videos from text or images: ```go videoModel, err := provider.VideoModel("fal-ai/stable-video-diffusion") // Generate from image initImage, _ := os.ReadFile("frame.png") fps := 10 result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{ Image: &ai.VideoPromptImage{Data: initImage}, }, FPS: &fps, ProviderOptions: map[string]interface{}{ "fal": map[string]interface{}{ "frames": 25, }, }, }) os.WriteFile("output.mp4", result.Video.Data, 0644) ``` ### Async Video And Webhooks `fal.VideoModel` implements `provider.VideoModelStarter` / `VideoModelStatusChecker`, so it also works with `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus`, and can complete via a webhook instead of polling: ```go started, err := ai.ExperimentalStartVideo(ctx, ai.StartVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{Text: "A cat playing piano"}, WebhookURL: "https://example.com/fal-webhook", }) ``` ### Speech and Transcription ```go speechModel, err := provider.SpeechModel("fal-ai/minimax/speech-02-hd") if err != nil { log.Fatal(err) } result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello from Fal.", }) ``` ```go transcriptionModel, err := provider.TranscriptionModel("fal-ai/wizper") if err != nil { log.Fatal(err) } transcript, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioBytes, }) ``` Fal does not support reranking — `provider.RerankingModel` returns an error. ## Examples ### Basic Image Generation ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/fal" ) func main() { provider := fal.New(fal.Config{ APIKey: os.Getenv("FAL_KEY"), }) model, err := provider.ImageModel("fal-ai/flux-schnell") if err != nil { log.Fatal(err) } start := time.Now() result, err := ai.GenerateImage(context.Background(), ai.GenerateImageOptions{ Model: model, Prompt: "A peaceful zen garden", }) if err != nil { log.Fatal(err) } fmt.Printf("Generated in %v\n", time.Since(start)) os.WriteFile("garden.png", result.Images[0].Data, 0644) } ``` ### Fast Batch Generation ```go prompts := []string{ "A red apple", "A green apple", "A yellow apple", } // Parallel generation var wg sync.WaitGroup for i, prompt := range prompts { wg.Add(1) go func(idx int, p string) { defer wg.Done() result, err := ai.GenerateImage(context.Background(), ai.GenerateImageOptions{ Model: model, Prompt: p, }) if err != nil { return } filename := fmt.Sprintf("apple_%d.png", idx) os.WriteFile(filename, result.Images[0].Data, 0644) }(i, prompt) } wg.Wait() ``` ## Best Practices 1. **Speed Optimization** - Use flux-schnell for rapid iteration - Parallel requests for batches - Leverage fast inference for real-time apps 2. **Quality vs Speed** - Use schnell for previews - Use pro for finals - Balance based on use case 3. **Cost Management** - Schnell is very cost-effective - Cache results - Use for high-volume applications ## Rate Limits & Pricing Check dashboard for current limits and pricing. ## Workflow Serialization Fal image, speech, and transcription models can cross a workflow boundary with `provider.SerializeImageModel` / `DeserializeImageModel`, `provider.SerializeSpeechModel` / `DeserializeSpeechModel`, and `provider.SerializeTranscriptionModel` / `DeserializeTranscriptionModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; video models are not yet serializable. ## See Also - [API Reference: GenerateImage](https://goaisdk.com/docs/reference/ai/generate-image.md) - [FAL Documentation](https://fal.ai/docs) --- # Luma Provider > Setup guide for Luma AI's asynchronous Photon image generation in the Go AI SDK, covering reference-image-guided generation and workflow serialization. Canonical URL: https://goaisdk.com/docs/providers/luma Documentation index: https://goaisdk.com/llms.txt Luma AI provides asynchronous image generation (Photon models) with reference-image-guided generation — image, style, character, and modify-image reference types. Luma has no language, embedding, speech, transcription, reranking, or video model in this SDK; only `ImageModel` is supported. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/luma" ) ``` ### Configuration ```go provider := luma.New(luma.Config{ APIKey: os.Getenv("LUMA_API_KEY"), }) model, err := provider.ImageModel(luma.ModelPhoton1) ``` `luma.Config` also accepts `BaseURL` (default `https://api.lumalabs.ai`) and `Headers`. `APIKey` falls back to `LUMA_API_KEY` when left empty. An empty model ID defaults to `luma.ModelPhoton1`. ### Get API Key ```bash export LUMA_API_KEY=... ``` ## Basic Usage `DoGenerate` submits an async generation, polls its status, and downloads the resulting image — the whole lifecycle happens inside a single `ai.GenerateImage` call: ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: "A neon-lit cyberpunk alley in the rain", }) if err != nil { log.Fatal(err) } os.WriteFile("output.png", result.Images[0].Data, 0644) ``` Model IDs: `luma.ModelPhoton1` (`photon-1`, default), `luma.ModelPhotonFlash1` (`photon-flash-1`). ## Reference Images Pass up to 4 reference images (as `Files` on the lower-level `provider.ImageGenerateOptions`) and select a reference type through `providerOptions.luma`: ```go referenceType := luma.ReferenceTypeStyle weight := 0.8 result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{ Prompt: "A portrait in the same style as the reference", Files: []provider.ImageFile{{ Type: "url", URL: "https://example.com/style-reference.jpg", }}, ProviderOptions: map[string]interface{}{ "luma": luma.ImageModelOptions{ ReferenceType: &referenceType, Images: []luma.ImageConfig{ {Weight: &weight}, }, }, }, }) ``` Reference types: `luma.ReferenceTypeImage` (default, up to 4 images), `ReferenceTypeStyle`, `ReferenceTypeCharacter` (up to 4 images, grouped by `ImageConfig.ID`), `ReferenceTypeModifyImage` (single input image). `ImageModelOptions` also accepts `PollIntervalMillis` (default 500) and `MaxPollAttempts` (default 120) to override the async polling behavior, and an `Additional map[string]interface{}` for any passthrough field not explicitly modeled. ## Workflow Serialization Luma image models can cross a workflow boundary with `provider.SerializeImageModel` / `DeserializeImageModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateImage](https://goaisdk.com/docs/reference/ai/generate-image.md) - [Luma AI Documentation](https://docs.lumalabs.ai) --- # ElevenLabs Provider > Setup guide for ElevenLabs text-to-speech in the Go AI SDK, covering the provider.SpeechModel interface, generating speech, provider options, and models. Canonical URL: https://goaisdk.com/docs/providers/elevenlabs Documentation index: https://goaisdk.com/llms.txt ElevenLabs provides text-to-speech models. The current Go provider implements the shared `provider.SpeechModel` interface and is used through `ai.GenerateSpeech`. ## Setup ```go import ( "context" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/elevenlabs" ) ``` ```go provider := elevenlabs.New(elevenlabs.Config{ APIKey: os.Getenv("ELEVENLABS_API_KEY"), }) model, err := provider.SpeechModel("eleven_multilingual_v2") if err != nil { log.Fatal(err) } ``` ## Generate Speech ```go result, err := ai.GenerateSpeech(context.Background(), ai.GenerateSpeechOptions{ Model: model, Text: "Welcome to the Go AI SDK.", Voice: "Rachel", OutputFormat: "mp3", }) if err != nil { log.Fatal(err) } if err := os.WriteFile("welcome.mp3", result.Audio.Data, 0644); err != nil { log.Fatal(err) } ``` ## Provider Options Pass ElevenLabs-specific request options under the `"elevenlabs"` key when the provider supports them. ```go result, err := ai.GenerateSpeech(context.Background(), ai.GenerateSpeechOptions{ Model: model, Text: "A more expressive delivery.", Voice: "Rachel", ProviderOptions: map[string]interface{}{ "elevenlabs": map[string]interface{}{ "voiceSettings": map[string]interface{}{ "stability": 0.3, "similarityBoost": 0.8, "style": 0.5, "useSpeakerBoost": true, }, }, }, }) ``` ## Models Common model IDs include: | Model ID | Best For | |---|---| | `eleven_multilingual_v2` | Multilingual speech | | `eleven_flash_v2_5` | Low-latency speech | | `eleven_turbo_v2_5` | Fast speech synthesis | ## Transcription `provider.TranscriptionModel(modelID)` supports both batch (`scribe_v1`, `scribe_v2`) and realtime (`scribe_v2_realtime`) models: ```go transcriptionModel, err := provider.TranscriptionModel("scribe_v2") if err != nil { log.Fatal(err) } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioBytes, }) ``` ### Streaming Transcription (Scribe v2 Realtime) `scribe_v2_realtime` is a streaming-only model — `DoTranscribe` (the batch path) rejects it. Use `ai.ExperimentalStreamTranscribe` instead: ```go liveModel, err := provider.TranscriptionModel("scribe_v2_realtime") if err != nil { log.Fatal(err) } stream, err := ai.ExperimentalStreamTranscribe(ctx, ai.StreamTranscribeOptions{ Model: liveModel, Audio: audioStream, }) ``` `ELEVENLABS_API_KEY` is read automatically when no API key is passed to `Config`. ## Workflow Serialization ElevenLabs speech and transcription models can cross a workflow boundary with `provider.SerializeSpeechModel` / `DeserializeSpeechModel` and `provider.SerializeTranscriptionModel` / `DeserializeTranscriptionModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateSpeech](https://goaisdk.com/docs/reference/ai/generate-speech.md) - [Speech Generation Guide](https://goaisdk.com/docs/ai-sdk-core/speech.md) - [ElevenLabs Documentation](https://elevenlabs.io/docs) --- # Deepgram Provider > Setup guide for Deepgram speech-to-text in the Go AI SDK, covering the provider.TranscriptionModel interface, timestamps, models, and speech synthesis. Canonical URL: https://goaisdk.com/docs/providers/deepgram Documentation index: https://goaisdk.com/llms.txt Deepgram provides speech-to-text transcription models. The current Go provider implements the shared `provider.TranscriptionModel` interface and is used through `ai.Transcribe`. ## Setup ```go import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/deepgram" ) ``` ```go provider := deepgram.New(deepgram.Config{ APIKey: os.Getenv("DEEPGRAM_API_KEY"), }) model, err := provider.TranscriptionModel("nova-2") if err != nil { log.Fatal(err) } ``` ## Transcribe Audio ```go audioData, err := os.ReadFile("recording.mp3") if err != nil { log.Fatal(err) } result, err := ai.Transcribe(context.Background(), ai.TranscribeOptions{ Model: model, Audio: audioData, MimeType: "audio/mpeg", ProviderOptions: map[string]interface{}{ "deepgram": map[string]interface{}{ "language": "en", }, }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` ## Timestamps Use Deepgram provider options to request word timing data. The high-level `ai.Transcribe` result exposes normalized segment fields. ```go result, err := ai.Transcribe(context.Background(), ai.TranscribeOptions{ Model: model, Audio: audioData, MimeType: "audio/mpeg", ProviderOptions: map[string]interface{}{ "deepgram": map[string]interface{}{ "utterances": true, }, }, }) if err != nil { log.Fatal(err) } for _, segment := range result.Segments { fmt.Printf("[%.2f-%.2f] %s\n", segment.StartSecond, segment.EndSecond, segment.Text) } ``` ## Models Common Deepgram model IDs include: | Model ID | Best For | |---|---| | `nova-2` | General transcription (default when the model ID is empty) | | `nova-2-phonecall` | Phone audio | | `nova-2-meeting` | Meetings | | `enhanced` | General-purpose transcription | | `base` | Cost-sensitive use cases | ## Speech Synthesis Deepgram also exposes Aura text-to-speech through `provider.SpeechModel(modelID)`: ```go speechModel, err := provider.SpeechModel("aura-2-thalia-en") if err != nil { log.Fatal(err) } result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello from Deepgram.", }) ``` Deepgram does not support reranking — `provider.RerankingModel` returns an error. ## Array Query Options Deepgram provider options that take a list (e.g. `keywords`) are sent as comma-joined query values, matching the Deepgram REST API's expected query string format. `DEEPGRAM_API_KEY` is read automatically when no API key is passed to `Config`. ## Workflow Serialization Deepgram speech and transcription models can cross a workflow boundary with `provider.SerializeSpeechModel` / `DeserializeSpeechModel` and `provider.SerializeTranscriptionModel` / `DeserializeTranscriptionModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: Transcribe](https://goaisdk.com/docs/reference/ai/transcribe.md) - [Transcription Guide](https://goaisdk.com/docs/ai-sdk-core/transcription.md) - [Deepgram Documentation](https://developers.deepgram.com) --- # AssemblyAI Provider > Setup guide for AssemblyAI transcription in the Go AI SDK, covering the provider.TranscriptionModel interface, timestamps, supported models, and usage. Canonical URL: https://goaisdk.com/docs/providers/assemblyai Documentation index: https://goaisdk.com/llms.txt AssemblyAI provides speech-to-text transcription models. The current Go provider implements the shared `provider.TranscriptionModel` interface and is used through `ai.Transcribe`. ## Setup ```go import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/assemblyai" ) ``` ```go provider := assemblyai.New(assemblyai.Config{ APIKey: os.Getenv("ASSEMBLYAI_API_KEY"), }) model, err := provider.TranscriptionModel("best") if err != nil { log.Fatal(err) } ``` ## Transcribe Audio Pass local audio bytes to `ai.Transcribe`; the AssemblyAI provider uploads the audio and polls until the transcript is complete. ```go audioData, err := os.ReadFile("recording.mp3") if err != nil { log.Fatal(err) } result, err := ai.Transcribe(context.Background(), ai.TranscribeOptions{ Model: model, Audio: audioData, MimeType: "audio/mpeg", ProviderOptions: map[string]interface{}{ "assemblyai": map[string]interface{}{ "languageCode": "en", }, }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` You can also provide a URL and let the SDK download it before invoking the provider: ```go result, err := ai.Transcribe(context.Background(), ai.TranscribeOptions{ Model: model, AudioURL: "https://example.com/recording.mp3", ProviderOptions: map[string]interface{}{ "assemblyai": map[string]interface{}{ "languageCode": "en", }, }, }) ``` ## Timestamps AssemblyAI word timings are returned as normalized segments when available. ```go result, err := ai.Transcribe(context.Background(), ai.TranscribeOptions{ Model: model, Audio: audioData, MimeType: "audio/mpeg", ProviderOptions: map[string]interface{}{ "assemblyai": map[string]interface{}{ "speakerLabels": true, }, }, }) if err != nil { log.Fatal(err) } for _, segment := range result.Segments { fmt.Printf("[%.2f-%.2f] %s\n", segment.StartSecond, segment.EndSecond, segment.Text) } ``` ## Models | Model ID | Best For | |---|---| | `best` | Highest-accuracy transcription | | `nano` | Lower-latency transcription | ## Workflow Serialization AssemblyAI transcription models can cross a workflow boundary with `provider.SerializeTranscriptionModel` / `DeserializeTranscriptionModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: Transcribe](https://goaisdk.com/docs/reference/ai/transcribe.md) - [Transcription Guide](https://goaisdk.com/docs/ai-sdk-core/transcription.md) - [AssemblyAI Documentation](https://www.assemblyai.com/docs) --- # Cartesia Provider > Setup and usage guide for Cartesia speech synthesis and transcription with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/cartesia Documentation index: https://goaisdk.com/llms.txt Cartesia provides low-latency text-to-speech (Sonic) and both batch and realtime speech-to-text (Ink) models. It does not offer language, embedding, or image models — `LanguageModel`, `EmbeddingModel`, and `ImageModel` all return errors. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/cartesia" ) ``` ### Configuration ```go provider := cartesia.New(cartesia.Config{ APIKey: os.Getenv("CARTESIA_API_KEY"), }) ``` `cartesia.Config` also accepts `BaseURL` (default `https://api.cartesia.ai`) and `APIVersion` (sent as the `Cartesia-Version` header, default `"2026-03-01"`). ### Get API Key ```bash export CARTESIA_API_KEY=... ``` ## Speech Synthesis ```go speechModel, err := provider.SpeechModel(cartesia.ModelSonic2) if err != nil { log.Fatal(err) } result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello from Cartesia.", ProviderOptions: map[string]interface{}{ "cartesia": map[string]interface{}{ "container": "mp3", "sampleRate": 44100, "speed": 1.0, }, }, }) ``` Model IDs: `cartesia.ModelSonic35` (`sonic-3.5`), `ModelSonic3` (`sonic-3`), `ModelSonic2` (`sonic-2`), `ModelSonicTurbo` (`sonic-turbo`), `ModelSonicLatest` (`sonic-latest`). ## Transcription ```go transcriptionModel, err := provider.TranscriptionModel(cartesia.ModelInkWhisper) if err != nil { log.Fatal(err) } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioBytes, ProviderOptions: map[string]interface{}{ "cartesia": map[string]interface{}{ "language": "en", }, }, }) ``` ### Streaming Transcription (Ink 2) `cartesia.ModelInk2` (and any `ink-2-*` variant) is streaming-only — the batch `DoTranscribe` path rejects it. Use `ai.ExperimentalStreamTranscribe` instead, which streams over Cartesia's Ink 2 WebSocket: ```go liveModel, err := provider.TranscriptionModel(cartesia.ModelInk2) if err != nil { log.Fatal(err) } stream, err := ai.ExperimentalStreamTranscribe(ctx, ai.StreamTranscribeOptions{ Model: liveModel, Audio: audioStream, ProviderOptions: map[string]interface{}{ "cartesia": map[string]interface{}{ "streaming": map[string]interface{}{ "turnDetection": true, }, }, }, }) ``` ## Provider Options | Option | Applies to | Type | Description | |--------|------------|------|-------------| | `container` | Speech | string | Output audio container: `raw`, `wav`, or `mp3` | | `encoding` | Speech | string | `pcm_f32le`, `pcm_s16le`, `pcm_mulaw`, or `pcm_alaw` | | `sampleRate` | Speech | int | Output sample rate in Hz | | `bitRate` | Speech | int | Bitrate for mp3 output | | `speed` | Speech | float64 | Speech speed, 0.6 to 1.5 | | `language` | Speech, Transcription | string | ISO 639-1 language code | | `timestampGranularities` | Transcription | []string | Currently only `["word"]` is supported | | `streaming.encoding` | Transcription (streaming) | string | Raw audio encoding sent to Ink 2 | | `streaming.turnDetection` | Transcription (streaming) | *bool | Use Cartesia's native turn detection (default true) | ## Workflow Serialization Cartesia speech and transcription models can cross a workflow boundary with `provider.SerializeSpeechModel` / `DeserializeSpeechModel` and `provider.SerializeTranscriptionModel` / `DeserializeTranscriptionModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateSpeech](https://goaisdk.com/docs/reference/ai/generate-speech.md) - [API Reference: Transcribe](https://goaisdk.com/docs/reference/ai/transcribe.md) - [Cartesia Documentation](https://docs.cartesia.ai) --- # Fish Audio Provider > Setup and usage guide for Fish Audio speech synthesis and transcription with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/fishaudio Documentation index: https://goaisdk.com/llms.txt Fish Audio provides text-to-speech and speech-to-text models. It does not offer language, embedding, or image models — `LanguageModel`, `EmbeddingModel`, and `ImageModel` all return errors. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/fishaudio" ) ``` ### Configuration ```go provider := fishaudio.New(fishaudio.Config{ APIKey: os.Getenv("FISH_AUDIO_API_KEY"), }) ``` `fishaudio.Config` also accepts `BaseURL` (default `https://api.fish.audio`). ### Get API Key ```bash export FISH_AUDIO_API_KEY=... ``` ## Speech Synthesis ```go speechModel, err := provider.SpeechModel(fishaudio.ModelS21Pro) if err != nil { log.Fatal(err) } result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello from Fish Audio.", ProviderOptions: map[string]interface{}{ "fishaudio": map[string]interface{}{ "referenceId": "some-voice-id", "latency": "normal", }, }, }) ``` Model IDs: `fishaudio.ModelS1` (`s1`), `ModelS2Pro` (`s2-pro`), `ModelS21Pro` (`s2.1-pro`, Fish Audio's recommended default), `ModelS21ProFree` (`s2.1-pro-free`, a free developer tier with no time-to-first-audio or data-processing guarantees — prefer `ModelS21Pro` for production). ## Transcription `POST /v1/asr` has no model selector, so the model ID is only a routing label; it defaults to `"transcribe-1"` when left empty: ```go transcriptionModel, err := provider.TranscriptionModel(fishaudio.ModelTranscribe1) if err != nil { log.Fatal(err) } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioBytes, }) ``` ## Provider Options Speech options are passed under the `"fishaudio"` `ProviderOptions` key: | Option | Type | Description | |--------|------|--------------| | `referenceId` | string or []string | Voice model; a slice enables S2-Pro multi-speaker dialogue. Takes precedence over the top-level `Voice` option. | | `sampleRate` | int | Output sample rate in Hz | | `mp3Bitrate` | int | mp3 bitrate in kbps: 64, 128, or 192 | | `opusBitrate` | int | opus bitrate in bps, or -1000 for automatic | | `latency` | string | `low`, `normal`, or `balanced` | | `volume` | float64 | Volume offset in dB | | `normalizeLoudness` | *bool | Loudness normalization (S2 family only) | | `temperature` | float64 | Expressiveness, 0 to 1 | | `topP` | float64 | Nucleus sampling diversity, 0 to 1 | | `chunkLength` | int | Text segment size, 100 to 300 | | `minChunkLength` | int | Minimum characters before a new chunk, 0 to 100 | | `normalize` | *bool | Text normalization for English and Chinese | | `maxNewTokens` | int | Maximum audio tokens per text chunk | ## Workflow Serialization Fish Audio speech and transcription models can cross a workflow boundary with `provider.SerializeSpeechModel` / `DeserializeSpeechModel` and `provider.SerializeTranscriptionModel` / `DeserializeTranscriptionModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateSpeech](https://goaisdk.com/docs/reference/ai/generate-speech.md) - [API Reference: Transcribe](https://goaisdk.com/docs/reference/ai/transcribe.md) - [Fish Audio Documentation](https://docs.fish.audio) --- # Hume Provider > Setup guide for Hume's expressive text-to-speech in the Go AI SDK, noting it only supports speech synthesis, with workflow serialization details. Canonical URL: https://goaisdk.com/docs/providers/hume Documentation index: https://goaisdk.com/llms.txt Hume provides expressive text-to-speech synthesis. It does not offer language, embedding, image, or transcription models — `LanguageModel`, `EmbeddingModel`, `ImageModel`, and `TranscriptionModel` all return errors. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/hume" ) ``` ### Configuration ```go provider := hume.New(hume.Config{ APIKey: os.Getenv("HUME_API_KEY"), }) ``` `hume.Config` also accepts `BaseURL` (default `https://api.hume.ai`). ### Get API Key ```bash export HUME_API_KEY=... ``` ## Speech Synthesis Hume has no selectable model ID — `SpeechModel`'s `modelID` argument is accepted for interface parity but ignored, matching the TypeScript SDK's `speech()` factory (which takes no model ID parameter): ```go speechModel, err := provider.SpeechModel("") if err != nil { log.Fatal(err) } result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello from Hume.", }) ``` ### Multi-Utterance Context Pass a list of utterances (each with its own voice, speed, description, and trailing silence), or continue a previous generation by ID, through `providerOptions.hume.context`: ```go result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "", // ignored when context.utterances is set ProviderOptions: map[string]interface{}{ "hume": hume.SpeechModelOptions{ Context: &hume.Context{ Utterances: []hume.Utterance{ { Text: "Welcome to the show.", Description: "warm, welcoming tone", Voice: &hume.UtteranceVoice{Name: "Ava Song", Provider: "HUME_AI"}, }, }, }, }, }, }) ``` To continue a previous generation instead of specifying utterances, set `Context.GenerationID` (leave `Utterances` empty). ## Workflow Serialization Hume speech models can cross a workflow boundary with `provider.SerializeSpeechModel` / `DeserializeSpeechModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateSpeech](https://goaisdk.com/docs/reference/ai/generate-speech.md) - [Hume Documentation](https://dev.hume.ai) --- # Rev.ai Provider > Setup guide for Rev.ai's asynchronous, job-based transcription in the Go AI SDK, noting only TranscriptionModel is supported, plus provider options. Canonical URL: https://goaisdk.com/docs/providers/revai Documentation index: https://goaisdk.com/llms.txt Rev.ai provides asynchronous, job-based speech-to-text transcription. It does not offer language, embedding, image, speech, or reranking models — only `TranscriptionModel` is supported. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/revai" ) ``` ### Configuration ```go provider := revai.New(revai.Config{ APIKey: os.Getenv("REVAI_API_KEY"), }) ``` `revai.Config` also accepts `BaseURL` (default `https://api.rev.ai`). ### Get API Key ```bash export REVAI_API_KEY=... ``` ## Transcription `ai.Transcribe` submits a job (`POST /speechtotext/v1/jobs`), polls its status, then fetches the transcript once complete — the whole async job lifecycle is handled for you: ```go transcriptionModel, err := provider.TranscriptionModel(revai.ModelMachine) if err != nil { log.Fatal(err) } result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioBytes, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` Model IDs (the "transcriber" job field): `revai.ModelMachine` (`"machine"`), `ModelLowCost` (`"low_cost"`), `ModelFusion` (`"fusion"`). ## Provider Options Rev.ai-specific job options are passed under the `"revai"` `ProviderOptions` key. Unlike most providers, these field names intentionally match Rev.ai's own snake_case job submission API rather than the SDK's usual camelCase convention: ```go verbatim := true result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioBytes, ProviderOptions: map[string]interface{}{ "revai": revai.TranscriptionModelOptions{ Verbatim: &verbatim, SkipDiarization: types.BoolPtr(false), Language: "en", }, }, }) ``` Notable options: `verbatim`, `rush`, `skip_diarization`, `skip_punctuation`, `remove_disfluencies`, `speakers_count`, `custom_vocabulary_id`, `summarization_config`, `translation_config`, `notification_config` (webhook on completion), and `delete_after_seconds`. See `pkg/providers/revai/transcription_model_options.go` for the complete field list. ## Workflow Serialization Rev.ai transcription models can cross a workflow boundary with `provider.SerializeTranscriptionModel` / `DeserializeTranscriptionModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: Transcribe](https://goaisdk.com/docs/reference/ai/transcribe.md) - [Rev.ai Documentation](https://docs.rev.ai) --- # Baseten Provider > Setup and usage guide for Baseten Model APIs and custom deployments with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/baseten Documentation index: https://goaisdk.com/llms.txt Baseten hosts an OpenAI-compatible Model APIs endpoint alongside custom model deployments. `pkg/providers/baseten` builds on the shared `openai` package (like the TypeScript provider) — there is no deployment orchestration API in this SDK; use the Baseten dashboard or CLI to deploy models, and this provider to call them. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/baseten" ) ``` ### Configuration ```go provider := baseten.New(baseten.Config{ APIKey: os.Getenv("BASETEN_API_KEY"), }) model, err := provider.LanguageModel("") // "" defaults to "chat" against the Model APIs ``` ### Get API Key ```bash export BASETEN_API_KEY=... ``` `Config.APIKey` falls back to `BASETEN_API_KEY` when left empty. ## Config ```go type Config struct { APIKey string // falls back to BASETEN_API_KEY BaseURL string // default: https://inference.baseten.co/v1 ModelURL string // custom deployment URL (see below) Headers map[string]string HTTPClient *http.Client } ``` ## Model APIs vs. Custom Deployments - **Model APIs (default):** leave `ModelURL` empty. `provider.LanguageModel(modelID)` calls Baseten's shared OpenAI-compatible Model APIs at `https://inference.baseten.co/v1` with `ChatProviderName: "baseten.chat"`. An empty `modelID` defaults to `"chat"`. - **Custom deployment:** set `Config.ModelURL` to your deployment's `/sync/v1` endpoint. An empty `modelID` then defaults to `"placeholder"` instead of `"chat"`, since dedicated single-model endpoints ignore the wire-level `model` field. A `/predict` URL is rejected — chat requires a `/sync/v1` (OpenAI-compatible Chat Completions) endpoint. ```go provider := baseten.New(baseten.Config{ APIKey: os.Getenv("BASETEN_API_KEY"), ModelURL: "https://model-xxxxxxx.api.baseten.co/environments/production/sync/v1", }) model, err := provider.LanguageModel("") if err != nil { log.Fatal(err) } ``` > **Breaking change (this cycle):** the default base URL changed from the > old `bridge.baseten.co` host to `https://inference.baseten.co/v1`, and > the chat model's provider name is now `"baseten.chat"`. With a custom > `/sync/v1` `ModelURL`, an empty model ID now defaults to `"placeholder"` > (was `"chat"`). ## Embeddings Baseten has no default embeddings endpoint on the Model APIs. `Config.ModelURL` must point at a `/sync` or `/sync/v1` deployment URL, or `EmbeddingModel` returns an error: ```go provider := baseten.New(baseten.Config{ APIKey: os.Getenv("BASETEN_API_KEY"), ModelURL: "https://model-xxxxxxx.api.baseten.co/environments/production/sync", }) embeddingModel, err := provider.EmbeddingModel("") if err != nil { log.Fatal(err) } ``` Baseten embeddings batch at most 128 inputs per call. ## Video Parts Chat requests sent through this provider set `AllowVideo: true` internally, so `video/*` file parts are sent as `video_url` content parts — see [OpenAI-compatible providers: Video Parts](https://goaisdk.com/docs/providers/openai-compatible.md#video-parts). ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/baseten" ) func main() { provider := baseten.New(baseten.Config{ APIKey: os.Getenv("BASETEN_API_KEY"), }) model, err := provider.LanguageModel("") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{Model: model, Prompt: "Explain Baseten"}) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Workflow Serialization Baseten embedding models can cross a workflow boundary with `provider.SerializeEmbeddingModel` / `DeserializeEmbeddingModel` (language models could already be serialized with `providerutils.SerializeModel` / `DeserializeModel`). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Baseten Documentation](https://docs.baseten.co) --- # Cerebras Provider > Setup and usage guide for Cerebras ultra-fast wafer-scale inference with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/cerebras Documentation index: https://goaisdk.com/llms.txt Cerebras provides ultra-fast AI inference using wafer-scale engines — ideal for latency-critical and high-throughput applications. ## Setup ### Installation Cerebras has its own dedicated package (not the generic `openai` package with a custom `BaseURL`): ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/cerebras" ) ``` ### Configuration ```go provider := cerebras.New(cerebras.Config{ APIKey: os.Getenv("CEREBRAS_API_KEY"), }) model, err := provider.LanguageModel(cerebras.ModelGPTOSS120B) ``` `cerebras.Config` also accepts `BaseURL` (default `https://api.cerebras.ai/v1`) and `HTTPClient`. ### Get API Key ```bash export CEREBRAS_API_KEY=... ``` ## Available Models | Model ID | Go constant | Best For | |----------|-------------|----------| | `gpt-oss-120b` | `cerebras.ModelGPTOSS120B` | General-purpose, high throughput | | `gemma-4-31b` | `cerebras.ModelGemma4_31B` | Fast, cost-effective | > **Removed this cycle:** `llama3.1-8b`, `qwen-3-235b-a22b-instruct-2507`, > `qwen-3-235b-a22b-thinking-2507`, `zai-glm-4.6`, and `zai-glm-4.7` were > retired from Cerebras and no longer have Go constants. Cerebras accepts > any model ID string, so passing a retired ID compiles but fails against > the live API — use `provider.LanguageModel(cerebras.ModelGPTOSS120B)` or > check [inference-docs.cerebras.ai/models](https://inference-docs.cerebras.ai/models/overview) > for the current catalog. ## Provider-Specific Features ### Measuring Throughput ```go start := time.Now() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a detailed 500-word essay about AI", }) if err != nil { log.Fatal(err) } elapsed := time.Since(start) tokensPerSec := float64(result.Usage.GetOutputTokens()) / elapsed.Seconds() fmt.Printf("Generated %d tokens in %v (%.0f tokens/sec)\n", result.Usage.GetOutputTokens(), elapsed, tokensPerSec) ``` ### Streaming ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a story"}) if err != nil { log.Fatal(err) } defer stream.Close() for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } ``` ### Video Parts Chat requests sent through this provider set `AllowVideo: true` internally, so `video/*` file parts are sent as `video_url` content parts — see [OpenAI-compatible providers: Video Parts](https://goaisdk.com/docs/providers/openai-compatible.md#video-parts). ### High Throughput Handle massive concurrent requests: ```go var wg sync.WaitGroup requests := 100 start := time.Now() for i := 0; i < requests; i++ { wg.Add(1) go func(idx int) { defer wg.Done() _, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Query %d", idx), }) if err != nil { log.Printf("Request %d failed: %v", idx, err) } }(i) } wg.Wait() elapsed := time.Since(start) fmt.Printf("Processed %d requests in %v\n", requests, elapsed) ``` ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/cerebras" ) func main() { provider := cerebras.New(cerebras.Config{ APIKey: os.Getenv("CEREBRAS_API_KEY"), }) model, err := provider.LanguageModel(cerebras.ModelGPTOSS120B) if err != nil { log.Fatal(err) } start := time.Now() result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Explain wafer-scale computing in one paragraph", }) if err != nil { log.Fatal(err) } elapsed := time.Since(start) tokensPerSec := float64(result.Usage.GetOutputTokens()) / elapsed.Seconds() fmt.Println(result.Text) fmt.Printf("\nGenerated %d tokens in %v (%.0f tokens/sec)\n", result.Usage.GetOutputTokens(), elapsed, tokensPerSec) } ``` ### Interactive Chat Application ```go func interactiveChat(ctx context.Context, model provider.LanguageModel) { scanner := bufio.NewScanner(os.Stdin) messages := []types.Message{} for { fmt.Print("You: ") scanner.Scan() userInput := scanner.Text() if userInput == "exit" { break } messages = append(messages, types.Message{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: userInput}, }, }) stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Messages: messages, }) if err != nil { log.Printf("Error: %v", err) continue } fmt.Print("AI: ") var response string for chunk := range stream.Chunks() { fmt.Print(chunk.Text) response += chunk.Text } fmt.Println() messages = append(messages, types.Message{ Role: types.RoleAssistant, Content: []types.ContentPart{ types.TextContent{Text: response}, }, }) } } ``` ### High-Volume Processing ```go func processHighVolume(ctx context.Context, model provider.LanguageModel, prompts []string) { results := make(chan string, len(prompts)) semaphore := make(chan struct{}, 50) // Limit concurrency start := time.Now() var wg sync.WaitGroup for i, prompt := range prompts { wg.Add(1) semaphore <- struct{}{} // Acquire go func(idx int, p string) { defer wg.Done() defer func() { <-semaphore }() // Release result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: p}) if err != nil { log.Printf("Failed %d: %v", idx, err) return } results <- result.Text }(i, prompt) } wg.Wait() close(results) elapsed := time.Since(start) fmt.Printf("Processed %d prompts in %v\n", len(prompts), elapsed) fmt.Printf("Average: %.2fs per prompt\n", elapsed.Seconds()/float64(len(prompts))) } ``` ## Best Practices 1. **Leverage Speed** - Build real-time interactive applications - Enable instant user feedback - Process high volumes efficiently 2. **Model Selection** - Check [inference-docs.cerebras.ai/models](https://inference-docs.cerebras.ai/models/overview) for current production models - Use `cerebras.ModelGPTOSS120B` for general-purpose tasks - Use `cerebras.ModelGemma4_31B` for cost-sensitive, high-throughput tasks 3. **Architecture** - Wafer-scale engine eliminates GPU bottlenecks - No memory bandwidth constraints - Consistent low latency 4. **Use Cases** - Real-time chat applications - Live content generation - High-volume batch processing - Interactive AI assistants ## Rate Limits & Pricing Check the [Cerebras dashboard](https://cloud.cerebras.ai) for current rate limits and per-model pricing — both vary by model and plan. ## Error Handling ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) if err != nil { if strings.Contains(err.Error(), "rate_limit") { log.Println("Rate limited") } log.Fatal(err) } ``` ## Workflow Serialization Cerebras language models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Cerebras Documentation](https://cerebras.ai/inference) - [Groq Provider](https://goaisdk.com/docs/providers/groq.md) - Alternative fast inference --- # DeepInfra Provider > Setup and usage guide for DeepInfra serverless GPU inference with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/deepinfra Documentation index: https://goaisdk.com/llms.txt DeepInfra provides serverless GPU inference for open-source models with excellent performance and competitive pricing. Access hundreds of models through a simple API with no infrastructure management. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/deepinfra" ) ``` ### Configuration ```go provider := deepinfra.New(deepinfra.Config{ APIKey: os.Getenv("DEEPINFRA_API_KEY"), }) model, err := provider.LanguageModel("meta-llama/Meta-Llama-3.1-70B-Instruct") ``` > **Note:** The DeepInfra provider automatically fixes token counting issues for Gemini/Gemma models. > See [Token Counting Fix](#token-counting-fix) section for details. ### Get API Key 1. Sign up at [deepinfra.com](https://deepinfra.com) 2. Get API key from dashboard 3. Set environment variable: ```bash export DEEPINFRA_API_KEY=... ``` ## Available Models ### Language Models | Model ID | Parameters | Best For | | ---------- | ----------- | ---------- | | meta-llama/Meta-Llama-3.1-70B-Instruct | 70B | General purpose | | meta-llama/Meta-Llama-3.1-8B-Instruct | 8B | Fast, cheap | | mistralai/Mixtral-8x7B-Instruct-v0.1 | 47B | Balanced | | Qwen/Qwen2.5-72B-Instruct | 72B | Multilingual | | microsoft/WizardLM-2-8x22B | 141B | Complex tasks | ### Embedding Models | Model ID | Dimensions | Best For | | ---------- | ----------- | ---------- | | BAAI/bge-large-en-v1.5 | 1024 | English embeddings | | sentence-transformers/all-MiniLM-L6-v2 | 384 | Fast embeddings | ### Vision Models | Model ID | Capability | Best For | | ---------- | ----------- | ---------- | | meta-llama/Llama-3.2-90B-Vision-Instruct | Vision | Image understanding | ## Provider-Specific Features ### Token Counting Fix The Go AI SDK automatically corrects token counting issues for Gemini and Gemma models on DeepInfra. **The Issue:** DeepInfra's API has a bug where `reasoning_tokens` are not included in `completion_tokens` for Gemini/Gemma models with thinking capabilities. This violates the OpenAI-compatible spec and can result in negative token counts. **Example of Incorrect API Response:** ```json { "completion_tokens": 84, // Text-only tokens "completion_tokens_details": { "reasoning_tokens": 1081 // Not included in completion_tokens! } } ``` This would incorrectly calculate: `text_tokens = 84 - 1081 = -997` ❌ **The Fix:** The Go AI SDK automatically detects and corrects this: ```go // When reasoning_tokens > completion_tokens, automatically add them // corrected_completion_tokens = 84 + 1081 = 1165 ✅ ``` **Affected Models:** - `google/gemini-2.5-flash` - `google/gemini-2.5-flash:free` - `google/gemma-2-9b-it` - Other Gemini/Gemma variants with reasoning support **Usage:** ```go provider := deepinfra.New(deepinfra.Config{ APIKey: os.Getenv("DEEPINFRA_API_KEY"), }) model, err := provider.LanguageModel("google/gemini-2.5-flash") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain quantum computing"}) // Token usage is automatically corrected if result.Usage.OutputDetails != nil { textTokens := result.Usage.OutputDetails.TextTokens // Correct text tokens reasoningTokens := result.Usage.OutputDetails.ReasoningTokens // Correct reasoning tokens totalOutput := result.Usage.OutputTokens // Correct total (text + reasoning) } ``` No configuration needed - the fix is automatic and transparent. See `examples/deepinfra-token-fix` for a complete example. ### Serverless Infrastructure No infrastructure management required: ```go // Automatic scaling and GPU allocation result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) // DeepInfra handles all infrastructure ``` ### Pay-Per-Use Pay only for actual inference time: ```go // No idle costs, no minimum commitments result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) // Charged only for tokens used ``` ### Wide Model Selection Access hundreds of open-source models: ```go // Browse available models models := []string{ "meta-llama/Meta-Llama-3.1-70B-Instruct", "mistralai/Mixtral-8x7B-Instruct-v0.1", "google/gemma-2-9b-it", "microsoft/Phi-3-medium-128k-instruct", "Qwen/Qwen2.5-72B-Instruct", } for _, modelID := range models { model, err := provider.LanguageModel(modelID) // Test different models easily } ``` ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/deepinfra" ) func main() { provider := deepinfra.New(deepinfra.Config{ APIKey: os.Getenv("DEEPINFRA_API_KEY"), }) model, err := provider.LanguageModel("meta-llama/Meta-Llama-3.1-70B-Instruct") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Explain serverless GPU inference", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) fmt.Printf("Cost: $%.6f\n", calculateCost(result.Usage.GetInputTokens(), result.Usage.GetOutputTokens())) } func calculateCost(inputTokens, outputTokens int64) float64 { inputCost := float64(inputTokens) * 0.52 / 1_000_000 outputCost := float64(outputTokens) * 0.75 / 1_000_000 return inputCost + outputCost } ``` ### Model Comparison ```go func compareModels(prompt string) { models := map[string]string{ "Llama 3.1 70B": "meta-llama/Meta-Llama-3.1-70B-Instruct", "Mixtral 8x7B": "mistralai/Mixtral-8x7B-Instruct-v0.1", "Qwen 2.5 72B": "Qwen/Qwen2.5-72B-Instruct", } for name, modelID := range models { model, err := provider.LanguageModel(modelID) if err != nil { log.Printf("Failed to load %s: %v", name, err) continue } start := time.Now() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) if err != nil { log.Printf("Failed %s: %v", name, err) continue } elapsed := time.Since(start) fmt.Printf("\n=== %s ===\n", name) fmt.Printf("Response: %s\n", result.Text) fmt.Printf("Time: %v\n", elapsed) fmt.Printf("Tokens: %d\n", result.Usage.GetTotalTokens()) } } ``` ### Embeddings ```go embeddingModel, err := provider.EmbeddingModel("BAAI/bge-large-en-v1.5") if err != nil { log.Fatal(err) } texts := []string{ "DeepInfra provides serverless AI", "GPU inference without infrastructure", "Open-source models on demand", } result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: texts, }) if err != nil { log.Fatal(err) } fmt.Printf("Generated %d embeddings of dimension %d\n", len(result.Embeddings), len(result.Embeddings[0])) // Use embeddings for semantic search, clustering, etc. ``` ### Streaming ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a detailed article"}) if err != nil { log.Fatal(err) } defer stream.Close() for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } ``` ## Best Practices 1. **Model Selection** - Use Llama 3.1 70B for high quality - Use Llama 3.1 8B for cost efficiency - Use Mixtral for balanced performance - Use Qwen for multilingual tasks 2. **Cost Optimization** - Compare prices across models - Use smaller models for simple tasks - Monitor token usage - Cache results when appropriate 3. **Performance** - Serverless = no cold starts to worry about - Automatic scaling for traffic spikes - Geographic distribution for low latency 4. **Reliability** - Implement retry logic - Handle rate limits gracefully - Monitor API status ## Rate Limits & Pricing ### Rate Limits Varies by plan: - Free tier: 10 requests/min - Pay-as-you-go: Higher limits - Enterprise: Custom limits ### Pricing Pricing varies by model. See the [DeepInfra pricing page](https://deepinfra.com/pricing) for current values. ## Error Handling ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) if err != nil { if strings.Contains(err.Error(), "rate_limit") { log.Println("Rate limited, implement backoff") time.Sleep(time.Second * 5) // Retry } else if strings.Contains(err.Error(), "model_not_found") { log.Fatal("Model not available on DeepInfra") } else if strings.Contains(err.Error(), "insufficient_credits") { log.Fatal("Add credits to account") } log.Fatal(err) } ``` ## Advanced Features ### Custom Deployments Deploy custom models: ```go // Contact DeepInfra for custom model deployments // Bring your own fine-tuned models ``` ### API Compatibility Works with OpenAI SDK: ```go // DeepInfra is OpenAI-compatible // Easy migration from OpenAI // Use same code, just change base URL ``` ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [DeepInfra Documentation](https://deepinfra.com/docs) - [DeepInfra Model Library](https://deepinfra.com/models) - [Together AI Provider](https://goaisdk.com/docs/providers/together.md) - Alternative - [Fireworks AI Provider](https://goaisdk.com/docs/providers/fireworks.md) - Alternative --- # GMI Cloud Provider > Setup and usage guide for GMI Cloud OpenAI-compatible inference with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/gmicloud Documentation index: https://goaisdk.com/llms.txt GMI Cloud provides GPU inference for open-weight models over an OpenAI-compatible chat API. It only supports language models — `EmbeddingModel`, `ImageModel`, `SpeechModel`, `TranscriptionModel`, and `RerankingModel` all return errors. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/gmicloud" ) ``` ### Configuration ```go provider := gmicloud.New(gmicloud.Config{ APIKey: os.Getenv("GMI_CLOUD_APIKEY"), }) model, err := provider.LanguageModel(gmicloud.ModelDeepSeekV4FlashChat) ``` `gmicloud.Config` also accepts `BaseURL` (default `https://api.gmi-serving.com/v1`), `Headers`, and `HTTPClient`. ### Get API Key ```bash export GMI_CLOUD_APIKEY=... ``` ## Available Models | Model ID | Go constant | |----------|-------------| | `deepseek-ai/DeepSeek-V4-Flash-0731` | `gmicloud.ModelDeepSeekV4FlashChat` | | `Qwen/Qwen3.8-Max` | `gmicloud.ModelQwen38Max` | | `moonshotai/kimi-k3` | `gmicloud.ModelKimiK3` | | `zai-org/GLM-5.2-FP8` | `gmicloud.ModelGLM52FP8` | | `MiniMaxAI/MiniMax-M3` | `gmicloud.ModelMiniMaxM3` | `ChatModel` is an alias for `LanguageModel`. ## Basic Usage ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain GPU inference in one paragraph", }) ``` ## Provider Options GMI Cloud-specific options are passed under the `"gmicloud"` `ProviderOptions` key, resolved through the same OpenAI-compatible option resolver used by the other OpenAI-compatible providers in this SDK (see [OpenAI-compatible providers](https://goaisdk.com/docs/providers/openai-compatible.md)). ## Video Parts Chat requests sent through this provider set `AllowVideo: true` internally, so `video/*` file parts are sent as `video_url` content parts — see [OpenAI-compatible providers: Video Parts](https://goaisdk.com/docs/providers/openai-compatible.md#video-parts). ## Workflow Serialization GMI Cloud language models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [GMI Cloud Documentation](https://docs.gmicloud.ai) --- # Z.AI Provider > Setup guide for Z.AI's GLM chat models over an OpenAI-compatible API in the Go AI SDK, covering available models, provider options, and basic usage. Canonical URL: https://goaisdk.com/docs/providers/zai Documentation index: https://goaisdk.com/llms.txt Z.AI provides GLM chat models over an OpenAI-compatible API. It only supports language models — `EmbeddingModel`, `ImageModel`, `SpeechModel`, `TranscriptionModel`, and `RerankingModel` all return errors. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/zai" ) ``` ### Configuration ```go provider := zai.New(zai.Config{ APIKey: os.Getenv("ZAI_API_KEY"), }) model, err := provider.LanguageModel(zai.ModelGLM53) ``` `zai.Config` also accepts `BaseURL` (default `https://api.z.ai/api/paas/v4`) and `Headers`. ### Get API Key ```bash export ZAI_API_KEY=... ``` ## Available Models | Model ID | Go constant | |----------|-------------| | `glm-5.3` | `zai.ModelGLM53` | | `glm-5.2` | `zai.ModelGLM52` | | `glm-5.1` | `zai.ModelGLM51` | | `glm-5-turbo` | `zai.ModelGLM5Turbo` | | `glm-5` | `zai.ModelGLM5` | | `glm-4.7` | `zai.ModelGLM47` | | `glm-4.7-flash` | `zai.ModelGLM47Flash` | | `glm-4.7-flashx` | `zai.ModelGLM47FlashX` | | `glm-4.6` | `zai.ModelGLM46` | | `glm-4.5` | `zai.ModelGLM45` | | `glm-4.5-air` | `zai.ModelGLM45Air` | See `pkg/providers/zai/model_ids.go` for the full catalog, including dated/legacy variants. `ChatModel` is an alias for `LanguageModel`. ## Basic Usage ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain the GLM model family in one paragraph", }) ``` ## Provider Options Z.AI-specific options are passed under the `"zai"` `ProviderOptions` key, resolved through the same OpenAI-compatible option resolver used by the other OpenAI-compatible providers in this SDK (see [OpenAI-compatible providers](https://goaisdk.com/docs/providers/openai-compatible.md)). ## Workflow Serialization Z.AI language models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Z.AI Documentation](https://docs.z.ai) --- # MiniMax Provider > Setup and usage guide for MiniMax chat and video models with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/minimax Documentation index: https://goaisdk.com/llms.txt MiniMax provides chat language models over an Anthropic-compatible endpoint, plus its own video generation API. Chat requests delegate to the shared Anthropic language model implementation (`Provider.Name()` is `"minimax"`), so `providerOptions.minimax` and `providerOptions.anthropic` are both read on every call. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/minimax" ) ``` ### Configuration ```go provider := minimax.New(minimax.Config{ APIKey: os.Getenv("MINIMAX_API_KEY"), }) model, err := provider.LanguageModel(minimax.ModelM3) ``` `minimax.Config` also accepts `BaseURL` (default `https://api.minimax.io/anthropic/v1`), `VideoBaseURL` (default `https://api.minimax.io`), and `Headers`. Video-request credentials are resolved fresh on every call, so rotating `MINIMAX_API_KEY` after `New()` still takes effect on the next request. ### Get API Key ```bash export MINIMAX_API_KEY=... ``` ## Available Chat Models | Model ID | Go constant | |----------|-------------| | `minimax-m3` | `minimax.ModelM3` | | `minimax-m2.7` | `minimax.ModelM27` | | `minimax-m2.7-highspeed` | `minimax.ModelM27HighSpeed` | | `minimax-m2.5` | `minimax.ModelM25` | | `minimax-m2.5-highspeed` | `minimax.ModelM25HighSpeed` | | `minimax-m2.1` | `minimax.ModelM21` | | `minimax-m2.1-highspeed` | `minimax.ModelM21HighSpeed` | | `minimax-m2` | `minimax.ModelM2` | `ChatModel` is an alias for `LanguageModel`. `EmbeddingModel`, `ImageModel`, `SpeechModel`, `TranscriptionModel`, and `RerankingModel` all return errors — MiniMax only supports chat and video here. ## Basic Usage ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain the MiniMax M3 model in one paragraph", }) ``` ## Video Generation ```go videoModel, err := provider.VideoModel(string(minimax.ModelH3)) if err != nil { log.Fatal(err) } result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{Text: "A neon-lit city street at night, rain reflections"}, }) ``` Model IDs: `minimax.ModelH3` (`MiniMax-H3`), `minimax.ModelH3Max` (`MiniMax-H3-Max`). Default resolution is `"2K"` and default aspect ratio is `"16:9"`. `VideoModel` also implements `provider.VideoModelStarter` / `VideoModelStatusChecker`, so it works with `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus`. ## Provider Options Since chat delegates to the Anthropic language model, MiniMax accepts the same per-call `providerOptions.anthropic` fields (thinking, context management, etc. — see the [Anthropic provider](https://goaisdk.com/docs/providers/anthropic.md)) under either the `"minimax"` or `"anthropic"` key. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Anthropic Provider](https://goaisdk.com/docs/providers/anthropic.md) - MiniMax chat shares this implementation - [MiniMax Documentation](https://www.minimax.io/platform_overview) --- # Moonshot AI Provider > Setup guide for Moonshot AI's Kimi and legacy Moonshot v1 chat models in the Go AI SDK, covering provider options, metadata, and workflow serialization. Canonical URL: https://goaisdk.com/docs/providers/moonshot Documentation index: https://goaisdk.com/llms.txt Moonshot AI provides the Kimi model family and the legacy Moonshot v1 chat models over an OpenAI-compatible API. It only supports language models — `EmbeddingModel`, `ImageModel`, `SpeechModel`, and `TranscriptionModel` all return errors. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/moonshot" ) ``` ### Configuration ```go provider := moonshot.New(moonshot.Config{ APIKey: os.Getenv("MOONSHOT_API_KEY"), }) model, err := provider.LanguageModel(moonshot.ModelKimiK26) ``` Unlike most providers, `moonshot.New` does **not** fall back to an environment variable on its own — pass `APIKey` explicitly, or use `moonshot.NewConfig(apiKey)` (empty string falls back to `MOONSHOT_API_KEY`, returning an error if that is empty too): ```go cfg, err := moonshot.NewConfig("") if err != nil { log.Fatal(err) } provider := moonshot.New(cfg) ``` Any model ID is accepted; an empty ID defaults to `moonshot-v1-32k`. ### Get API Key ```bash export MOONSHOT_API_KEY=... ``` > **Breaking change (this cycle):** the default base URL is now > `https://api.moonshot.ai/v1` (was the China-region > `https://api.moonshot.cn/v1`). Set `Config.BaseURL` explicitly to the > `.cn` host if you relied on the previous default. ## Available Models | Model ID | Go constant | |----------|-------------| | `kimi-k2.7-code` | `moonshot.ModelKimiK27Code` | | `kimi-k2.6` | `moonshot.ModelKimiK26` | | `kimi-k2.5` | `moonshot.ModelKimiK25` | | `kimi-k2-thinking` | `moonshot.ModelKimiK2Thinking` | | `kimi-k2-thinking-turbo` | `moonshot.ModelKimiK2ThinkingTurbo` | | `kimi-k2-turbo` | `moonshot.ModelKimiK2Turbo` | | `kimi-k2-0905` | `moonshot.ModelKimiK20905` | | `kimi-k2` | `moonshot.ModelKimiK2` | | `moonshot-v1-auto` | `moonshot.ModelMoonshotV1Auto` | | `moonshot-v1-128k` | `moonshot.ModelMoonshotV1128k` | | `moonshot-v1-32k` | `moonshot.ModelMoonshotV132k` (default) | | `moonshot-v1-8k` | `moonshot.ModelMoonshotV18k` | | `moonshot-v1-128k-vision-preview` | `moonshot.ModelMoonshotV1128kVisionPreview` | | `moonshot-v1-32k-vision-preview` | `moonshot.ModelMoonshotV132kVisionPreview` | | `moonshot-v1-8k-vision-preview` | `moonshot.ModelMoonshotV18kVisionPreview` | ## Basic Usage ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain the Kimi K2 model family in one paragraph", }) ``` ## Provider Options And Metadata > **Breaking change (this cycle):** result metadata moved from > `ProviderMetadata["moonshot"]` to `ProviderMetadata["moonshotai"]` > (matching the TypeScript SDK's key). `providerOptions.moonshotai` is now > honored alongside the package's own `providerOptions.moonshot` key. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Solve this step by step: ...", ProviderOptions: map[string]interface{}{ "moonshotai": map[string]interface{}{ "reasoningEffort": "high", }, }, }) if meta, ok := result.ProviderMetadata["moonshotai"].(map[string]interface{}); ok { fmt.Println(meta) } ``` ## Workflow Serialization Moonshot language models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Moonshot AI Documentation](https://platform.moonshot.ai/docs) --- # Alibaba Cloud Provider > Setup and usage guide for Alibaba Cloud Qwen and Wan models with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/alibaba Documentation index: https://goaisdk.com/llms.txt Alibaba Cloud provides powerful Qwen language models with strong Chinese language capabilities and Wan video generation models. Known for cost-effective prompt caching, reasoning capabilities, and innovative video generation features. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/alibaba" ) ``` ### Configuration ```go provider := alibaba.New(alibaba.Config{ APIKey: os.Getenv("ALIBABA_API_KEY"), }) // Chat model model, err := provider.LanguageModel("qwen-plus") if err != nil { log.Fatal(err) } // Video model videoModel, err := provider.VideoModel("wan2.6-t2v") if err != nil { log.Fatal(err) } ``` ### Get API Key 1. Sign up at [Alibaba Cloud DashScope](https://dashscope.console.aliyun.com/) 2. Navigate to API Keys section 3. Create a new API key 4. Set environment variable: ```bash export ALIBABA_API_KEY=sk-... ``` ## Available Models ### Qwen Language Models | Model ID | Context | Best For | |----------|---------|----------| | qwen-plus | 32K | Balanced performance and cost | | qwen-turbo | 8K | Fast responses, economical | | qwen3-max | 32K | Most capable, highest quality | | qwq-32b | 32K | Complex reasoning with thinking | | qwen-vl-max | 32K | Vision + language understanding | ### Wan Video Models `Provider.VideoModel` accepts any model ID (there is no fixed allowlist); the constants below (`pkg/providers/alibaba/video_model_ids.go`) cover the documented catalog. An empty model ID defaults to `wan2.6-t2v`. | Model ID | Go constant | Type | Best For | |----------|-------------|------|----------| | wan2.5-t2v-preview | `VideoModelWan2_5T2VPreview` | Text-to-video | Preview quality (720p) | | wan2.6-t2v | `VideoModelWan2_6T2V` | Text-to-video | High quality generation | | wan2.7-t2v | `VideoModelWan2_7T2V` | Text-to-video | Latest generation | | wan2.6-i2v | `VideoModelWan2_6I2V` | Image-to-video | Animate static images | | wan2.6-i2v-flash | `VideoModelWan2_6I2VFlash` | Image-to-video | Faster generation | | wan2.6-r2v | `VideoModelWan2_6R2V` | Reference-to-video | Style transfer from reference | | wan2.6-r2v-flash | `VideoModelWan2_6R2VFlash` | Reference-to-video | Faster style transfer | | wan2.7-r2v | `VideoModelWan2_7R2V` | Reference-to-video | Latest reference-to-video | | wan3.0-video | `VideoModelWan3Video` | All-in-one | One model ID serves text-, image-, and reference-to-video, and supports custom aspect ratios | `wan2.7` and `wan3` are new this cycle — `wan3.0-video` in particular replaces separate t2v/i2v/r2v model IDs with one model that branches on which inputs (`Prompt` only, `Image`, or `InputReferences`) you pass. ### Embedding Models Alibaba DashScope text embeddings are exposed through `Provider.EmbeddingModel` and `Provider.Embedding`. | Model ID | Go constant | Notes | |----------|-------------|-------| | text-embedding-v4 | `alibaba.AlibabaEmbeddingTextV4` | Dense, sparse, or dense+sparse output | | text-embedding-v3 | `alibaba.AlibabaEmbeddingTextV3` | Dense text embeddings | ```go embeddingModel, err := provider.EmbeddingModel(alibaba.AlibabaEmbeddingTextV4) if err != nil { log.Fatal(err) } dimension := 1024.0 result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: embeddingModel, Inputs: []string{"first document", "second document"}, ProviderOptions: map[string]interface{}{ "alibaba": alibaba.AlibabaEmbeddingModelOptions{ TextType: "document", Dimension: &dimension, OutputType: alibaba.AlibabaEmbeddingOutputDense, }, }, }) ``` `OutputType` accepts `dense`, `sparse`, or `dense&sparse`. Sparse-only responses are reported as unsupported for the SDK's dense embedding result contract; dense+sparse responses preserve sparse vectors in provider metadata. ## Provider-Specific Features ### Thinking/Reasoning Enable reasoning for complex problems with token tracking: ```go result, err := model.DoGenerate(ctx, &provider.GenerateOptions{ Prompt: types.Prompt{ Text: "Solve this logic puzzle: ...", }, ProviderOptions: map[string]interface{}{ "alibaba": map[string]interface{}{ "enableThinking": true, "thinkingBudget": 1000, // Max reasoning tokens }, }, }) // Access reasoning tokens if result.Usage.OutputDetails != nil && result.Usage.OutputDetails.ReasoningTokens != nil { fmt.Printf("Reasoning tokens used: %d\n", *result.Usage.OutputDetails.ReasoningTokens) } ``` ### Prompt Caching Automatic caching for repeated content with cost savings: ```go // System prompts are automatically cached prompt := types.Prompt{ System: "You are an expert software architect...", // Cached Text: "Design a microservices system", } result, err := model.DoGenerate(ctx, &provider.GenerateOptions{ Prompt: prompt, }) // Check cache usage if result.Usage.InputDetails != nil { if result.Usage.InputDetails.CacheReadTokens != nil { fmt.Printf("Cache hit: %d tokens saved!\n", *result.Usage.InputDetails.CacheReadTokens) } if result.Usage.InputDetails.CacheWriteTokens != nil { fmt.Printf("Cached: %d tokens for future use\n", *result.Usage.InputDetails.CacheWriteTokens) } } ``` ### Vision Capabilities Image understanding with Qwen VL models: ```go messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.FileContent{MediaType: "image/jpeg", URL: "https://example.com/photo.jpg"}, types.TextContent{Text: "What's in this image?"}, }, }, } visionModel, _ := provider.LanguageModel("qwen-vl-max") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: visionModel, Messages: messages, }) ``` ### Tool Calling Function calling with single or parallel execution: ```go tools := []types.Tool{ { Name: "get_weather", Description: "Get current weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "City name", }, }, "required": []string{"location"}, }, }, } result, err := model.DoGenerate(ctx, &provider.GenerateOptions{ Prompt: types.Prompt{Text: "What's the weather in Paris?"}, Tools: tools, ProviderOptions: map[string]interface{}{ "alibaba": map[string]interface{}{ "parallelToolCalls": true, // Enable parallel execution }, }, }) // Handle tool calls if len(result.ToolCalls) > 0 { // Execute tools and send results back } ``` ### Video Generation #### Text-to-Video ```go videoModel, _ := provider.VideoModel("wan2.6-t2v") duration := 5.0 result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "A golden retriever playing in a park, sunny day", AspectRatio: "16:9", Duration: &duration, }) if len(result.Videos) > 0 { fmt.Printf("Video URL: %s\n", result.Videos[0].URL) } ``` #### Image-to-Video Animate static images: ```go videoModel, _ := provider.VideoModel("wan2.6-i2v-flash") duration := 6.0 result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "Add gentle camera movement and natural motion", AspectRatio: "16:9", Duration: &duration, Image: &provider.VideoModelV3File{ Type: "url", URL: "https://example.com/landscape.jpg", }, }) ``` #### Reference-to-Video (Style Transfer) Apply style from reference image to new content: ```go videoModel, _ := provider.VideoModel("wan2.6-r2v") duration := 5.0 result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "A cat walking through a garden", AspectRatio: "16:9", Duration: &duration, Image: &provider.VideoModelV3File{ Type: "url", URL: "https://example.com/anime-style.jpg", // Reference style }, }) ``` #### Async Video `VideoModel` also implements `provider.VideoModelStarter` / `VideoModelStatusChecker`, so it works with `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus` instead of the polling `ai.GenerateVideo`. The default poll interval/timeout are 5s / 10 minutes, and the provider-returned task ID is path-encoded when polling status. ```go started, err := ai.ExperimentalStartVideo(ctx, ai.StartVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{Text: "A golden retriever playing in a park"}, }) if err != nil { log.Fatal(err) } status, err := ai.ExperimentalGetVideoStatus(ctx, videoModel, ai.GetVideoStatusOptions{ Operation: started.Operation, }) ``` ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/alibaba" ) func main() { cfg, err := alibaba.NewConfig(os.Getenv("ALIBABA_API_KEY")) if err != nil { log.Fatal(err) } prov := alibaba.New(cfg) model, _ := prov.LanguageModel("qwen-plus") result, err := model.DoGenerate(context.Background(), &provider.GenerateOptions{ Prompt: types.Prompt{ Text: "Write a haiku about artificial intelligence.", }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Reasoning with Thinking ```go model, _ := provider.LanguageModel("qwq-32b") result, err := model.DoGenerate(ctx, &provider.GenerateOptions{ Prompt: types.Prompt{ Text: `Solve this logic puzzle: Three friends - Alice, Bob, and Carol - each have different pets. - Alice doesn't have a dog - Bob is allergic to cats - Carol lives in an apartment that doesn't allow fish Who has which pet?`, }, ProviderOptions: map[string]interface{}{ "alibaba": map[string]interface{}{ "enableThinking": true, "thinkingBudget": 1000, }, }, }) ``` ### Video Generation ```go videoModel, _ := provider.VideoModel("wan2.6-t2v") duration := 5.0 result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "A person dancing in the rain", AspectRatio: "9:16", // Vertical for mobile Duration: &duration, }) ``` ## Best Practices 1. **Use Prompt Caching** - Cache long system prompts for cost savings - Second request with same system prompt uses cached tokens - Significant cost reduction for repeated queries 2. **Choose the Right Model** - `qwen-turbo`: Simple tasks, fast responses - `qwen-plus`: Balanced for most use cases - `qwen3-max`: Complex reasoning, highest quality - `qwq-32b`: Logic puzzles, math problems 3. **Video Generation** - Text-to-video takes 1-2 minutes - Use flash variants when speed > quality - Aspect ratios: 16:9 (landscape), 9:16 (portrait), 1:1 (square) - Reference-to-video for consistent visual style 4. **Thinking Budget** - Set limits to control costs - Monitor reasoning token usage - Higher budgets for complex problems ## Rate Limits & Pricing ### Rate Limits Check [Alibaba Cloud DashScope documentation](https://help.aliyun.com/zh/dashscope/) for current limits. ### Cost Optimization - Enable prompt caching for repeated content - Use `qwen-turbo` for simple tasks - Set thinking budget limits for reasoning models - Use flash video variants when appropriate ## API Endpoints - **Chat API**: `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` (OpenAI-compatible) - **Video API**: `https://dashscope-intl.aliyuncs.com` (DashScope native) ## Complete Examples See the [Alibaba provider examples](https://github.com/digitallysavvy/go-ai/tree/main/examples/providers/alibaba) directory for 8 complete examples: 1. Basic chat 2. Thinking/reasoning 3. Prompt caching 4. Text-to-video 5. Image-to-video 6. Vision chat 7. Tool calling 8. Reference-to-video ## Workflow Serialization Alibaba embedding models can cross a workflow boundary with `provider.SerializeEmbeddingModel` / `DeserializeEmbeddingModel` (language models could already be serialized with `providerutils.SerializeModel` / `DeserializeModel`). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; video models are not yet serializable. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [API Reference: GenerateVideo](https://goaisdk.com/docs/ai-sdk-core/video-generation.md) - [Alibaba Cloud DashScope Documentation](https://help.aliyun.com/zh/dashscope/) - [Qwen Models](https://qwen.readthedocs.io/) --- # KlingAI Provider > Setup guide for KlingAI's video generation in the Go AI SDK, covering text-to-video, image-to-video, motion control, async polling, and error handling. Canonical URL: https://goaisdk.com/docs/providers/klingai Documentation index: https://goaisdk.com/llms.txt KlingAI specializes in high-quality video generation with advanced features including text-to-video, image-to-video, start/end frame control, and motion control. Known for competitive pricing and professional-grade video output. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/klingai" ) ``` ### Configuration The recommended path is a single API key, sent directly as a bearer token: ```go provider, err := klingai.New(klingai.Config{ APIKey: os.Getenv("KLINGAI_API_KEY"), }) model, err := provider.VideoModel("kling-v2.6-t2v") ``` ```bash export KLINGAI_API_KEY=... ``` The legacy access-key/secret-key pair (which signs a short-lived JWT per request) is still supported and used when `APIKey` is unset: ```go provider, err := klingai.New(klingai.Config{ AccessKey: os.Getenv("KLINGAI_ACCESS_KEY"), SecretKey: os.Getenv("KLINGAI_SECRET_KEY"), }) ``` ```bash export KLINGAI_ACCESS_KEY=your-access-key export KLINGAI_SECRET_KEY=your-secret-key ``` Credentials are resolved on every request rather than once in `New()`, so changing the environment variables takes effect without rebuilding the provider. ## Available Models ### Text-to-Video Models | Model ID | Quality | Speed | Duration | Best For | |----------|---------|-------|----------|----------| | kling-v2.6-t2v | Excellent | Medium | 5-10s | Latest quality | | kling-v2.5-turbo-t2v | Good | Fast | 5-10s | Speed optimized | | kling-v2.1-master-t2v | High | Medium | 5-10s | Stable version | | kling-v1.6-t2v | Good | Fast | 5s | Legacy support | ### Image-to-Video Models | Model ID | Quality | Speed | Features | Best For | |----------|---------|-------|----------|----------| | kling-v2.6-i2v | Excellent | Medium | Start/end control | Professional work | | kling-v2.5-turbo-i2v | Good | Fast | Basic animation | Quick animations | | kling-v2.1-master-i2v | High | Medium | Advanced control | Production quality | ### Motion Control Models | Model ID | Quality | Speed | Max Reference | Best For | |----------|---------|-------|---------------|----------| | kling-v2.6-motion-control | High | Slow | 30s (pro) / 10s (std) | Motion transfer | | kling-v3.0-motion-control | Highest | Slow | 30s (pro) / 10s (std) | Latest motion transfer | ## Provider-Specific Features ### Text-to-Video Generation Create videos from text descriptions: ```go model, _ := provider.VideoModel("kling-v2.6-t2v") duration := 10.0 result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: ai.VideoPrompt{Text: "A sunrise over mountains in 4K"}, Duration: &duration, AspectRatio: "16:9", ProviderOptions: map[string]interface{}{ "klingai": map[string]interface{}{ "mode": "pro", "negativePrompt": "low quality, blurry", }, }, }) ``` ### Image-to-Video Animation Animate static images with camera control: ```go model, _ := provider.VideoModel("kling-v2.6-i2v") result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: ai.VideoPrompt{ Image: &ai.VideoPromptImage{Data: imageData}, Text: "The cat slowly turns its head and blinks", }, Duration: &duration, ProviderOptions: map[string]interface{}{ "klingai": map[string]interface{}{ "mode": "pro", "cameraControl": map[string]interface{}{ "type": "forward_up", "config": map[string]interface{}{ "horizontal": 0.5, "vertical": 0.3, "zoom": 0.2, }, }, }, }, }) ``` ### Start/End Frame Control Precise control over video start and end states: ```go model, _ := provider.VideoModel("kling-v2.6-i2v") result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: ai.VideoPrompt{ Image: &ai.VideoPromptImage{Data: startImage}, Text: "Transform cat into dog with smooth transition", }, ProviderOptions: map[string]interface{}{ "klingai": map[string]interface{}{ "mode": "pro", // Required for start/end control "imageTail": endImageURL, "sound": "on", }, }, }) ``` ### Motion Control Transfer motion from reference videos to your characters: ```go model, _ := provider.VideoModel("kling-v2.6-motion-control") result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: ai.VideoPrompt{ Image: &ai.VideoPromptImage{Data: characterImage}, Text: "Character performs smooth dance move", }, ProviderOptions: map[string]interface{}{ "klingai": map[string]interface{}{ "videoUrl": "https://example.com/reference-motion.mp4", "characterOrientation": "video", // "image" or "video" "mode": "pro", "keepOriginalSound": "yes", "pollTimeoutMs": 600000, // 10 minutes for pro mode }, }, }) ``` ## Provider Options Configure generation with `ProviderOptions.klingai`: ### Common Options | Option | Type | Description | Modes | |--------|------|-------------|-------| | `mode` | string | `"std"` or `"pro"` quality mode | All | | `pollIntervalMs` | int | Polling interval (default: 5000ms) | All | | `pollTimeoutMs` | int | Max wait time (default: 600000ms) | All | ### Text-to-Video & Image-to-Video Options | Option | Type | Description | |--------|------|-------------| | `negativePrompt` | string | What to avoid (max 2500 chars) | | `sound` | string | `"on"` or `"off"` (V2.6+ pro only) | | `cfgScale` | float64 | Prompt adherence 0.0-1.0 (V1.x only) | ### Camera Control ```go "cameraControl": map[string]interface{}{ "type": "forward_up", // simple, down_back, right_turn_forward, left_turn_forward "config": map[string]interface{}{ "horizontal": 0.5, // -1.0 to 1.0 "vertical": 0.3, // -1.0 to 1.0 "zoom": 0.2, // -1.0 to 1.0 }, } ``` ### Image-to-Video Specific | Option | Type | Description | |--------|------|-------------| | `imageTail` | string | End frame image URL or base64 | | `staticMask` | string | Static brush mask image | | `dynamicMasks` | array | Dynamic brush configurations | ### Motion Control Specific | Option | Type | Description | |--------|------|-------------| | `videoUrl` | string | Reference video URL (required) | | `characterOrientation` | string | `"image"` or `"video"` (required) | | `keepOriginalSound` | string | `"yes"` or `"no"` | | `watermarkEnabled` | bool | Enable watermark | ## Examples ### Basic Text-to-Video ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/klingai" ) func main() { prov, err := klingai.New(klingai.Config{ APIKey: os.Getenv("KLINGAI_API_KEY"), }) if err != nil { log.Fatal(err) } model, err := prov.VideoModel("kling-v2.6-t2v") if err != nil { log.Fatal(err) } duration := 5.0 result, err := model.DoGenerate(context.Background(), &provider.VideoModelV3CallOptions{ Prompt: "A sunrise over mountains", Duration: &duration, AspectRatio: "16:9", }) if err != nil { log.Fatal(err) } fmt.Printf("Video URL: %s\n", result.Videos[0].URL) } ``` See `examples/providers/klingai/` for more complete examples including: - Image-to-video with camera control - Start/end frame control - Motion control (standard and pro) ## Async Video `VideoModel` also implements `provider.VideoModelStarter` / `VideoModelStatusChecker`, so it works with `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus` in addition to the polling `ai.GenerateVideo`. ## Polling Behavior KlingAI video generation is asynchronous with automatic polling: - **Default interval:** 5 seconds - **Default timeout:** 10 minutes - **Configurable** via provider options ```go ProviderOptions: map[string]interface{}{ "klingai": map[string]interface{}{ "pollIntervalMs": 3000, // Poll every 3 seconds "pollTimeoutMs": 300000, // 5 minute timeout }, } ``` ## Error Handling Common error scenarios: ```go result, err := model.DoGenerate(ctx, opts) if err != nil { if klingErr, ok := err.(*klingai.Error); ok { switch klingErr.Code { case 401: log.Println("Authentication failed") case 504: log.Println("Generation timeout") default: log.Printf("KlingAI error %d: %s", klingErr.Code, klingErr.Message) } } return err } ``` ## Best Practices 1. **Use Pro Mode for Quality** - Better visual quality - Required for start/end frame control - Supports audio generation (V2.6+) 2. **Optimize Polling** - Use longer intervals for pro mode - Set appropriate timeouts (5-10 minutes) - Handle timeout errors gracefully 3. **Motion Control Tips** - Use clear, well-lit reference videos - Match character orientation correctly - Pro mode supports longer references (up to 30s) 4. **Prompting** - Be specific and detailed - Use negative prompts to avoid unwanted elements - Test different camera controls for best results ## Limitations - **Resolution:** Not configurable (determined by model/image) - **FPS:** Not configurable - **Seed:** Not supported for deterministic generation - **Max videos per call:** 1 - **Max duration:** 10 seconds (standard), varies by model Attempting unsupported options will generate warnings but not fail. ## Rate Limits & Pricing ### Rate Limits - **Token Generation:** 100 requests/minute - **Video Generation:** 10 concurrent tasks per account - **Polling:** No explicit limit (use backoff) ### Pricing Contact KlingAI for current pricing information. Pricing varies by: - Model version (V1, V2, V2.5, V2.6) - Quality mode (standard vs pro) - Video duration (5s vs 10s) ## See Also - [API Reference: GenerateVideo](https://goaisdk.com/docs/ai-sdk-core/video-generation.md) - [Examples: KlingAI](https://github.com/digitallysavvy/go-ai/tree/main/examples/providers/klingai) - [KlingAI API Documentation](https://app.klingai.com/global/dev/document-api) - [KlingAI Model Capabilities](https://app.klingai.com/global/dev/document-api/apiReference/model/skillsMap) --- # AI Gateway > Use Vercel's AI Gateway for unified multi-provider access, automatic failover, and cost optimization. Canonical URL: https://goaisdk.com/docs/providers/gateway Documentation index: https://goaisdk.com/llms.txt The AI Gateway provider enables you to access multiple AI providers through a unified interface with automatic provider selection, failover, and cost optimization. ## Features - **Unified Interface**: Single API for multiple AI providers - **Automatic Provider Selection**: Gateway selects the best available provider - **Failover Handling**: Automatic failover to alternative providers - **Cost Optimization**: Intelligentrouting based on cost and performance - **Zero Data Retention**: Option to not log or retain request data - **Observability**: Built-in metrics and monitoring - **Rate Limiting**: Automatic rate limit handling across providers ## Installation The Gateway provider is included in the core Go AI SDK: ```bash go get github.com/digitallysavvy/go-ai ``` ## Setup ### Environment Variables ```bash export AI_GATEWAY_API_KEY="your-api-key-here" ``` `AI_GATEWAY_API_KEY` may contain either an AI Gateway API key or a Vercel access token. When a Vercel access token can access multiple teams, set `TeamIDOrSlug`; the provider sends it as `x-vercel-ai-gateway-team`. ### Basic Configuration ```go package main import ( "log" "os" "github.com/digitallysavvy/go-ai/pkg/providers/gateway" ) func main() { // Create provider with API key from environment provider, err := gateway.New(gateway.Config{ APIKey: os.Getenv("AI_GATEWAY_API_KEY"), }) if err != nil { log.Fatal(err) } _ = provider } ``` ### Advanced Configuration ```go provider, err := gateway.New(gateway.Config{ APIKey: "your-api-key-or-vercel-access-token", TeamIDOrSlug: "team-slug-or-id", BaseURL: "https://ai-gateway.vercel.sh/v4/ai", // Custom base URL Headers: map[string]string{ "X-Custom-Header": "value", }, ZeroDataRetention: true, // Enable zero data retention DisallowPromptTraining: true, QuotaEntityID: "tenant-123", MetadataCacheRefreshMillis: 300000, // 5 minutes }) ``` ## Text Generation ```go import ( "context" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/gateway" ) // Create provider provider, _ := gateway.New(gateway.Config{ APIKey: os.Getenv("AI_GATEWAY_API_KEY"), }) // Create language model // Gateway routes to the best available provider model, _ := provider.LanguageModel(string(gateway.GatewayLanguageModelAnthropicClaudeSonnet45)) // Generate text maxTokens := 200 result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing in simple terms.", MaxTokens: &maxTokens, }) ``` ### Model ID catalog `pkg/providers/gateway/model_ids.go` is generated from the upstream Gateway catalog (`go generate ./pkg/providers/gateway`) and gives compile-time constants such as `gateway.GatewayLanguageModelAnthropicClaudeSonnet45`, or you can pass the raw Gateway model ID string directly. > **Breaking change:** all `xai/*` model IDs were renamed to `spacexai/*` > upstream (`GatewayLanguageModelXai*` constants are now > `GatewayLanguageModelSpacexai*`, e.g. > `gateway.GatewayLanguageModelSpacexaiGrok420MultiAgent` / > `"spacexai/grok-4.20-multi-agent"`), and 44 language, 4 image, and 4 video > IDs that no longer exist upstream were removed from the generated file. ## Video Generation The Gateway provider supports video generation with automatic provider selection: ```go // Create video model videoModel, err := provider.VideoModel("google/veo-3.1-generate-001") if err != nil { log.Fatal(err) } // Generate video duration := 4.0 result, err := ai.GenerateVideo(context.Background(), ai.GenerateVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{ Text: "A cat playing with a ball of yarn", }, AspectRatio: "16:9", Duration: &duration, }) ``` ### Video Generation Features - **Text-to-Video**: Generate videos from text descriptions - **Image-to-Video**: Animate static images - **Batch Generation**: Generate multiple videos in one request - **Provider Selection**: Automatic routing to best video provider - **Custom Parameters**: Aspect ratio, duration, FPS, resolution ### Supported Video Providers The gateway can route video requests to: - Google Vertex AI (Veo) - FAL AI - Replicate - Other configured providers ### Video Generation Example ```go videoModel, _ := provider.VideoModel("google/veo-3.1-generate-001") duration := 4.0 fps := 30 result, err := ai.GenerateVideo(context.Background(), ai.GenerateVideoOptions{ Model: videoModel, Prompt: ai.VideoPrompt{ Text: "A serene ocean sunset with waves", }, AspectRatio: "16:9", Duration: &duration, FPS: &fps, ProviderOptions: map[string]interface{}{ "enhancePrompt": true, }, }) // Access generated videos for _, video := range result.Videos { fmt.Printf("Video URL: %s\n", video.URL) fmt.Printf("Media Type: %s\n", video.MediaType) } ``` ## Image Generation ```go imageModel, _ := provider.ImageModel(string(gateway.GatewayImageModelOpenaiGptImage2)) n := 1 result, err := ai.GenerateImage(context.Background(), ai.GenerateImageOptions{ Model: imageModel, Prompt: "A futuristic cityscape at sunset", N: &n, Size: "1024x1024", }) // The Gateway image model sets its own retryability signal (*bool) on a // failed/empty generation; ai.GenerateImage reads it internally through // the retry loop, so this does not appear as a field on the returned // GenerateImageResult. ``` ## Embeddings ```go embeddingModel, _ := provider.EmbeddingModel(string(gateway.GatewayEmbeddingModelOpenaiTextEmbedding3Small)) result, err := ai.EmbedMany(context.Background(), ai.EmbedManyOptions{ Model: embeddingModel, Inputs: []string{ "The cat sat on the mat", "A dog ran through the park", }, }) ``` ## Provider Options Customize provider selection and behavior: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello, world!", ProviderOptions: map[string]interface{}{ "gateway": map[string]interface{}{ "only": []string{"anthropic", "openai", "google"}, "tags": []string{"production", "fast"}, "order": []string{"openai", "anthropic"}, "sort": "cost", }, }, }) ``` ### Provider Option Fields - **gateway.only**: Restrict candidate providers. - **gateway.order**: Explicit provider order. - **gateway.sort**: Gateway sort strategy. - **gateway.tags**: Routing tags. - **gateway.models**: Ordered fallback model list. Build entries with `gateway.GatewayModel(id)` for a plain fallback, or `gateway.GatewayConditionalModelFallback(id, when)` for an evaluation-only conditional fallback (only the first entry may be conditional). JSON also accepts a plain model-ID string in place of either helper. - **gateway.has**: Restrict routing to models satisfying capability tags (`gateway.GatewayHasImplicitCaching`, `GatewayHasReasoning`, `GatewayHasStructuredOutput`, `GatewayHasToolUse`, `GatewayHasVision`, or a weight-format condition from `GatewayHasQuantization` / `GatewayHasNotQuantization`). - **gateway.zeroDataRetention**: Per-call zero retention. - **gateway.disallowPromptTraining**: Restrict to no-prompt-training providers. - **gateway.quotaEntityId**: Quota identity for tenant/account-level tracking. - **gateway.serviceTier**: Unified service tier intent. Supported values are `"flex"` and `"priority"`; Gateway translates this to the provider-specific tier and overrides lower-level provider tier options. > **Breaking change:** `gateway.hipaaCompliant` / `Config.HIPAACompliant` were > removed; the Gateway API no longer accepts this field. ## Reranking ```go import goprovider "github.com/digitallysavvy/go-ai/pkg/provider" reranker, _ := gatewayProvider.RerankingModel("cohere/rerank-v3.5") topN := 3 res, err := reranker.DoRerank(ctx, &goprovider.RerankOptions{ Query: "gateway routing", Documents: []string{ "Provider fallback order can reduce latency variance.", "Embedding vectors power semantic search.", "HIPAA compliance restricts provider choice.", }, TopN: &topN, }) ``` ## Timeout Handling The Gateway SDK provides enhanced timeout error handling with troubleshooting guidance. ### Typed Error Classification Gateway response errors are decoded into typed errors under `pkg/providers/gateway/errors`. The common `GatewayError` interface exposes status code, public type, generation ID, and retryability. Unknown Gateway error types are classified the same way as the TypeScript SDK while preserving the raw type through `GatewayErrorDetails`. ```go var gatewayErr gatewayerrors.GatewayError if errors.As(err, &gatewayErr) { fmt.Println(gatewayErr.GetType(), gatewayErr.GetStatusCode(), gatewayErr.IsRetryable()) } var details gatewayerrors.GatewayErrorDetails if errors.As(err, &details) { fmt.Println(details.GetRawType(), details.GetCode(), details.GetParam()) } ``` Inline `[]byte` file data in Gateway language-model requests is base64-encoded exactly once for file, reasoning-file, and tool-result file parts. URL, provider-reference, and text file data are preserved in their provider-native form. ## Realtime Runtime Client Secrets Gateway realtime is exposed through `provider.ExperimentalRealtime(modelID)`. Use `GetRealtimeToken` or `MintRealtimeClientSecret` on a server to create the short-lived client secret used by browser WebSocket clients. ```go gw, err := gateway.New(gateway.Config{ APIKey: os.Getenv("AI_GATEWAY_API_KEY"), TeamIDOrSlug: os.Getenv("VERCEL_TEAM_ID"), }) if err != nil { return err } expiresAfterSeconds := 600 secret, err := gw.GetRealtimeToken(ctx, provider.RealtimeFactoryGetTokenOptions{ Model: "openai/gpt-realtime", ClientSecretOptions: provider.ClientSecretOptions{ ExpiresAfterSeconds: &expiresAfterSeconds, }, }) if err != nil { return err } model := gw.ExperimentalRealtime("openai/gpt-realtime") ws := model.GetWebSocketConfig(secret.Token, secret.URL) fmt.Println(ws.URL, ws.Protocols) ``` `MintRealtimeClientSecret` sends `POST /v1/realtime/client-secrets` at the Gateway origin, not under `/v4/ai`. The realtime WebSocket protocols include `ai-gateway-realtime.v1`, `ai-gateway-auth.`, and, when configured, a base64url encoded `ai-gateway-team.` entry. Realtime event parsing, client event serialization, and session config building are identity mappings for provider-native realtime payloads. ### Setting Timeouts ```go import ( "context" "time" gatewayerrors "github.com/digitallysavvy/go-ai/pkg/providers/gateway/errors" ) // Text generation: 30-60 seconds ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() options.Model = model result, err := ai.GenerateText(ctx, options) // Video generation: 5-10 minutes ctx, cancel := context.WithTimeout(context.Background(), 10*time.Minute) defer cancel() options.Model = videoModel videoResult, err := ai.GenerateVideo(ctx, options) ``` ### Detecting Timeout Errors ```go options.Model = videoModel result, err := ai.GenerateVideo(ctx, options) if err != nil { if gatewayerrors.IsGatewayTimeoutError(err) { // This is a client-side timeout // Error message includes troubleshooting guidance fmt.Printf("Timeout error: %v\n", err) } else { // Other error fmt.Printf("Error: %v\n", err) } } ``` ### Timeout Best Practices **Text Generation:** - Simple prompts: 30-60 seconds - Complex generation: 2-5 minutes **Video Generation:** - Short videos (2-4s): 5-10 minutes - Long videos: 15-30 minutes - Batch generation: 20-40 minutes **Image Generation:** - Single image: 1-2 minutes - Batch images: 3-5 minutes ## Zero Data Retention Enable zero data retention mode to ensure requests are not logged: ```go provider, err := gateway.New(gateway.Config{ APIKey: os.Getenv("AI_GATEWAY_API_KEY"), ZeroDataRetention: true, }) ``` ## Observability The Gateway automatically includes observability headers when running on Vercel: ```go // Automatically includes: // - ai-o11y-deployment-id // - ai-o11y-environment // - ai-o11y-region // - ai-o11y-request-id ``` ## Available Models Query available models and providers: ```go metadata, err := provider.GetAvailableModels(context.Background()) if err != nil { log.Fatal(err) } for _, model := range metadata.Models { fmt.Printf("Provider: %s\n", model.Specification.Provider) fmt.Printf(" - %s (%s)\n", model.Name, model.ID) fmt.Printf(" Specification: %v\n", model.Specification) } ``` ## Credits Check your Gateway credit usage: ```go credits, err := provider.GetCredits(context.Background()) if err != nil { log.Fatal(err) } fmt.Printf("Balance: %s\n", credits.Balance) fmt.Printf("Total Used: %s\n", credits.TotalUsed) ``` `/v1/credits` reads the team from a `teamId`/`slug` query parameter rather than the `x-vercel-ai-gateway-team` header every other endpoint uses. When `Config.TeamIDOrSlug` (or a per-call `x-vercel-ai-gateway-team` header) is set, `GetCredits` forwards it as that query parameter too, so credentials that can access multiple teams (such as Vercel access tokens) are scoped correctly instead of getting a 401. ## Evaluation model fallbacks `EvaluationFallbackCondition.Question` is optional. Without it, `ConfidenceBelow` checks every Choice/Score question and `ProbabilityBetween` checks every Boolean question, instead of a single named question: ```go confidenceBelow := 0.6 when := gateway.EvaluationFallbackCondition{ConfidenceBelow: &confidenceBelow} fallback := gateway.GatewayConditionalModelFallback("openai/gpt-5.6-sol", when) ``` A condition still needs a check: `{}` or a condition with only `Question` set is rejected. ## Error Handling The Gateway provides detailed error types: ```go import gatewayerrors "github.com/digitallysavvy/go-ai/pkg/providers/gateway/errors" options.Model = model result, err := ai.GenerateText(ctx, options) if err != nil { switch { case gatewayerrors.IsGatewayTimeoutError(err): // Handle timeout with troubleshooting guidance fmt.Printf("Request timed out: %v\n", err) case providererrors.IsRateLimitError(err): // Handle rate limiting fmt.Printf("Rate limited: %v\n", err) case providererrors.IsProviderError(err): // Handle provider errors fmt.Printf("Provider error: %v\n", err) default: // Handle other errors fmt.Printf("Error: %v\n", err) } } ``` ## Configuration Reference ### Config ```go type Config struct { // APIKey is the AI Gateway API key or Vercel access token (required) // Can also be set via AI_GATEWAY_API_KEY environment variable APIKey string // TeamIDOrSlug scopes Vercel access-token requests to a team. TeamIDOrSlug string // BaseURL is the base URL for the AI Gateway API // Default: https://ai-gateway.vercel.sh/v4/ai BaseURL string // Headers are custom headers to include in requests Headers map[string]string // MetadataCacheRefreshMillis is how frequently to refresh metadata cache // Default: 300000 (5 minutes) MetadataCacheRefreshMillis int64 // HTTPClient is a custom HTTP client HTTPClient *http.Client // ZeroDataRetention enables zero data retention mode // When true, requests are not logged or retained ZeroDataRetention bool // ProjectID is forwarded as the "ai-o11y-project-id" header. Can also be // set via VERCEL_PROJECT_ID. ProjectID *string // DisallowPromptTraining filters routing to providers that do not train on prompts DisallowPromptTraining bool // QuotaEntityID identifies the entity against which quota is tracked QuotaEntityID string } ``` ## Workflow Serialization Gateway embedding and image models can cross a workflow boundary with `provider.SerializeEmbeddingModel` / `DeserializeEmbeddingModel` and `provider.SerializeImageModel` / `DeserializeImageModel` (language models could already be serialized with `providerutils.SerializeModel` / `DeserializeModel`). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; speech, transcription, and video models are not yet serializable. ## Limitations - Speech synthesis and transcription not directly supported (use specific providers) - Some provider-specific features may not be available through Gateway ## Examples - [Basic Gateway Usage](https://github.com/digitallysavvy/go-ai/tree/main/examples/providers/gateway/basic) - [Video Generation](https://github.com/digitallysavvy/go-ai/tree/main/examples/providers/gateway/video-generation) - [Timeout Handling](https://github.com/digitallysavvy/go-ai/tree/main/examples/providers/gateway/timeout-handling) - [Zero Data Retention](https://github.com/digitallysavvy/go-ai/tree/main/examples/providers/gateway/zero-retention) - [Parallel Search](https://github.com/digitallysavvy/go-ai/tree/main/examples/providers/gateway/parallel-search) ## Related Documentation - [Video Generation Guide](https://goaisdk.com/docs/ai-sdk-core/video-generation.md) - [Error Handling](https://goaisdk.com/docs/troubleshooting/common-errors.md) - [Provider Overview](https://goaisdk.com/docs/providers/overview.md) ## Learn More - [AI Gateway Documentation](https://vercel.com/docs/ai-gateway) - [Video Generation Guide](https://vercel.com/docs/ai-gateway/capabilities/video-generation) - [Gateway API Reference](https://vercel.com/docs/ai-gateway/api-reference) ## May 2026 parity updates ### Reranking Gateway supports reranking through `provider.RerankingModel` and the `/reranking-model` endpoint. ```go ranker, err := gatewayProvider.RerankingModel("cohere/rerank-v3.5") topN := 3 result, err := ranker.DoRerank(ctx, &provider.RerankOptions{ Query: "go concurrency", Documents: []string{"Goroutines are lightweight.", "SQL joins tables."}, TopN: &topN, }) ``` ### Routing, quota, and compliance `gateway.GatewayProviderOptions` includes `Sort`, `QuotaEntityID`, `DisallowPromptTraining`, `ZeroDataRetention`, `ServiceTier`, `Has`, provider order, model filters, and BYOK options. ```go zdr := true opts := gateway.GatewayProviderOptions{ Sort: "cost", QuotaEntityID: "tenant_123", ServiceTier: "priority", ZeroDataRetention: &zdr, Models: []gateway.GatewayModelFallback{ gateway.GatewayModel("openai/gpt-5.1"), gateway.GatewayModel("anthropic/claude-sonnet-4-6"), }, } ``` ### Unknown model types Gateway metadata parsing is resilient to unknown model types and recognizes `reranking` as a valid model type. --- # Vercel AI Provider > Setup and usage guide for the Vercel AI OpenAI-compatible endpoint with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/vercel Documentation index: https://goaisdk.com/llms.txt The Vercel provider (`pkg/providers/vercel`) is a thin wrapper around the `openai` package pointed at Vercel's OpenAI-compatible chat endpoint. It is **Go-only** — there is no equivalent package in the upstream TypeScript AI SDK. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/vercel" ) ``` ### Configuration ```go provider := vercel.New(vercel.Config{ APIKey: os.Getenv("VERCEL_API_KEY"), }) model, err := provider.LanguageModel("your-model-id") ``` `vercel.Config` accepts `APIKey` and `BaseURL` (default `https://api.vercel.com/v1`) — there is no environment-variable fallback built into `New()`, so pass the key explicitly. `LanguageModel` calls the Chat Completions endpoint (`/chat/completions`), not the OpenAI Responses API — OpenAI-compatible hosts generally don't implement Responses. ## Basic Usage ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain how Vercel's AI infrastructure works", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` Because `vercel.Provider` embeds `*openai.Provider`, embedding, image, speech, and transcription factory methods are also available and behave like the `openai` package's — see the [OpenAI provider](https://goaisdk.com/docs/providers/openai.md) for their full option surface. ## See Also - [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - [OpenAI Provider](https://goaisdk.com/docs/providers/openai.md) - this provider's underlying implementation --- # Gladia Provider > Setup and usage guide for Gladia speech-to-text transcription with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/gladia Documentation index: https://goaisdk.com/llms.txt Gladia provides speech-to-text transcription. The current Go provider implements the shared `provider.TranscriptionModel` interface and is used through `ai.Transcribe`. ## Setup ```go import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/gladia" ) ``` ```go provider := gladia.New(gladia.Config{ APIKey: os.Getenv("GLADIA_API_KEY"), }) model, err := provider.TranscriptionModel("") if err != nil { log.Fatal(err) } ``` Gladia has a single transcription model, not a selectable catalog — the model ID argument is kept for interface parity with other providers but is not sent in the request; pass an empty string. ## Transcribe Audio ```go audioData, err := os.ReadFile("meeting.mp3") if err != nil { log.Fatal(err) } result, err := ai.Transcribe(context.Background(), ai.TranscribeOptions{ Model: model, Audio: audioData, MimeType: "audio/mpeg", ProviderOptions: map[string]interface{}{ "gladia": map[string]interface{}{ "language": "en", }, }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) if result.DurationInSeconds != nil { fmt.Printf("Duration: %.2f seconds\n", *result.DurationInSeconds) } ``` ## Timestamps ```go result, err := ai.Transcribe(context.Background(), ai.TranscribeOptions{ Model: model, Audio: audioData, MimeType: "audio/mpeg", ProviderOptions: map[string]interface{}{ "gladia": map[string]interface{}{ "detectLanguage": true, }, }, }) if err != nil { log.Fatal(err) } for _, segment := range result.Segments { fmt.Printf("[%.2f-%.2f] %s\n", segment.StartSecond, segment.EndSecond, segment.Text) } ``` ## Workflow Serialization Gladia transcription models can cross a workflow boundary with `provider.SerializeTranscriptionModel` / `DeserializeTranscriptionModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: Transcribe](https://goaisdk.com/docs/reference/ai/transcribe.md) - [Transcription Guide](https://goaisdk.com/docs/ai-sdk-core/transcription.md) - [Gladia Documentation](https://www.gladia.io/docs) --- # QuiverAI Provider > Covers generating SVGs and vectorizing raster images with QuiverAI through the Go AI SDK image model interface, including provider options and setup. Canonical URL: https://goaisdk.com/docs/providers/quiverai Documentation index: https://goaisdk.com/llms.txt QuiverAI provides SVG generation and raster-to-SVG vectorization through the Go AI SDK image model interface. ## Setup ```go import "github.com/digitallysavvy/go-ai/pkg/providers/quiverai" qprovider := quiverai.New(quiverai.Config{ APIKey: "your-api-key", }) ``` `APIKey` defaults to `QUIVERAI_API_KEY`. `BaseURL` defaults to `https://api.quiver.ai/v1` and can be overridden with `QUIVERAI_BASE_URL`. ## SVG Generation ```go model, _ := qprovider.ImageModel(quiverai.ModelArrow11) result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{ Prompt: "minimal rocket logo", }) ``` Generated images are returned as SVG bytes, use `image/svg+xml` as the MIME type, and include token usage when QuiverAI reports it. ## Vectorization ```go result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{ Files: []provider.ImageFile{{ Type: "url", URL: "https://example.com/logo.png", }}, ProviderOptions: map[string]interface{}{ "quiverai": map[string]interface{}{ "operation": quiverai.OperationVectorize, "autoCrop": true, "targetSize": 512, }, }, }) ``` ## Provider Options | Option | Type | Description | | --- | --- | --- | | `operation` | `generate` or `vectorize` | Selects text-to-SVG generation or raster vectorization. | | `instructions` | string | Extra style guidance for generation. | | `temperature` | number | Sampling temperature. | | `topP` | number | Nucleus sampling value. | | `presencePenalty` | number | Presence penalty. | | `maxOutputTokens` | number | Maximum output tokens. | | `autoCrop` | bool | Vectorization-only auto-crop setting. | | `targetSize` | number | Vectorization-only target canvas size in pixels. | ## Language Models The Arrow 2 / Arrow 2 Telos models are available through the standard language model interface, reusing the Open Responses transport with QuiverAI's own request policy layered on top: ```go model, _ := qprovider.LanguageModel(quiverai.ModelArrow2) // or quiverai.ModelArrow2Telos result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Suggest a minimal color palette for a weather app icon set.", ProviderOptions: map[string]interface{}{ "quiverai": map[string]interface{}{ "reasoningEffort": "medium", // low | medium | high | xhigh "reasoningSummary": "auto", }, }, }) ``` QuiverAI also supports a caller-executed "custom" tool (`ProviderID: "quiverai.custom"`), whose input the model streams as raw text (e.g. SVG markup) instead of JSON function arguments. Build one with `qprovider.Tools().CustomTool(...)`: ```go import "github.com/digitallysavvy/go-ai/pkg/providers/openresponses" svgTool := qprovider.Tools().CustomTool("write_svg", openresponses.CustomToolOptions{ Description: "Return raw SVG markup for the requested icon.", Format: &openresponses.CustomToolFormat{Type: "text"}, }) ``` See [`examples/providers/quiverai`](https://github.com/digitallysavvy/go-ai/tree/main/examples/providers/quiverai) for full runnable examples. ## Workflow Serialization QuiverAI image models can cross a workflow boundary with `provider.SerializeImageModel` / `DeserializeImageModel` (Open Responses language models can too, via `providerutils.SerializeModel` / `DeserializeModel`, gated by `provider.SerializableModelStrict` and refused when extensions are registered). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. --- # LMNT Provider > Setup and usage guide for LMNT text-to-speech synthesis with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/lmnt Documentation index: https://goaisdk.com/llms.txt LMNT provides text-to-speech synthesis with voice selection and speed control. The current Go provider implements the shared `provider.SpeechModel` interface and is used through `ai.GenerateSpeech`. ## Setup ```go import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/lmnt" ) ``` ```go provider := lmnt.New(lmnt.Config{ APIKey: os.Getenv("LMNT_API_KEY"), }) model, err := provider.SpeechModel("default") if err != nil { log.Fatal(err) } ``` ## Generate Speech ```go speed := 1.2 result, err := ai.GenerateSpeech(context.Background(), ai.GenerateSpeechOptions{ Model: model, Text: "Welcome to the Go AI SDK.", Voice: "aurora", Speed: &speed, }) if err != nil { log.Fatal(err) } if err := os.WriteFile("output.mp3", result.Audio.Data, 0644); err != nil { log.Fatal(err) } fmt.Printf("Generated %d bytes as %s\n", len(result.Audio.Data), result.Audio.MediaType) ``` ## Voices Common voices include `aurora`, `lily`, `harper`, and `sage`. `Voice` defaults to `"ava"` when left empty. Check the [LMNT documentation](https://docs.lmnt.com/) for the complete voice list. ## See Also - [API Reference: GenerateSpeech](https://goaisdk.com/docs/reference/ai/generate-speech.md) - [Speech Generation Guide](https://goaisdk.com/docs/ai-sdk-core/speech.md) - [LMNT Documentation](https://docs.lmnt.com/) --- # Prodia Provider > Setup and usage guide for Prodia image and video generation with the Go AI SDK. Canonical URL: https://goaisdk.com/docs/providers/prodia Documentation index: https://goaisdk.com/llms.txt [Prodia](https://prodia.com/) is a fast inference platform for generative AI, offering high-speed image generation with FLUX and Stable Diffusion models, plus video generation with Wan 2.2 Lightning models. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/providers/prodia" ) ``` ### Configuration ```go provider := prodia.New(prodia.Config{ APIKey: os.Getenv("PRODIA_TOKEN"), }) ``` You can use the following optional settings to customize the Prodia provider instance: - **APIKey** `string` — API key sent as a Bearer token in the `Authorization` header. Defaults to the `PRODIA_TOKEN` environment variable. - **BaseURL** `string` — Custom URL prefix for API calls. Defaults to `https://inference.prodia.com/v2`. - **Headers** `map[string]string` — Custom headers to include in all requests. ### Get API keys 1. Sign up at [prodia.com](https://prodia.com) 2. Navigate to your API settings 3. Generate an API token 4. Set the environment variable: ```bash export PRODIA_TOKEN=your-api-key ``` ## Image model Prodia's image model handles text-to-image and image-to-image generation. You create an image model using the `ImageModel()` factory method. ### Basic usage ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/prodia" ) func main() { prov := prodia.New(prodia.Config{}) model, err := prov.ImageModel("inference.flux-fast.schnell.txt2img.v2") if err != nil { log.Fatal(err) } result, err := model.DoGenerate(context.Background(), &provider.ImageGenerateOptions{ Prompt: "A cat wearing an intricate robe", Size: "1024x1024", }) if err != nil { log.Fatal(err) } fmt.Printf("Generated image: %d bytes, type: %s\n", len(result.Image), result.MimeType) } ``` ### Model capabilities | Model ID | Description | |----------|-------------| | `inference.flux-fast.schnell.txt2img.v2` | Fast FLUX Schnell model for text-to-image generation | | `inference.flux.schnell.txt2img.v2` | FLUX Schnell model for text-to-image generation | > **Note:** Prodia supports additional model IDs. Check the [Prodia documentation](https://docs.prodia.com/) for the full list of available models. ## Language Model Prodia's language model handles img2img generation via multipart form-data. It accepts a text prompt and an optional input image and returns a transformed image along with an optional text description. Because Prodia does not expose SSE for this endpoint, `DoStream` wraps `DoGenerate` and emits the result as a synchronous chunk sequence. ### Model IDs | Model ID | Description | |----------|-------------| | `inference.nano-banana.img2img.v2` | Nano Banana image-to-image (default) | ### DoGenerate Example ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/prodia" ) func main() { prov := prodia.New(prodia.Config{}) model, err := prov.LanguageModel("inference.nano-banana.img2img.v2") if err != nil { log.Fatal(err) } imageBytes, err := os.ReadFile("input.png") if err != nil { log.Fatal(err) } result, err := model.DoGenerate(context.Background(), &provider.GenerateOptions{ Prompt: types.Prompt{ Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Transform this image into a watercolor painting"}, types.ImageContent{Image: imageBytes, MimeType: "image/png"}, }, }, }, }, ProviderOptions: map[string]interface{}{ "prodia": map[string]interface{}{ "aspectRatio": "16:9", }, }, }) if err != nil { log.Fatal(err) } fmt.Printf("Finish reason: %s\n", result.FinishReason) fmt.Printf("Content parts: %d\n", len(result.Content)) } ``` The language model does not support temperature, topP, topK, maxTokens, stopSequences, tools, toolChoice, responseFormat, or reasoning. Unsupported options emit warnings but do not cause failures. ## Video Models Prodia supports text-to-video (T2V) and image-to-video (I2V) generation with Wan 2.2 Lightning models. Text-to-video requests are sent as JSON; image-to-video requests are sent as multipart/form-data. ### Model IDs | Model ID | Type | Description | |----------|------|-------------| | `inference.wan2-2.lightning.txt2vid.v0` | T2V | Wan 2.2 Lightning text-to-video (default) | | `inference.wan2-2.lightning.img2vid.v0` | I2V | Wan 2.2 Lightning image-to-video | ### Text-to-Video Example ```go model, err := prov.VideoModel("inference.wan2-2.lightning.txt2vid.v0") if err != nil { log.Fatal(err) } result, err := model.DoGenerate(context.Background(), &provider.VideoModelV3CallOptions{ Prompt: "A sunrise over mountains in 4K", AspectRatio: "16:9", }) if err != nil { log.Fatal(err) } fmt.Printf("Video MIME: %s\n", result.Videos[0].MediaType) fmt.Printf("Video bytes: %d\n", len(result.Videos[0].Binary)) ``` ### Image-to-Video Example ```go model, err := prov.VideoModel("inference.wan2-2.lightning.img2vid.v0") if err != nil { log.Fatal(err) } result, err := model.DoGenerate(context.Background(), &provider.VideoModelV3CallOptions{ Prompt: "The cat slowly turns its head", AspectRatio: "9:16", Image: &provider.VideoModelV3File{ Type: "file", Data: imageBytes, MediaType: "image/png", }, }) if err != nil { log.Fatal(err) } fmt.Printf("Video bytes: %d\n", len(result.Videos[0].Binary)) ``` Max videos per call is 1. ## Image size You can specify the image size using the `Size` field in `WIDTHxHEIGHT` format. Width and height must be between 256 and 1920 pixels. ```go result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{ Prompt: "A serene mountain landscape at sunset", Size: "1024x768", }) ``` ## Provider options Prodia image models support additional options through the `ProviderOptions` map under the `"prodia"` key: ```go steps := 4 stylePreset := "cinematic" result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{ Prompt: "A cat wearing an intricate robe", Size: "1024x768", ProviderOptions: map[string]interface{}{ "prodia": map[string]interface{}{ "steps": steps, "stylePreset": stylePreset, }, }, }) ``` ### Image model options | Option | Type | Description | |--------|------|-------------| | `steps` | `int` | Number of computational iterations (1-4). More steps produce higher quality results. | | `width` | `int` | Output width in pixels (256-1920). Overrides width from `Size`. | | `height` | `int` | Output height in pixels (256-1920). Overrides height from `Size`. | | `stylePreset` | `string` | Visual theme for the output image (see style presets below). | | `loras` | `[]string` | Up to 3 LoRA model references for output augmentation. | | `progressive` | `bool` | Return a progressive JPEG when using JPEG output. | ### Style presets Supported values: `3d-model`, `analog-film`, `anime`, `cinematic`, `comic-book`, `digital-art`, `enhance`, `fantasy-art`, `isometric`, `line-art`, `low-poly`, `neon-punk`, `origami`, `photographic`, `pixel-art`, `texture`, `craft-clay`. ### Language and video model options | Option | Type | Applies to | Description | |--------|------|------------|-------------| | `aspectRatio` | `string` | Language, Video | One of: `1:1`, `2:3`, `3:2`, `4:5`, `5:4`, `4:7`, `7:4`, `9:16`, `16:9`, `9:21`, `21:9` | | `resolution` | `string` | Video only | Output resolution (e.g. `"480p"`, `"720p"`) | ## Seed You can use the `Seed` parameter to get reproducible results: ```go seed := 12345 result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{ Prompt: "A serene mountain landscape at sunset", Seed: &seed, }) ``` The seed used for generation is available in the provider metadata (see below), which is useful when you do not set a specific seed and want to reproduce a result later. ## Provider metadata The response includes provider-specific metadata in `ProviderMetadata["prodia"]`. The nesting depends on the model type: - **Image models**: metadata is nested under `"images"` as a slice — `ProviderMetadata["prodia"]["images"][0]` - **Video models**: metadata is nested under `"videos"` as a slice — `ProviderMetadata["prodia"]["videos"][0]` - **Language models**: metadata is flat — `ProviderMetadata["prodia"]` directly contains the fields Each metadata object may contain: | Field | Type | Description | |-------|------|-------------| | `jobId` | `string` | Unique identifier for the generation job | | `seed` | `float64` | Seed used for generation (useful for reproducing results) | | `elapsed` | `float64` | Generation time in seconds | | `iterationsPerSecond` | `float64` | Processing speed metric | | `createdAt` | `string` | Timestamp when the job was created | | `updatedAt` | `string` | Timestamp when the job was last updated | | `dollars` | `float64` | Cost of the generation in USD | ```go result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{ Prompt: "A serene mountain landscape at sunset", }) if err != nil { log.Fatal(err) } // Access provider metadata if prodiaMeta, ok := result.ProviderMetadata["prodia"].(map[string]interface{}); ok { if images, ok := prodiaMeta["images"].([]interface{}); ok && len(images) > 0 { if img, ok := images[0].(map[string]interface{}); ok { fmt.Println("Job ID:", img["jobId"]) fmt.Println("Seed:", img["seed"]) fmt.Println("Elapsed:", img["elapsed"]) } } } ``` ## Workflow Serialization Prodia image models can cross a workflow boundary with `provider.SerializeImageModel` / `DeserializeImageModel` (language models could already be serialized with `providerutils.SerializeModel` / `DeserializeModel`). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; video models are not yet serializable. ## See also - [Provider options](https://goaisdk.com/docs/foundations/provider-options.md) — how to pass provider-specific configuration - [KlingAI provider](https://goaisdk.com/docs/providers/klingai.md) — another video generation provider - [Prodia documentation](https://docs.prodia.com/) --- # Voyage Provider > Setup guide for Voyage embedding and reranking models in the Go AI SDK, covering model IDs, retrieval-optimized usage, and May 2026 parity updates. Canonical URL: https://goaisdk.com/docs/providers/voyage Documentation index: https://goaisdk.com/llms.txt Voyage provides high-quality embedding and reranking models optimized for retrieval workflows. ## Setup ```go import ( goprovider "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/voyage" ) voyageProvider := voyage.New(voyage.Config{ APIKey: os.Getenv("VOYAGE_API_KEY"), }) ``` ## Embeddings ```go embeddingModel, _ := voyageProvider.EmbeddingModel(voyage.ModelVoyage3Large) result, err := embeddingModel.DoEmbedMany(ctx, []string{ "Go is a statically typed language", "Vector search uses embeddings", }, nil) ``` ### Embedding Options Voyage-specific options are passed under the `voyage` provider-options key and serialized to the Voyage API field names. ```go result, err := embeddingModel.DoEmbedMany(ctx, []string{"search text"}, &goprovider.EmbedModelOptions{ ProviderOptions: map[string]interface{}{ "voyage": map[string]interface{}{ "inputType": "query", "truncation": true, "outputDimension": 1024, "outputDtype": "float", }, }, }) ``` ## Reranking ```go import ( goprovider "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/voyage" ) voyageProvider := voyage.New(voyage.Config{ APIKey: os.Getenv("VOYAGE_API_KEY"), }) rerankModel, _ := voyageProvider.RerankingModel(voyage.ModelRerank25) topN := 3 result, err := rerankModel.DoRerank(ctx, &goprovider.RerankOptions{ Query: "go concurrency", Documents: []string{ "Goroutines are lightweight threads.", "SQL joins combine rows from tables.", "Channels coordinate goroutines.", }, TopN: &topN, }) ``` ### Reranking Options ```go result, err := rerankModel.DoRerank(ctx, &goprovider.RerankOptions{ Query: "go concurrency", Documents: []string{"Goroutines are lightweight.", "SQL joins tables."}, TopN: &topN, ProviderOptions: map[string]interface{}{ "voyage": map[string]interface{}{ "returnDocuments": true, "truncation": true, }, }, }) ``` ## Model IDs ### Embeddings - `voyage.ModelVoyage4Large` (`voyage-4-large`) - `voyage.ModelVoyage4` (`voyage-4`) - `voyage.ModelVoyage4Lite` (`voyage-4-lite`) - `voyage.ModelVoyage4Nano` (`voyage-4-nano`) - `voyage.ModelVoyageCode35` (`voyage-code-3.5`) - `voyage.ModelVoyageCode3` (`voyage-code-3`) - `voyage.ModelVoyage3Large` (`voyage-3-large`) - `voyage.ModelVoyage35` (`voyage-3.5`) - `voyage.ModelVoyage35Lite` (`voyage-3.5-lite`) - `voyage.ModelVoyage3` (`voyage-3`) - `voyage.ModelVoyage3Lite` (`voyage-3-lite`) - `voyage.ModelVoyageFinance2` (`voyage-finance-2`) - `voyage.ModelVoyageLaw2` (`voyage-law-2`) - `voyage.ModelVoyageMultilingual2` (`voyage-multilingual-2`) - `voyage.ModelVoyageCode2` (`voyage-code-2`) - `voyage.ModelVoyage2` (`voyage-2`) ### Reranking - `voyage.ModelRerank25` (`rerank-2.5`) - `voyage.ModelRerank25Lite` (`rerank-2.5-lite`) - `voyage.ModelRerank2` (`rerank-2`) - `voyage.ModelRerank2Lite` (`rerank-2-lite`) - `voyage.ModelRerank1` (`rerank-1`) - `voyage.ModelRerankLite1` (`rerank-lite-1`) ## Notes - Default base URL: `https://api.voyageai.com/v1` - Auth header: `Authorization: Bearer $VOYAGE_API_KEY` - Embedding endpoint: `POST /embeddings` - Reranking endpoint: `POST /rerank` - Provider identities match the TypeScript SDK: `voyage.embedding` and `voyage.reranking` - Embedding requests are limited to 128 inputs per call to match the TypeScript SDK. ## Workflow Serialization Voyage embedding models can cross a workflow boundary with `provider.SerializeEmbeddingModel` / `DeserializeEmbeddingModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism; reranking models are not yet serializable. ## May 2026 parity updates Voyage is a full provider as of v0.5.0, with embedding and reranking support. Provider-specific options live under the `voyage` provider-options key. Embedding options include `inputType`, `truncation`, `outputDimension`, and `outputDtype`; reranking options include `returnDocuments` and `truncation`. --- # TypeSafe AI Provider > Setup and usage guide for TypeSafe AI evaluation models with Go AI SDK Canonical URL: https://goaisdk.com/docs/providers/typesafeai Documentation index: https://goaisdk.com/llms.txt TypeSafe AI provides evaluation models for use with the experimental `ai.ExperimentalEvaluate` API — judging free-form output against a set of questions (choice, score, or boolean) rather than generating text. It's the only capability this provider implements: `LanguageModel`, `EmbeddingModel`, `ImageModel`, `SpeechModel`, and `TranscriptionModel` all return errors. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" goprovider "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/typesafeai" ) ``` ### Configuration ```go provider := typesafeai.New(typesafeai.Config{ APIKey: os.Getenv("TYPESAFE_AI_API_KEY"), }) ``` `typesafeai.Config` also accepts `BaseURL` (default `https://api.typesafe.ai/v1`), `Headers`, and `HTTPClient`. If `APIKey` is left empty, it is read from `TYPESAFE_AI_API_KEY` lazily on each request (so setting the environment variable after `New()` still works). ### Get API Key ```bash export TYPESAFE_AI_API_KEY=... ``` ## Evaluation ```go evalModel, err := provider.EvaluationModel("jev-latest") if err != nil { log.Fatal(err) } result, err := ai.ExperimentalEvaluate(ctx, ai.EvaluateOptions{ Model: evalModel, State: "The capital of France is Paris.", Questions: map[string]goprovider.EvaluationQuestion{ "factual": { Type: "boolean", Instructions: "Is this statement factually correct?", }, "helpfulness": { Type: "score", Instructions: "How helpful is this response?", Criteria: []interface{}{"poor", "fair", "good", "excellent"}, }, }, }) if err != nil { log.Fatal(err) } for id, answer := range result.Answers { fmt.Printf("%s: %v\n", id, answer) } ``` `ai.ExperimentalEvaluate` also accepts any `provider.EvaluationModel` instance (or a model ID string, resolved via `Provider`, falling back to the AI Gateway when `Provider` is nil) — you are not limited to TypeSafe AI as the judge model. ## Workflow Serialization TypeSafe AI evaluation models can cross a workflow boundary with `provider.SerializeEvaluationModel` / `DeserializeEvaluationModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [TypeSafe AI Documentation](https://typesafe.ai) --- # OpenAI-compatible providers > Configure OpenAI-compatible endpoints with Go AI SDK provider options and warnings. Canonical URL: https://goaisdk.com/docs/providers/openai-compatible Documentation index: https://goaisdk.com/llms.txt Go does not expose a separate `openai-compatible` provider package. OpenAI-compatible request behavior is shared by provider implementations such as Groq, Together, Moonshot, Alibaba, and DeepSeek; use each provider's package and provider-specific option key. ## Setup Configure the provider package that matches the target API. Providers that use the OpenAI-compatible option resolver accept camelCase provider options and keep deprecated fallback keys only for compatibility warnings. ## May 2026 parity updates Provider options should use camelCase keys to match the TypeScript SDK. Prefer the concrete provider key, for example `"deepseek"` or `"groq"`, rather than a generic `"openai-compatible"` key. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a short status update.", ProviderOptions: map[string]interface{}{ "deepseek": map[string]interface{}{ "reasoningEffort": "low", }, }, }) ``` ## Assistant Content Serialization Assistant messages with tool calls no longer use one shape for every OpenAI-compatible provider — each matches its own API's expectations: - **Groq** sends the text content verbatim, including an empty string (`""`), never `null`. - **Cohere** omits the `content` key entirely on a tool-call assistant turn. - Other OpenAI-compatible providers keep the previous `null`-content shape for tool-call turns. ## Video Parts `video/*` file parts can now be sent as OpenAI-compatible `video_url` content parts. This is controlled internally by `prompt.ToOpenAIMessagesOptions.AllowVideo` (default `false` in the plain `openai` package, since not every OpenAI-compatible API accepts video content): Fireworks, GMI Cloud, Z.AI, Together, Baseten, Cerebras, and DeepInfra all set it to `true` for their own requests, so passing a `video/*` `types.FileContent` to those providers just works — ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, // a Fireworks, GMI Cloud, Z.AI, Together, Baseten, Cerebras, or DeepInfra model Messages: []types.Message{{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What happens in this clip?"}, types.FileContent{MediaType: "video/mp4", URL: "https://example.com/clip.mp4"}, }, }}, }) ``` — there is no per-call or per-provider `AllowVideo` setting to opt into; the plain `openai` package used directly against a custom `BaseURL` still omits video parts unless the endpoint is one of the providers above. ## Multipart Tool Result Content By default, a tool result's structured `content` output (text and file blocks, e.g. from `toModelOutput`) is sent as a single JSON-stringified `content` field, since not every OpenAI-compatible endpoint accepts structured tool-result content. Endpoints documented to extend Chat Completions with structured tool-result content can opt in internally via `prompt.ToOpenAIMessagesOptions.SupportsMultiPartToolContent`, which instead sends an array of content parts (text/image_url/video_url/input_audio/file) so a multimodal model can see an image a tool returned, rather than its JSON-stringified representation: ```go toolMessages := prompt.ToOpenAIMessages(messages, prompt.ToOpenAIMessagesOptions{ SupportsMultiPartToolContent: true, }) ``` There is no per-call setting; a provider implementation sets this only for an endpoint it has confirmed accepts the structured shape. ## Streams Truncated Without A Finish Reason If the underlying SSE stream ends without ever sending a `finish_reason` (a truncated or misbehaving upstream), the shared `OpenAICompatStream` base (used by Fireworks, Together, Mistral, Ollama, XAI, Azure, Perplexity, and others) now emits an `InvalidResponseDataError` error chunk ("Response stream ended without a finish reason.") followed by a finish chunk with reason `"error"`, instead of ending the stream silently as in v0.4.0. ## Trailing Usage Chunks Streamed responses on providers built on the shared base now report token usage: the trailing `stream_options.include_usage` chunk that v0.4.0 dropped is applied to the final usage totals. ## Unreliable Streamed Tool-Call Labels Some OpenAI-compatible gateways omit, blank, repeat, or change streamed tool-call `id` and `index` values mid-stream, and may strip the `type` field from individual deltas. The shared tool-call tracker (used by OpenAI, OpenAI-compatible providers, Groq, DeepSeek, Alibaba, Mistral, and MoonshotAI) now correlates fragments using the wire ID, index, function name, incremental argument structure, and explicit-start evidence together, rather than treating any single label as authoritative: - Tool calls that reuse the same `index` (or the same `id`, or both) are kept distinct once another signal — a new function name, or a fresh JSON object/array start — shows the provider has moved on to a new call. - A continuation fragment with a conflicting or unattributable label (no matching `id`/`index`, and more than one call it could belong to) is dropped instead of being merged into the wrong call. - A tool call whose function name only arrives on a later, correlated delta (rather than on the delta that first introduces its `id`/`index`) is still reported correctly, with its buffered arguments included. This requires no code changes — it is a correctness fix to the existing `tool-input-start` / `tool-input-delta` / `tool-input-end` / `tool-call` chunk sequence for well-formed streams. A tool call whose function name never arrives by the end of the stream still surfaces as an `InvalidResponseDataError` error chunk, exactly as before. --- # OpenResponses Provider > Explains the OpenResponses provider for Responses-compatible APIs like LMStudio, Ollama, and LocalAI, covering custom tools and streaming text parts. Canonical URL: https://goaisdk.com/docs/providers/openresponses Documentation index: https://goaisdk.com/llms.txt The OpenResponses provider targets APIs that implement the OpenAI Responses wire shape (self-hosted servers like LMStudio, Ollama, or LocalAI, or any other `/responses`-compatible endpoint). Use it when a provider exposes `/responses` semantics instead of chat completions. It only supports language models — `EmbeddingModel`, `ImageModel`, `SpeechModel`, `TranscriptionModel`, and `RerankingModel` all return errors. [QuiverAI](https://goaisdk.com/docs/providers/quiverai.md) is built on this same transport. ## Setup ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openresponses" ) provider := openresponses.New(openresponses.Config{ BaseURL: "http://localhost:1234/v1", // e.g. LMStudio APIKey: os.Getenv("OPENRESPONSES_API_KEY"), // optional; many local servers need none }) model, err := provider.LanguageModel("responses-model") if err != nil { log.Fatal(err) } ``` `openresponses.Config` also accepts `Headers`, `Name` (overrides the provider identity, default `"open-responses"`), and `StrictResponseInput` (controls how assistant history without a known item ID is replayed). ## Custom Tools `provider.Tools().CustomTool(name, options)` builds a caller-executed "custom" tool whose input the model streams as raw text (not JSON function arguments) — useful for tools like `apply_patch` or free-form code/markup generation. `custom_tool_call` / `custom_tool_call_output` items are replayed correctly in message history, and the default custom tool ID is `"open-responses.custom"` (`Config.CustomToolID` overrides it; QuiverAI overrides it to `"quiverai.custom"`). ```go patchTool := provider.Tools().CustomTool("apply_patch", openresponses.CustomToolOptions{ Description: "Apply a unified diff patch to the workspace.", Format: &openresponses.CustomToolFormat{Type: "text"}, }) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Fix the off-by-one bug in main.go.", Tools: []types.Tool{patchTool}, }) ``` ## Streaming And Text Parts Streaming emits `text-start` / `text-end` chunk boundaries around text content, and text parts carry the Responses `itemId` and any annotations in their provider metadata. ## Extensions `openresponses.Config.Extensions` lets a provider that embeds the Open Responses spec plug in its own tool/item/event codecs via `openresponses.Extension`, without the core language model needing to know about them ahead of time. By default a codec's wire types must be namespaced (`":"`, e.g. `"lmstudio:code_execution"`), which keeps extensions portable across implementations. For an implementation that documents a bare (non-namespaced) wire discriminator and can't be changed, set `AllowBareTypes: true` and use `BareToolType`/`BareItemTypes`/`BareEventTypes` instead of `ToolType`/`ItemTypes`/`EventTypes`. A bare type must be non-empty, must not contain a colon, and must not collide with a core Responses wire type (e.g. `"function"`, `"message"`, `"response.completed"`); registering one is an explicit, non-portable opt-in. ```go p := openresponses.New(openresponses.Config{ BaseURL: "https://api.example.com/v1", Extensions: []openresponses.Extension{{ ID: "acme.legacy_tool", AllowBareTypes: true, BareToolType: "legacy_tool", EncodeTool: func(name string, args map[string]interface{}) (map[string]interface{}, error) { return args, nil }, }}, }) ``` ## Workflow Serialization OpenResponses language models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel`, gated by `provider.SerializableModelStrict` — a model is refused for serialization when extensions are registered on it (since extension codecs can't be reconstructed from a serialized config alone). See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the general mechanism. ## May 2026 parity updates ### Reasoning summary OpenResponses supports the `reasoningSummary` provider option. The provider serializes it to the Responses `reasoning.summary` request field for APIs that accept `auto`, `concise`, or `detailed` summaries. ```go p := openresponses.New(openresponses.Config{ APIKey: os.Getenv("OPENRESPONSES_API_KEY"), BaseURL: "https://api.example.com/v1", }) model, err := p.LanguageModel("responses-model") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Summarize the reasoning at a high level.", ProviderOptions: map[string]interface{}{ "openresponses": map[string]interface{}{ "reasoningSummary": "concise", }, }, }) ``` ### System messages System messages are converted to the Responses instructions shape by the provider. At the core SDK layer, system-role messages inside `Messages` are rejected by default; use the top-level `System` field or opt in with `AllowSystemMessages`. --- # Topaz Provider > Setup and usage guide for Topaz Labs image and video enhancement with the Go AI SDK. Canonical URL: https://goaisdk.com/docs/providers/topaz Documentation index: https://goaisdk.com/llms.txt [Topaz Labs](https://www.topazlabs.com) enhances media the caller already has — upscale, denoise, sharpen, restore — rather than generating from a text prompt. `topaz.image("wonder-3.5")` enhances a single input image; `topaz.video("proteus")` and `topaz.video("starlight-precise-2.6")` enhance a single input video through Topaz's async operation protocol (`DoStart`/`DoStatus`). Topaz has no language, embedding, speech, transcription, or reranking model. ## Setup ### Installation ```go import ( "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/topaz" ) ``` ### Configuration ```go prov := topaz.New(topaz.Config{ APIKey: os.Getenv("TOPAZ_API_KEY"), }) model, err := prov.ImageModel(topaz.ImageModelWonder35) ``` `topaz.Config` accepts: | Field | Type | Description | |-------|------|-------------| | `APIKey` | `string` | Topaz Labs API key. Falls back to the `TOPAZ_API_KEY` environment variable when empty. | | `BaseURL` | `string` | Base URL for the Topaz API. Defaults to `https://api.topazlabs.com`. | | `Headers` | `map[string]string` | Custom headers merged into every request. | | `VideoPollIntervalMillis` | `int` | Interval between status checks used by `VideoModel.DoGenerate`'s synchronous start+poll wrapper. Defaults to 2000ms. Only relevant for a direct `DoGenerate` call — `ai.GenerateVideo` always prefers the `DoStart`/`DoStatus` pair below. | | `VideoPollTimeoutMillis` | `int` | Total timeout for that same wrapper. Defaults to 600000ms (10 minutes). | `topaz.CreateTopaz(cfg)` is an alias for `topaz.New(cfg)`, mirroring the TypeScript SDK's `createTopaz` export. ### Get an API key ```bash export TOPAZ_API_KEY=... ``` See [Topaz's API key setup guide](https://developer.topazlabs.com/getting-started/api-key-setup). ## Image model: Wonder 3.5 `topaz.ImageModelWonder35` (`"wonder-3.5"`) submits an async enhance job, polls it, and downloads the result. Pass the input image through `Files` (exactly one image is processed per call — `MaxImagesPerCall() == 1`); a `url`-type file is forwarded to Topaz as `source_url` (Topaz fetches it itself), and a `file`-type file is uploaded as multipart form data. ```go package main import ( "context" "log" "os" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/topaz" ) func main() { prov := topaz.New(topaz.Config{APIKey: os.Getenv("TOPAZ_API_KEY")}) model, err := prov.ImageModel(topaz.ImageModelWonder35) if err != nil { log.Fatal(err) } result, err := model.DoGenerate(context.Background(), &provider.ImageGenerateOptions{ Files: []provider.ImageFile{ {Type: "url", URL: "https://example.com/input.png"}, }, Size: "4000x3000", }) if err != nil { log.Fatal(err) } os.WriteFile("enhanced.png", result.Image, 0o644) } ``` Unsupported call options (`Prompt`, `AspectRatio`, `Seed`, `Mask`, and `N` above 1) surface as `result.Warnings` rather than errors — Wonder 3.5 enhances one existing image per call and does not take a text prompt. ### Image model options Pass these under `ProviderOptions["topaz"]` as a `topaz.ImageModelOptions` value (or an equivalent `map[string]interface{}`). Documented request fields are sent to Topaz in snake_case; model-specific settings are sent in camelCase, matching the Topaz API itself — the Go field names and JSON tags below are camelCase either way, since that's the user-facing option surface. | Field | Type | Description | |-------|------|-------------| | `EnhancementStrength` | `*string` | `"low"`, `"medium"`, or `"high"` (default `"high"`). | | `Grain` | `*bool` | Add grain to the output (default `false`). | | `GrainDensity` | `*float64` | Grain intensity, 0.0–1.0 (default 0.5). | | `GrainModel` | `*string` | `"silver"`, `"gaussian"`, or `"grey"` (default `"silver"`). | | `GrainSize` | `*float64` | Grain particle size, 1–5 (default 1). | | `GrainStrength` | `*float64` | Grain effect strength, 0.0–1.0 (default 0.5). | | `InputWidth` / `InputHeight` | `*int` | Input image dimensions. Topaz infers these from the upload when omitted. | | `OutputWidth` / `OutputHeight` | `*int` | Output dimensions, 1–32000. Take precedence over the `Size` call option. | | `OutputFormat` | `*string` | `"jpeg"`, `"jpg"`, `"png"`, `"tiff"`, or `"tif"`. | | `CropToFill` | `*bool` | Crop the output to fill the requested dimensions (default `false`). | | `WebhookURL` | `*string` | URL to receive job-status webhooks. | | `PollIntervalMillis` | `*int` | Status-poll interval, in milliseconds (default 2000). | | `PollTimeoutMillis` | `*int` | Total poll timeout, in milliseconds, before the job is canceled and the call fails (default 600000, 10 minutes). | ```go strength := "medium" result, err := model.DoGenerate(context.Background(), &provider.ImageGenerateOptions{ Files: []provider.ImageFile{{Type: "url", URL: "https://example.com/input.png"}}, ProviderOptions: map[string]interface{}{ "topaz": topaz.ImageModelOptions{ EnhancementStrength: &strength, }, }, }) ``` ### Image provider metadata `result.ProviderMetadata["topaz"]` holds `{"images": [...]}`, with one entry (`processId`, and when Topaz reports them, `credits`, `width`, `height`, `format`) for the single enhanced image — `credits` is the number of credits Topaz billed for the completed job. ## Video models: Proteus and Starlight Precise 2.6 `topaz.VideoModelProteus` (`"proteus"`) and `topaz.VideoModelStarlightPrecise26` (`"starlight-precise-2.6"`) implement Topaz's express async operation protocol through `DoStart` and `DoStatus` (`MaxVideosPerCall() == 1`). Pass the input video through `InputReferences`; a `url`-type reference is forwarded to Topaz as `source.external` (Topaz fetches it itself), and a `file`-type reference is uploaded to the URL Topaz returns from `DoStart`. `Resolution` (or the `Output` provider option) sets the output resolution, which Topaz requires one way or another. ```go package main import ( "context" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/topaz" ) func main() { prov := topaz.New(topaz.Config{APIKey: os.Getenv("TOPAZ_API_KEY")}) model, err := prov.VideoModel(topaz.VideoModelProteus) if err != nil { log.Fatal(err) } result, err := ai.GenerateVideo(context.Background(), ai.GenerateVideoOptions{ Model: model, InputReferences: []ai.VideoReferenceInput{ {Data: ai.VideoPromptImage{URL: "https://example.com/input.mp4", MediaType: "video/mp4"}}, }, Resolution: "1920x1080", }) if err != nil { log.Fatal(err) } for _, video := range result.Videos { log.Println(video.URL) } } ``` `ai.GenerateVideo` uses `DoStart`/`DoStatus` directly (Topaz implements `provider.VideoModelStarter` and `provider.VideoModelStatusChecker`), so it also works with `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus` for a decoupled start-then-poll flow. `VideoModel.DoGenerate` exists only because Go's `provider.VideoModelV3` interface requires it directly (the TypeScript provider has no `doGenerate` of its own); calling it submits the request and polls to completion in one call, governed by `Config.VideoPollIntervalMillis` / `VideoPollTimeoutMillis`. Unsupported call options (`Prompt` when non-empty, `AspectRatio`, `Seed`, `Duration`, `GenerateAudio`, `FrameImages`, and `N` above 1) surface as warnings rather than errors. ### Source metadata Starlight models require source metadata (`width`, `height`, `duration`, `frameRate` — Topaz prices them from these values); other models accept it optionally, where it lets Topaz estimate the cost up front. The AI SDK never inspects the video bytes, so these values must come from the caller. Setting any one of them requires all four: ```go width, height, duration, frameRate := 1920, 1080, 10.0, 30.0 startResult, err := ai.ExperimentalStartVideo(context.Background(), ai.StartVideoOptions{ Model: model, InputReferences: []ai.VideoReferenceInput{ {Data: ai.VideoPromptImage{Data: videoBytes, MediaType: "video/mp4"}}, }, Resolution: "1920x1080", ProviderOptions: map[string]interface{}{ "topaz": topaz.VideoModelOptions{ Source: &topaz.VideoSource{ Width: &width, Height: &height, Duration: &duration, FrameRate: &frameRate, }, }, }, }) ``` ### Video model options Pass these under `ProviderOptions["topaz"]` as a `topaz.VideoModelOptions` value. | Field | Type | Description | |-------|------|-------------| | `Source` | `*topaz.VideoSource` | Input video metadata (`Width`, `Height`, `Duration`, `FrameRate`, `FrameCount`, `Container`). See above. | | `Output` | `*topaz.VideoOutput` | Output settings: `Width`/`Height` (precede `Resolution`), `FrameRate` (precedes `FPS`), `AudioCodec` (`"AAC"`/`"AC3"`/`"PCM"`, default `"AAC"`), `AudioBitrate`, `AudioTransfer` (`"Copy"`/`"Convert"`/`"None"`, default `"Copy"`), `VideoEncoder` (`"AV1"`/`"H264"`/`"H265"`/`"ProRes"`/`"VP9"`; `"ProRes"` forces a `mov` container, `"AV1"`/`"VP9"` force `mp4`), `VideoProfile`, `VideoBitrate`, `DynamicCompressionLevel` (`"Low"`/`"Mid"`/`"High"`), `CropToFit`, `Container` (`"mp4"`/`"mov"`/`"mkv"`/`"avi"`/`"webm"`). | | `AdditionalFilters` | `[]map[string]interface{}` | Extra `filters[]` entries alongside the model's own filter, e.g. a frame-interpolation filter. Each entry must include a `"model"` key. | | `Filter` | `map[string]interface{}` | Escape hatch merged into the model's filter entry, taking precedence over the typed settings below. | Proteus-specific settings (sent as filter fields): `VideoType`, `Auto`, `FieldOrder`, `FocusFixLevel`, `Compression`, `Details`, `Prenoise`, `Noise`, `Halo`, `Preblur`, `Blur`, `Grain`, `GrainSigma`, `GrainSize`, `GrainType`, `RecoverOriginalDetailValue`. Starlight Precise-specific settings: `Sharpness` (1.0–5.0, default 5.0), `VideoBitDepth`, `VideoCodec` (`"ffv1"`/`"prores"`/`"vp9"`), `VideoProfile` (`"420"`/`"422"`/`"444"`), `Watermark` (default `false`). See the [Proteus](https://developer.topazlabs.com/video-models/proteus/proteus-1) and [Starlight Precise 2.6](https://developer.topazlabs.com/video-models/starlight/starlight-precise-2.6) model references for the full field semantics. ### Video provider metadata - `DoStart`'s result carries `ProviderMetadata["topaz"]["requestId"]` and, when Topaz can estimate the cost up front (source metadata was supplied), `["estimatedCredits"]` — a `[lowerBound, upperBound]` pair. - A completed `DoStatus` result carries `requestId`, `credits` (the lower bound of Topaz's post-upload `estimates.cost`, which is what Topaz invoices), the full `estimatedCredits` range, `outputSize`, and `expiresAt`. - A failed or canceled `DoStatus` result's `Error` includes Topaz's `errorCode` (e.g. `CREDIT_DIFFERENCE`, which Topaz refunds) when present, and the metadata map carries `requestId` and `errorCode`. ## Errors Both models report Topaz API errors (`{"message": ...}` on the image API, `{"message": ..., "errorCode": ...}` on the video API, or the FastAPI `{"detail": [...]}` validation-error shape) as a `providererrors.ProviderError`, with Topaz's `errorCode` (when present) appended to the message in parentheses and any field-validation messages joined after a colon. ## Workflow Serialization Topaz image and video models can cross a workflow boundary with `providerutils.SerializeModel` / `DeserializeModel`. See [Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the mechanism. ## See Also - [API Reference: GenerateImage](https://goaisdk.com/docs/reference/ai/generate-image.md) - [API Reference: GenerateVideo](https://goaisdk.com/docs/ai-sdk-core/video-generation.md) - [Topaz Labs developer documentation](https://developer.topazlabs.com) --- # Providers > Complete guide to the 49 AI providers supported by the Go AI SDK, covering provider categories, comparison, selection guidance, and configuration patterns. Canonical URL: https://goaisdk.com/docs/providers/ Documentation index: https://goaisdk.com/llms.txt The Go AI SDK supports 49 AI providers (`pkg/providers/*`), giving you access to the latest language models, embedding models, image generation, video generation, speech synthesis, and transcription services. Five providers — **LMNT**, **Ollama**, **Stability AI**, **Vercel**, and **You.com** — are Go-only additions with no upstream TypeScript AI SDK equivalent. ## Provider Categories ### Top Providers The most popular and widely-used AI providers: - [OpenAI](https://goaisdk.com/docs/providers/openai.md) - GPT-6, GPT-5, o-series models - [Anthropic](https://goaisdk.com/docs/providers/anthropic.md) - Claude Sonnet, Opus, Haiku - [Google](https://goaisdk.com/docs/providers/google.md) - Gemini models - [Azure OpenAI](https://goaisdk.com/docs/providers/azure.md) - Enterprise OpenAI deployment - [AWS Bedrock](https://goaisdk.com/docs/providers/bedrock.md) - Multi-model AWS platform - AWS Bedrock Anthropic (`pkg/providers/anthropicaws`, "anthropic-aws") - Claude on Bedrock, rebuilt on the shared Anthropic language model ### Popular Providers High-performance providers with specialized capabilities: - [Cohere](https://goaisdk.com/docs/providers/cohere.md) - Command and Embed models, reranking - [Mistral AI](https://goaisdk.com/docs/providers/mistral.md) - Open-source language models - [Groq](https://goaisdk.com/docs/providers/groq.md) - Ultra-fast inference - [xAI](https://goaisdk.com/docs/providers/xai.md) - Grok models (Responses API, image, video, speech) - [DeepSeek](https://goaisdk.com/docs/providers/deepseek.md) - Advanced reasoning models - [Perplexity](https://goaisdk.com/docs/providers/perplexity.md) - Search-augmented Agent API ### Open Source & Serving Platforms for running open-source models: - [Together AI](https://goaisdk.com/docs/providers/together.md) - Open-source model hosting, reranking - [Fireworks AI](https://goaisdk.com/docs/providers/fireworks.md) - Fast open-source inference - [Replicate](https://goaisdk.com/docs/providers/replicate.md) - Run any open-source model - [HuggingFace](https://goaisdk.com/docs/providers/huggingface.md) - Router Responses API for HF models - [Ollama](https://goaisdk.com/docs/providers/ollama.md) - Local model deployment (**Go-only**) - [Google Vertex AI](https://goaisdk.com/docs/providers/google-vertex.md) - Enterprise AI platform - [GMI Cloud](https://goaisdk.com/docs/providers/gmicloud.md) - OpenAI-compatible chat hosting - [Z.AI](https://goaisdk.com/docs/providers/zai.md) - OpenAI-compatible chat hosting - [MiniMax](https://goaisdk.com/docs/providers/minimax.md) - Chat language models - [Moonshot AI](https://goaisdk.com/docs/providers/moonshot.md) - Kimi language models - [DeepInfra](https://goaisdk.com/docs/providers/deepinfra.md) - Serverless GPU inference - [Baseten](https://goaisdk.com/docs/providers/baseten.md) - ML model deployment - [Cerebras](https://goaisdk.com/docs/providers/cerebras.md) - Ultra-fast inference ### Specialized Providers Providers for specific use cases: - [Alibaba Cloud](https://goaisdk.com/docs/providers/alibaba.md) - Qwen language models & Wan video generation - [KlingAI](https://goaisdk.com/docs/providers/klingai.md) - Professional video generation with motion control - [Stability AI](https://goaisdk.com/docs/providers/stability.md) - Image generation (Stable Diffusion) (**Go-only**) - [Black Forest Labs](https://goaisdk.com/docs/providers/bfl.md) - FLUX image and video models - [FAL](https://goaisdk.com/docs/providers/fal.md) - Image, video, speech, and transcription - [QuiverAI](https://goaisdk.com/docs/providers/quiverai.md) - SVG generation and image vectorization - [Prodia](https://goaisdk.com/docs/providers/prodia.md) - Fast FLUX and Stable Diffusion image generation - [ByteDance](https://github.com/digitallysavvy/go-ai/tree/main/pkg/providers/bytedance) - Seedance/Seedream video and image (Volcengine Ark) - [Luma](https://goaisdk.com/docs/providers/luma.md) - Async image generation - Replicate video, xAI video - see their provider pages ### Speech, Transcription & Audio - [ElevenLabs](https://goaisdk.com/docs/providers/elevenlabs.md) - Speech synthesis and transcription - [Deepgram](https://goaisdk.com/docs/providers/deepgram.md) - Speech transcription - [AssemblyAI](https://goaisdk.com/docs/providers/assemblyai.md) - Audio intelligence - [Cartesia](https://goaisdk.com/docs/providers/cartesia.md) - Speech and transcription - [Fish Audio](https://goaisdk.com/docs/providers/fishaudio.md) - Speech and transcription - [Hume](https://goaisdk.com/docs/providers/hume.md) - Speech synthesis - [Rev.ai](https://goaisdk.com/docs/providers/revai.md) - Transcription with validated job polling - [Gladia](https://goaisdk.com/docs/providers/gladia.md) - Transcription - [LMNT](https://goaisdk.com/docs/providers/lmnt.md) - Speech synthesis (**Go-only**) ### Embeddings & Reranking - [Voyage](https://goaisdk.com/docs/providers/voyage.md) - Embeddings and reranking models ### Infrastructure & Tools - [AI Gateway](https://goaisdk.com/docs/providers/gateway.md) - Unified multi-provider routing, failover, batch, video, realtime - [Vercel](https://goaisdk.com/docs/providers/vercel.md) - OpenAI-compatible Vercel AI API (**Go-only**) - [OpenResponses](https://goaisdk.com/docs/providers/openresponses.md) - Open Responses API protocol (custom tools, QuiverAI's transport) - [OpenAI-compatible providers](https://goaisdk.com/docs/providers/openai-compatible.md) - shared request/response conventions - [TypeSafe AI](https://goaisdk.com/docs/providers/typesafeai.md) - Evaluation models for `ai.ExperimentalEvaluate` - [You.com](https://goaisdk.com/docs/providers/youcom-tools.md) - Search, research, and contents tools (not a model provider) (**Go-only**) ## Quick Start ### 1. Install the SDK ```bash go get github.com/digitallysavvy/go-ai ``` ### 2. Choose a Provider ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, err := provider.LanguageModel("gpt-6-astra") ``` ### 3. Generate Text ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "What is the Go AI SDK?"}) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` ## Provider Comparison ### Language Models | Provider | Best Models | Context Length | Streaming | Tools | Reasoning | |----------|-------------|----------------|-----------|-------|-----------| | OpenAI | GPT-6, o3, o4-mini | 128K+ | Yes | Yes | Yes (o3/o4-mini) | | Anthropic | Claude Opus 4.5 | 200K | Yes | Yes | Yes | | Google | Gemini 2.0 Flash | 1M+ | Yes | Yes | Yes | | Alibaba | Qwen Max, QwQ | 32K | Yes | Yes | Yes (QwQ) | | Groq | Llama 3.3 | 8K-128K | Yes | Yes | No | | DeepSeek | DeepSeek-V3 | 64K | Yes | Yes | Yes | ### Image & Video Generation | Provider | Best Models | Speed | Quality | Features | |----------|-------------|-------|---------|----------| | Alibaba Cloud | Wan 2.6 | Medium | High | Text/Image/Reference-to-video | | Stability AI | SDXL, SD3 | Fast | High | Flexible | | Black Forest Labs | FLUX.1 Pro | Medium | Excellent | Photorealistic | | FAL | FLUX, SD | Very Fast | High | Video support | | Prodia | FLUX, SD | Very Fast | High | Image + video | ### Speech & Audio | Provider | Capability | Quality | Speed | Features | |----------|-----------|---------|-------|----------| | ElevenLabs | TTS | Excellent | Fast | Voice cloning | | Deepgram | STT | Excellent | Real-time | Diarization | | AssemblyAI | STT | Excellent | Fast | Summarization | ## Provider Selection Guide ### For Production Applications **Best Overall**: OpenAI, Anthropic, Google - Enterprise-grade reliability - Excellent documentation - Strong safety measures ### For Cost-Effective Solutions **Best Value**: Groq, Together AI, DeepInfra - Lower pricing - Good performance - Open-source model access ### For Specialized Tasks **Video Generation**: KlingAI, Alibaba Cloud (Wan), FAL **Image Generation**: Stability AI, Black Forest Labs, FAL, Prodia, QuiverAI **Speech Synthesis**: ElevenLabs **Transcription**: Deepgram, AssemblyAI **Fast Inference**: Groq, Cerebras **Local Deployment**: Ollama **Chinese Language**: Alibaba Cloud (Qwen) ### For Enterprise **Best Choice**: Azure OpenAI, AWS Bedrock, Google Vertex AI - Private deployment options - Compliance features - SLA guarantees ## Configuration Patterns ### Environment Variables ```go provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) ``` ### Custom Base URL ```go provider := openai.New(openai.Config{ APIKey: os.Getenv("API_KEY"), BaseURL: "https://custom-endpoint.com/v1", }) ``` ### Regional Configuration ```go provider, err := azure.New(azure.Config{ ResourceName: os.Getenv("AZURE_RESOURCE_NAME"), APIKey: os.Getenv("AZURE_API_KEY"), UseDeploymentBasedURLs: true, }) if err != nil { log.Fatal(err) } ``` ## Common Features ### Streaming All major providers support streaming responses: ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a story"}) for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } ``` ### Tool Calling Most providers support function calling: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather?", Tools: []types.Tool{weatherTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) ``` ### Structured Output Generate JSON conforming to a schema: ```go personSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]string{"type": "string"}, "age": map[string]string{"type": "number"}, }, "required": []string{"name", "age"}, }) result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: personSchema, Prompt: "Describe a person", }) ``` ## Error Handling ### Rate Limits ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) && providerErr.StatusCode == 429 { time.Sleep(time.Second * 5) // Retry } } ``` ### Provider-Specific Errors Each provider has unique error codes and handling - see individual provider documentation. ## Best Practices 1. **Use Environment Variables**: Never hardcode API keys 2. **Handle Rate Limits**: Implement exponential backoff 3. **Monitor Costs**: Track token usage across providers 4. **Test Locally**: Use Ollama for development 5. **Fallback Providers**: Have backup providers configured 6. **Validate Models**: Check model availability before deployment ## Next Steps - Explore the [Overview Guide](https://goaisdk.com/docs/providers/overview.md) for detailed provider concepts - Read individual provider documentation for specific features - Check [API Reference](https://goaisdk.com/docs/reference/ai/generate-text.md) for complete API details - Review [Examples](https://github.com/digitallysavvy/go-ai/tree/main/examples) for implementation patterns ## See Also - [Quick Start Guide](https://goaisdk.com/docs/getting-started.md) - [Core Concepts](https://goaisdk.com/docs/foundations.md) - [API Reference](https://goaisdk.com/docs/reference.md) --- # You.com Tools > Reference for pkg/providers/youcom, a Go port of the You.com AI SDK plugin providing search, research, and content-extraction tools for agents. Canonical URL: https://goaisdk.com/docs/providers/youcom-tools Documentation index: https://goaisdk.com/llms.txt The `pkg/providers/youcom` package mirrors the TypeScript `@youdotcom-oss/ai-sdk-plugin` exports: ```go tools := []types.Tool{ youcom.YouSearch(), youcom.YouResearch(), youcom.YouContents(), } ``` These tools execute locally through the You.com API, use `YDC_API_KEY` by default, and carry `ProviderName="you.com"` plus provider metadata keyed by `you`. --- # Prompt Engineering > Advanced guide to prompt engineering for LLMs in Go, covering what a prompt is, a slogan-generator example, advanced techniques, and common pitfalls. Canonical URL: https://goaisdk.com/docs/advanced/prompt-engineering Documentation index: https://goaisdk.com/llms.txt ## What is a Large Language Model (LLM)? A Large Language Model is essentially a prediction engine that takes a sequence of words as input and aims to predict the most likely sequence to follow. It does this by assigning probabilities to potential next sequences and then selecting one. The model continues to generate sequences until it meets a specified stopping criterion. These models learn by training on massive text corpuses, which means they will be better suited to some use cases than others. For example, a model trained on GitHub data would understand the probabilities of sequences in source code particularly well. However, it's crucial to understand that the generated sequences, while often seeming plausible, can sometimes be random and not grounded in reality. As these models become more accurate, many surprising abilities and applications emerge. ## What is a prompt? Prompts are the starting points for LLMs. They are the inputs that trigger the model to generate text. The scope of prompt engineering involves not just crafting these prompts but also understanding related concepts such as system prompts, tokens, token limits, and context management. ## Why is prompt engineering needed? Prompt engineering currently plays a pivotal role in shaping the responses of LLMs. It allows us to tweak the model to respond more effectively to a broader range of queries. This includes the use of techniques like few-shot learning, chain-of-thought prompting, and structured output generation. The performance, context window, and cost of LLMs varies between models and model providers which adds further constraints to the mix. For example, GPT-4 is more expensive than GPT-4o-mini and significantly slower, but it can also be more effective at certain tasks. Like many things in software engineering, there are trade-offs between cost and performance. ## Example: Build a Slogan Generator ### Start with an instruction Imagine you want to build a slogan generator for marketing campaigns. Creating catchy slogans isn't always straightforward! First, you'll need a prompt that makes it clear what you want. Let's start with a basic instruction: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Create a slogan for a coffee shop.", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) // Output example: "Brew Your Day the Right Way" } ``` Not bad! Now, try making your instruction more specific: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Create a slogan for an organic coffee shop.", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) // Output example: "Naturally Brewed, Ethically Sourced" ``` Introducing a single descriptive term to our prompt influences the completion. Essentially, crafting your prompt is the means by which you "instruct" or "program" the model. ### Include examples Clear instructions are key for quality outcomes, but that might not always be enough. Let's try to enhance your instruction further: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Create three slogans for a coffee shop with live music.", }) ``` These slogans are fine, but could be even better. Often, it's beneficial to both demonstrate and tell the model your requirements. Incorporating examples in your prompt (few-shot prompting) can aid in conveying patterns or subtleties: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") prompt := `Create three slogans for a business with unique features. Business: Bookstore with cats Slogans: "Purr-fect Pages", "Books and Whiskers", "Novels and Nuzzles" Business: Gym with rock climbing Slogans: "Peak Performance", "Reach New Heights", "Climb Your Way Fit" Business: Coffee shop with live music Slogans:` result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) // Output example: "Sip, Savor, and Swing", "Beans and Beats", "Melody in Every Cup" } ``` Great! Incorporating examples of expected output for a certain input prompted the model to generate the kind of names we aimed for. ### Tweak your settings Apart from designing prompts, you can influence completions by tweaking model settings. A crucial setting is the **temperature**. With temperature set to 0, you'll get consistent, deterministic outputs: ```go temperature := 0.0 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Temperature: &temperature, Prompt: "Create a slogan for a coffee shop.", }) // Running this multiple times gives the same result ``` With temperature set higher, you get more varied and creative outputs: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") temperature := 1.0 prompt := `Create three slogans for a business with unique features. Business: Bookstore with cats Slogans: "Purr-fect Pages", "Books and Whiskers", "Novels and Nuzzles" Business: Gym with rock climbing Slogans: "Peak Performance", "Reach New Heights", "Climb Your Way Fit" Business: Coffee shop with live music Slogans:` // Run multiple times to see variation for i := 0; i < 3; i++ { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Temperature: &temperature, Prompt: prompt, }) if err != nil { log.Fatal(err) } fmt.Printf("Try %d: %s\n\n", i+1, result.Text) } } ``` Notice the difference? With a temperature above 0, the same prompt delivers varied completions each time. Keep in mind that the model forecasts the text most likely to follow the preceding text. Temperature, a value from 0 to 1, essentially governs the model's confidence level in making these predictions. A lower temperature implies lesser risks, leading to more precise and deterministic completions. A higher temperature yields a broader range of completions. For your slogan generator, you might want a large pool of name suggestions. A moderate temperature of 0.6 should serve well: ```go temperature := 0.6 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Temperature: &temperature, Prompt: prompt, }) ``` ## Advanced Prompting Techniques ### System Prompts Use system prompts to define the model's behavior and personality: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: `You are a professional marketing copywriter with 10 years of experience. You specialize in creating memorable, punchy slogans. Always think about: - Brand voice and personality - Target audience - Emotional impact - Memorability`, Prompt: "Create three slogans for an eco-friendly coffee shop that sources directly from farmers.", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Chain-of-Thought Prompting Encourage the model to think step-by-step: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") prompt := `Create a slogan for a sustainable fashion brand. Think through this step by step: 1. First, identify the key values of sustainable fashion 2. Consider the target audience and their motivations 3. Think about what makes this brand unique 4. Combine these elements into a memorable slogan 5. Present your final slogan Let's begin:` result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Role-Based Prompting Assign specific roles to the model: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") systemPrompt := `You are a brand strategist at a top advertising agency. Your role is to: - Analyze the brand's core values - Understand the competitive landscape - Create positioning statements that differentiate the brand - Craft slogans that resonate emotionally with the target audience` result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: systemPrompt, Prompt: "Create a positioning statement and three slogans for a premium plant-based protein powder targeted at athletes.", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Structured Output with Prompting Guide the model to return structured information: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") prompt := `Create marketing slogans for a vegan restaurant. Return your response in this exact format: PRIMARY SLOGAN: [main slogan] TAGLINE: [supporting tagline] SOCIAL MEDIA: [shorter version for social media] EMOTIONAL APPEAL: [one-word emotion the slogans evoke] Example: PRIMARY SLOGAN: "Plant Power, Every Hour" TAGLINE: "Deliciously sustainable dining" SOCIAL MEDIA: "🌱 Power Up Plant-Based" EMOTIONAL APPEAL: Energetic` result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Iterative Refinement Use multiple calls to refine the output: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") // Step 1: Generate initial slogans result1, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Create 5 slogans for a tech startup that makes AI-powered productivity tools.", }) if err != nil { log.Fatal(err) } fmt.Println("Initial slogans:") fmt.Println(result1.Text) fmt.Println() // Step 2: Refine the best ones refinePrompt := fmt.Sprintf(`Here are some slogans for an AI productivity tool startup: %s Pick the best 2 and refine them to be: - More punchy and memorable - Focused on the benefit to users - Under 6 words each`, result1.Text) result2, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: refinePrompt, }) if err != nil { log.Fatal(err) } fmt.Println("Refined slogans:") fmt.Println(result2.Text) } ``` ## Best Practices ### 1. Be Specific ```text // Bad - Vague "Write about coffee" // Good - Specific "Write a 3-sentence product description for organic Ethiopian coffee beans emphasizing fruity notes and sustainability" ``` ### 2. Use Examples (Few-Shot Learning) ```go // Show the pattern you want prompt := `Convert casual text to professional: Casual: "Hey, wanna meet up?" Professional: "Would you be available for a meeting?" Casual: "Got your email, thx!" Professional: "Thank you for your email. I have received it." Casual: "Can't make it tomorrow" Professional:` ``` ### 3. Set Constraints ```go // Include explicit constraints prompt := `Create a slogan for a coffee shop. Requirements: - Maximum 5 words - Include the word "brew" - Sound energetic and welcoming - Avoid cliches like "wake up" or "smell the coffee"` ``` ### 4. Specify Format ```go // Tell the model exactly how to format output prompt := `Generate 3 coffee shop slogans. Format each as: 1. [Slogan] - [Brief explanation of why it works] Example: 1. "Grounds for Celebration" - Uses wordplay on "grounds" which means both coffee and reason` ``` ### 5. Use System Prompts for Behavior ```go // Define consistent behavior with system prompts system := `You are a professional copywriter. For every slogan: - Keep it under 6 words - Make it memorable and punchy - Avoid overused marketing cliches - Explain your creative rationale` ``` ### 6. Iterate and Refine ```go // Don't expect perfection on the first try // Generate multiple options temperature := 0.8 for i := 0; i < 5; i++ { result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Temperature: &temperature, Prompt: prompt, }) fmt.Printf("Option %d: %s\n", i+1, result.Text) } ``` ### 7. Test Different Temperatures ```go // Low temperature (0.0-0.3): Focused, deterministic, factual temperature := 0.2 // Medium temperature (0.4-0.7): Balanced creativity and consistency temperature := 0.6 // High temperature (0.8-1.0): Creative, varied, exploratory temperature := 0.9 ``` ## Common Pitfalls ### 1. Ambiguous Instructions ```go // Bad "Make it better" // Good "Rewrite this slogan to be more energetic and under 5 words" ``` ### 2. Conflicting Instructions ```go // Bad - Contradictory "Be creative and unique, but follow these 10 exact examples" // Good - Clear expectations "Use these examples as inspiration for style, but create original slogans" ``` ### 3. Assuming Context ```text // Bad - No context "What should I do?" // Good - Full context "I'm creating a marketing campaign for a vegan restaurant targeting young professionals. What social media platforms should I focus on?" ``` ## Recommended Resources Prompt Engineering is evolving rapidly, with new methods and research papers surfacing regularly. Here are some resources we recommend for learning about and experimenting with prompt engineering: - [Anthropic Prompt Engineering Guide](https://docs.anthropic.com/claude/docs/prompt-engineering) - [OpenAI Prompt Engineering Guide](https://platform.openai.com/docs/guides/prompt-engineering) - [Prompt Engineering Guide by Dair AI](https://www.promptingguide.ai/) - [Brex Prompt Engineering](https://github.com/brexhq/prompt-engineering) ## Next Steps - Learn about [caching](https://goaisdk.com/docs/advanced/caching.md) to optimize repeated prompts - Explore [rate limiting](https://goaisdk.com/docs/advanced/rate-limiting.md) for production applications - See [model as router](https://goaisdk.com/docs/advanced/model-as-router.md) for dynamic model selection --- # Provider Architecture > How the Go AI SDK's provider abstraction layer works, and how to implement your own provider. Canonical URL: https://goaisdk.com/docs/advanced/provider-architecture Documentation index: https://goaisdk.com/llms.txt The Go AI SDK uses a layered abstraction to decouple your application code from specific AI provider APIs. This guide explains the architecture, the interfaces involved, and how to build a custom provider or apply middleware. ## Overview ``` Your Application │ ▼ ┌─────────────┐ │ ai package │ GenerateText, StreamText, GenerateImage, ... └──────┬──────┘ │ uses ▼ ┌──────────────────┐ │ provider.Provider │ LanguageModel(), ImageModel(), ... └──────┬───────────┘ │ implemented by ▼ ┌──────────────────────────────────────────────┐ │ providers/openai, providers/anthropic, ... │ │ (concrete provider implementations) │ └──────────────────────────────────────────────┘ ``` The key insight: the `ai` package operates entirely against the `provider` interfaces — it never imports a concrete provider. This makes providers interchangeable and testable. ## The Provider Interface `provider.Provider` is the entry point for every AI provider: ```go type Provider interface { Name() string LanguageModel(modelID string) (LanguageModel, error) EmbeddingModel(modelID string) (EmbeddingModel, error) ImageModel(modelID string) (ImageModel, error) SpeechModel(modelID string) (SpeechModel, error) TranscriptionModel(modelID string) (TranscriptionModel, error) RerankingModel(modelID string) (RerankingModel, error) } ``` Each method returns a model interface (or an error if the provider doesn't support that model type). Concrete providers like `openai.Provider` and `anthropic.Provider` implement this interface. ## The LanguageModel Interface `provider.LanguageModel` is the core interface for text generation: ```go type LanguageModel interface { // Metadata SpecificationVersion() string // e.g., "v3" Provider() string // e.g., "openai", "anthropic" ModelID() string // e.g., "gpt-6-astra", "claude-sonnet-4-5" // Capabilities SupportsTools() bool SupportsStructuredOutput() bool SupportsImageInput() bool // Generation DoGenerate(ctx context.Context, opts *GenerateOptions) (*types.GenerateResult, error) DoStream(ctx context.Context, opts *GenerateOptions) (TextStream, error) } ``` The `DoGenerate` and `DoStream` methods accept a `GenerateOptions` struct that standardizes all generation parameters across providers. Provider-specific options are passed via `ProviderOptions`. ## The ImageModel Interface `provider.ImageModel` covers image generation: ```go type ImageModel interface { SpecificationVersion() string Provider() string ModelID() string DoGenerate(ctx context.Context, opts *ImageGenerateOptions) (*types.ImageResult, error) } ``` ## ProviderOptions: Passing Provider-Specific Settings When you need features unique to one provider (e.g., Anthropic prompt caching, OpenAI reasoning effort, Google image size), use `ProviderOptions`. The key is the provider name returned by `Provider()`: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing.", ProviderOptions: map[string]interface{}{ "anthropic": map[string]interface{}{ "thinking": map[string]interface{}{ "type": "enabled", "budgetTokens": 8000, }, }, }, }) ``` ```go // Google image size control result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "A mountain landscape at dawn", ProviderOptions: map[string]interface{}{ "google": map[string]interface{}{ "imageSize": google.ImageSize2K, }, }, }) ``` ```go // Vertex AI — the ProviderOptions key is "googleVertex" (checked first), // with "vertex" and then "google" honored as fallbacks for compatibility. // "google-vertex" (with a hyphen) is the provider's identity string // (Provider().Name(), used for model IDs like "google-vertex:gemini-..."), // not a ProviderOptions key — it is never read here. result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: vertexModel, Prompt: "Summarize this document.", ProviderOptions: map[string]interface{}{ "googleVertex": map[string]interface{}{ "safetySettings": []map[string]interface{}{ {"category": "HARM_CATEGORY_DANGEROUS_CONTENT", "threshold": "BLOCK_ONLY_HIGH"}, }, }, }, }) ``` The provider's `DoGenerate` method extracts its own key from `ProviderOptions` and ignores others — so the same call can carry options for multiple providers safely. OpenAI-compatible providers use camelCase provider option keys, matching the TypeScript SDK public API. Deprecated kebab-case or snake_case keys are still accepted for migration compatibility, but they emit warnings so callers can move to camelCase. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: groqModel, Prompt: "Explain tool calling.", ProviderOptions: map[string]interface{}{ "groq": map[string]interface{}{ "reasoningEffort": "high", // preferred "parallelToolCalls": true, }, }, }) ``` ## Implementing a Custom Provider To add a new provider, implement the `provider.Provider` and `provider.LanguageModel` interfaces: ```go package myprovider import ( "context" "fmt" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) // Provider implements provider.Provider for MyAI type Provider struct { apiKey string baseURL string } // New creates a new MyAI provider func New(apiKey string) *Provider { return &Provider{ apiKey: apiKey, baseURL: "https://api.myai.example/v1", } } func (p *Provider) Name() string { return "myai" } func (p *Provider) LanguageModel(modelID string) (provider.LanguageModel, error) { if modelID == "" { return nil, fmt.Errorf("model ID required") } return &LanguageModel{provider: p, modelID: modelID}, nil } // Implement remaining Provider methods (can return errors for unsupported types) func (p *Provider) EmbeddingModel(modelID string) (provider.EmbeddingModel, error) { return nil, fmt.Errorf("myai does not support embeddings") } func (p *Provider) ImageModel(modelID string) (provider.ImageModel, error) { return nil, fmt.Errorf("myai does not support image generation") } func (p *Provider) SpeechModel(modelID string) (provider.SpeechModel, error) { return nil, fmt.Errorf("myai does not support speech synthesis") } func (p *Provider) TranscriptionModel(modelID string) (provider.TranscriptionModel, error) { return nil, fmt.Errorf("myai does not support transcription") } func (p *Provider) RerankingModel(modelID string) (provider.RerankingModel, error) { return nil, fmt.Errorf("myai does not support reranking") } // LanguageModel wraps a MyAI model type LanguageModel struct { provider *Provider modelID string } func (m *LanguageModel) SpecificationVersion() string { return "v3" } func (m *LanguageModel) Provider() string { return "myai" } func (m *LanguageModel) ModelID() string { return m.modelID } func (m *LanguageModel) SupportsTools() bool { return true } func (m *LanguageModel) SupportsStructuredOutput() bool { return true } func (m *LanguageModel) SupportsImageInput() bool { return false } // DoGenerate calls the MyAI API for text generation func (m *LanguageModel) DoGenerate(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { // Build request from opts reqBody := buildMyAIRequest(opts) // Call your API resp, err := callMyAIAPI(ctx, m.provider.baseURL, m.provider.apiKey, m.modelID, reqBody) if err != nil { return nil, fmt.Errorf("myai generate: %w", err) } // Convert response to standard types.GenerateResult return convertMyAIResponse(resp), nil } // DoStream calls the MyAI API for streaming text generation func (m *LanguageModel) DoStream(ctx context.Context, opts *provider.GenerateOptions) (provider.TextStream, error) { // Similar to DoGenerate but returns a provider.TextStream // See existing providers for streaming patterns return nil, fmt.Errorf("streaming not yet implemented") } // myAIRequest and myAIResponse mirror your API's wire format. type myAIRequest struct { Model string `json:"model"` Prompt string `json:"prompt"` } type myAIResponse struct { Text string `json:"text"` } func buildMyAIRequest(opts *provider.GenerateOptions) *myAIRequest { return &myAIRequest{Prompt: opts.Prompt.Text} } // callMyAIAPI sends the request to your API. Replace the body with a real HTTP call. func callMyAIAPI(ctx context.Context, baseURL, apiKey, modelID string, req *myAIRequest) (*myAIResponse, error) { req.Model = modelID return nil, fmt.Errorf("not implemented") } func convertMyAIResponse(resp *myAIResponse) *types.GenerateResult { return &types.GenerateResult{ Text: resp.Text, FinishReason: types.FinishReasonStop, } } ``` Once implemented, your provider works anywhere the SDK accepts a `provider.LanguageModel`: ```go myProvider := myprovider.New(os.Getenv("MYAI_API_KEY")) model, _ := myProvider.LanguageModel("my-model-v1") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello from my custom provider!", }) ``` ## Middleware Language model middleware wraps a model to intercept and modify generation calls — without changing your application code. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // loggingMiddleware logs every generation request var loggingMiddleware = &middleware.LanguageModelMiddleware{ WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { log.Printf("generating with prompt: %q", params.Prompt.Text) result, err := doGenerate() if err == nil { log.Printf("generated %d tokens", result.Usage.GetTotalTokens()) } return result, err }, } func main() { ctx := context.Background() p := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) base, _ := p.LanguageModel("gpt-6-astra") // Wrap the model with middleware wrapped := middleware.WrapLanguageModel(base, []*middleware.LanguageModelMiddleware{loggingMiddleware}, nil, nil) result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: wrapped, Prompt: "What is the capital of France?", }) fmt.Println(result.Text) } ``` ### Middleware Execution Order When multiple middlewares are applied, they execute from first to last on the way in, and last to first on the way out (like standard middleware chains): ```go wrapped := middleware.WrapLanguageModel( base, []*middleware.LanguageModelMiddleware{ rateLimitMiddleware, // executes first cachingMiddleware, // executes second loggingMiddleware, // executes third }, nil, nil, ) // Call order: rateLimitMiddleware → cachingMiddleware → loggingMiddleware → base ``` The SDK ships several built-in middlewares: | Middleware | Package | Purpose | |-----------|---------|---------| | Rate limiting | `middleware` | Throttle requests per second | | Caching | `middleware` | Cache identical prompts | | Retry | `middleware` | Automatic retries with backoff | | Logging | `middleware` | Request/response logging | | Telemetry | `middleware` | OpenTelemetry tracing | See the [Language Model Middleware](https://goaisdk.com/docs/ai-sdk-core/middleware.md) guide for full details on built-in and custom middlewares. ## Testing with Mock Providers To test code that uses the SDK without hitting real APIs, implement a minimal `provider.LanguageModel` that returns controlled responses: ```go type MockLanguageModel struct { responses []string idx int } func (m *MockLanguageModel) SpecificationVersion() string { return "v3" } func (m *MockLanguageModel) Provider() string { return "mock" } func (m *MockLanguageModel) ModelID() string { return "mock-model" } func (m *MockLanguageModel) SupportsTools() bool { return true } func (m *MockLanguageModel) SupportsStructuredOutput() bool { return true } func (m *MockLanguageModel) SupportsImageInput() bool { return false } func (m *MockLanguageModel) DoGenerate(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { text := m.responses[m.idx%len(m.responses)] m.idx++ return &types.GenerateResult{Text: text}, nil } func (m *MockLanguageModel) DoStream(ctx context.Context, opts *provider.GenerateOptions) (provider.TextStream, error) { return nil, fmt.Errorf("streaming not supported in mock") } // In tests: func TestMyFeature(t *testing.T) { model := &MockLanguageModel{ responses: []string{"Hello from mock!", "Another response"}, } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Say hello", }) require.NoError(t, err) assert.Equal(t, "Hello from mock!", result.Text) } ``` ## See Also - [Language Model Middleware](https://goaisdk.com/docs/ai-sdk-core/middleware.md) — built-in middleware reference - [Provider Management](https://goaisdk.com/docs/ai-sdk-core/provider-management.md) — switching between providers at runtime - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) — provider error types - [Providers](https://goaisdk.com/docs/providers.md) — available provider implementations --- # Stopping Streams > Shows how to cancel streaming AI responses in Go using context cancellation, timeouts, HTTP request cancellation, and graceful shutdown patterns. Canonical URL: https://goaisdk.com/docs/advanced/stopping-streams Documentation index: https://goaisdk.com/llms.txt Cancelling ongoing streams is often needed in production applications. For example, users might want to stop a stream when they realize that the response is not what they want, or you may need to implement timeout policies. The Go AI SDK uses Go's standard `context.Context` for cancellation, providing a native and idiomatic way to stop streams. ## Context-Based Cancellation All AI SDK functions accept a `context.Context` parameter that can be used to cancel operations. When the context is cancelled, the stream will stop immediately and clean up resources. ### Basic Example ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { // Create a cancellable context ctx, cancel := context.WithCancel(context.Background()) provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a very long story...", }) if err != nil { log.Fatal(err) } // Cancel the stream after 5 seconds go func() { time.Sleep(5 * time.Second) fmt.Println("\n[Cancelling stream...]") cancel() }() // Process the stream for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } // Check why the stream ended if err := stream.Err(); err != nil { if ctx.Err() == context.Canceled { fmt.Println("\nStream was cancelled successfully") } else { log.Printf("Stream error: %v", err) } } } ``` ## Timeout-Based Cancellation Use `context.WithTimeout` to automatically cancel a stream after a specified duration: ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { // Create a context with 10-second timeout ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second) defer cancel() provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a detailed technical explanation...", }) if err != nil { log.Fatal(err) } for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } // Check if timeout occurred if err := stream.Err(); err != nil { if ctx.Err() == context.DeadlineExceeded { fmt.Println("\n\nStream stopped due to timeout") } else { log.Printf("Stream error: %v", err) } } } ``` ## HTTP Request Cancellation In web servers, you can forward the request context to automatically cancel streams when clients disconnect: ```go package main import ( "context" "fmt" "log" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func streamHandler(w http.ResponseWriter, r *http.Request) { // Use request context - automatically cancelled when client disconnects ctx := r.Context() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a comprehensive guide...", }) if err != nil { http.Error(w, err.Error(), http.StatusInternalServerError) return } // Set up SSE headers w.Header().Set("Content-Type", "text/event-stream") w.Header().Set("Cache-Control", "no-cache") w.Header().Set("Connection", "keep-alive") flusher, ok := w.(http.Flusher) if !ok { http.Error(w, "Streaming not supported", http.StatusInternalServerError) return } // Stream to client for chunk := range stream.Chunks() { select { case <-ctx.Done(): // Client disconnected log.Println("Client disconnected, cleaning up") return default: fmt.Fprintf(w, "data: %s\n\n", chunk.Text) flusher.Flush() } } // Log cancellation reason if err := stream.Err(); err != nil { if ctx.Err() == context.Canceled { log.Println("Stream cancelled by client disconnect") } else { log.Printf("Stream error: %v", err) } } } func main() { http.HandleFunc("/stream", streamHandler) fmt.Println("Server listening on :8080") fmt.Println("Try: curl http://localhost:8080/stream") fmt.Println("Press Ctrl+C in curl to test cancellation") log.Fatal(http.ListenAndServe(":8080", nil)) } ``` ## Handling Cancellation with GenerateText Non-streaming operations also respect context cancellation: ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { // Create a context with 5-second timeout ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a comprehensive analysis of...", }) if err != nil { if ctx.Err() == context.DeadlineExceeded { log.Println("Operation timed out after 5 seconds") } else { log.Printf("Error: %v", err) } return } fmt.Println(result.Text) } ``` ## Graceful Shutdown For long-running services, implement graceful shutdown with context cancellation: ```go package main import ( "context" "fmt" "log" "os" "os/signal" "syscall" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func processStream(ctx context.Context) error { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate a report...", }) if err != nil { return err } for chunk := range stream.Chunks() { fmt.Print(chunk.Text) // Check context periodically select { case <-ctx.Done(): fmt.Println("\n[Shutting down gracefully...]") return ctx.Err() default: } } return stream.Err() } func main() { // Create root context ctx, cancel := context.WithCancel(context.Background()) defer cancel() // Handle OS signals for graceful shutdown sigChan := make(chan os.Signal, 1) signal.Notify(sigChan, os.Interrupt, syscall.SIGTERM) go func() { sig := <-sigChan fmt.Printf("\nReceived signal: %v\n", sig) cancel() // Cancel all operations }() // Process stream with cancellation support if err := processStream(ctx); err != nil { if ctx.Err() == context.Canceled { log.Println("Operation cancelled successfully") } else { log.Printf("Error: %v", err) } } } ``` ## Best Practices ### 1. Always Defer Cancel ```go // Good - Always defer cancel to prevent context leaks ctx, cancel := context.WithCancel(parent) defer cancel() result, err := ai.GenerateText(ctx, options) ``` ### 2. Check Context in Long Operations ```go for chunk := range stream.Chunks() { // Check context regularly select { case <-ctx.Done(): return ctx.Err() default: process(chunk) } } ``` ### 3. Distinguish Cancellation from Errors ```go if err := stream.Err(); err != nil { switch ctx.Err() { case context.Canceled: log.Println("User cancelled the operation") case context.DeadlineExceeded: log.Println("Operation timed out") default: log.Printf("Stream error: %v", err) } } ``` ### 4. Use Request Context in HTTP Handlers ```go func handler(w http.ResponseWriter, r *http.Request) { // Use request context - automatic cancellation on disconnect ctx := r.Context() stream, err := ai.StreamText(ctx, options) // Stream will automatically stop if client disconnects } ``` ### 5. Set Appropriate Timeouts ```go // Short timeout for user-facing requests ctx, cancel := context.WithTimeout(r.Context(), 30*time.Second) defer cancel() // Longer timeout for batch processing ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute) defer cancel() ``` ## Context Propagation Contexts propagate cancellation through your application: ```go func processRequest(ctx context.Context, prompt string) error { // Cancellation from parent context propagates here return generateAndProcess(ctx, prompt) } func generateAndProcess(ctx context.Context, prompt string) error { // Cancellation propagates to AI SDK calls stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: prompt, }) // ... Stream will stop if parent context is cancelled } ``` ## Comparison with JavaScript ### JavaScript (AbortSignal) ```typescript const controller = new AbortController(); const result = streamText({ model, prompt: "...", abortSignal: controller.signal, }); // Cancel the stream controller.abort(); ``` ### Go (Context) ```go ctx, cancel := context.WithCancel(context.Background()) stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "...", }) // Cancel the stream cancel() ``` Go's context pattern is more powerful because: - Automatic propagation through call chains - Built-in deadline and timeout support - Native integration with standard library (HTTP, database, etc.) - Type-safe cancellation reason checking ## Next Steps - Learn about [Backpressure](https://goaisdk.com/docs/advanced/backpressure.md) for managing stream flow - Explore [Rate Limiting](https://goaisdk.com/docs/advanced/rate-limiting.md) for production deployments - Review [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) for robust applications --- # Backpressure > How to handle backpressure and cancellation when working with the Go AI SDK Canonical URL: https://goaisdk.com/docs/advanced/backpressure Documentation index: https://goaisdk.com/llms.txt This page focuses on understanding backpressure and cancellation when working with streams. You do not need to know this information to use the Go AI SDK, but for those interested, it offers a deeper dive on why and how the SDK optimally streams responses. In the following sections, we'll explore backpressure and cancellation in the context of Go channels and goroutines. We'll discuss the issues that can arise from an eager approach and demonstrate how Go's natural blocking behavior resolves them. ## Backpressure and Cancellation with Channels Let's begin by setting up a simple example program using Go channels: ```go package main import ( "fmt" "time" ) // A generator that will yield positive integers func integers(ch chan<- int) { i := 1 for { fmt.Printf("yielding %d\n", i) ch <- i i++ time.Sleep(100 * time.Millisecond) } } // Consume data from channel func main() { // Create a channel ch := make(chan int) // Start the generator go integers(ch) // Read values from the channel for i := 0; i < 10; i++ { value := <-ch fmt.Printf("read %d\n", value) time.Sleep(1 * time.Second) } } ``` In this example, we create a generator that yields positive integers through a channel, and a consumer which reads values from that channel. The generator logs `"yielding %d"` and the consumer logs `"read %d"`. The generator sleeps for 100ms between values, while the consumer takes 1 second to process each value. ## Backpressure with Unbuffered Channels If you run the program above with an **unbuffered channel**, you'll notice something important: the "yield" and "read" logs are paired. The generator blocks on `ch <- i` until the consumer reads with `<-ch`. This is **natural backpressure** - the producer automatically waits for the consumer. ```go // Unbuffered channel - natural backpressure ch := make(chan int) // Blocks until read ``` Output: ``` yielding 1 read 1 yielding 2 read 2 ... ``` The producer and consumer are synchronized. Go's unbuffered channels provide automatic backpressure! ## Problem: Buffered Channels Without Cancellation Now let's see what happens with a buffered channel: ```go package main import ( "fmt" "time" ) func integers(ch chan<- int) { i := 1 for { fmt.Printf("yielding %d\n", i) ch <- i i++ time.Sleep(100 * time.Millisecond) } } func main() { // Buffered channel - can store multiple values ch := make(chan int, 100) go integers(ch) // Only read 3 values for i := 0; i < 3; i++ { value := <-ch fmt.Printf("read %d\n", value) time.Sleep(1 * time.Second) } // Main function exits, but integers() goroutine continues! fmt.Println("Done reading") time.Sleep(2 * time.Second) // See the problem } ``` Output: ``` yielding 1 read 1 yielding 2 yielding 3 yielding 4 ... read 2 yielding 5 yielding 6 ... read 3 Done reading yielding 7 yielding 8 ... ``` The goroutine continues producing values even after we stop consuming! This wastes resources and can cause memory leaks. ## Solution: Context Cancellation Go provides `context.Context` for cancellation: ```go package main import ( "context" "fmt" "time" ) func integers(ctx context.Context, ch chan<- int) { i := 1 for { select { case <-ctx.Done(): fmt.Println("Generator stopped") return case ch <- i: fmt.Printf("yielding %d\n", i) i++ } time.Sleep(100 * time.Millisecond) } } func main() { ctx, cancel := context.WithCancel(context.Background()) defer cancel() // Clean up when main exits ch := make(chan int, 100) go integers(ctx, ch) // Read 3 values for i := 0; i < 3; i++ { value := <-ch fmt.Printf("read %d\n", value) time.Sleep(1 * time.Second) } fmt.Println("Done reading") cancel() // Signal generator to stop time.Sleep(500 * time.Millisecond) // Give time to see cleanup } ``` Output: ``` yielding 1 read 1 yielding 2 read 2 yielding 3 read 3 Done reading Generator stopped ``` Perfect! The generator stops cleanly when we signal cancellation. ## Backpressure in the Go AI SDK The Go AI SDK uses these principles for efficient streaming: ### Chunks - Natural Backpressure ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Count from 1 to 100 slowly, explaining each number.", }) if err != nil { log.Fatal(err) } // Slow consumer - the producer automatically slows down for chunk := range stream.Chunks() { fmt.Print(chunk.Text) // Simulate slow processing time.Sleep(100 * time.Millisecond) // The LLM API won't send more data faster than we can process! } if err := stream.Err(); err != nil { log.Fatal(err) } } ``` The stream chunk channel is consumed as chunks arrive, so: - Each chunk blocks until read - The SDK doesn't buffer unlimited data - Memory usage stays constant - Natural backpressure is maintained ### Context Cancellation for Early Termination ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { // Create context with timeout ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a very long story that never ends...", }) if err != nil { log.Fatal(err) } // Read for a while for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } // Check if cancelled if err := stream.Err(); err != nil { if ctx.Err() == context.DeadlineExceeded { fmt.Println("\n\nStream stopped after timeout - resources freed!") } else { log.Fatal(err) } } } ``` When the context times out or is cancelled: 1. The SDK stops fetching from the LLM API 2. The channel closes cleanly 3. All resources are freed 4. No memory leaks ### User-Initiated Cancellation ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx, cancel := context.WithCancel(context.Background()) provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Tell me a long story...", }) if err != nil { log.Fatal(err) } // Simulate user clicking "Stop" button after 2 seconds go func() { time.Sleep(2 * time.Second) fmt.Println("\n[User clicked Stop button]") cancel() // Stop the stream immediately }() // Process chunks chunkCount := 0 for chunk := range stream.Chunks() { fmt.Print(chunk.Text) chunkCount++ } fmt.Printf("\n\nReceived %d chunks before stopping\n", chunkCount) // Check for cancellation if ctx.Err() == context.Canceled { fmt.Println("Stream cancelled by user - all resources freed") } } ``` ## Real-World Example: HTTP Streaming with Cancellation Here's how backpressure and cancellation work in a web server: ```go package main import ( "context" "fmt" "log" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func streamHandler(w http.ResponseWriter, r *http.Request) { // Use request context - automatically cancelled if client disconnects ctx := r.Context() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long technical explanation...", }) if err != nil { http.Error(w, err.Error(), http.StatusInternalServerError) return } // Set up SSE headers w.Header().Set("Content-Type", "text/event-stream") w.Header().Set("Cache-Control", "no-cache") w.Header().Set("Connection", "keep-alive") flusher, ok := w.(http.Flusher) if !ok { http.Error(w, "Streaming not supported", http.StatusInternalServerError) return } // Stream to client for chunk := range stream.Chunks() { select { case <-ctx.Done(): // Client disconnected - context cancelled log.Println("Client disconnected, cleaning up") return default: // Send chunk to client fmt.Fprintf(w, "data: %s\n\n", chunk.Text) flusher.Flush() } } if err := stream.Err(); err != nil { if ctx.Err() == context.Canceled { log.Println("Stream cancelled, resources freed") } else { log.Printf("Stream error: %v", err) } } } func main() { http.HandleFunc("/stream", streamHandler) fmt.Println("Server listening on :8080") fmt.Println("Try: curl http://localhost:8080/stream") fmt.Println("Press Ctrl+C in curl to see cancellation") log.Fatal(http.ListenAndServe(":8080", nil)) } ``` Key benefits: - **Client disconnects**: Context automatically cancelled - **No memory leaks**: LLM API connection closed immediately - **Natural backpressure**: Client read speed controls data flow - **Clean shutdown**: All goroutines properly terminated ## Best Practices ### 1. Always Use Context ```go // Good - Always pass context func makeRequest(ctx context.Context) error { stream, err := ai.StreamText(ctx, options) // ... } // Bad - Missing context func makeRequest() error { stream, err := ai.StreamText(context.Background(), options) // No way to cancel! } ``` ### 2. Set Timeouts for Long Operations ```go // Good - Timeout prevents hanging ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() stream, err := ai.StreamText(ctx, options) ``` ### 3. Use Unbuffered Channels ```go // Good - Natural backpressure ch := make(chan string) // Bad - Can accumulate unbounded data ch := make(chan string, 10000) ``` ### 4. Check Context in Loops ```go // Good - Check for cancellation for chunk := range stream.Chunks() { select { case <-ctx.Done(): return ctx.Err() default: process(chunk) } } // Better - Context check is automatic with range for chunk := range stream.Chunks() { process(chunk) } // Channel closes when context cancelled ``` ### 5. Defer Cancel Calls ```go // Good - Ensures cleanup ctx, cancel := context.WithCancel(parent) defer cancel() // Always called result, err := doWork(ctx) return result, err ``` ### 6. Handle Errors Properly ```go for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } // Always check for errors after stream completes if err := stream.Err(); err != nil { if ctx.Err() == context.Canceled { log.Println("Cancelled by user") } else if ctx.Err() == context.DeadlineExceeded { log.Println("Timeout") } else { log.Printf("Stream error: %v", err) } } ``` ## Comparison: JavaScript vs Go ### JavaScript (Streams API) ```javascript // Requires explicit pull() implementation for backpressure const stream = new ReadableStream({ async pull(controller) { const { value, done } = await iterator.next(); if (done) controller.close(); else controller.enqueue(value); } }); ``` ### Go (Channels) ```go // Backpressure is automatic with unbuffered channels ch := make(chan string) go func() { for value := range source { ch <- value // Blocks until read } close(ch) }() ``` Go's channel model makes backpressure natural and automatic. Combined with context cancellation, Go provides elegant solutions to streaming challenges. ## Next Steps - Learn about [caching](https://goaisdk.com/docs/advanced/caching.md) to optimize repeated requests - Explore [rate limiting](https://goaisdk.com/docs/advanced/rate-limiting.md) for production deployments - See [error handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) for robust stream handling --- # Caching > Explains caching AI responses in the Go AI SDK using language model middleware, different storage backends, the OnFinish callback, and streaming caches. Canonical URL: https://goaisdk.com/docs/advanced/caching Documentation index: https://goaisdk.com/llms.txt Depending on the type of application you're building, you may want to cache the responses you receive from your AI provider, at least temporarily. ## Using Language Model Middleware (Recommended) The recommended approach to caching responses is using [language model middleware](https://goaisdk.com/docs/ai-sdk-core/middleware.md). Middleware allows you to intercept and modify calls to the language model, making it perfect for caching. Let's see how you can use language model middleware to cache responses: ```go skip-compile package main import ( "context" "crypto/sha256" "encoding/hex" "encoding/json" "fmt" "log" "time" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/testutil" "github.com/redis/go-redis/v9" ) // CacheMiddleware creates middleware that caches LLM responses func CacheMiddleware(client *redis.Client, ttl time.Duration) *middleware.LanguageModelMiddleware { return &middleware.LanguageModelMiddleware{ WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), options *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { // Create cache key from options cacheKey := createCacheKey(options) // Check cache cached, err := client.Get(ctx, cacheKey).Result() if err == nil { // Cache hit var result types.GenerateResult if err := json.Unmarshal([]byte(cached), &result); err == nil { fmt.Println("Cache hit for generate") return &result, nil } } // Cache miss - call the model result, err := doGenerate() if err != nil { return nil, err } // Store in cache data, _ := json.Marshal(result) client.Set(ctx, cacheKey, data, ttl) return result, nil }, WrapStream: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), options *provider.GenerateOptions, model provider.LanguageModel, ) (provider.TextStream, error) { // Create cache key cacheKey := createCacheKey(options) // Check cache cached, err := client.Get(ctx, cacheKey).Result() if err == nil { // Cache hit - return simulated stream var chunks []provider.StreamChunk if err := json.Unmarshal([]byte(cached), &chunks); err == nil { fmt.Println("Cache hit for stream") return testutil.NewMockTextStream(chunks), nil } } // Cache miss - call the model stream, err := doStream() if err != nil { return nil, err } return stream, nil }, } } func createCacheKey(options *provider.GenerateOptions) string { // Create a consistent cache key from options data, _ := json.Marshal(options) hash := sha256.Sum256(data) return "cache:" + hex.EncodeToString(hash[:]) } func main() { ctx := context.Background() // Set up Redis client redisClient := redis.NewClient(&redis.Options{ Addr: "localhost:6379", }) defer redisClient.Close() // Test connection if err := redisClient.Ping(ctx).Err(); err != nil { log.Fatalf("Failed to connect to Redis: %v", err) } // Create model with caching middleware baseModel, _ := getBaseModel() cachedModel := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{ CacheMiddleware(redisClient, 1*time.Hour), }, nil, nil, ) // First call - cache miss fmt.Println("First call:") result1, err := cachedModel.DoGenerate(ctx, &provider.GenerateOptions{ Prompt: types.Prompt{ Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What is 2+2?"}, }, }, }, }, }) if err != nil { log.Fatal(err) } fmt.Println("Response:", result1.Text) // Second call - cache hit fmt.Println("\nSecond call:") result2, err := cachedModel.DoGenerate(ctx, &provider.GenerateOptions{ Prompt: types.Prompt{ Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What is 2+2?"}, }, }, }, }, }) if err != nil { log.Fatal(err) } fmt.Println("Response:", result2.Text) } func getBaseModel() (provider.LanguageModel, error) { // Import your provider and create model // provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) // return provider.LanguageModel("gpt-6-astra") return nil, nil } ``` This example uses Redis to store and retrieve responses, but you can use any key-value storage provider you prefer. ## Caching with Different Storage Backends ### In-Memory Cache For simple use cases, you can use an in-memory cache: ```go package main import ( "context" "crypto/sha256" "encoding/hex" "encoding/json" "sync" "time" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) type CacheEntry struct { Data []byte ExpiresAt time.Time } type MemoryCache struct { mu sync.RWMutex cache map[string]CacheEntry } func NewMemoryCache() *MemoryCache { return &MemoryCache{ cache: make(map[string]CacheEntry), } } func (c *MemoryCache) Get(key string) ([]byte, bool) { c.mu.RLock() defer c.mu.RUnlock() entry, exists := c.cache[key] if !exists || time.Now().After(entry.ExpiresAt) { return nil, false } return entry.Data, true } func (c *MemoryCache) Set(key string, data []byte, ttl time.Duration) { c.mu.Lock() defer c.mu.Unlock() c.cache[key] = CacheEntry{ Data: data, ExpiresAt: time.Now().Add(ttl), } } func (c *MemoryCache) Cleanup() { c.mu.Lock() defer c.mu.Unlock() now := time.Now() for key, entry := range c.cache { if now.After(entry.ExpiresAt) { delete(c.cache, key) } } } func MemoryCacheMiddleware(cache *MemoryCache, ttl time.Duration) *middleware.LanguageModelMiddleware { // Periodic cleanup go func() { ticker := time.NewTicker(5 * time.Minute) defer ticker.Stop() for range ticker.C { cache.Cleanup() } }() return &middleware.LanguageModelMiddleware{ WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), options *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { cacheKey := createCacheKey(options) // Check cache if data, found := cache.Get(cacheKey); found { var result types.GenerateResult if err := json.Unmarshal(data, &result); err == nil { return &result, nil } } // Call model result, err := doGenerate() if err != nil { return nil, err } // Cache result data, _ := json.Marshal(result) cache.Set(cacheKey, data, ttl) return result, nil }, } } func createCacheKey(options *provider.GenerateOptions) string { data, _ := json.Marshal(options) hash := sha256.Sum256(data) return hex.EncodeToString(hash[:]) } ``` ### File-Based Cache For persistent caching without external dependencies: ```go package main import ( "context" "crypto/sha256" "encoding/hex" "encoding/json" "os" "path/filepath" "time" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) type FileCacheEntry struct { Data json.RawMessage `json:"data"` ExpiresAt time.Time `json:"expiresAt"` } func FileCacheMiddleware(cacheDir string, ttl time.Duration) *middleware.LanguageModelMiddleware { // Ensure cache directory exists os.MkdirAll(cacheDir, 0755) return &middleware.LanguageModelMiddleware{ WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), options *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { cacheKey := createCacheKey(options) cacheFile := filepath.Join(cacheDir, cacheKey+".json") // Check cache file if data, err := os.ReadFile(cacheFile); err == nil { var entry FileCacheEntry if err := json.Unmarshal(data, &entry); err == nil { // Check expiration if time.Now().Before(entry.ExpiresAt) { var result types.GenerateResult if err := json.Unmarshal(entry.Data, &result); err == nil { return &result, nil } } } } // Call model result, err := doGenerate() if err != nil { return nil, err } // Cache result resultData, _ := json.Marshal(result) entry := FileCacheEntry{ Data: resultData, ExpiresAt: time.Now().Add(ttl), } entryData, _ := json.Marshal(entry) os.WriteFile(cacheFile, entryData, 0644) return result, nil }, } } func createCacheKey(options *provider.GenerateOptions) string { data, _ := json.Marshal(options) hash := sha256.Sum256(data) return hex.EncodeToString(hash[:]) } ``` ## Using OnFinish Callback Alternatively, you can use the `OnFinish` callback to cache responses: ```go skip-compile package main import ( "context" "crypto/sha256" "encoding/hex" "encoding/json" "fmt" "log" "net/http" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/redis/go-redis/v9" ) func chatHandler(w http.ResponseWriter, r *http.Request) { ctx := r.Context() var req struct { Messages []string `json:"messages"` } if err := json.NewDecoder(r.Body).Decode(&req); err != nil { http.Error(w, err.Error(), http.StatusBadRequest) return } // Create cache key cacheKey := createCacheKeyFromMessages(req.Messages) // Set up Redis redisClient := redis.NewClient(&redis.Options{ Addr: os.Getenv("REDIS_ADDR"), }) defer redisClient.Close() // Check cache if cached, err := redisClient.Get(ctx, cacheKey).Result(); err == nil { fmt.Println("Cache hit") w.Header().Set("Content-Type", "text/plain") w.WriteHeader(http.StatusOK) w.Write([]byte(cached)) return } // Set up provider provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Convert messages messages := convertToMessages(req.Messages) // Call model with OnFinish callback result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: messages, OnFinish: func(ctx context.Context, result *ai.GenerateTextResult, userContext interface{}) { // Cache the response redisClient.Set(context.Background(), cacheKey, result.Text, 1*time.Hour) fmt.Println("Cached response") }, }) if err != nil { http.Error(w, err.Error(), http.StatusInternalServerError) return } // Return response w.Header().Set("Content-Type", "text/plain") w.WriteHeader(http.StatusOK) w.Write([]byte(result.Text)) } func createCacheKeyFromMessages(messages []string) string { data, _ := json.Marshal(messages) hash := sha256.Sum256(data) return "chat:" + hex.EncodeToString(hash[:]) } func convertToMessages(msgs []string) []types.Message { // Convert string messages to Message types return nil } func main() { http.HandleFunc("/api/chat", chatHandler) fmt.Println("Server listening on :8080") log.Fatal(http.ListenAndServe(":8080", nil)) } ``` ## Streaming with Cache For streaming responses, cache the complete stream after it finishes: ```go skip-compile package main import ( "context" "encoding/json" "fmt" "log" "net/http" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/redis/go-redis/v9" ) func streamHandler(w http.ResponseWriter, r *http.Request) { ctx := r.Context() var req struct { Prompt string `json:"prompt"` } json.NewDecoder(r.Body).Decode(&req) cacheKey := "stream:" + req.Prompt redisClient := redis.NewClient(&redis.Options{ Addr: "localhost:6379", }) defer redisClient.Close() // Check cache if cached, err := redisClient.Get(ctx, cacheKey).Result(); err == nil { fmt.Println("Cache hit - sending cached stream") w.Header().Set("Content-Type", "text/event-stream") w.Header().Set("Cache-Control", "no-cache") w.Header().Set("Connection", "keep-alive") flusher, _ := w.(http.Flusher) // Send cached response as if streaming for _, char := range cached { fmt.Fprintf(w, "data: %c\n\n", char) flusher.Flush() } return } // No cache - stream from model provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: req.Prompt, OnFinish: func(result *ai.StreamTextResult) { // Cache complete response redisClient.Set(context.Background(), cacheKey, result.Text(), 1*time.Hour) fmt.Println("Cached stream") }, }) if err != nil { http.Error(w, err.Error(), http.StatusInternalServerError) return } // Stream to client w.Header().Set("Content-Type", "text/event-stream") w.Header().Set("Cache-Control", "no-cache") w.Header().Set("Connection", "keep-alive") flusher, _ := w.(http.Flusher) for chunk := range stream.Chunks() { fmt.Fprintf(w, "data: %s\n\n", chunk.Text) flusher.Flush() } if err := stream.Err(); err != nil { log.Printf("Stream error: %v", err) } } func main() { http.HandleFunc("/stream", streamHandler) log.Fatal(http.ListenAndServe(":8080", nil)) } ``` ## Best Practices ### 1. Use Appropriate Cache Keys ```go // Good - Deterministic cache key func createCacheKey(options *provider.GenerateOptions) string { // Sort and normalize to ensure consistency normalized := normalizeOptions(options) data, _ := json.Marshal(normalized) hash := sha256.Sum256(data) return hex.EncodeToString(hash[:]) } // Bad - Non-deterministic func badCacheKey(options *provider.GenerateOptions) string { return fmt.Sprintf("%v", options) // String representation may vary } ``` ### 2. Set Appropriate TTL ```go // Different TTLs for different use cases const ( ShortTTL = 5 * time.Minute // Frequently changing data MediumTTL = 1 * time.Hour // Stable queries LongTTL = 24 * time.Hour // Static content ) ``` ### 3. Handle Cache Errors Gracefully ```go // Check cache if cached, err := cache.Get(key); err == nil { return cached, nil } // If cache fails, continue with normal operation result, err := model.DoGenerate(ctx, options) if err != nil { return nil, err } // Try to cache, but don't fail if caching fails _ = cache.Set(key, result, ttl) return result, nil ``` ### 4. Cache Invalidation ```go // Implement cache clearing for specific patterns func InvalidateCache(cache *redis.Client, pattern string) error { ctx := context.Background() iter := cache.Scan(ctx, 0, pattern, 0).Iterator() for iter.Next(ctx) { cache.Del(ctx, iter.Val()) } return iter.Err() } // Usage InvalidateCache(redisClient, "chat:*") // Clear all chat caches ``` ### 5. Monitor Cache Performance ```go type CacheStats struct { Hits int64 Misses int64 mu sync.Mutex } func (s *CacheStats) RecordHit() { s.mu.Lock() defer s.mu.Unlock() s.Hits++ } func (s *CacheStats) RecordMiss() { s.mu.Lock() defer s.mu.Unlock() s.Misses++ } func (s *CacheStats) HitRate() float64 { s.mu.Lock() defer s.mu.Unlock() total := s.Hits + s.Misses if total == 0 { return 0 } return float64(s.Hits) / float64(total) } ``` ## Next Steps - Learn about [rate limiting](https://goaisdk.com/docs/advanced/rate-limiting.md) for production deployments - Explore [middleware](https://goaisdk.com/docs/ai-sdk-core/middleware.md) for other use cases - See [model as router](https://goaisdk.com/docs/advanced/model-as-router.md) for dynamic model selection --- # Multiple Concurrent Streams > Patterns for managing multiple concurrent AI streams in Go with goroutines and channels, including fan-out/fan-in and rate-limited streaming approaches. Canonical URL: https://goaisdk.com/docs/advanced/multiple-streamables Documentation index: https://goaisdk.com/llms.txt This guide covers patterns for managing multiple concurrent streams using Go's goroutines, channels, and synchronization primitives. ## Basic Concurrent Streaming ### Multiple Streams in Parallel ```go package main import ( "context" "fmt" "os" "sync" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type StreamResult struct { ID int Text string Error error } func streamConcurrently(ctx context.Context, model provider.LanguageModel, prompts []string) []StreamResult { results := make([]StreamResult, len(prompts)) var wg sync.WaitGroup for i, prompt := range prompts { wg.Add(1) go func(index int, p string) { defer wg.Done() var fullText string stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: p, }) if err != nil { results[index] = StreamResult{ ID: index, Error: err, } return } // Collect all chunks for chunk := range stream.Chunks() { fullText += chunk.Text fmt.Printf("[Stream %d] %s", index, chunk.Text) } // Check for streaming errors if err := stream.Err(); err != nil { results[index] = StreamResult{ ID: index, Text: fullText, Error: err, } return } results[index] = StreamResult{ ID: index, Text: fullText, } fmt.Printf("\n[Stream %d] Completed\n", index) }(i, prompt) } wg.Wait() return results } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-5.4-mini") prompts := []string{ "Explain AI in one sentence", "Explain ML in one sentence", "Explain DL in one sentence", "Explain NLP in one sentence", "Explain CV in one sentence", } results := streamConcurrently(ctx, model, prompts) fmt.Println("\n=== Results ===") for _, result := range results { if result.Error != nil { fmt.Printf("Stream %d: ERROR - %v\n", result.ID, result.Error) } else { fmt.Printf("Stream %d: %s\n", result.ID, result.Text) } } } ``` ## Fan-Out/Fan-In Pattern ### Distributing Work and Collecting Results ```go package main import ( "context" "fmt" "log" "os" "sync" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type Request struct { ID int Prompt string } type Response struct { ID int Text string Error error } // Fan-out: Distribute prompts to workers func fanOut(ctx context.Context, prompts []Request, numWorkers int) <-chan Request { requestChan := make(chan Request, len(prompts)) go func() { defer close(requestChan) for _, prompt := range prompts { select { case requestChan <- prompt: case <-ctx.Done(): return } } }() return requestChan } // Worker: Process requests func worker(ctx context.Context, id int, model provider.LanguageModel, requests <-chan Request, responses chan<- Response) { for req := range requests { log.Printf("Worker %d processing request %d", id, req.ID) var fullText string stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: req.Prompt, }) if err != nil { responses <- Response{ ID: req.ID, Error: err, } continue } for chunk := range stream.Chunks() { fullText += chunk.Text } if err := stream.Err(); err != nil { responses <- Response{ ID: req.ID, Text: fullText, Error: err, } continue } responses <- Response{ ID: req.ID, Text: fullText, } } } // Fan-in: Collect results from workers func fanIn(ctx context.Context, numWorkers int, model provider.LanguageModel, requests <-chan Request) <-chan Response { responses := make(chan Response, numWorkers) var wg sync.WaitGroup // Start workers for i := 0; i < numWorkers; i++ { wg.Add(1) go func(workerID int) { defer wg.Done() worker(ctx, workerID, model, requests, responses) }(i) } // Close responses channel when all workers done go func() { wg.Wait() close(responses) }() return responses } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-5.4-mini") // Create requests requests := []Request{ {ID: 1, Prompt: "What is AI?"}, {ID: 2, Prompt: "What is ML?"}, {ID: 3, Prompt: "What is DL?"}, {ID: 4, Prompt: "What is NLP?"}, {ID: 5, Prompt: "What is CV?"}, {ID: 6, Prompt: "What is RL?"}, {ID: 7, Prompt: "What is GANs?"}, {ID: 8, Prompt: "What is transformers?"}, } // Fan-out to workers requestChan := fanOut(ctx, requests, 3) // Fan-in from workers responseChan := fanIn(ctx, 3, model, requestChan) // Collect results results := make(map[int]Response) for response := range responseChan { results[response.ID] = response } // Print results in order fmt.Println("\n=== Results ===") for i := 1; i <= len(requests); i++ { if resp, ok := results[i]; ok { if resp.Error != nil { fmt.Printf("%d: ERROR - %v\n", resp.ID, resp.Error) } else { fmt.Printf("%d: %s\n", resp.ID, resp.Text[:min(100, len(resp.Text))]) } } } } func min(a, b int) int { if a < b { return a } return b } ``` ## Rate-Limited Concurrent Streaming ### Control Concurrency and Rate Limits ```go package main import ( "context" "fmt" "log" "os" "sync" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "golang.org/x/time/rate" ) type RateLimitedStreamer struct { model provider.LanguageModel limiter *rate.Limiter semaphore chan struct{} } func NewRateLimitedStreamer(model provider.LanguageModel, requestsPerMinute float64, maxConcurrent int) *RateLimitedStreamer { return &RateLimitedStreamer{ model: model, limiter: rate.NewLimiter(rate.Limit(requestsPerMinute/60.0), 1), semaphore: make(chan struct{}, maxConcurrent), } } func (s *RateLimitedStreamer) Stream(ctx context.Context, prompt string, resultChan chan<- string, errorChan chan<- error) { // Acquire semaphore (limit concurrency) select { case s.semaphore <- struct{}{}: defer func() { <-s.semaphore }() case <-ctx.Done(): errorChan <- ctx.Err() return } // Wait for rate limiter if err := s.limiter.Wait(ctx); err != nil { errorChan <- err return } // Start streaming stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: s.model, Prompt: prompt, }) if err != nil { errorChan <- err return } var fullText string for chunk := range stream.Chunks() { fullText += chunk.Text } if err := stream.Err(); err != nil { errorChan <- err return } resultChan <- fullText } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-5.4-mini") // Create streamer: 30 requests per minute, max 5 concurrent streamer := NewRateLimitedStreamer(model, 30, 5) prompts := []string{ "Explain AI", "Explain ML", "Explain DL", "Explain NLP", "Explain CV", "Explain RL", "Explain GANs", "Explain transformers", } resultChan := make(chan string, len(prompts)) errorChan := make(chan error, len(prompts)) var wg sync.WaitGroup start := time.Now() // Start all streams for i, prompt := range prompts { wg.Add(1) go func(index int, p string) { defer wg.Done() log.Printf("Starting stream %d", index) streamer.Stream(ctx, p, resultChan, errorChan) log.Printf("Completed stream %d", index) }(i, prompt) } // Wait for completion go func() { wg.Wait() close(resultChan) close(errorChan) }() // Collect results var results []string var errors []error for resultChan != nil || errorChan != nil { select { case result, ok := <-resultChan: if !ok { resultChan = nil } else { results = append(results, result) } case err, ok := <-errorChan: if !ok { errorChan = nil } else { errors = append(errors, err) } } } elapsed := time.Since(start) fmt.Printf("\n=== Summary ===\n") fmt.Printf("Total time: %v\n", elapsed) fmt.Printf("Successful: %d\n", len(results)) fmt.Printf("Errors: %d\n", len(errors)) for i, result := range results { fmt.Printf("\nResult %d:\n%s\n", i+1, result[:min(100, len(result))]) } } func min(a, b int) int { if a < b { return a } return b } ``` ## Real-Time Stream Aggregation ### Combine Multiple Streams in Real-Time ```go package main import ( "context" "fmt" "log" "os" "sync" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type StreamChunk struct { SourceID int Text string } func aggregateStreams(ctx context.Context, model provider.LanguageModel, prompts []string) <-chan StreamChunk { aggregated := make(chan StreamChunk, 100) var wg sync.WaitGroup for i, prompt := range prompts { wg.Add(1) go func(sourceID int, p string) { defer wg.Done() stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: p, }) if err != nil { log.Printf("Stream %d error: %v", sourceID, err) return } // Forward chunks to aggregated channel for chunk := range stream.Chunks() { select { case aggregated <- StreamChunk{ SourceID: sourceID, Text: chunk.Text, }: case <-ctx.Done(): return } } if err := stream.Err(); err != nil { log.Printf("Stream %d error: %v", sourceID, err) } }(i, prompt) } // Close aggregated channel when all streams complete go func() { wg.Wait() close(aggregated) }() return aggregated } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-5.4-mini") prompts := []string{ "List 3 programming languages", "List 3 database systems", "List 3 cloud providers", } // Get aggregated stream aggregated := aggregateStreams(ctx, model, prompts) // Process chunks as they arrive from any stream streamTexts := make(map[int]string) for chunk := range aggregated { streamTexts[chunk.SourceID] += chunk.Text fmt.Printf("[Stream %d] %s", chunk.SourceID, chunk.Text) } fmt.Println("\n\n=== Final Results ===") for id, text := range streamTexts { fmt.Printf("\nStream %d:\n%s\n", id, text) } } ``` ## Error Handling Across Streams ### Robust Error Handling ```go package main import ( "context" "fmt" "log" "os" "sync" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type StreamStatus struct { ID int Prompt string Text string Error error StartTime time.Time EndTime time.Time ChunkCount int } func streamWithMonitoring(ctx context.Context, id int, model provider.LanguageModel, prompt string) StreamStatus { status := StreamStatus{ ID: id, Prompt: prompt, StartTime: time.Now(), } stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: prompt, }) if err != nil { status.Error = err status.EndTime = time.Now() return status } for chunk := range stream.Chunks() { status.Text += chunk.Text status.ChunkCount++ // Check for context cancellation select { case <-ctx.Done(): status.Error = ctx.Err() status.EndTime = time.Now() return status default: } } status.Error = stream.Err() status.EndTime = time.Now() return status } func main() { // Create context with timeout ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute) defer cancel() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-5.4-mini") prompts := []string{ "Explain quantum computing", "Explain artificial intelligence", "Explain blockchain technology", "Explain machine learning", "Explain neural networks", } statusChan := make(chan StreamStatus, len(prompts)) var wg sync.WaitGroup // Start all streams with monitoring for i, prompt := range prompts { wg.Add(1) go func(id int, p string) { defer wg.Done() status := streamWithMonitoring(ctx, id, model, p) statusChan <- status }(i, prompt) } // Wait and close go func() { wg.Wait() close(statusChan) }() // Collect and analyze statuses var statuses []StreamStatus successCount := 0 errorCount := 0 for status := range statusChan { statuses = append(statuses, status) if status.Error != nil { errorCount++ log.Printf("Stream %d FAILED: %v", status.ID, status.Error) } else { successCount++ duration := status.EndTime.Sub(status.StartTime) log.Printf("Stream %d SUCCESS: %d chunks in %v", status.ID, status.ChunkCount, duration) } } // Print summary fmt.Println("\n=== Summary ===") fmt.Printf("Total streams: %d\n", len(statuses)) fmt.Printf("Successful: %d\n", successCount) fmt.Printf("Failed: %d\n", errorCount) // Print detailed results for _, status := range statuses { duration := status.EndTime.Sub(status.StartTime) fmt.Printf("\n[Stream %d] %v\n", status.ID, duration) if status.Error != nil { fmt.Printf(" Error: %v\n", status.Error) } else { fmt.Printf(" Chunks: %d\n", status.ChunkCount) fmt.Printf(" Text length: %d\n", len(status.Text)) } } } ``` ## Best Practices ### 1. Always Use WaitGroups ```go var wg sync.WaitGroup for _, prompt := range prompts { wg.Add(1) go func(p string) { defer wg.Done() // Process stream }(prompt) } wg.Wait() ``` ### 2. Limit Concurrency ```go sem := make(chan struct{}, maxConcurrent) // Acquire: sem <- struct{}{} // Release: <-sem ``` ### 3. Use Context for Cancellation ```go ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute) defer cancel() ``` ### 4. Handle Errors Gracefully ```go if err := stream.Err(); err != nil { log.Printf("Stream error: %v", err) // Handle error, don't panic } ``` ### 5. Monitor Stream Health ```go log.Printf("Stream %d: %d chunks, %v elapsed", id, chunkCount, time.Since(start)) ``` ## See Also - [Streaming Guide](https://goaisdk.com/docs/foundations/streaming.md) - [Rate Limiting](https://goaisdk.com/docs/advanced/rate-limiting.md) - [Sequential Generations](https://goaisdk.com/docs/advanced/sequential-generations.md) - [Backpressure](https://goaisdk.com/docs/advanced/backpressure.md) --- # Rate Limiting > Shows how to rate limit Go AI SDK applications using a token bucket, a Redis sliding window, middleware, and per-user versus global strategies. Canonical URL: https://goaisdk.com/docs/advanced/rate-limiting Documentation index: https://goaisdk.com/llms.txt Rate limiting helps you protect your APIs from abuse. It involves setting a maximum threshold on the number of requests a client can make within a specified timeframe. This simple technique acts as a gatekeeper, preventing excessive usage that can degrade service performance and incur unnecessary costs. ## Rate Limiting with Token Bucket (Standard Library) Go's `golang.org/x/time/rate` package provides a simple token bucket rate limiter: ```go package main import ( "encoding/json" "fmt" "log" "net/http" "os" "sync" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "golang.org/x/time/rate" ) // RateLimiter manages rate limits per client type RateLimiter struct { limiters map[string]*rate.Limiter mu sync.RWMutex r rate.Limit b int } // NewRateLimiter creates a new rate limiter // r: requests per second // b: burst size func NewRateLimiter(r rate.Limit, b int) *RateLimiter { return &RateLimiter{ limiters: make(map[string]*rate.Limiter), r: r, b: b, } } // GetLimiter returns the rate limiter for a client func (rl *RateLimiter) GetLimiter(clientID string) *rate.Limiter { rl.mu.Lock() defer rl.mu.Unlock() limiter, exists := rl.limiters[clientID] if !exists { limiter = rate.NewLimiter(rl.r, rl.b) rl.limiters[clientID] = limiter } return limiter } // Global rate limiter: 5 requests per 30 seconds var rateLimiter = NewRateLimiter(rate.Limit(5.0/30.0), 5) func generateHandler(w http.ResponseWriter, r *http.Request) { // Get client identifier (IP address or user ID) clientIP := r.RemoteAddr if forwarded := r.Header.Get("X-Forwarded-For"); forwarded != "" { clientIP = forwarded } // Check rate limit limiter := rateLimiter.GetLimiter(clientIP) if !limiter.Allow() { http.Error(w, "Rate limit exceeded", http.StatusTooManyRequests) return } // Process request var req struct { Prompt string `json:"prompt"` } if err := json.NewDecoder(r.Body).Decode(&req); err != nil { http.Error(w, err.Error(), http.StatusBadRequest) return } provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(r.Context(), ai.GenerateTextOptions{ Model: model, Prompt: req.Prompt, }) if err != nil { http.Error(w, err.Error(), http.StatusInternalServerError) return } w.Header().Set("Content-Type", "application/json") json.NewEncoder(w).Encode(map[string]string{ "text": result.Text, }) } func main() { http.HandleFunc("/api/generate", generateHandler) fmt.Println("Server listening on :8080") log.Fatal(http.ListenAndServe(":8080", nil)) } ``` ## Rate Limiting with Redis (Sliding Window) For distributed systems, use Redis for shared rate limiting: ```go skip-compile package main import ( "context" "fmt" "net/http" "os" "strconv" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/redis/go-redis/v9" ) type RedisRateLimiter struct { client *redis.Client limit int window time.Duration } func NewRedisRateLimiter(client *redis.Client, limit int, window time.Duration) *RedisRateLimiter { return &RedisRateLimiter{ client: client, limit: limit, window: window, } } // Allow checks if request is allowed and increments counter func (rl *RedisRateLimiter) Allow(ctx context.Context, clientID string) (bool, int, error) { key := fmt.Sprintf("ratelimit:%s", clientID) now := time.Now().Unix() windowStart := now - int64(rl.window.Seconds()) // Use Lua script for atomic operations script := redis.NewScript(` -- Remove old entries redis.call('ZREMRANGEBYSCORE', KEYS[1], '-inf', ARGV[1]) -- Count current entries local count = redis.call('ZCARD', KEYS[1]) if count < tonumber(ARGV[3]) then -- Add new entry redis.call('ZADD', KEYS[1], ARGV[2], ARGV[2]) redis.call('EXPIRE', KEYS[1], ARGV[4]) return {1, tonumber(ARGV[3]) - count - 1} else return {0, 0} end `) result, err := script.Run(ctx, rl.client, []string{key}, windowStart, // ARGV[1]: window start now, // ARGV[2]: current timestamp rl.limit, // ARGV[3]: limit int(rl.window.Seconds())+1, // ARGV[4]: expiration ).Result() if err != nil { return false, 0, err } values := result.([]interface{}) allowed := values[0].(int64) == 1 remaining := int(values[1].(int64)) return allowed, remaining, nil } func rateLimitHandler(w http.ResponseWriter, r *http.Request) { ctx := r.Context() // Set up Redis client redisClient := redis.NewClient(&redis.Options{ Addr: "localhost:6379", }) defer redisClient.Close() // Create rate limiter: 5 requests per 30 seconds limiter := NewRedisRateLimiter(redisClient, 5, 30*time.Second) // Get client ID clientIP := r.RemoteAddr // Check rate limit allowed, remaining, err := limiter.Allow(ctx, clientIP) if err != nil { http.Error(w, "Rate limit check failed", http.StatusInternalServerError) return } // Set rate limit headers w.Header().Set("X-RateLimit-Limit", "5") w.Header().Set("X-RateLimit-Remaining", strconv.Itoa(remaining)) if !allowed { w.Header().Set("Retry-After", "30") http.Error(w, "Rate limit exceeded", http.StatusTooManyRequests) return } // Process request provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello, world!", }) if err != nil { http.Error(w, err.Error(), http.StatusInternalServerError) return } fmt.Fprint(w, result.Text) } func main() { http.HandleFunc("/api/generate", rateLimitHandler) fmt.Println("Server listening on :8080") http.ListenAndServe(":8080", nil) } ``` ## Rate Limiting Middleware Create reusable middleware for rate limiting: ```go package main import ( "fmt" "net/http" "sync" "time" "golang.org/x/time/rate" ) // RateLimitMiddleware creates middleware that rate limits requests func RateLimitMiddleware(requestsPerSecond float64, burst int) func(http.Handler) http.Handler { type client struct { limiter *rate.Limiter lastSeen time.Time } var ( mu sync.Mutex clients = make(map[string]*client) ) // Clean up old clients periodically go func() { for { time.Sleep(time.Minute) mu.Lock() for ip, client := range clients { if time.Since(client.lastSeen) > 3*time.Minute { delete(clients, ip) } } mu.Unlock() } }() return func(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { // Get client IP ip := r.RemoteAddr if forwarded := r.Header.Get("X-Forwarded-For"); forwarded != "" { ip = forwarded } // Get or create limiter for this client mu.Lock() if _, exists := clients[ip]; !exists { clients[ip] = &client{ limiter: rate.NewLimiter(rate.Limit(requestsPerSecond), burst), } } clients[ip].lastSeen = time.Now() limiter := clients[ip].limiter mu.Unlock() // Check rate limit if !limiter.Allow() { http.Error(w, "Rate limit exceeded", http.StatusTooManyRequests) return } // Continue to next handler next.ServeHTTP(w, r) }) } } func generateHandler(w http.ResponseWriter, r *http.Request) { // Your handler logic here fmt.Fprintf(w, "Request processed successfully") } func main() { // Wrap handler with rate limiting middleware // 5 requests per 30 seconds = 5/30 requests per second handler := RateLimitMiddleware(5.0/30.0, 5)(http.HandlerFunc(generateHandler)) http.Handle("/api/generate", handler) fmt.Println("Server listening on :8080") http.ListenAndServe(":8080", nil) } ``` ## Different Rate Limiting Strategies ### Fixed Window Fixed time window - simple but can have burst at window boundaries: ```go type FixedWindowLimiter struct { client *redis.Client limit int window time.Duration } func (l *FixedWindowLimiter) Allow(ctx context.Context, key string) (bool, error) { now := time.Now() windowKey := fmt.Sprintf("ratelimit:%s:%d", key, now.Unix()/int64(l.window.Seconds())) count, err := l.client.Incr(ctx, windowKey).Result() if err != nil { return false, err } if count == 1 { l.client.Expire(ctx, windowKey, l.window) } return count <= int64(l.limit), nil } ``` ### Sliding Window Log More accurate but uses more memory: ```go type SlidingWindowLogLimiter struct { client *redis.Client limit int window time.Duration } func (l *SlidingWindowLogLimiter) Allow(ctx context.Context, key string) (bool, error) { now := time.Now() windowStart := now.Add(-l.window).Unix() // Remove old entries l.client.ZRemRangeByScore(ctx, key, "-inf", fmt.Sprintf("%d", windowStart)) // Count current entries count, err := l.client.ZCard(ctx, key).Result() if err != nil { return false, err } if count < int64(l.limit) { // Add new entry l.client.ZAdd(ctx, key, redis.Z{ Score: float64(now.Unix()), Member: fmt.Sprintf("%d", now.UnixNano()), }) l.client.Expire(ctx, key, l.window) return true, nil } return false, nil } ``` ### Token Bucket (Per User) ```go type TokenBucketLimiter struct { client *redis.Client capacity int refillRate float64 // tokens per second } func (l *TokenBucketLimiter) Allow(ctx context.Context, key string) (bool, error) { script := redis.NewScript(` local capacity = tonumber(ARGV[1]) local refillRate = tonumber(ARGV[2]) local now = tonumber(ARGV[3]) local lastRefill = redis.call('HGET', KEYS[1], 'lastRefill') local tokens = redis.call('HGET', KEYS[1], 'tokens') if not lastRefill then lastRefill = now tokens = capacity else lastRefill = tonumber(lastRefill) tokens = tonumber(tokens) -- Calculate tokens to add local elapsed = now - lastRefill local tokensToAdd = elapsed * refillRate tokens = math.min(capacity, tokens + tokensToAdd) end if tokens >= 1 then redis.call('HSET', KEYS[1], 'tokens', tokens - 1) redis.call('HSET', KEYS[1], 'lastRefill', now) redis.call('EXPIRE', KEYS[1], 60) return {1, tokens - 1} else redis.call('HSET', KEYS[1], 'lastRefill', now) return {0, tokens} end `) result, err := script.Run(ctx, l.client, []string{key}, l.capacity, l.refillRate, time.Now().Unix(), ).Result() if err != nil { return false, err } values := result.([]interface{}) allowed := values[0].(int64) == 1 return allowed, nil } ``` ## Per-User and Global Rate Limiting Implement both per-user and global limits: ```go package main import ( "fmt" "net/http" "sync" "golang.org/x/time/rate" ) type MultiTierRateLimiter struct { perUser map[string]*rate.Limiter userRate rate.Limit userBurst int global *rate.Limiter mu sync.RWMutex } func NewMultiTierRateLimiter(globalRate rate.Limit, globalBurst int, userRate rate.Limit, userBurst int) *MultiTierRateLimiter { return &MultiTierRateLimiter{ perUser: make(map[string]*rate.Limiter), userRate: userRate, userBurst: userBurst, global: rate.NewLimiter(globalRate, globalBurst), } } func (m *MultiTierRateLimiter) Allow(userID string) (bool, string) { // Check global limit first if !m.global.Allow() { return false, "global" } // Check per-user limit m.mu.Lock() if _, exists := m.perUser[userID]; !exists { m.perUser[userID] = rate.NewLimiter(m.userRate, m.userBurst) } limiter := m.perUser[userID] m.mu.Unlock() if !limiter.Allow() { return false, "user" } return true, "" } func rateLimitMiddleware(limiter *MultiTierRateLimiter) func(http.Handler) http.Handler { return func(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { // Get user ID from auth or use IP userID := r.Header.Get("X-User-ID") if userID == "" { userID = r.RemoteAddr } allowed, reason := limiter.Allow(userID) if !allowed { if reason == "global" { http.Error(w, "Global rate limit exceeded", http.StatusTooManyRequests) } else { http.Error(w, "User rate limit exceeded", http.StatusTooManyRequests) } return } next.ServeHTTP(w, r) }) } } func main() { // Global: 100 requests per second // Per user: 10 requests per minute limiter := NewMultiTierRateLimiter(100, 100, rate.Limit(10.0/60.0), 10) mux := http.NewServeMux() mux.HandleFunc("/api/generate", func(w http.ResponseWriter, r *http.Request) { fmt.Fprintf(w, "Request processed") }) handler := rateLimitMiddleware(limiter)(mux) http.ListenAndServe(":8080", handler) } ``` ## Rate Limiting for Streaming Responses Handle rate limits for streaming AI responses: ```go package main import ( "fmt" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "golang.org/x/time/rate" ) var streamLimiter = rate.NewLimiter(rate.Limit(5.0/60.0), 5) // 5 per minute func streamHandler(w http.ResponseWriter, r *http.Request) { // Check rate limit if !streamLimiter.Allow() { http.Error(w, "Rate limit exceeded", http.StatusTooManyRequests) return } ctx := r.Context() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long story...", }) if err != nil { http.Error(w, err.Error(), http.StatusInternalServerError) return } w.Header().Set("Content-Type", "text/event-stream") w.Header().Set("Cache-Control", "no-cache") w.Header().Set("Connection", "keep-alive") flusher, _ := w.(http.Flusher) for chunk := range stream.Chunks() { fmt.Fprintf(w, "data: %s\n\n", chunk.Text) flusher.Flush() } if err := stream.Err(); err != nil { fmt.Fprintf(w, "data: error: %v\n\n", err) } } func main() { http.HandleFunc("/stream", streamHandler) http.ListenAndServe(":8080", nil) } ``` ## Best Practices ### 1. Return Rate Limit Headers ```go w.Header().Set("X-RateLimit-Limit", "5") w.Header().Set("X-RateLimit-Remaining", strconv.Itoa(remaining)) w.Header().Set("X-RateLimit-Reset", strconv.FormatInt(resetTime.Unix(), 10)) if !allowed { w.Header().Set("Retry-After", "30") http.Error(w, "Rate limit exceeded", http.StatusTooManyRequests) } ``` ### 2. Use Different Limits for Different Endpoints ```go var ( generateLimiter = rate.NewLimiter(5.0/60.0, 5) // 5 per minute streamLimiter = rate.NewLimiter(2.0/60.0, 2) // 2 per minute embedLimiter = rate.NewLimiter(50.0/60.0, 50) // 50 per minute ) ``` ### 3. Implement Graceful Degradation ```go if !limiter.Allow() { // Instead of hard failure, use a simpler/faster model model, _ = provider.LanguageModel("gpt-5.4-mini") } else { model, _ = provider.LanguageModel("gpt-6-astra") } ``` ### 4. Clean Up Old Entries ```go // Periodically clean up old rate limit entries go func() { ticker := time.NewTicker(1 * time.Minute) defer ticker.Stop() for range ticker.C { mu.Lock() for key, client := range clients { if time.Since(client.lastSeen) > 5*time.Minute { delete(clients, key) } } mu.Unlock() } }() ``` ### 5. Monitor Rate Limit Metrics ```go type RateLimitMetrics struct { Allowed int64 Rejected int64 mu sync.Mutex } func (m *RateLimitMetrics) RecordAllowed() { m.mu.Lock() defer m.mu.Unlock() m.Allowed++ } func (m *RateLimitMetrics) RecordRejected() { m.mu.Lock() defer m.mu.Unlock() m.Rejected++ } func (m *RateLimitMetrics) Report() (int64, int64, float64) { m.mu.Lock() defer m.mu.Unlock() total := m.Allowed + m.Rejected rejectionRate := float64(m.Rejected) / float64(total) * 100 return m.Allowed, m.Rejected, rejectionRate } ``` ## Next Steps - Learn about [caching](https://goaisdk.com/docs/advanced/caching.md) to reduce API calls - Explore [middleware](https://goaisdk.com/docs/ai-sdk-core/middleware.md) for other cross-cutting concerns - See [error handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) for robust error management --- # Language Models as Routers > Use language models to dynamically route requests to appropriate handlers Canonical URL: https://goaisdk.com/docs/advanced/model-as-router Documentation index: https://goaisdk.com/llms.txt Language models can act as intelligent routers, using their reasoning capabilities to determine which function or handler should process a user's request. This allows you to build applications where the model routes requests to appropriate logic based on user intent. ## Probabilistic Routing with Deterministic Outputs Language models can route requests deterministically by using function calling. When provided with a set of function definitions, the model will: - Execute the function most relevant to the user query - Not execute any function if the query is out of scope ```go package main import ( "context" "fmt" "log" "math/rand" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Define available routes as tools weatherTool := types.Tool{ Name: "getWeather", Description: "Get the weather in a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "The location to get the weather for", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location := input["location"].(string) temperature := 72 + rand.Intn(21) - 10 return map[string]interface{}{ "location": location, "temperature": temperature, }, nil }, } // Test different queries queries := []string{ "What is the weather in San Francisco?", // getWeather called "What is the weather in New York?", // getWeather called "What events are happening in London?", // No function called } for _, query := range queries { fmt.Printf("\nQuery: %s\n", query) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a friendly weather assistant!", Prompt: query, Tools: []types.Tool{weatherTool}, }) if err != nil { log.Fatal(err) } if len(result.ToolCalls) > 0 { fmt.Printf("Routed to: %s\n", result.ToolCalls[0].ToolName) } else { fmt.Println("No routing - general response") } fmt.Printf("Response: %s\n", result.Text) } } ``` This emergent ability to choose whether a function should be executed is the model performing "reasoning". Combined with function calling, this enables language models to act as routers. ## Language Models as Application Routers Traditionally, developers write explicit routing logic: ```go // Traditional routing mux := http.NewServeMux() mux.HandleFunc("/login", handleLogin) mux.HandleFunc("/user/{username}", handleUserProfile) mux.HandleFunc("/api/events", handleEvents) ``` With language models as routers, the model determines which handler to invoke based on user intent: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) // Define application routes as tools func createRoutes() []types.Tool { return []types.Tool{ { Name: "getUserProfile", Description: "Get a user's profile information", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "username": map[string]interface{}{ "type": "string", "description": "The username to look up", }, }, "required": []string{"username"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { username := input["username"].(string) return map[string]interface{}{ "username": username, "name": "John Doe", "email": "john@example.com", "joined": "2023-01-15", }, nil }, }, { Name: "searchEvents", Description: "Search for events", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{ "type": "string", "description": "Search query", }, "limit": map[string]interface{}{ "type": "number", "description": "Maximum number of results", }, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := input["query"].(string) limit := 10 if l, ok := input["limit"].(float64); ok { limit = int(l) } return map[string]interface{}{ "query": query, "results": []string{"Event 1", "Event 2", "Event 3"}, "total": limit, }, nil }, }, { Name: "login", Description: "Handle user authentication", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{ "loginUrl": "/login", "message": "Please log in to continue", }, nil }, }, } } func routeRequest(ctx context.Context, userQuery string) (*ai.GenerateTextResult, error) { provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") routes := createRoutes() return ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are an intelligent router. Call the appropriate function based on the user's request.", Prompt: userQuery, Tools: routes, }) } func main() { ctx := context.Background() queries := []string{ "Show me John's profile", "Find upcoming tech events", "I need to log in", } for _, query := range queries { fmt.Printf("\n=== User Query: %s ===\n", query) result, err := routeRequest(ctx, query) if err != nil { log.Fatal(err) } if len(result.ToolCalls) > 0 { fmt.Printf("Routed to: %s\n", result.ToolCalls[0].ToolName) fmt.Printf("Parameters: %v\n", result.ToolCalls[0].Arguments) } fmt.Printf("Response: %s\n", result.Text) } } ``` ## Routing by Parameters Language models can extract parameters from natural language and route accordingly: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func searchArtworks(ctx context.Context, artist string) ([]string, error) { // Simulated database query artworks := map[string][]string{ "van gogh": {"Starry Night", "Sunflowers", "Irises"}, "picasso": {"Guernica", "Les Demoiselles d'Avignon", "The Weeping Woman"}, "monet": {"Water Lilies", "Impression, Sunrise", "Haystacks"}, } if works, exists := artworks[artist]; exists { return works, nil } return []string{}, fmt.Errorf("no artworks found for %s", artist) } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") searchTool := types.Tool{ Name: "searchArtworks", Description: "Search for artworks by a specific artist", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "artist": map[string]interface{}{ "type": "string", "description": "The artist's name", }, }, "required": []string{"artist"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { artist := input["artist"].(string) works, err := searchArtworks(ctx, artist) if err != nil { return nil, err } return map[string]interface{}{ "artist": artist, "artworks": works, }, nil }, } queries := []string{ "Show me paintings by Van Gogh", "Find Picasso's work", "What did Monet paint?", } for _, query := range queries { fmt.Printf("\n=== Query: %s ===\n", query) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are an art search assistant.", Prompt: query, Tools: []types.Tool{searchTool}, }) if err != nil { log.Fatal(err) } if len(result.ToolCalls) > 0 { fmt.Printf("Routed to: %s with artist=%v\n", result.ToolCalls[0].ToolName, result.ToolCalls[0].Arguments["artist"]) } fmt.Printf("Response: %s\n", result.Text) } } ``` ## Routing by Sequence Language models can perform multi-step routing for complex tasks: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") // Define sequential operations as tools tools := []types.Tool{ { Name: "lookupCalendar", Description: "Check the user's calendar for availability", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "date": map[string]interface{}{ "type": "string", "description": "Date to check (YYYY-MM-DD)", }, }, "required": []string{"date"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{ "available": []string{"5:00 PM", "6:00 PM", "7:00 PM"}, }, nil }, }, { Name: "lookupFriendsCalendar", Description: "Check friends' calendars for availability", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "friends": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, }, "date": map[string]interface{}{ "type": "string", }, }, "required": []string{"friends", "date"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{ "mutuallyAvailable": []string{"6:00 PM", "7:00 PM"}, }, nil }, }, { Name: "searchNearbyPlaces", Description: "Search for nearby happy hour spots", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "category": map[string]interface{}{ "type": "string", }, "time": map[string]interface{}{ "type": "string", }, }, "required": []string{"category"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{ "places": []map[string]interface{}{ {"name": "The Rooftop Bar", "address": "123 Main St"}, {"name": "Happy Hour Spot", "address": "456 Oak Ave"}, }, }, nil }, }, { Name: "createEvent", Description: "Create a calendar event and send invites", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "title": map[string]interface{}{ "type": "string", }, "date": map[string]interface{}{ "type": "string", }, "time": map[string]interface{}{ "type": "string", }, "location": map[string]interface{}{ "type": "string", }, "attendees": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, }, }, "required": []string{"title", "date", "time", "location", "attendees"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{ "eventId": "evt_123", "status": "created", "invites": "sent", }, nil }, }, } // Create agent that routes through the sequence eventAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful calendar assistant. Use the available tools to help plan events.", Tools: tools, MaxSteps: 20, OnToolCall: func(toolCall types.ToolCall) { fmt.Printf("Step: Calling %s\n", toolCall.ToolName) }, }) result, err := eventAgent.Execute(ctx, "Schedule a happy hour with Sarah and Mike for tomorrow evening. Find a good spot nearby.") if err != nil { log.Fatal(err) } fmt.Printf("\n=== Agent Response ===\n%s\n", result.Text) fmt.Printf("\nCompleted in %d steps\n", len(result.Steps)) } ``` Output: ``` Step: Calling lookupCalendar Step: Calling lookupFriendsCalendar Step: Calling searchNearbyPlaces Step: Calling createEvent === Agent Response === I've scheduled a happy hour for tomorrow at 6:00 PM at The Rooftop Bar (123 Main St). Invites have been sent to Sarah and Mike. Completed in 4 steps ``` ## Intent-Based Routing Route to different application logic based on user intent: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type RouteHandler func(context.Context, map[string]interface{}, types.ToolExecutionOptions) (interface{}, error) type Route struct { Name string Description string Parameters map[string]interface{} Handler RouteHandler } func createApplicationRoutes() []types.Tool { routes := []Route{ { Name: "processPayment", Description: "Process a payment transaction", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "amount": map[string]interface{}{ "type": "number", "description": "Payment amount", }, "currency": map[string]interface{}{ "type": "string", "description": "Currency code (USD, EUR, etc.)", }, }, "required": []string{"amount", "currency"}, }, Handler: func(ctx context.Context, params map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { amount := params["amount"].(float64) currency := params["currency"].(string) // Process payment return map[string]interface{}{ "transactionId": "txn_123", "amount": amount, "currency": currency, "status": "success", }, nil }, }, { Name: "getOrderStatus", Description: "Check the status of an order", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "orderId": map[string]interface{}{ "type": "string", "description": "Order ID to check", }, }, "required": []string{"orderId"}, }, Handler: func(ctx context.Context, params map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { orderId := params["orderId"].(string) return map[string]interface{}{ "orderId": orderId, "status": "shipped", "eta": "2024-01-15", }, nil }, }, { Name: "requestRefund", Description: "Request a refund for an order", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "orderId": map[string]interface{}{ "type": "string", }, "reason": map[string]interface{}{ "type": "string", }, }, "required": []string{"orderId", "reason"}, }, Handler: func(ctx context.Context, params map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { orderId := params["orderId"].(string) reason := params["reason"].(string) return map[string]interface{}{ "refundId": "ref_456", "orderId": orderId, "reason": reason, "status": "processing", }, nil }, }, } // Convert routes to tools var tools []types.Tool for _, route := range routes { r := route // Capture for closure tools = append(tools, types.Tool{ Name: r.Name, Description: r.Description, Parameters: r.Parameters, Execute: types.ToolExecutor(r.Handler), }) } return tools } func handleUserQuery(ctx context.Context, query string) error { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") tools := createApplicationRoutes() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "You are a customer service assistant. Route user requests to the appropriate function.", Prompt: query, Tools: tools, }) if err != nil { return err } fmt.Printf("Query: %s\n", query) if len(result.ToolCalls) > 0 { fmt.Printf("Routed to: %s\n", result.ToolCalls[0].ToolName) fmt.Printf("Parameters: %v\n", result.ToolCalls[0].Arguments) } fmt.Printf("Response: %s\n\n", result.Text) return nil } func main() { ctx := context.Background() queries := []string{ "I need to pay $50 in USD", "Where is my order #12345?", "I want to return order #67890 because it's damaged", } for _, query := range queries { if err := handleUserQuery(ctx, query); err != nil { log.Fatal(err) } } } ``` ## Best Practices ### 1. Clear Tool Descriptions ```go // Good - Clear and specific types.Tool{ Name: "searchProducts", Description: "Search for products in the catalog by name, category, or description. Returns up to 10 matching products.", } // Bad - Vague types.Tool{ Name: "search", Description: "Search for things", } ``` ### 2. Validate Parameters ```go Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { amount, ok := input["amount"].(float64) if !ok || amount <= 0 { return nil, fmt.Errorf("invalid amount") } // Process with validated parameter return processPayment(ctx, amount) }, ``` ### 3. Handle Routing Fallbacks ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: query, Tools: routes, }) if len(result.ToolCalls) == 0 { // No route matched - handle gracefully return handleGeneralQuery(ctx, query) } ``` ### 4. Log Routing Decisions ```go OnToolCall: func(toolCall types.ToolCall) { log.Printf("Routing: %s -> %s (params: %v)", userQuery, toolCall.ToolName, toolCall.Arguments) }, ``` ### 5. Monitor Routing Accuracy ```go type RoutingMetrics struct { TotalRequests int64 RoutedCorrectly int64 RoutedIncorrectly int64 } func (m *RoutingMetrics) RecordRouting(correct bool) { m.TotalRequests++ if correct { m.RoutedCorrectly++ } else { m.RoutedIncorrectly++ } } ``` ## Next Steps - Learn about [agents](https://goaisdk.com/docs/agents/overview.md) for multi-step routing - Explore [workflow patterns](https://goaisdk.com/docs/agents/workflows.md) for complex routing scenarios - See [tools and tool calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) for more on function calling --- # Sequential Generations > Learn how to implement sequential generations ("chains") with the Go AI SDK Canonical URL: https://goaisdk.com/docs/advanced/sequential-generations Documentation index: https://goaisdk.com/llms.txt When working with the Go AI SDK, you may want to create sequences of generations (often referred to as "chains" or "pipes"), where the output of one becomes the input for the next. This can be useful for creating more complex AI-powered workflows or for breaking down larger tasks into smaller, more manageable steps. ## Basic Sequential Generations In a sequential chain, the output of one generation is directly used as input for the next generation. This allows you to create a series of dependent generations, where each step builds upon the previous one. Here's a basic example: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func sequentialActions(ctx context.Context) error { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Step 1: Generate blog post ideas fmt.Println("=== Step 1: Generating Ideas ===") ideasResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate 10 ideas for a blog post about making spaghetti.", }) if err != nil { return err } fmt.Println("Generated Ideas:") fmt.Println(ideasResult.Text) fmt.Println() // Step 2: Pick the best idea fmt.Println("=== Step 2: Selecting Best Idea ===") bestIdeaResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf(`Here are some blog post ideas about making spaghetti: %s Pick the best idea from the list above and explain why it's the best.`, ideasResult.Text), }) if err != nil { return err } fmt.Println("Best Idea:") fmt.Println(bestIdeaResult.Text) fmt.Println() // Step 3: Generate an outline fmt.Println("=== Step 3: Creating Outline ===") outlineResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf(`We've chosen the following blog post idea about making spaghetti: %s Create a detailed outline for a blog post based on this idea.`, bestIdeaResult.Text), }) if err != nil { return err } fmt.Println("Blog Post Outline:") fmt.Println(outlineResult.Text) return nil } func main() { ctx := context.Background() if err := sequentialActions(ctx); err != nil { log.Fatal(err) } } ``` In this example, we: 1. Generate ideas for a blog post 2. Pick the best idea from the generated list 3. Create an outline based on the selected idea Each step uses the output from the previous step as input for the next generation. ## Sequential Refinement Pattern Iteratively refine content through multiple generations: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func refineContent(ctx context.Context, initialContent string, iterations int) (string, error) { provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-4-5") currentContent := initialContent for i := 0; i < iterations; i++ { fmt.Printf("\n=== Refinement Iteration %d ===\n", i+1) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf(`Review and improve this text: %s Make it more concise, clear, and impactful. Return only the improved version.`, currentContent), }) if err != nil { return "", err } currentContent = result.Text fmt.Println(currentContent) } return currentContent, nil } func main() { ctx := context.Background() initialText := `The process of making spaghetti is something that many people enjoy doing. It involves boiling water in a pot, adding salt to the water, putting the spaghetti noodles into the boiling water, and then waiting for them to cook for the appropriate amount of time before draining the water from the pot.` fmt.Println("=== Initial Text ===") fmt.Println(initialText) finalText, err := refineContent(ctx, initialText, 3) if err != nil { log.Fatal(err) } fmt.Println("\n=== Final Refined Text ===") fmt.Println(finalText) } ``` ## Extract, Transform, Load (ETL) Pattern Chain generations for data processing pipelines: ```go package main import ( "context" "encoding/json" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/schema" ) type ExtractedData struct { Title string `json:"title"` KeyPoints []string `json:"keyPoints"` Entities []string `json:"entities"` } type TransformedData struct { Summary string `json:"summary"` Category string `json:"category"` Sentiment string `json:"sentiment"` } func extractData(ctx context.Context, model provider.LanguageModel, rawText string) (*ExtractedData, error) { fmt.Println("=== Step 1: Extract Data ===") result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "title": map[string]interface{}{ "type": "string", "description": "The main title or topic", }, "keyPoints": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, "description": "Key points from the text", }, "entities": map[string]interface{}{ "type": "array", "items": map[string]string{"type": "string"}, "description": "Important entities mentioned", }, }, "required": []string{"title", "keyPoints", "entities"}, }), Prompt: fmt.Sprintf("Extract structured data from this text:\n\n%s", rawText), }) if err != nil { return nil, err } var extracted ExtractedData extractedBytes, err := json.Marshal(result.Object) if err != nil { return nil, err } if err := json.Unmarshal(extractedBytes, &extracted); err != nil { return nil, err } fmt.Printf("Extracted: %+v\n\n", extracted) return &extracted, nil } func transformData(ctx context.Context, model provider.LanguageModel, extracted *ExtractedData) (*TransformedData, error) { fmt.Println("=== Step 2: Transform Data ===") extractedJSON, _ := json.Marshal(extracted) result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "summary": map[string]interface{}{ "type": "string", "description": "A one-sentence summary", }, "category": map[string]interface{}{ "type": "string", "description": "The category this belongs to", }, "sentiment": map[string]interface{}{ "type": "string", "enum": []string{"positive", "neutral", "negative"}, }, }, "required": []string{"summary", "category", "sentiment"}, }), Prompt: fmt.Sprintf("Transform this extracted data into a summary:\n\n%s", string(extractedJSON)), }) if err != nil { return nil, err } var transformed TransformedData transformedBytes, err := json.Marshal(result.Object) if err != nil { return nil, err } if err := json.Unmarshal(transformedBytes, &transformed); err != nil { return nil, err } fmt.Printf("Transformed: %+v\n\n", transformed) return &transformed, nil } func loadData(ctx context.Context, extracted *ExtractedData, transformed *TransformedData) error { fmt.Println("=== Step 3: Load Data ===") // Simulate saving to database fmt.Printf("Saving to database:\n") fmt.Printf(" Title: %s\n", extracted.Title) fmt.Printf(" Summary: %s\n", transformed.Summary) fmt.Printf(" Category: %s\n", transformed.Category) fmt.Printf(" Sentiment: %s\n", transformed.Sentiment) fmt.Printf(" Key Points: %v\n", extracted.KeyPoints) return nil } func etlPipeline(ctx context.Context, rawText string) error { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Extract extracted, err := extractData(ctx, model, rawText) if err != nil { return err } // Transform transformed, err := transformData(ctx, model, extracted) if err != nil { return err } // Load return loadData(ctx, extracted, transformed) } func main() { ctx := context.Background() rawText := `Breaking: Major tech company announces new AI breakthrough. The revolutionary system achieves unprecedented accuracy in natural language understanding. Researchers say this could transform how we interact with technology. Industry experts are calling it a game-changer for artificial intelligence.` if err := etlPipeline(ctx, rawText); err != nil { log.Fatal(err) } } ``` ## Multi-Model Sequential Pipeline Use different models for different steps: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func multiModelPipeline(ctx context.Context, userQuery string) error { // Use fast model for classification openaiProvider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) fastModel, _ := openaiProvider.LanguageModel("gpt-5.4-mini") fmt.Println("=== Step 1: Classify Query (Fast Model) ===") classifyResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: fastModel, Prompt: fmt.Sprintf("Classify this query as 'technical', 'creative', or 'analytical':\n\n%s", userQuery), }) if err != nil { return err } classification := classifyResult.Text fmt.Printf("Classification: %s\n\n", classification) // Use appropriate model based on classification anthropicProvider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) var processingModel provider.LanguageModel if classification == "technical" { processingModel, _ = openaiProvider.LanguageModel("gpt-6-astra") } else { processingModel, _ = anthropicProvider.LanguageModel("claude-sonnet-4-5") } fmt.Println("=== Step 2: Process Query (Specialized Model) ===") processResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: processingModel, Prompt: userQuery, }) if err != nil { return err } fmt.Printf("Response: %s\n\n", processResult.Text) // Use fast model for summary fmt.Println("=== Step 3: Summarize (Fast Model) ===") summaryResult, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: fastModel, Prompt: fmt.Sprintf("Summarize this in one sentence:\n\n%s", processResult.Text), }) if err != nil { return err } fmt.Printf("Summary: %s\n", summaryResult.Text) return nil } func main() { ctx := context.Background() if err := multiModelPipeline(ctx, "Explain quantum entanglement and its applications in computing"); err != nil { log.Fatal(err) } } ``` ## Error Handling in Chains Handle errors gracefully in sequential pipelines: ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type PipelineStep struct { Name string Execute func(context.Context, string) (string, error) RetryCount int SkipOnError bool } func runPipeline(ctx context.Context, steps []PipelineStep, initialInput string) (string, error) { currentInput := initialInput for i, step := range steps { fmt.Printf("\n=== Step %d: %s ===\n", i+1, step.Name) var result string var err error // Retry logic for attempt := 0; attempt <= step.RetryCount; attempt++ { if attempt > 0 { fmt.Printf("Retry attempt %d/%d\n", attempt, step.RetryCount) } result, err = step.Execute(ctx, currentInput) if err == nil { break } if attempt == step.RetryCount { if step.SkipOnError { fmt.Printf("Error (skipping): %v\n", err) result = currentInput // Use previous input break } else { return "", fmt.Errorf("step %s failed after %d retries: %w", step.Name, step.RetryCount, err) } } } currentInput = result fmt.Printf("Output: %s\n", result) } return currentInput, nil } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") steps := []PipelineStep{ { Name: "Generate Draft", RetryCount: 2, Execute: func(ctx context.Context, input string) (string, error) { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Write a draft about: %s", input), }) if err != nil { return "", err } return result.Text, nil }, }, { Name: "Add Examples", RetryCount: 1, SkipOnError: true, // Skip if this fails Execute: func(ctx context.Context, input string) (string, error) { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Add 2 examples to this text:\n\n%s", input), }) if err != nil { return "", err } return result.Text, nil }, }, { Name: "Polish", RetryCount: 2, Execute: func(ctx context.Context, input string) (string, error) { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Polish and improve this text:\n\n%s", input), }) if err != nil { return "", err } return result.Text, nil }, }, } finalResult, err := runPipeline(ctx, steps, "How to make spaghetti") if err != nil { log.Fatal(err) } fmt.Println("\n=== Final Result ===") fmt.Println(finalResult) } ``` ## Parallel Sequential Branches Execute multiple sequential chains in parallel: ```go package main import ( "context" "fmt" "log" "os" "sync" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type ChainResult struct { Name string Output string Error error } func executeChain(ctx context.Context, name string, steps []string) ChainResult { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") currentOutput := "" for i, step := range steps { var prompt string if i == 0 { prompt = step } else { prompt = fmt.Sprintf("%s\n\nPrevious output:\n%s", step, currentOutput) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return ChainResult{Name: name, Error: err} } currentOutput = result.Text } return ChainResult{Name: name, Output: currentOutput} } func main() { ctx := context.Background() chains := map[string][]string{ "Technical": { "Generate a technical article about quantum computing", "Add code examples to the article", "Write a technical summary", }, "Marketing": { "Generate marketing copy for a tech product", "Add emotional appeal", "Create a call to action", }, "Educational": { "Explain quantum computing for beginners", "Add visual descriptions", "Create quiz questions", }, } var wg sync.WaitGroup results := make(chan ChainResult, len(chains)) // Execute chains in parallel for name, steps := range chains { wg.Add(1) go func(chainName string, chainSteps []string) { defer wg.Done() fmt.Printf("Starting chain: %s\n", chainName) result := executeChain(ctx, chainName, chainSteps) results <- result }(name, steps) } // Wait and collect results go func() { wg.Wait() close(results) }() // Process results fmt.Println("\n=== Results ===") for result := range results { if result.Error != nil { log.Printf("Chain %s failed: %v", result.Name, result.Error) } else { fmt.Printf("\n--- %s Chain ---\n%s\n", result.Name, result.Output) } } } ``` ## Best Practices ### 1. Clear Step Descriptions ```go // Good - Clear purpose fmt.Println("=== Step 1: Extract Key Information ===") fmt.Println("=== Step 2: Analyze Sentiment ===") fmt.Println("=== Step 3: Generate Summary ===") ``` ### 2. Validate Intermediate Results ```go result, err := ai.GenerateText(ctx, options) if err != nil { return fmt.Errorf("step failed: %w", err) } if result.Text == "" { return fmt.Errorf("step produced empty output") } ``` ### 3. Use Context for Cancellation ```go ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute) defer cancel() // All steps respect the timeout step1, err := generateStep1(ctx, input) step2, err := generateStep2(ctx, step1) ``` ### 4. Log Progress ```go log.Printf("Pipeline started with input: %s", initialInput) log.Printf("Step 1 completed: %d tokens used", result1.Usage.GetTotalTokens()) log.Printf("Step 2 completed: %d tokens used", result2.Usage.GetTotalTokens()) log.Printf("Pipeline completed: total %d tokens", totalTokens) ``` ### 5. Consider Caching ```go // Cache intermediate results to avoid re-computation cacheKey := fmt.Sprintf("step1:%s", hash(input)) if cached, found := cache.Get(cacheKey); found { return cached, nil } result, err := executeStep(ctx, input) cache.Set(cacheKey, result, 1*time.Hour) return result, err ``` ## Next Steps - Learn about [agents](https://goaisdk.com/docs/agents/overview.md) for automated sequential workflows - Explore [workflow patterns](https://goaisdk.com/docs/agents/workflows.md) for complex multi-step processes - See [caching](https://goaisdk.com/docs/advanced/caching.md) to optimize repeated sequential operations --- # Deployment Guide > Guide to deploying Go AI SDK applications to production, covering binary compilation, Docker containerization, cloud deployment, and health checks. Canonical URL: https://goaisdk.com/docs/advanced/deployment-guide Documentation index: https://goaisdk.com/llms.txt This comprehensive guide covers deploying Go AI SDK applications to production, including binary compilation, containerization, and cloud deployment. ## Binary Compilation ### Building Production Binaries ```bash # Build for current platform go build -o ai-app main.go # Build with optimizations go build -ldflags="-s -w" -o ai-app main.go # Cross-compile for different platforms GOOS=linux GOARCH=amd64 go build -o ai-app-linux-amd64 main.go GOOS=darwin GOARCH=amd64 go build -o ai-app-darwin-amd64 main.go GOOS=windows GOARCH=amd64 go build -o ai-app-windows-amd64.exe main.go # Build with version information VERSION=$(git describe --tags --always --dirty) BUILD_TIME=$(date -u +"%Y-%m-%dT%H:%M:%SZ") go build -ldflags="-X main.Version=$VERSION -X main.BuildTime=$BUILD_TIME" -o ai-app main.go ``` ### Application Structure ```go package main import ( "context" "flag" "log" "os" "os/signal" "syscall" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) var ( Version = "dev" BuildTime = "unknown" ) type Config struct { APIKey string Model string Port int LogLevel string Environment string } func loadConfig() *Config { cfg := &Config{} flag.StringVar(&cfg.APIKey, "api-key", os.Getenv("OPENAI_API_KEY"), "OpenAI API key") flag.StringVar(&cfg.Model, "model", "gpt-5.4-mini", "Model to use") flag.IntVar(&cfg.Port, "port", 8080, "Port to listen on") flag.StringVar(&cfg.LogLevel, "log-level", "info", "Log level (debug, info, warn, error)") flag.StringVar(&cfg.Environment, "env", "production", "Environment (development, staging, production)") flag.Parse() if cfg.APIKey == "" { log.Fatal("API key is required") } return cfg } func main() { cfg := loadConfig() log.Printf("Starting AI service v%s (built %s)", Version, BuildTime) log.Printf("Environment: %s", cfg.Environment) // Setup graceful shutdown ctx, cancel := context.WithCancel(context.Background()) defer cancel() sigCh := make(chan os.Signal, 1) signal.Notify(sigCh, os.Interrupt, syscall.SIGTERM) go func() { <-sigCh log.Println("Received shutdown signal...") cancel() }() // Initialize provider provider := openai.New(openai.Config{ APIKey: cfg.APIKey, }) model, err := provider.LanguageModel(cfg.Model) if err != nil { log.Fatalf("Failed to initialize model: %v", err) } // Start application if err := run(ctx, model, cfg); err != nil { log.Fatalf("Application error: %v", err) } log.Println("Shutdown complete") } func run(ctx context.Context, model interface{}, cfg *Config) error { // Your application logic here log.Printf("Application running on port %d", cfg.Port) <-ctx.Done() return nil } ``` ## Docker Deployment ### Multi-Stage Dockerfile ```dockerfile # Build stage FROM golang:1.22-alpine AS builder # Install build dependencies RUN apk add --no-cache git make WORKDIR /app # Copy go mod files COPY go.mod go.sum ./ RUN go mod download # Copy source code COPY . . # Build binary RUN CGO_ENABLED=0 GOOS=linux go build \ -ldflags="-s -w -X main.Version=${VERSION:-dev} -X main.BuildTime=$(date -u +%Y-%m-%dT%H:%M:%SZ)" \ -o ai-app \ ./cmd/server # Runtime stage FROM alpine:3.19 # Install CA certificates for HTTPS RUN apk --no-cache add ca-certificates # Create non-root user RUN addgroup -g 1000 appuser && \ adduser -D -u 1000 -G appuser appuser WORKDIR /app # Copy binary from builder COPY --from=builder /app/ai-app . # Change ownership RUN chown -R appuser:appuser /app # Switch to non-root user USER appuser # Expose port EXPOSE 8080 # Health check HEALTHCHECK --interval=30s --timeout=3s --start-period=5s --retries=3 \ CMD ["/app/ai-app", "health"] || exit 1 # Run application ENTRYPOINT ["/app/ai-app"] CMD ["serve"] ``` ### Docker Compose ```yaml version: '3.8' services: ai-app: build: context: . dockerfile: Dockerfile args: VERSION: ${VERSION:-latest} image: ai-app:${VERSION:-latest} container_name: ai-app restart: unless-stopped ports: - "8080:8080" environment: - OPENAI_API_KEY=${OPENAI_API_KEY} - MODEL=gpt-5.4-mini - LOG_LEVEL=info - ENVIRONMENT=production env_file: - .env.production volumes: - ./data:/app/data networks: - ai-network healthcheck: test: ["CMD", "/app/ai-app", "health"] interval: 30s timeout: 10s retries: 3 start_period: 40s logging: driver: "json-file" options: max-size: "10m" max-file: "3" # Optional: Redis for caching redis: image: redis:7-alpine container_name: ai-redis restart: unless-stopped ports: - "6379:6379" volumes: - redis-data:/data networks: - ai-network networks: ai-network: driver: bridge volumes: redis-data: ``` ### Build and Run ```bash # Build image docker build -t ai-app:latest . # Run container docker run -d \ --name ai-app \ -p 8080:8080 \ -e OPENAI_API_KEY=$OPENAI_API_KEY \ ai-app:latest # Run with docker-compose docker-compose up -d # View logs docker logs -f ai-app # Stop docker-compose down ``` ## Cloud Deployment ### AWS Deployment (EC2) ```bash #!/bin/bash # deploy-aws.sh # Build binary for Linux GOOS=linux GOARCH=amd64 go build -o ai-app main.go # Create deployment package tar -czf deploy.tar.gz ai-app .env.production # Upload to EC2 scp deploy.tar.gz ec2-user@your-instance:/home/ec2-user/ # SSH and deploy ssh ec2-user@your-instance << 'EOF' cd /home/ec2-user tar -xzf deploy.tar.gz # Stop existing service sudo systemctl stop ai-app # Replace binary sudo mv ai-app /opt/ai-app/ sudo chmod +x /opt/ai-app/ai-app # Start service sudo systemctl start ai-app sudo systemctl status ai-app EOF ``` #### Systemd Service File ```ini # /etc/systemd/system/ai-app.service [Unit] Description=AI Application Service After=network.target [Service] Type=simple User=appuser Group=appuser WorkingDirectory=/opt/ai-app ExecStart=/opt/ai-app/ai-app serve Restart=always RestartSec=10 StandardOutput=journal StandardError=journal SyslogIdentifier=ai-app # Environment Environment="OPENAI_API_KEY=your-key-here" EnvironmentFile=/opt/ai-app/.env.production # Security NoNewPrivileges=true PrivateTmp=true ProtectSystem=strict ProtectHome=true ReadWritePaths=/opt/ai-app/data [Install] WantedBy=multi-user.target ``` ### AWS ECS (Fargate) ```json { "family": "ai-app", "networkMode": "awsvpc", "requiresCompatibilities": ["FARGATE"], "cpu": "256", "memory": "512", "containerDefinitions": [ { "name": "ai-app", "image": "your-account.dkr.ecr.us-east-1.amazonaws.com/ai-app:latest", "portMappings": [ { "containerPort": 8080, "protocol": "tcp" } ], "environment": [ { "name": "ENVIRONMENT", "value": "production" } ], "secrets": [ { "name": "OPENAI_API_KEY", "valueFrom": "arn:aws:secretsmanager:us-east-1:account:secret:openai-key" } ], "logConfiguration": { "logDriver": "awslogs", "options": { "awslogs-group": "/ecs/ai-app", "awslogs-region": "us-east-1", "awslogs-stream-prefix": "ecs" } }, "healthCheck": { "command": ["CMD-SHELL", "/app/ai-app health || exit 1"], "interval": 30, "timeout": 5, "retries": 3, "startPeriod": 60 } } ] } ``` ### Google Cloud Run ```bash # Build and push image gcloud builds submit --tag gcr.io/your-project/ai-app # Deploy to Cloud Run gcloud run deploy ai-app \ --image gcr.io/your-project/ai-app \ --platform managed \ --region us-central1 \ --allow-unauthenticated \ --set-env-vars "ENVIRONMENT=production" \ --set-secrets "OPENAI_API_KEY=openai-key:latest" \ --memory 512Mi \ --cpu 1 \ --min-instances 0 \ --max-instances 10 \ --concurrency 80 \ --timeout 300 ``` ### Azure Container Apps ```bash # Create resource group az group create --name ai-app-rg --location eastus # Create container app environment az containerapp env create \ --name ai-app-env \ --resource-group ai-app-rg \ --location eastus # Deploy container app az containerapp create \ --name ai-app \ --resource-group ai-app-rg \ --environment ai-app-env \ --image your-registry.azurecr.io/ai-app:latest \ --target-port 8080 \ --ingress external \ --secrets openai-key=$OPENAI_API_KEY \ --env-vars "ENVIRONMENT=production" "OPENAI_API_KEY=secretref:openai-key" \ --cpu 0.5 \ --memory 1.0Gi \ --min-replicas 1 \ --max-replicas 5 ``` ### Fly.io Deployment ```toml # fly.toml app = "ai-app" primary_region = "sjc" [build] dockerfile = "Dockerfile" [env] ENVIRONMENT = "production" MODEL = "gpt-5.4-mini" [[services]] internal_port = 8080 protocol = "tcp" [[services.ports]] handlers = ["http"] port = 80 [[services.ports]] handlers = ["tls", "http"] port = 443 [services.concurrency] type = "connections" hard_limit = 25 soft_limit = 20 [[services.tcp_checks]] interval = "15s" timeout = "2s" grace_period = "1s" restart_limit = 0 [http_service] internal_port = 8080 force_https = true auto_stop_machines = true auto_start_machines = true min_machines_running = 0 ``` ```bash # Deploy to Fly.io fly launch fly secrets set OPENAI_API_KEY=$OPENAI_API_KEY fly deploy fly status fly logs ``` ## Configuration Management ### Environment Variables ```go package config import ( "fmt" "os" "strconv" "time" ) type Config struct { // API Configuration APIKey string Model string Provider string // Server Configuration Port int ReadTimeout time.Duration WriteTimeout time.Duration // Application Configuration Environment string LogLevel string Debug bool // Rate Limiting RateLimit int RateLimitWindow time.Duration // Caching CacheEnabled bool CacheTTL time.Duration } func Load() (*Config, error) { cfg := &Config{ APIKey: getEnv("OPENAI_API_KEY", ""), Model: getEnv("MODEL", "gpt-5.4-mini"), Provider: getEnv("PROVIDER", "openai"), Port: getEnvInt("PORT", 8080), ReadTimeout: getEnvDuration("READ_TIMEOUT", 30*time.Second), WriteTimeout: getEnvDuration("WRITE_TIMEOUT", 30*time.Second), Environment: getEnv("ENVIRONMENT", "production"), LogLevel: getEnv("LOG_LEVEL", "info"), Debug: getEnvBool("DEBUG", false), RateLimit: getEnvInt("RATE_LIMIT", 100), RateLimitWindow: getEnvDuration("RATE_LIMIT_WINDOW", time.Minute), CacheEnabled: getEnvBool("CACHE_ENABLED", true), CacheTTL: getEnvDuration("CACHE_TTL", 5*time.Minute), } if err := cfg.Validate(); err != nil { return nil, err } return cfg, nil } func (c *Config) Validate() error { if c.APIKey == "" { return fmt.Errorf("API key is required") } if c.Port < 1 || c.Port > 65535 { return fmt.Errorf("invalid port: %d", c.Port) } return nil } func getEnv(key, defaultValue string) string { if value := os.Getenv(key); value != "" { return value } return defaultValue } func getEnvInt(key string, defaultValue int) int { if value := os.Getenv(key); value != "" { if intValue, err := strconv.Atoi(value); err == nil { return intValue } } return defaultValue } func getEnvBool(key string, defaultValue bool) bool { if value := os.Getenv(key); value != "" { if boolValue, err := strconv.ParseBool(value); err == nil { return boolValue } } return defaultValue } func getEnvDuration(key string, defaultValue time.Duration) time.Duration { if value := os.Getenv(key); value != "" { if duration, err := time.ParseDuration(value); err == nil { return duration } } return defaultValue } ``` ## Monitoring and Health Checks ### Health Check Endpoint ```go package main import ( "context" "encoding/json" "net/http" "time" ) type HealthStatus struct { Status string `json:"status"` Version string `json:"version"` Uptime string `json:"uptime"` Checks map[string]CheckResult `json:"checks"` } type CheckResult struct { Status string `json:"status"` Message string `json:"message,omitempty"` } var ( // Version is set via -ldflags "-X main.Version=..." at build time, // as shown in the "Application Structure" section above. Version = "dev" startTime = time.Now() ) func healthHandler(w http.ResponseWriter, r *http.Request) { ctx, cancel := context.WithTimeout(r.Context(), 5*time.Second) defer cancel() status := HealthStatus{ Status: "healthy", Version: Version, Uptime: time.Since(startTime).String(), Checks: make(map[string]CheckResult), } // Check AI provider connection if err := checkAIProvider(ctx); err != nil { status.Checks["ai_provider"] = CheckResult{ Status: "unhealthy", Message: err.Error(), } status.Status = "unhealthy" } else { status.Checks["ai_provider"] = CheckResult{Status: "healthy"} } // Check database (if applicable) // if err := checkDatabase(ctx); err != nil { ... } // Set response status code statusCode := http.StatusOK if status.Status != "healthy" { statusCode = http.StatusServiceUnavailable } w.Header().Set("Content-Type", "application/json") w.WriteHeader(statusCode) json.NewEncoder(w).Encode(status) } func checkAIProvider(ctx context.Context) error { // Test AI provider connectivity // Return error if unhealthy return nil } ``` ## Production Best Practices ### 1. Use Environment-Specific Configuration ```bash # .env.production ENVIRONMENT=production LOG_LEVEL=info DEBUG=false OPENAI_API_KEY=sk-... # .env.staging ENVIRONMENT=staging LOG_LEVEL=debug DEBUG=true ``` ### 2. Implement Graceful Shutdown ```go sigCh := make(chan os.Signal, 1) signal.Notify(sigCh, os.Interrupt, syscall.SIGTERM) <-sigCh log.Println("Shutting down gracefully...") // Stop accepting new requests // Complete in-flight requests // Close connections ``` ### 3. Add Structured Logging ```go skip-compile import "go.uber.org/zap" logger, _ := zap.NewProduction() defer logger.Sync() logger.Info("request_processed", zap.String("request_id", requestID), zap.Int("status_code", statusCode), zap.Duration("duration", duration), ) ``` ### 4. Monitor Performance ```go skip-compile import "github.com/prometheus/client_golang/prometheus" var ( requestDuration = prometheus.NewHistogramVec( prometheus.HistogramOpts{ Name: "ai_request_duration_seconds", Help: "AI request duration in seconds", }, []string{"model", "status"}, ) ) ``` ### 5. Implement Rate Limiting ```go import "golang.org/x/time/rate" limiter := rate.NewLimiter(rate.Limit(100), 10) if !limiter.Allow() { http.Error(w, "Rate limit exceeded", http.StatusTooManyRequests) return } ``` ## See Also - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [Testing](https://goaisdk.com/docs/ai-sdk-core/testing.md) - [Rate Limiting](https://goaisdk.com/docs/advanced/rate-limiting.md) - [Caching](https://goaisdk.com/docs/advanced/caching.md) --- # Memory Optimization > Explains how retention settings in the Go AI SDK reduce memory consumption by 50-80%, covering usage, streaming support, and performance impact. Canonical URL: https://goaisdk.com/docs/advanced/memory-optimization Documentation index: https://goaisdk.com/llms.txt Learn how to reduce memory consumption by 50-80% using retention settings. ## Overview The Go AI SDK includes retention settings that control what data is retained from LLM requests and responses. By excluding request and response bodies from results, you can significantly reduce memory usage while maintaining all essential metadata. ## When to Use Retention Settings ### Image Processing Workloads When sending images in prompts (vision models), the base64-encoded images can be 1-5 MB each. By excluding the request body, you avoid storing these large images in memory: ```go retention := &types.RetentionSettings{ RequestBody: types.BoolPtr(false), // Don't store the image } result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: imagePrompt, // Contains large image ExperimentalRetention: retention, }) ``` **Memory savings**: 50-80% per request ### Large Context Windows When analyzing documents, code repositories, or long conversations, the context can be 10-100 KB per request: ```go retention := &types.RetentionSettings{ RequestBody: types.BoolPtr(false), // Don't store large context ResponseBody: types.BoolPtr(false), // Don't store large responses } result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: longConversation, // Large message history ExperimentalRetention: retention, }) ``` **Memory savings**: 40-60% per request ### Long-Running Services For services that process thousands of requests, memory accumulation can lead to: - High memory usage - Frequent garbage collection - Potential out-of-memory errors Retention settings prevent memory accumulation over time. ### Privacy-Sensitive Applications Exclude sensitive data from being stored: - Confidential user inputs - Personal information - Proprietary content ```go retention := &types.RetentionSettings{ RequestBody: types.BoolPtr(false), // Don't store user data ResponseBody: types.BoolPtr(false), // Don't store AI responses } ``` ## Usage ### Basic Usage ```go import ( "context" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) // Create retention settings retention := &types.RetentionSettings{ RequestBody: types.BoolPtr(false), ResponseBody: types.BoolPtr(false), } // Use with GenerateText result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, ExperimentalRetention: retention, }) ``` ### Retention Options The `RetentionSettings` struct has two fields: ```go type RetentionSettings struct { // RequestBody controls whether to retain the request body // nil = default (retain), false = exclude, true = explicitly retain RequestBody *bool // ResponseBody controls whether to retain the response body // nil = default (retain), false = exclude, true = explicitly retain ResponseBody *bool } ``` **Helper function**: ```go types.BoolPtr(false) // Creates *bool pointing to false ``` ### Default Behavior By default (when `ExperimentalRetention` is `nil`), **all data is retained** for backwards compatibility: ```go result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, // ExperimentalRetention is nil - retains everything }) // result.RawRequest and result.RawResponse are present ``` ### Exclude Request Only Useful when you want to save the response for debugging but not the input: ```go retention := &types.RetentionSettings{ RequestBody: types.BoolPtr(false), // Exclude request // ResponseBody is nil, so response is retained } ``` ### Exclude Response Only Useful when you want to audit inputs but not outputs: ```go retention := &types.RetentionSettings{ // RequestBody is nil, so request is retained ResponseBody: types.BoolPtr(false), // Exclude response } ``` ### Exclude Both Maximum memory savings: ```go retention := &types.RetentionSettings{ RequestBody: types.BoolPtr(false), ResponseBody: types.BoolPtr(false), } ``` ## What's Still Available Even with retention disabled, you still get: - ✅ Generated text content - ✅ Token usage information (`Usage`) - ✅ Finish reason (`FinishReason`) - ✅ Tool calls and results - ✅ Warnings - ✅ All metadata Only the raw request/response bodies (`RawRequest`, `RawResponse`) are excluded. ## Streaming Support Retention settings work with `StreamText` as well: ```go stream, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: prompt, ExperimentalRetention: retention, }) ``` ## Performance Impact ### Memory Savings Typical memory reduction for different workloads: | Workload Type | Without Retention | With Retention | Savings | | ------------------- | ----------------- | -------------- | ------- | | Text-only | 0.1-0.5 MB | 0.05-0.1 MB | 40-50% | | Single image | 2-5 MB | 0.1-0.5 MB | 80-90% | | Multiple images | 10-20 MB | 0.2-0.8 MB | 90-95% | | Large documents | 1-3 MB | 0.1-0.3 MB | 70-80% | | 1000 requests (avg) | 500 MB | 100 MB | 80% | ### No Performance Overhead Retention settings have **zero performance overhead**: - No extra processing when enabled - No slowdown in generation - Only affects memory storage ### Garbage Collection Benefits Lower memory usage leads to: - Fewer GC cycles - Shorter GC pause times - Better overall application performance ## Best Practices ### 1. Enable in Production For production applications, enable retention by default: ```go // config.go var DefaultRetention = &types.RetentionSettings{ RequestBody: types.BoolPtr(false), ResponseBody: types.BoolPtr(false), } // usage result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, ExperimentalRetention: config.DefaultRetention, }) ``` ### 2. Disable in Development During development, retain data for debugging: ```go func getRetentionSettings() *types.RetentionSettings { if os.Getenv("ENV") == "production" { return &types.RetentionSettings{ RequestBody: types.BoolPtr(false), ResponseBody: types.BoolPtr(false), } } return nil // Retain everything in development } ``` ### 3. Conditional Retention Apply retention based on prompt characteristics: ```go func getRetentionForPrompt(hasImages bool) *types.RetentionSettings { if hasImages { // Images are large, exclude request return &types.RetentionSettings{ RequestBody: types.BoolPtr(false), } } return nil // Text-only prompts, no retention needed } ``` ### 4. Monitor Memory Usage Track memory usage to verify savings: ```go import "runtime" func trackMemory() { var m runtime.MemStats runtime.ReadMemStats(&m) fmt.Printf("Heap: %v MB\n", m.HeapAlloc / 1024 / 1024) fmt.Printf("Total Alloc: %v MB\n", m.TotalAlloc / 1024 / 1024) } ``` ### 5. Log Retention Policy Make retention policies visible in logs: ```go logger.Info("generating text", "prompt_size", len(prompt), "retention_enabled", retention != nil, ) ``` ## Debugging Considerations ### Trade-offs When retention is enabled, you lose access to: - Raw request payloads (for debugging API calls) - Raw response payloads (for inspecting provider responses) ### Debugging Strategies **1. Conditional Retention**: ```go retention := &types.RetentionSettings{ RequestBody: types.BoolPtr(!debugMode), ResponseBody: types.BoolPtr(!debugMode), } ``` **2. Separate Debug Requests**: ```go if needsDebugging { // Make a debug request without retention debugResult, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, // No retention - keep raw data }) logDebugInfo(debugResult) } // Normal request with retention result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, ExperimentalRetention: retention, }) ``` **3. Log Before Generation**: ```go logger.Debug("about to generate", "prompt", prompt) result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, ExperimentalRetention: retention, }) logger.Debug("generation complete", "text", result.Text) ``` ## Compliance and Privacy ### GDPR Compliance Retention settings help with GDPR compliance: - Don't store personal data longer than necessary - Reduce attack surface for data breaches - Simplify data retention policies ```go // For EU customers, always exclude bodies if customer.Region == "EU" { retention = &types.RetentionSettings{ RequestBody: types.BoolPtr(false), ResponseBody: types.BoolPtr(false), } } ``` ### Audit Trails Keep audit logs without storing full content: ```go auditLog := AuditEntry{ Timestamp: time.Now(), UserID: user.ID, Model: model.ModelID(), TokensUsed: result.Usage.GetTotalTokens(), FinishReason: result.FinishReason, // No sensitive content stored } ``` ## Examples See the [retention examples](https://github.com/digitallysavvy/go-ai/tree/main/examples/features/retention) for complete working code: - **[Basic Usage](https://github.com/digitallysavvy/go-ai/tree/main/examples/features/retention/basic)** - Simple examples of retention settings - **[Memory Benchmark](https://github.com/digitallysavvy/go-ai/tree/main/examples/features/retention/benchmark)** - Measure memory savings ## API Reference See RetentionSettings API Reference for complete type documentation. ## Related Topics - [Context Management](https://goaisdk.com/docs/advanced/prompt-engineering.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [Testing](https://goaisdk.com/docs/ai-sdk-core/testing.md) ## Summary Retention settings provide: - ✅ 50-80% memory reduction - ✅ Privacy protection - ✅ Better garbage collection - ✅ No performance overhead - ✅ Backwards compatible (opt-in) - ✅ Production-ready Enable retention settings in your production applications to reduce memory usage and improve performance. --- # MLflow Observability > Shows how to track and debug Go AI SDK operations with MLflow Tracing over OpenTelemetry, covering configuration, telemetry settings, and troubleshooting. Canonical URL: https://goaisdk.com/docs/advanced/mlflow-observability Documentation index: https://goaisdk.com/llms.txt MLflow Tracing provides automatic observability for applications built with the Go AI SDK via OpenTelemetry. When enabled, MLflow records prompts, responses, token usage, latencies, and exceptions for comprehensive experiment tracking and debugging. ## What is MLflow? [MLflow](https://mlflow.org/) is an open-source platform for managing the ML lifecycle, including experimentation, reproducibility, deployment, and a central model registry. MLflow Tracing enables tracking of LLM applications with detailed observability. ## What Gets Tracked When MLflow observability is enabled, the following information is automatically captured: - **Prompts and Messages**: Input text and conversation history - **Generated Responses**: Complete model outputs - **Token Usage**: Input, output, and total tokens (with cache details when available) - **Latencies**: Request duration and call hierarchy - **Exceptions**: Errors and stack traces - **Metadata**: Model IDs, parameters, and explicitly included context ## Quick Start ### Prerequisites 1. **Start MLflow Tracking Server:** ```bash mlflow server --backend-store-uri sqlite:///mlruns.db --port 5000 ``` For production, use a proper database backend. See [MLflow Setup Guide](https://mlflow.org/docs/latest/genai/getting-started/connect-environment/). 2. **Install Go AI SDK:** ```bash go get github.com/digitallysavvy/go-ai ``` ### Basic Setup ```go package main import ( "context" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/observability/mlflow" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/telemetry" ) func main() { // Create MLflow tracker tracker, err := mlflow.New(mlflow.Config{ TrackingURI: "http://localhost:5000", ExperimentName: "my-ai-app", ServiceName: "openai-service", Insecure: true, // For local development }) if err != nil { log.Fatal(err) } defer tracker.Shutdown(context.Background()) // Create provider and model provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Send spans to MLflow: register an integration built with the MLflow tracer telemetry.RegisterTelemetryIntegration( telemetry.NewLegacyOpenTelemetry(telemetry.LegacyOpenTelemetryOptions{Tracer: tracker.Tracer()}), ) // Configure telemetry telemetrySettings := &telemetry.Options{ IsEnabled: telemetry.Bool(true), RecordInputs: true, RecordOutputs: true, FunctionID: "text-generation", } // Generate text with telemetry result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Explain machine learning in simple terms", Telemetry: telemetrySettings, }) if err != nil { log.Fatal(err) } log.Printf("Response: %s\n", result.Text) // Flush traces to MLflow tracker.ForceFlush(context.Background()) log.Println("✅ Traces sent to MLflow at http://localhost:5000") } ``` ### View Traces 1. Open MLflow UI: `http://localhost:5000` 2. Navigate to: **Experiments** > **my-ai-app** > **Traces** 3. Click on individual traces to see detailed information ## Configuration Options ### MLflow Config ```go type Config struct { // TrackingURI is the MLflow tracking server endpoint (required) // Examples: "http://localhost:5000", "https://mlflow.example.com" TrackingURI string // ExperimentName is the MLflow experiment to log to (optional) // Default: "default" ExperimentName string // ExperimentID is the MLflow experiment ID (optional) // Takes precedence over ExperimentName if both are provided ExperimentID string // ServiceName for OpenTelemetry (optional) // Default: "go-ai-sdk" ServiceName string // Insecure controls HTTP vs HTTPS (optional) // Set to true for local development // Default: false (uses HTTPS) Insecure bool // Headers for additional authentication (optional) // Example: map[string]string{"Authorization": "Bearer token"} Headers map[string]string } ``` ### Examples **Local Development:** ```go tracker, err := mlflow.New(mlflow.Config{ TrackingURI: "http://localhost:5000", ExperimentName: "dev-experiment", Insecure: true, }) ``` **Production with HTTPS:** ```go tracker, err := mlflow.New(mlflow.Config{ TrackingURI: "https://mlflow.company.com", ExperimentID: "12345", ServiceName: "production-ai-service", Headers: map[string]string{ "Authorization": "Bearer " + os.Getenv("MLFLOW_TOKEN"), }, }) ``` **Multiple Experiments:** ```go // Experiment 1: Development devTracker, _ := mlflow.New(mlflow.Config{ TrackingURI: "http://localhost:5000", ExperimentName: "development", }) // Experiment 2: Production prodTracker, _ := mlflow.New(mlflow.Config{ TrackingURI: "https://mlflow.company.com", ExperimentName: "production", }) ``` ## Telemetry Settings Configure what gets tracked: ```go telemetrySettings := &telemetry.Options{ // Enable/disable telemetry IsEnabled: telemetry.Bool(true), // Control what gets recorded RecordInputs: true, // Record prompts and messages RecordOutputs: true, // Record generated responses // Function identifier for grouping traces FunctionID: "chatbot-generation", // Explicit context inclusion for telemetry IncludeRuntimeContext: map[string]bool{"request_id": true}, } // The MLflow tracer is configured on the registered integration // (telemetry.NewLegacyOpenTelemetry), not on telemetry.Options. ``` ### Privacy and Security Control sensitive data recording: ```go // Disable input/output recording for sensitive data telemetrySettings := &telemetry.Options{ IsEnabled: telemetry.Bool(true), RecordInputs: false, // Don't log user prompts RecordOutputs: false, // Don't log responses } // Token usage and latency still tracked ``` ## Advanced Usage ### Streaming with MLflow ```go stream, err := ai.StreamText(context.Background(), ai.StreamTextOptions{ Model: model, Prompt: "Write a long article about AI", // Telemetry automatically tracks streaming }) for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } if err := stream.Err(); err != nil { log.Printf("stream error: %v", err) } // Final token usage is tracked when stream completes ``` ### Multi-Step Tool Calling MLflow automatically tracks tool calling loops: ```go result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in San Francisco?", Tools: []types.Tool{ weatherTool, }, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, // Each tool call step is tracked as a span }) // MLflow shows: // - Initial prompt // - Model requests tool call // - Tool execution // - Final response with tool result ``` ### Custom Spans Add custom tracing for your application logic: ```go import "go.opentelemetry.io/otel/trace" func processWithTracing(ctx context.Context, tracker *mlflow.Tracker) { tracer := tracker.Tracer() ctx, span := tracer.Start(ctx, "custom-processing") defer span.End() // Your processing logic result, err := ai.GenerateText(ctx, options) // Add custom attributes span.SetAttributes( attribute.String("result_type", "summary"), attribute.Int64("result_length", int64(len(result.Text))), ) } ``` ## Best Practices ### 1. Experiment Organization ```go // Group related runs by experiment experiments := map[string]string{ "development": "dev-experiment", "staging": "staging-experiment", "production": "prod-experiment", "feature-test": "feature-xyz-experiment", } env := os.Getenv("ENVIRONMENT") tracker, _ := mlflow.New(mlflow.Config{ TrackingURI: mlflowURI, ExperimentName: experiments[env], }) ``` ### 2. Include Safe Context ```go telemetrySettings.IncludeRuntimeContext = map[string]bool{ "request_id": true, "feature_flag": true, "ab_test_group": true, } ``` ### 3. Graceful Shutdown ```go func main() { tracker, _ := mlflow.New(config) // Ensure traces are flushed on shutdown defer func() { ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() if err := tracker.ForceFlush(ctx); err != nil { log.Printf("Error flushing traces: %v", err) } if err := tracker.Shutdown(ctx); err != nil { log.Printf("Error shutting down tracker: %v", err) } }() // Your application logic } ``` ### 4. Error Tracking ```go result, err := ai.GenerateText(ctx, options) if err != nil { // Errors are automatically captured in traces log.Printf("Generation failed: %v", err) return } ``` ## Troubleshooting ### Traces Not Appearing 1. **Check MLflow server is running:** ```bash curl http://localhost:5000/api/2.0/mlflow/experiments/list ``` 2. **Verify experiment exists:** ```bash mlflow experiments list ``` 3. **Check for flush errors:** ```go if err := tracker.ForceFlush(ctx); err != nil { log.Printf("Flush error: %v", err) } ``` 4. **Enable debug logging:** ```go import "go.opentelemetry.io/otel" import "go.opentelemetry.io/otel/sdk/trace" // Add this before creating tracker otel.SetErrorHandler(otel.ErrorHandlerFunc(func(err error) { log.Printf("OTEL error: %v", err) })) ``` ### Connection Issues ```go // Ensure correct URL format tracker, err := mlflow.New(mlflow.Config{ TrackingURI: "http://localhost:5000", // Include http:// Insecure: true, // For local HTTP }) if err != nil { log.Fatalf("Failed to create tracker: %v", err) } ``` ### Network Timeouts ```go // Create context with timeout ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second) defer cancel() // Use context for operations result, err := ai.GenerateText(ctx, options) ``` ## Comparison with TypeScript SDK The Go implementation provides the same MLflow integration as the TypeScript AI SDK: **TypeScript:** ```typescript import { init } from 'mlflow-tracing'; import { generateText } from 'ai'; init(); // Auto-configures from environment const result = await generateText({ model: openai('gpt-6-astra'), prompt: 'Hello', experimental_telemetry: { isEnabled: true }, }); ``` **Go:** ```go import "github.com/digitallysavvy/go-ai/pkg/observability/mlflow" tracker, _ := mlflow.New(mlflow.Config{ TrackingURI: os.Getenv("MLFLOW_TRACKING_URI"), ExperimentName: "my-app", }) defer tracker.Shutdown(context.Background()) result, _ := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Hello", // Telemetry configured via tracer }) ``` ## See Also - [MLflow Official Documentation](https://mlflow.org/docs/latest/) - [MLflow Tracing Guide](https://mlflow.org/docs/latest/genai/tracing) - [OpenTelemetry Go SDK](https://opentelemetry.io/docs/instrumentation/go/) - [Telemetry Package Reference](https://goaisdk.com/docs/ai-sdk-core/telemetry.md) - [Example: observability-mlflow](https://github.com/digitallysavvy/go-ai/tree/main/examples/observability-mlflow) --- # Advanced > Landing page for advanced Go AI SDK topics, covering concurrency primitives, caching, rate limiting, deployment, and other production-oriented guides. Canonical URL: https://goaisdk.com/docs/advanced/ Documentation index: https://goaisdk.com/llms.txt This section covers advanced topics and concepts for the Go AI SDK. Working with LLMs often requires a different mental model compared to traditional software development, and Go's concurrency primitives provide powerful patterns for building production-ready AI applications. After reading these topics, you should have a better understanding of the paradigms behind the Go AI SDK and how to use them to build robust, scalable AI applications. ## Topics - **[Prompt Engineering](https://goaisdk.com/docs/advanced/prompt-engineering.md)** - Learn advanced techniques for crafting effective prompts. - **[Provider Architecture](https://goaisdk.com/docs/advanced/provider-architecture.md)** - Understand the provider abstraction layer: LanguageModel and ImageModel interfaces, ProviderOptions keys, implementing custom providers, middleware, and mock providers for testing. - **[Backpressure](https://goaisdk.com/docs/advanced/backpressure.md)** - Learn how the SDK handles backpressure and cancellation with Go channels and contexts. - **[Caching](https://goaisdk.com/docs/advanced/caching.md)** - Learn how to implement caching strategies to optimize performance and reduce costs. - **[Rate Limiting](https://goaisdk.com/docs/advanced/rate-limiting.md)** - Learn how to implement rate limiting for production deployments. - **[Model as Router](https://goaisdk.com/docs/advanced/model-as-router.md)** - Learn how to use a language model as an intelligent router to select the best tool or model for each request. - **[Sequential Generations](https://goaisdk.com/docs/advanced/sequential-generations.md)** - Learn how to chain multiple AI generations together for complex workflows. ## Key Concepts ### Concurrency and Control Flow Go's goroutines, channels, and contexts provide natural patterns for managing AI workloads: - **Channels** provide automatic backpressure for streaming responses - **Contexts** enable clean cancellation and timeout handling - **Select statements** allow responsive concurrent operations ### Production Patterns Building production AI applications requires: - Robust error handling and retry logic - Caching to reduce latency and costs - Rate limiting to prevent quota exhaustion - Monitoring and observability ## Next Steps - Review [AI SDK Core](https://goaisdk.com/docs/ai-sdk-core.md) for API reference - Learn about [Building Agents](https://goaisdk.com/docs/agents.md) for autonomous workflows - Explore [Foundations](https://goaisdk.com/docs/foundations.md) for core concepts --- # Agent UI Stream Helpers > API reference for pkg/agent UI stream helpers that bridge ToolLoopAgent streams to UI-transport-friendly forms, including CreateAgentUIStream. Canonical URL: https://goaisdk.com/docs/reference/ai/agent-ui-stream-helpers Documentation index: https://goaisdk.com/llms.txt These helper APIs live in `pkg/agent` and bridge `ToolLoopAgent` streams to UI transport-friendly forms. ## `CreateAgentUIStream` ```go func CreateAgentUIStream(ctx context.Context, agent *agent.ToolLoopAgent, opts agent.AgentStreamOptions) (<-chan ai.UIMessageChunk, <-chan error, error) ``` Creates UI message chunks from an agent stream. ## `CreateAgentUIStreamResponse` ```go func CreateAgentUIStreamResponse(ctx context.Context, agent *agent.ToolLoopAgent, opts agent.AgentStreamOptions) (*http.Response, error) ``` Returns a Server-Sent Events (`text/event-stream`) HTTP response for agent UI chunks. ## `PipeAgentUIStreamToResponse` ```go func PipeAgentUIStreamToResponse(ctx context.Context, agent *agent.ToolLoopAgent, opts agent.AgentStreamOptions, w io.Writer) error ``` Writes agent UI chunks to a writer as Server-Sent Events. ## Serving `useChat` from UI messages A chat frontend built with `useChat` posts its whole message history as UI messages. These helpers take those messages directly. They validate them against the agent's tools, convert them to model messages (including any tool approvals the user just gave), run the agent, and stream the reply back in the UI message stream protocol. They mirror TS `createAgentUIStream`, `createAgentUIStreamResponse` and `pipeAgentUIStreamToResponse`, which take `uiMessages`. ### `CreateAgentUIStreamFromUIMessages` ```go func CreateAgentUIStreamFromUIMessages(ctx context.Context, agent *agent.ToolLoopAgent, opts agent.CreateAgentUIStreamFromUIMessagesOptions) (<-chan ai.UIMessageChunk, <-chan error, error) ``` Returns the reply as a UI message chunk stream. `opts.UIMessages` accepts `[]ai.UIMessage`, decoded JSON, or raw JSON bytes or a string. It returns an error before streaming when the messages fail validation. ### `PipeAgentUIStreamFromUIMessagesToResponse` ```go func PipeAgentUIStreamFromUIMessagesToResponse(ctx context.Context, agent *agent.ToolLoopAgent, opts agent.CreateAgentUIStreamFromUIMessagesOptions, w io.Writer) error ``` Writes the reply to `w` as Server-Sent Events. On an `http.ResponseWriter` it sets the UI message stream headers and status first. The agent stops when writing fails, for example when the client disconnects. ```go package main import ( "encoding/json" "log" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { model, err := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}).LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } assistant := agent.NewToolLoopAgent(agent.AgentConfig{Model: model, System: "You are a helpful assistant."}) http.HandleFunc("POST /api/chat", func(w http.ResponseWriter, r *http.Request) { var body struct { Messages json.RawMessage `json:"messages"` } if err := json.NewDecoder(r.Body).Decode(&body); err != nil { http.Error(w, "invalid body", http.StatusBadRequest) return } err := agent.PipeAgentUIStreamFromUIMessagesToResponse(r.Context(), assistant, agent.CreateAgentUIStreamFromUIMessagesOptions{UIMessages: []byte(body.Messages)}, w) if err != nil { log.Printf("chat: %v", err) } }) log.Fatal(http.ListenAndServe(":8080", nil)) } ``` Messages that fail validation return an error before anything is written. To answer those with a 400, start the stream yourself and pipe it once it succeeds: ```go chunks, _, err := agent.CreateAgentUIStreamFromUIMessages(r.Context(), assistant, agent.CreateAgentUIStreamFromUIMessagesOptions{UIMessages: []byte(body.Messages)}) if err != nil { http.Error(w, err.Error(), http.StatusBadRequest) return } if err := ai.PipeUIMessageChunksToResponse(chunks, w, nil); err != nil { log.Printf("chat: %v", err) } ``` Once streaming has started, errors reach the client as `error` chunks. ### `CreateAgentUIStreamResponseFromUIMessages` ```go func CreateAgentUIStreamResponseFromUIMessages(ctx context.Context, agent *agent.ToolLoopAgent, opts agent.CreateAgentUIStreamFromUIMessagesOptions) (*http.Response, error) ``` Returns an `*http.Response` that streams the reply, with the UI message stream status and headers already set. ## Options {/* gen:fields agent.CreateAgentUIStreamFromUIMessagesOptions */} | Field | Type | Description | | --- | --- | --- | | `UIMessages` | `interface{}` | UIMessages are the input UI messages (for example []ai.UIMessage, decoded JSON, or raw JSON bytes/string). Validated against the agent's tools via ai.ValidateUIMessagesForAgent before conversion. | | `ExperimentalRefineToolInput` | `map[string]ai.ToolInputRefiner` | ExperimentalRefineToolInput reconstructs approved tool input from an approval's inputSchemaInput during validation. Optional. | | `ConvertDataPart` | `func(part ai.DataUIPart) types.ContentPart` | ConvertDataPart converts custom UI data parts to text or file model message parts. Data parts are ignored when nil or when the callback returns nil. Mirrors TS createAgentUIStream/createAgentUIStreamResponse/ pipeAgentUIStreamToResponse's `convertDataPart` option (TS #21816). | | `AgentOptions` | `AgentStreamOptions` | AgentOptions carries the remaining agent stream options (RuntimeContext, ToolsContext, CallOptions, AbortSignal-equivalents, etc). Its Prompt and Messages fields are ignored: the converted UI messages are used instead. | | `UIMessageStream` | `ai.UIMessageStreamResultOptions` | UIMessageStream carries ToUIMessageStream/CreateUIMessageStream options (message metadata, ID generation, callbacks, etc). Tools defaults to agent.Tools() when nil. OriginalMessages defaults to the validated UI messages (converted to UIMessageChunk) when nil, kept for backwards compatibility with callers that already computed their own chunks (TS `originalMessages` escape hatch). | {/* /gen:fields */} ## Tool approval signing When `agent.AgentConfig.ExperimentalToolApprovalSecret` is set, these helpers sign each approval request the agent issues and verify each approval response in the posted UI messages before a tool runs. A response with a missing or wrong signature fails with `*ai.InvalidToolApprovalSignatureError`. See [Tool approval helpers](https://goaisdk.com/docs/reference/ai/tool-approvals.md). --- # Agent > API reference for agent.Agent, the Go AI SDK interface implemented by autonomous agents that use tools and make decisions to complete tasks. Canonical URL: https://goaisdk.com/docs/reference/ai/agent Documentation index: https://goaisdk.com/llms.txt `agent.Agent` is the interface implemented by autonomous agents in the Go AI SDK. It lives in `github.com/digitallysavvy/go-ai/pkg/agent`, not `pkg/ai`. The SDK's concrete implementation is [`ToolLoopAgent`](https://goaisdk.com/docs/reference/ai/tool-loop-agent.md), created with `agent.NewToolLoopAgent`. ## Signature ```go type Agent interface { // Version returns the agent interface specification version. Version() string // ID returns the optional agent identifier. ID() string // Tools returns the tools that the agent can use. Tools() []types.Tool // Generate runs the agent with per-call options and returns the core // GenerateText result. Generate(ctx context.Context, opts AgentGenerateOptions) (*ai.GenerateTextResult, error) // Stream streams the agent with per-call options. Stream(ctx context.Context, opts AgentStreamOptions) (*ai.StreamTextResult, error) // Execute runs the agent with the given prompt and returns the final result Execute(ctx context.Context, prompt string) (*AgentResult, error) // ExecuteWithMessages runs the agent with a message history ExecuteWithMessages(ctx context.Context, messages []types.Message) (*AgentResult, error) } ``` ## Return Value ### AgentResult | Field | Type | Description | |-------|------|-------------| | Text | string | Final text output | | Output | interface{} | Parsed/structured final output when an output strategy is configured | | Steps | []types.StepResult | All steps taken by the agent | | ToolResults | []types.ToolResult | All tool results from agent execution | | Delegations | []agent.SubagentDelegation | Subagent delegations during execution | | FinishReason | types.FinishReason | Final finish reason | | StopReason | string | Reason string from the `StopCondition` that stopped the loop; empty if the agent ended naturally | | Usage | types.Usage | Total usage across all steps | | Warnings | []types.Warning | Warnings from any step | ## Examples ### Basic Agent ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { p := openai.New(openai.Config{ APIKey: "your-api-key", }) model, _ := p.LanguageModel("gpt-6-astra") searchTool := types.Tool{ Name: "search", Description: "Search the web for information", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{ "type": "string", "description": "Search query", }, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := input["query"].(string) return fmt.Sprintf("Search results for: %s", query), nil }, } var a agent.Agent = agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: []types.Tool{searchTool}, }) result, err := a.Execute(context.Background(), "Find the current weather in Tokyo") if err != nil { log.Fatal(err) } fmt.Printf("Result: %s\n", result.Text) fmt.Printf("Steps taken: %d\n", len(result.Steps)) } ``` ## Call options (generated) These tables are generated from the `agent.AgentGenerateOptions` and `agent.AgentStreamOptions` source. ### AgentGenerateOptions {/* gen:fields agent.AgentGenerateOptions */} | Field | Type | Description | | --- | --- | --- | | `Prompt` | `string` | | | `Messages` | `[]types.Message` | | | `System` | `string` | | | `Instructions` | `*string` | | | `AllowSystemInMessages` | `bool` | AllowSystemInMessages controls whether system-role messages in Messages are accepted instead of requiring Instructions/System. | | `RuntimeContext` | `interface{}` | | | `ToolsContext` | `map[string]interface{}` | | | `CallOptions` | `interface{}` | | | `Model` | `provider.LanguageModel` | | | `Tools` | `[]types.Tool` | | | `ToolChoice` | `types.ToolChoice` | | | `ToolOrder` | `[]string` | | | `ActiveTools` | `[]string` | | | `StopWhen` | `[]ai.StopCondition` | | | `MaxSteps` | `int` | | | `ExperimentalToolCallers` | `ai.ExperimentalToolCallers` | ExperimentalToolCallers configures which tools may call which other tools, and which tools stay hidden until discovered via ai.ToolSearch. Forwarded to the underlying GenerateText/StreamText call. | | `Temperature` | `*float64` | | | `MaxTokens` | `*int` | | | `TopP` | `*float64` | | | `TopK` | `*int` | | | `FrequencyPenalty` | `*float64` | | | `PresencePenalty` | `*float64` | | | `StopSequences` | `[]string` | | | `Seed` | `*int` | | | `Headers` | `map[string]string` | | | `Reasoning` | `*types.ReasoningLevel` | | | `SendReasoning` | `*bool` | | | `ProviderOptions` | `map[string]interface{}` | | | `Output` | `interface{}` | | | `Telemetry` | `*ai.TelemetrySettings` | | | `ExperimentalTelemetry` | `*ai.TelemetrySettings` | | | `SensitiveRuntimeContext` | `bool` | SensitiveRuntimeContext omits runtime context from telemetry payloads. | | `Include` | `*ai.IncludeOptions` | | | `ToolApproval` | `types.ToolApprovalConfig` | | | `Internal` | `*ai.InternalOptions` | | | `ExperimentalSandbox` | `interface{}` | | | `ExperimentalRefineToolInput` | `map[string]ai.ToolInputRefiner` | | | `ExperimentalDownload` | `ai.DownloadFunction` | | | `HarnessSession` | `interface{}` | HarnessSession is the extension point a harness.HarnessAgent (see pkg/harness) uses to receive the \*harness.AgentSession its caller created via HarnessAgent.CreateSession. Typed as interface\{\} (rather than a pkg/harness type) to avoid an import cycle: pkg/harness imports this package so HarnessAgent can satisfy the Agent interface. Other Agent implementations ignore this field. Mirrors the way TS `AgentCallParameters` carries a `session` for HarnessAgent's generate/stream (state/parity/sep_23_2026/harness.md §1.3, "the session is passed via options"). See harness.SessionFromCallOptions. | | `InstructionMessages` | `[]types.Message` | InstructionMessages supplies this call's instructions as system messages. Takes precedence over Instructions/System/Prompt when non-empty. See AgentConfig.InstructionMessages. | | `PrepareStep` | `func(ctx context.Context, step ai.PrepareStepOptions) ai.PrepareStepOptions` | PrepareStep overrides AgentConfig.PrepareStep for this call. | | `RepairToolCall` | `ai.ToolCallRepairFunction` | RepairToolCall overrides AgentConfig.RepairToolCall for this call. | | `ExperimentalRepairToolCall` | `ai.ToolCallRepairFunction` | Deprecated: use RepairToolCall. | | `OnLanguageModelCallStart` | `ai.OnLanguageModelCallStartCallback` | OnLanguageModelCallStart overrides AgentConfig.OnLanguageModelCallStart for this call. | | `ExperimentalOnLanguageModelCallStart` | `ai.OnLanguageModelCallStartCallback` | Deprecated: use OnLanguageModelCallStart. | | `OnLanguageModelCallEnd` | `ai.OnLanguageModelCallEndCallback` | OnLanguageModelCallEnd overrides AgentConfig.OnLanguageModelCallEnd for this call. | | `ExperimentalOnLanguageModelCallEnd` | `ai.OnLanguageModelCallEndCallback` | Deprecated: use OnLanguageModelCallEnd. | | `OnStart` | `func(ctx context.Context, e ai.OnStartEvent)` | | | `OnStepStart` | `func(ctx context.Context, e ai.OnStepStartEvent)` | | | `OnToolExecutionStart` | `func(ctx context.Context, e ai.OnToolCallStartEvent)` | | | `OnToolExecutionEnd` | `func(ctx context.Context, e ai.OnToolCallFinishEvent)` | | | `OnToolCallStart` | `func(ctx context.Context, e ai.OnToolCallStartEvent)` | Deprecated: use OnToolExecutionStart. | | `OnToolCallFinish` | `func(ctx context.Context, e ai.OnToolCallFinishEvent)` | Deprecated: use OnToolExecutionEnd. | | `OnStepEnd` | `func(ctx context.Context, e ai.OnStepFinishEvent)` | | | `OnStepFinish` | `func(ctx context.Context, e ai.OnStepFinishEvent)` | Deprecated: use OnStepEnd. | | `OnEnd` | `func(ctx context.Context, e ai.OnFinishEvent)` | | | `OnFinish` | `func(ctx context.Context, e ai.OnFinishEvent)` | Deprecated: use OnEnd. | {/* /gen:fields */} ### AgentStreamOptions {/* gen:fields agent.AgentStreamOptions */} | Field | Type | Description | | --- | --- | --- | | (embedded) | `AgentGenerateOptions` | Embedded. | | `OnChunk` | `func(chunk provider.StreamChunk)` | | | `InitialStreamChunks` | `[]provider.StreamChunk` | | | `ExperimentalTransform` | `[]ai.StreamTransformFunc` | ExperimentalTransform is an ordered list of transforms applied to stream chunks before they reach OnChunk and the returned StreamTextResult, forwarded to ai.StreamTextOptions.ExperimentalTransform. | {/* /gen:fields */} Set `ExperimentalToolApprovalSecret` on `agent.AgentConfig` (or per call on `agent.PrepareCallConfig`) to sign approval requests and verify approval responses. See [ToolLoopAgent](https://goaisdk.com/docs/reference/ai/tool-loop-agent.md#tool-approval-signing). ## See Also - [ToolLoopAgent](https://goaisdk.com/docs/reference/ai/tool-loop-agent.md) - The SDK's concrete `Agent` implementation - [Tool](https://goaisdk.com/docs/reference/ai/tool.md) - Tool definition - [Agents Guide](https://goaisdk.com/docs/agents/overview.md) --- # Batch API > Reference for the experimental batch functions in the Go AI SDK: start, status, results, cancel and list, with their option and result types. Canonical URL: https://goaisdk.com/docs/reference/ai/batch Documentation index: https://goaisdk.com/llms.txt Batch functions submit many requests to a provider for asynchronous processing, then poll for status and read the results. The functions are experimental and carry the `Experimental` prefix. A `Provider` field accepts a `provider.BatchV4` value, a provider that exposes `ExperimentalBatch()`, or nil. When it is nil, the batch runs through the AI Gateway. ## Functions | Function | Returns | | --- | --- | | `ai.ExperimentalStartBatch(ctx, StartBatchOptions)` | `*StartBatchResult`. Submits the batch. | | `ai.ExperimentalGetBatchStatus(ctx, GetBatchStatusOptions)` | `*Batch`. The batch with its latest status. | | `ai.ExperimentalGetBatchResults(ctx, GetBatchResultsOptions)` | `*BatchResultsStream`. Streams the terminal results. | | `ai.ExperimentalCancelBatch(ctx, CancelBatchOptions)` | `*CancelBatchResult`. Requests cancellation. | | `ai.ExperimentalListBatches(ctx, ListBatchesOptions)` | `*ListBatchesResult`. One page of batches. | ## BatchReference and Batch `ai.BatchReference` is the value to persist. It holds `Version`, `ID` and `Provider`, which is enough to check status, read results or cancel from another process. `ai.Batch` embeds `BatchReference` and `provider.BatchV4Status`. ## StartBatchOptions {/* gen:fields ai.StartBatchOptions */} | Field | Type | Description | | --- | --- | --- | | `Provider` | `interface{}` | Provider is a provider.BatchV4 instance, a provider.BatchProvider (a Provider exposing ExperimentalBatch()), or nil. When nil, batches are processed through the AI Gateway (mirrors TypeScript: `globalThis.AI_SDK_DEFAULT_PROVIDER ?? gateway`). | | `Requests` | `[]BatchRequest` | | | `ProviderOptions` | `map[string]interface{}` | | | `WebhookURL` | `string` | WebhookURL, when set, is the URL the provider notifies when the batch reaches a terminal state. Providers that do not support completion webhooks return an unsupported-functionality warning instead of erroring. | | `MaxRetries` | `*int` | | | `Headers` | `map[string]string` | | | `Timeout` | `*time.Duration` | | {/* /gen:fields */} ### StartBatchResult {/* gen:fields ai.StartBatchResult */} | Field | Type | Description | | --- | --- | --- | | (embedded) | `Batch` | Embedded. | | `Warnings` | `[]provider.BatchV4Warning` | | {/* /gen:fields */} ### BatchImageRequest An image request inside a batch. {/* gen:fields ai.BatchImageRequest */} | Field | Type | Description | | --- | --- | --- | | `ID` | `string` | | | `Model` | `string` | | | `Prompt` | `string` | | | `N` | `*int` | | | `Size` | `string` | | | `AspectRatio` | `string` | | | `Seed` | `*int` | | | `Files` | `[]provider.ImageFile` | | | `Mask` | `*provider.ImageFile` | | | `ProviderOptions` | `map[string]interface{}` | | {/* /gen:fields */} ## GetBatchStatusOptions {/* gen:fields ai.GetBatchStatusOptions */} | Field | Type | Description | | --- | --- | --- | | `Provider` | `interface{}` | | | `Batch` | `BatchReference` | | | `ProviderOptions` | `map[string]interface{}` | | | `MaxRetries` | `*int` | | | `Headers` | `map[string]string` | | | `Timeout` | `*time.Duration` | | {/* /gen:fields */} ## GetBatchResultsOptions {/* gen:fields ai.GetBatchResultsOptions */} | Field | Type | Description | | --- | --- | --- | | `Provider` | `interface{}` | | | `Batch` | `BatchReference` | | | `ProviderOptions` | `map[string]interface{}` | | | `MaxRetries` | `*int` | | | `Headers` | `map[string]string` | | | `Timeout` | `*time.Duration` | | {/* /gen:fields */} ### BatchResultsStream `Next()` returns the next `*ai.BatchItemResult` and `io.EOF` when the stream is complete. `Err()` returns the first error. `Close()` releases the stream. The shape matches `provider.TextStream`. {/* gen:fields ai.BatchItemResult */} | Field | Type | Description | | --- | --- | --- | | `Type` | `provider.BatchRequestType` | | | `ID` | `string` | | | `Status` | `provider.BatchItemStatus` | | | `Text` | `*types.GenerateResult` | Text is populated when Type is text and Status is succeeded. | | `Image` | `*types.ImageResult` | Image is populated when Type is image and Status is succeeded. | | `Error` | `*provider.BatchError` | | | `ProviderMetadata` | `map[string]interface{}` | | {/* /gen:fields */} ## CancelBatchOptions {/* gen:fields ai.CancelBatchOptions */} | Field | Type | Description | | --- | --- | --- | | `Provider` | `interface{}` | | | `Batch` | `BatchReference` | | | `ProviderOptions` | `map[string]interface{}` | | | `Headers` | `map[string]string` | | | `Timeout` | `*time.Duration` | | {/* /gen:fields */} {/* gen:fields ai.CancelBatchResult */} | Field | Type | Description | | --- | --- | --- | | `ProviderMetadata` | `map[string]interface{}` | | {/* /gen:fields */} ## ListBatchesOptions {/* gen:fields ai.ListBatchesOptions */} | Field | Type | Description | | --- | --- | --- | | `Provider` | `interface{}` | | | `ProviderOptions` | `map[string]interface{}` | | | `Limit` | `*int` | Limit is optional (nil means unset), mirroring TypeScript's `limit?: number`. An explicit 0 is forwarded to the provider rather than treated as "not set". | | `Cursor` | `string` | | | `MaxRetries` | `*int` | | | `Headers` | `map[string]string` | | | `Timeout` | `*time.Duration` | | {/* /gen:fields */} ### ListBatchesResult {/* gen:fields ai.ListBatchesResult */} | Field | Type | Description | | --- | --- | --- | | `Batches` | `[]Batch` | | | `NextCursor` | `string` | | | `ProviderMetadata` | `map[string]interface{}` | | {/* /gen:fields */} --- # Callback Events > Structured lifecycle events for observability in GenerateText, StreamText, and ToolLoopAgent Canonical URL: https://goaisdk.com/docs/reference/ai/callback-events Documentation index: https://goaisdk.com/llms.txt The Go AI SDK fires structured lifecycle events during text generation and agent execution. Each event is a typed struct delivered via callback fields on `GenerateTextOptions`, `StreamTextOptions`, and `AgentConfig`. `GenerateTextOptions`, `StreamTextOptions`, `AgentGenerateOptions`, and `AgentConfig` use the TypeScript-aligned tool callback names `OnToolExecutionStart` and `OnToolExecutionEnd`. The event structs keep their Go names, `OnToolCallStartEvent` and `OnToolCallFinishEvent`. All callbacks are **optional** and **panic-safe** — a panicking callback never aborts generation. ## Event Types ### OnStartEvent Fired once at the beginning of a `GenerateText` or `StreamText` call, before any LLM request is made. Each struct below shows selected fields; see `go doc ./pkg/ai ` for the complete list. ```go type OnStartEvent struct { CallID string // correlates this event with other events for the same call ModelProvider string ModelID string Prompt string System string // Deprecated: use Instructions Messages []types.Message Tools []types.Tool // ... generation parameters, ProviderOptions, RuntimeContext, etc. } ``` | Field | Description | |-------|-------------| | CallID | Unique ID for this generation call; correlates `OnStart`/`OnStepStart`/`OnStepFinish`/`OnFinish` events | | ModelProvider | Provider name (e.g., `"openai"`, `"anthropic"`) | | ModelID | Model identifier (e.g., `"gpt-6-astra"`) | | Prompt | The user prompt string | | System | The system prompt string (deprecated alias for `Instructions`) | | Messages | Initial conversation messages | | Tools | Tools available for the call | --- ### OnStepStartEvent Fired at the beginning of each generation step (each LLM call in a multi-step tool loop). ```go type OnStepStartEvent struct { CallID string StepNumber int // 0-indexed Messages []types.Message Tools []types.Tool // ... ToolChoice, ActiveTools, Steps (prior step results), etc. } ``` | Field | Description | |-------|-------------| | CallID | Correlates this event with the other events for this call | | StepNumber | 0-indexed step counter | | Messages | Messages sent to the LLM for this step | | Tools | Tools available for this step | --- ### OnToolCallStartEvent Fired before each individual tool is executed. Use this event with `OnToolExecutionStart`. ```go type OnToolCallStartEvent struct { CallID string ToolCallID string ToolName string ToolCall types.ToolCall Messages []types.Message // Args and StepNumber are deprecated; use ToolCall.Arguments and // correlate with step events via CallID instead. } ``` | Field | Description | |-------|-------------| | CallID | Correlates this event with the generation call that triggered it | | ToolCallID | Unique identifier for this tool call | | ToolName | Name of the tool being called | | ToolCall | The full tool call, including `Arguments` | --- ### OnToolCallFinishEvent Fired after each tool completes (success or failure). Use this event with `OnToolExecutionEnd`. ```go type OnToolCallFinishEvent struct { CallID string ToolCallID string ToolName string ToolCall types.ToolCall ToolOutput types.ToolResult // exactly one of ToolOutput.Result / ToolOutput.Error is meaningful ToolExecutionMs int64 Messages []types.Message // Result, Error, Args, StepNumber, DurationMs are deprecated aliases // for ToolOutput.Result, ToolOutput.Error, ToolCall.Arguments, // CallID-based correlation, and ToolExecutionMs respectively. } ``` | Field | Description | |-------|-------------| | ToolOutput | Tool result; `ToolOutput.Result` is non-nil on success, `ToolOutput.Error` on failure | | ToolExecutionMs | Tool execution time in milliseconds | | Result (deprecated) | Tool output on success (`nil` on failure) | | Error (deprecated) | Execution error on failure (`nil` on success) | --- ### OnStepFinishEvent Fired after each generation step completes, including after all tools for that step have been executed. ```go type OnStepFinishEvent struct { CallID string StepNumber int // 0-indexed Text string ToolCalls []types.ToolCall ToolResults []types.ToolResult FinishReason types.FinishReason Usage types.Usage Warnings []types.Warning // ... Reasoning, Sources, Files, ProviderMetadata, Request, Response, etc. } ``` | Field | Description | |-------|-------------| | StepNumber | 0-indexed step counter | | Text | LLM text output for this step | | ToolCalls | Tool calls made during this step | | ToolResults | Results from tool executions in this step | | FinishReason | Why the LLM stopped (e.g., `ToolCalls`, `Stop`) | | Usage | Token usage for this step | --- ### OnFinishEvent Fired once when generation completes (after all steps). ```go type OnFinishEvent struct { CallID string Text string ToolCalls []types.ToolCall ToolResults []types.ToolResult FinishReason types.FinishReason Steps []types.StepResult Usage types.Usage // usage of the final step TotalUsage types.Usage // sum across all steps Output interface{} // parsed structured output, if configured // ... Reasoning, Sources, Files, Warnings, ProviderMetadata, etc. } ``` | Field | Description | |-------|-------------| | Text | Final generated text | | ToolResults | Tool results from the final step | | FinishReason | Final finish reason | | Steps | All steps taken during generation | | TotalUsage | Aggregated token usage across all steps | --- ## Canonical event names Each lifecycle event has a canonical name that matches the TypeScript SDK. The older names are type aliases for the same types, so both compile. | Canonical name | Alias for | | --- | --- | | `ai.GenerateTextStartEvent` | `ai.OnStartEvent` | | `ai.GenerateTextStepStartEvent` | `ai.OnStepStartEvent` | | `ai.GenerateTextStepEndEvent` | `ai.OnStepFinishEvent` | | `ai.GenerateTextEndEvent` | `ai.OnFinishEvent` | | `ai.ToolExecutionStartEvent` | `ai.OnToolCallStartEvent` | | `ai.ToolExecutionEndEvent` | `ai.OnToolCallFinishEvent` | | `ai.GenerateObjectStartEvent` | `ai.ObjectOnStartEvent` | | `ai.GenerateObjectStepStartEvent` | `ai.ObjectOnStepStartEvent` | | `ai.GenerateObjectStepEndEvent`, `ai.ObjectOnStepEndEvent` | `ai.ObjectOnStepFinishEvent` | | `ai.GenerateObjectEndEvent` | `ai.ObjectOnFinishEvent` | | `ai.EmbedStartEvent` | `ai.EmbedOnStartEvent` | | `ai.EmbedEndEvent` | `ai.EmbedOnFinishEvent` | | `ai.RerankStartEvent` | `ai.RerankOnStartEvent` | | `ai.RerankEndEvent` | `ai.RerankOnFinishEvent` | ## Language model call events These events wrap each provider call inside a step. ### LanguageModelCallStartEvent {/* gen:fields ai.LanguageModelCallStartEvent */} | Field | Type | Description | | --- | --- | --- | | `CallID` | `string` | CallID correlates this event with the other events of the call. | | `Provider` | `string` | Provider and ModelID identify the model that is called. | | `ModelID` | `string` | | | `Instructions` | `string` | Instructions are the step's text instructions. | | `InstructionMessages` | `[]types.Message` | InstructionMessages are the step's instructions when given as system messages. | | `Messages` | `[]types.Message` | Messages are the step's input messages. | | `Tools` | `[]types.Tool` | Tools are the prepared tool definitions sent to the model. | | `MaxOutputTokens` | `*int` | Call settings. | | `Temperature` | `*float64` | | | `TopP` | `*float64` | | | `TopK` | `*int` | | | `PresencePenalty` | `*float64` | | | `FrequencyPenalty` | `*float64` | | | `StopSequences` | `[]string` | | | `Seed` | `*int` | | | `Reasoning` | `*types.ReasoningLevel` | | {/* /gen:fields */} ### LanguageModelCallEndEvent {/* gen:fields ai.LanguageModelCallEndEvent */} | Field | Type | Description | | --- | --- | --- | | `CallID` | `string` | CallID correlates this event with the other events of the call. | | `Provider` | `string` | Provider is the provider identifier. | | `ModelID` | `string` | ModelID is the model that produced the response (the response model ID when the provider reports one, otherwise the requested model). | | `FinishReason` | `types.FinishReason` | FinishReason is the unified finish reason of the model call. | | `Usage` | `types.Usage` | Usage is the token usage reported by the model call. | | `Content` | `[]types.ContentPart` | Content are the (parsed) content parts produced by the model call. | | `ResponseID` | `string` | ResponseID is the provider-returned response id. | | `ProviderMetadata` | `map[string]interface{}` | ProviderMetadata is optional provider-specific metadata. | | `Performance` | `LanguageModelCallPerformance` | Performance contains performance metrics for the model call. | {/* /gen:fields */} ## Evaluate events `ai.ExperimentalEvaluate` emits `ai.EvaluateOnStartEvent` and `ai.EvaluateOnEndEvent`. ### EvaluateOnStartEvent {/* gen:fields ai.EvaluateOnStartEvent */} | Field | Type | Description | | --- | --- | --- | | `CallID` | `string` | CallID is a unique identifier for this evaluate call. | | `OperationID` | `string` | OperationID is the canonical operation name ("ai.evaluate"). | | `RuntimeContext` | `interface{}` | RuntimeContext is the user-defined runtime context passed via options. | | `Provider` | `string` | | | `ModelID` | `string` | | | `State` | `interface{}` | | | `Questions` | `map[string]provider.EvaluationQuestion` | | | `MaxRetries` | `int` | | | `Headers` | `map[string]string` | | | `ProviderOptions` | `map[string]interface{}` | | {/* /gen:fields */} ### EvaluateOnEndEvent {/* gen:fields ai.EvaluateOnEndEvent */} | Field | Type | Description | | --- | --- | --- | | (embedded) | `EvaluateOnStartEvent` | Embedded. | | `Answers` | `map[string]provider.EvaluationAnswer` | | | `Usage` | `EvaluateUsage` | | | `Warnings` | `[]types.Warning` | | | `Rounding` | `*provider.EvaluationRounding` | | | `ProviderMetadata` | `map[string]interface{}` | | | `Response` | `EvaluateResponse` | | {/* /gen:fields */} ## Notify Utility The SDK exposes the `Notify` function used internally for safe event dispatch: ```go // Listener is a function that receives a typed event. type Listener[E any] func(ctx context.Context, event E) // Notify safely dispatches event to each listener. // Panics in any listener are recovered and do not propagate. func Notify[E any](ctx context.Context, event E, listeners ...Listener[E]) ``` `Notify` is useful for fan-out scenarios where you want to dispatch the same event to multiple independent listeners. --- ## Usage ### With GenerateText ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What is 2+2?", Tools: []types.Tool{calculatorTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, OnStart: func(ctx context.Context, e ai.OnStartEvent) { log.Printf("[start] model=%s/%s prompt=%q", e.ModelProvider, e.ModelID, e.Prompt) }, OnStepStart: func(ctx context.Context, e ai.OnStepStartEvent) { log.Printf("[step %d start] messages=%d", e.StepNumber, len(e.Messages)) }, OnToolExecutionStart: func(ctx context.Context, e ai.OnToolCallStartEvent) { log.Printf("[tool start] %s args=%v", e.ToolName, e.Args) }, OnToolExecutionEnd: func(ctx context.Context, e ai.OnToolCallFinishEvent) { if e.Error != nil { log.Printf("[tool error] %s: %v", e.ToolName, e.Error) } else { log.Printf("[tool finish] %s result=%v", e.ToolName, e.Result) } }, OnStepFinishEvent: func(ctx context.Context, e ai.OnStepFinishEvent) { log.Printf("[step %d finish] finish_reason=%s tokens=%d", e.StepNumber, e.FinishReason, safeTokens(e.Usage.TotalTokens)) }, OnFinishEvent: func(ctx context.Context, e ai.OnFinishEvent) { log.Printf("[finish] steps=%d total_tokens=%d text=%q", len(e.Steps), safeTokens(e.TotalUsage.TotalTokens), e.Text) }, }) ``` ### With AgentConfig (ToolLoopAgent) Callbacks on `AgentConfig` fire for every call to `Execute`. They are merged with any per-call callbacks: ```go agent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(10)}, OnStart: func(ctx context.Context, e ai.OnStartEvent) { log.Printf("Agent starting: model=%s", e.ModelID) }, OnToolExecutionStart: func(ctx context.Context, e ai.OnToolCallStartEvent) { log.Printf("Tool call: %s", e.ToolName) }, OnToolExecutionEnd: func(ctx context.Context, e ai.OnToolCallFinishEvent) { if e.Error != nil { log.Printf("Tool failed: %s: %v", e.ToolName, e.Error) } }, OnFinishEvent: func(ctx context.Context, e ai.OnFinishEvent) { log.Printf("Agent done: %d steps", len(e.Steps)) }, }) ``` --- ## Event Ordering For a single `GenerateText` call with one tool call, the event order is: ``` OnStart └── Step 1 ├── OnStepStart ├── OnToolExecutionStart (for each tool execution) ├── OnToolExecutionEnd (for each tool execution) └── OnStepFinish └── Step 2 ├── OnStepStart └── OnStepFinish OnFinish ``` --- ## Callback Merging in Agents When using `ToolLoopAgent`, callbacks set in `AgentConfig` are merged with any per-call callbacks. Both fire in order: **settings-level first, then call-level**. This is useful for attaching persistent observability (logging, metrics) at the agent level while allowing per-request overrides. --- ## See Also - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) — Full reference for `GenerateText` - [StreamText](https://goaisdk.com/docs/reference/ai/stream-text.md) — Full reference for `StreamText` - [ToolLoopAgent](https://goaisdk.com/docs/reference/ai/tool-loop-agent.md) — Full reference for `ToolLoopAgent` --- # Code Mode > API reference for pkg/codemode, a Go port of @ai-sdk/code-mode letting a model write JavaScript that calls host tools instead of per-call tool use. Canonical URL: https://goaisdk.com/docs/reference/ai/code-mode Documentation index: https://goaisdk.com/llms.txt `pkg/codemode` is a Go port of the TypeScript AI SDK's `@ai-sdk/code-mode` package. It lets a model write JavaScript that calls host tools programmatically instead of making one tool call per model turn ("code mode"). The model's source runs in an isolated QuickJS sandbox compiled to WebAssembly and executed with [wazero](https://wazero.io/) — no cgo required. **Experimental:** this package is experimental, mirroring the TypeScript package's own `experimental_` naming convention on every exported entry point. Its API may change in a minor version. Go names drop the `experimental_` prefix. See [Known differences: code-mode never-settling promises and in-flight limits](https://goaisdk.com/docs/migration-guides/known-differences.md#code-mode-never-settling-promises-and-in-flight-limits) for the behavior differences from the TypeScript (V8-based) runtime. ## Package ```go import "github.com/digitallysavvy/go-ai/pkg/codemode" ``` ## Entry points ### RunCodeMode ```go func RunCodeMode(ctx context.Context, input RunInput) (interface{}, error) ``` Runs code-mode JavaScript directly, without wrapping it as an AI SDK tool. The source runs as the body of an async function (top-level `await`/`return` are supported) inside a fresh QuickJS sandbox with `input.Options`' execution limits applied. Returns the sandboxed program's JSON-decoded return value on completion, a `*Interrupt` (as the returned `interface{}`, with a nil error) if execution paused on a host tool call, or an error. ```go type RunInput struct { JS string Tools ToolSet // map[string]types.Tool, reachable as tools.(input) ToolExecutionOptions *types.ToolExecutionOptions Options *Options Continuation *Continuation // resumes a prior interrupt InterruptResolution *InterruptResolution // resolution for Continuation's next pending interrupt } ``` ### CreateCodeModeTool ```go func CreateCodeModeTool(tools ToolSet, options Options) types.Tool ``` Creates an AI SDK `types.Tool` (name `"codeMode"`) that executes code-mode JavaScript in an isolated sandbox, with `tools` reachable inside the sandbox as `tools.(input)`. Drop the result straight into `GenerateTextOptions.Tools`. ### CodeModeTool (tool caller) ```go func CodeModeTool(options ToolCallerOptions) types.Tool ``` Creates a code-mode tool for use with `ai.ExperimentalToolCallers`, so other tools can invoke code-mode as a caller rather than the model calling it directly. Messages a tool caller adds to the conversation (for example the code-mode tool catalog with `ToolDiscoveryConversation`) persist across steps in `GenerateText`, `StreamText` and `ToolLoopAgent`, and `PrepareStep`'s `InitialMessages` stays the raw prompt, as in TS. ## TypeScript in model-written code Before running, code-mode strips TypeScript annotations from the source with Node `stripTypeScriptTypes` semantics, as TS code-mode does: - type annotations, generics (including generic arrow functions such as `(x: T) => x`) and default type parameters - `as` and `satisfies` expressions, and non-null assertions (`x!`) - `interface` and `type` declarations, `import type` / `export type` - class modifiers (`public`, `private`, `protected`, `readonly`, `override`, `declare`, `abstract`), `declare` statements, and a function's `this` parameter Syntax that has no type-only erasure, such as `enum`, isn't rewritten. If stripping fails for any reason, the source runs unchanged, and the JavaScript engine reports a syntax error if it isn't valid JavaScript. Plain JavaScript is never modified. ## Tool catalog order The tool catalog and prompt that code-mode generates list tools in declaration order, matching TS. With `CodeModeTool` and `ai.ExperimentalToolCallers`, that is the order of the tools passed to the call: `types.ToolCallerDefinition.Bind` and `PrepareModelMessage` receive an ordered `[]types.Tool`. `CreateCodeModeTool`, `BuildCodeModeToolDescription` and `BuildCodeModeToolCatalogMessage` take a `ToolSet` (a Go map, which has no order), so they list tools alphabetically. ## Interrupts and approvals (Code-mode part 2) Host tool calls made from inside the sandbox can pause execution (an interrupt) or require approval before running. These functions manage that lifecycle; see the package's Go doc comments for full semantics. ```go func RequestCodeModeInterrupt(payload InterruptPayload) error func IsCodeModeInterrupt(value interface{}, security ...ContinuationSecurityOptions) bool func GetCodeModeInterrupt(result interface{}, security ...ContinuationSecurityOptions) *Interrupt func UnwrapCodeModeResult(result interface{}, security ...ContinuationSecurityOptions) UnwrappedResult func ContinueCodeModeInterrupt(ctx context.Context, interrupt Interrupt, resolution interface{}, tools ToolSet, options *Options, toolExecutionOptions *types.ToolExecutionOptions) (interface{}, error) func IsCodeModeApprovalInterrupt(value interface{}, security ...ContinuationSecurityOptions) bool func ContinueCodeModeApproval(ctx context.Context, interrupt ApprovalInterrupt, response ApprovalResponse, tools ToolSet, options *Options, toolExecutionOptions *types.ToolExecutionOptions) (interface{}, error) func GetCodeModeApprovalResponse(messages []types.Message, interrupt ApprovalInterrupt) *ApprovalResponse func ToCodeModeApprovalMessages(interrupt ApprovalInterrupt) []types.Message func SetCodeModeContinuationSigningKey(key []byte, maxAge time.Duration) error ``` ### Concurrent approvals (`Promise.all`) With `Options.Approval` set to `&ApprovalOptions{Mode: ApprovalModeInterrupt}`, tool calls that need approval and start concurrently (for example inside one `Promise.all`) are batched, as in TS. Each such call stays pending until the sandbox's job queue is idle, and then all of them are returned together: - `RunCodeMode` (or the tool) returns an approval `*Interrupt` whose `Continuation.PendingInterruptions` lists every call in the batch. - Resolve the entries one at a time with `ContinueCodeModeApproval`, passing `ApprovalResponse{ApprovalID: interrupt.InterruptID, Approved: ...}`. Each call returns the next pending `*Interrupt` until the batch is complete. - No tool in the batch runs until every entry is resolved. If any entry is denied, none of them run, and the result is a `*ToolApprovalDeniedError`. Calls that don't need approval run immediately and aren't part of the batch. `MaxInFlightBridgeRequests` counts the calls in a batch. A promise that never settles while no tool call is pending fails fast with a `*ProtocolError` instead of waiting for the execution timeout. Continuations are signed (HMAC-SHA256, TS-compatible canonical envelope) so they can safely round-trip through client-visible state between requests. `SetCodeModeContinuationSigningKey` sets the process-wide signing key; a random key is generated at process start if it's never called. ## Vendored dependency `pkg/codemode` vendors a patched copy of `fastschema/qjs` v0.0.6 at `pkg/internal/third_party/qjs` (MIT license) to fix two memory-read bugs and to add a job-queue quiescence hook (used to batch concurrent approval requests) and an option that disables module imports in the engine (code-mode sandboxes run with it on, so `import()` and the native `qjs:*` modules are unavailable). `qjs.wasm` is rebuilt from pinned, checksum-verified sources by `pkg/internal/third_party/qjs/build/build.sh`. `github.com/tetratelabs/wazero` is a direct dependency. ## See Also - [Known differences from the TypeScript AI SDK](https://goaisdk.com/docs/migration-guides/known-differences.md) - [Migrating from v0.4.x to v0.5.0](https://goaisdk.com/docs/migration-guides/from-v0.4-to-v0.5.md) --- # CosineSimilarity > API reference for CosineSimilarity, a Go AI SDK function that measures similarity between two embedding vectors, returning a value from -1 to 1. Canonical URL: https://goaisdk.com/docs/reference/ai/cosine-similarity Documentation index: https://goaisdk.com/llms.txt Calculates the cosine similarity between two embedding vectors. Returns a value between -1 (opposite) and 1 (identical). ## Signature ```go func CosineSimilarity(a, b []float64) (float64, error) ``` ## Parameters | Parameter | Type | Description | |-----------|------|-------------| | a | []float64 | First embedding vector | | b | []float64 | Second embedding vector | ## Return Value Returns a float64 between -1 and 1: - 1.0: Vectors are identical - 0.0: Vectors are orthogonal (no similarity) - -1.0: Vectors are opposite ### Zero and empty vectors A zero vector (all-zero components) or an empty (`len == 0`) vector on either side returns `(0, nil)` — not an error — matching the TypeScript AI SDK. Only a length mismatch between `a` and `b` returns an error, a typed `*errors.InvalidArgumentError` with message `"Vectors must have the same length (vector1Length=…, vector2Length=…)"`. ## Examples ### Basic Similarity Calculation ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { provider := openai.New(openai.Config{ APIKey: "your-api-key", }) model, _ := provider.EmbeddingModel("text-embedding-3-small") // Embed two texts result1, _ := ai.Embed(context.Background(), ai.EmbedOptions{ Model: model, Input: "machine learning", }) result2, _ := ai.Embed(context.Background(), ai.EmbedOptions{ Model: model, Input: "artificial intelligence", }) // Calculate similarity similarity, err := ai.CosineSimilarity(result1.Embedding, result2.Embedding) if err != nil { log.Fatal(err) } fmt.Printf("Similarity: %.4f\n", similarity) } ``` ### Semantic Search ```go // Query embedding queryResult, _ := ai.Embed(ctx, ai.EmbedOptions{ Model: model, Input: "neural networks", }) // Document embeddings documents := []string{ "Deep learning uses neural networks", "The weather is sunny today", "Artificial neural networks are computational models", } // Find most similar var bestMatch string var bestScore float64 for _, doc := range documents { docResult, _ := ai.Embed(ctx, ai.EmbedOptions{ Model: model, Input: doc, }) similarity, _ := ai.CosineSimilarity(queryResult.Embedding, docResult.Embedding) if similarity > bestScore { bestScore = similarity bestMatch = doc } } fmt.Printf("Best match (%.4f): %s\n", bestScore, bestMatch) ``` ### Similarity Threshold ```go similarity, err := ai.CosineSimilarity(embedding1, embedding2) if err != nil { log.Fatal(err) } switch { case similarity > 0.9: fmt.Println("Highly similar") case similarity > 0.7: fmt.Println("Moderately similar") case similarity > 0.5: fmt.Println("Somewhat similar") default: fmt.Println("Not similar") } ``` ### Pairwise Similarity Matrix ```go embeddings := [][]float64{ embedding1, embedding2, embedding3, } fmt.Println("Similarity Matrix:") for i := 0; i < len(embeddings); i++ { for j := 0; j < len(embeddings); j++ { sim, _ := ai.CosineSimilarity(embeddings[i], embeddings[j]) fmt.Printf("%.3f ", sim) } fmt.Println() } ``` ## Error Handling `CosineSimilarity` only returns an error for a length mismatch; a zero or empty vector is not an error condition. ```go similarity, err := ai.CosineSimilarity(a, b) if err != nil { var invalidArg *providererrors.InvalidArgumentError if errors.As(err, &invalidArg) { log.Printf("Dimension mismatch: %v", invalidArg) } else { log.Println("Unknown error:", err) } return } // similarity is 0 (not an error) if either vector is all-zero or empty. ``` ## See Also - [Embed](https://goaisdk.com/docs/reference/ai/embed.md) - Generate embeddings - [EmbedMany](https://goaisdk.com/docs/reference/ai/embed-many.md) - Batch embeddings - EuclideanDistance - Alternative distance metric - DotProduct - Dot product similarity - [RAG Guide](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) --- # Dynamic Tools > API reference for dynamic tools in the Go AI SDK: registering tools at runtime with types.Tool.Type set to types.ToolTypeDynamic instead of a static, typed definition. Canonical URL: https://goaisdk.com/docs/reference/ai/dynamic-tool Documentation index: https://goaisdk.com/llms.txt The Go AI SDK does not have a `DynamicTool` helper function. Instead, a tool is marked as **dynamic** by setting its `Type` field to `types.ToolTypeDynamic`. Dynamic tools are registered and typed at runtime (for example, tools loaded from an MCP server or built from user-supplied configuration) rather than being known at compile time. The SDK tracks dynamic tool calls and results separately from statically defined ones. ## Tool.Type ```go const ( ToolTypeFunction = "function" // default: locally executed function tool ToolTypeDynamic = "dynamic" // dynamic tool whose schema can vary at runtime ToolTypeProviderDefined = "provider" // provider-defined native tool ToolTypeProviderExecuted = "provider-executed" ) ``` Setting `Type: types.ToolTypeDynamic` on a `types.Tool` tells the SDK this tool's shape isn't known statically. It otherwise has the same fields as any other tool: `Name`, `Description`, `Parameters`, and `Execute`. ## Tracking dynamic tool calls and results Results that come from dynamic tools are split out from static ones: - `ai.GenerateTextResult.DynamicToolCalls []types.ToolCall` - `ai.GenerateTextResult.DynamicToolResults []types.ToolResult` - `ai.StreamTextResult.DynamicToolCalls() []types.ToolCall` - `ai.StreamTextResult.DynamicToolResults() []types.ToolResult` - `types.ToolCall.Dynamic bool` and `types.ToolResult.Dynamic bool` ## Examples ### Registering a dynamic tool ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { p := openai.New(openai.Config{APIKey: "your-api-key"}) model, _ := p.LanguageModel("gpt-6-astra") searchTool := types.Tool{ Name: "search_db", Type: types.ToolTypeDynamic, Description: "Search the database", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{"type": "string"}, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := input["query"].(string) return []string{"result for " + query}, nil }, } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Search the database for invoices", Tools: []types.Tool{searchTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Fatal(err) } fmt.Printf("Dynamic tool calls: %d\n", len(result.DynamicToolCalls)) } ``` ### Checking whether a result came from a dynamic tool ```go for _, tr := range result.ToolResults { if tr.Dynamic { fmt.Printf("dynamic tool %s returned %v\n", tr.ToolName, tr.Result) } } ``` ## See Also - [Tool](https://goaisdk.com/docs/reference/ai/tool.md) - Static tool definition - [Tool Calling Guide](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) --- # EmbedMany > API reference for EmbedMany, a Go AI SDK function that generates vector embeddings for multiple text inputs in a single batched operation. Canonical URL: https://goaisdk.com/docs/reference/ai/embed-many Documentation index: https://goaisdk.com/llms.txt Generates vector embeddings for multiple text inputs in a single batch operation for better performance. ## Signature ```go func EmbedMany(ctx context.Context, opts EmbedManyOptions) (*EmbedManyResult, error) ``` ## Parameters ### EmbedManyOptions | Field | Type | Required | Description | |-------|------|----------|-------------| | Model | provider.EmbeddingModel | Yes | Embedding model to use | | Inputs | []string | Yes | Texts to embed (at least one required) | ## Return Value ### EmbedManyResult | Field | Type | Description | |-------|------|-------------| | Embeddings | [][]float64 | Vector embeddings (one per input) | | Usage | types.EmbeddingUsage | Token usage information | ## Examples ### Basic Batch Embedding ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { provider := openai.New(openai.Config{ APIKey: "your-api-key", }) model, err := provider.EmbeddingModel("text-embedding-3-small") if err != nil { log.Fatal(err) } texts := []string{ "The cat sat on the mat", "Dogs are loyal companions", "Birds can fly in the sky", } result, err := ai.EmbedMany(context.Background(), ai.EmbedManyOptions{ Model: model, Inputs: texts, }) if err != nil { log.Fatal(err) } fmt.Printf("Generated %d embeddings\n", len(result.Embeddings)) for i, emb := range result.Embeddings { fmt.Printf("Text %d: %d dimensions\n", i+1, len(emb)) } fmt.Printf("Total tokens used: %d\n", result.Usage.TotalTokens) } ``` ### Document Indexing ```go // Prepare documents for indexing documents := []string{ "Introduction to machine learning and AI", "Natural language processing techniques", "Computer vision and image recognition", "Reinforcement learning algorithms", "Deep neural network architectures", } // Batch embed all documents result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: model, Inputs: documents, }) if err != nil { log.Fatal(err) } // Store embeddings in database or vector store for i, embedding := range result.Embeddings { fmt.Printf("Indexing document %d with embedding of size %d\n", i, len(embedding)) // store(documents[i], embedding) } ``` ### Similarity Matrix ```go texts := []string{ "machine learning", "artificial intelligence", "deep learning", "weather forecast", } result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: model, Inputs: texts, }) if err != nil { log.Fatal(err) } // Calculate pairwise similarities fmt.Println("Similarity Matrix:") for i := 0; i < len(result.Embeddings); i++ { for j := 0; j < len(result.Embeddings); j++ { sim, _ := ai.CosineSimilarity( result.Embeddings[i], result.Embeddings[j], ) fmt.Printf("%.3f ", sim) } fmt.Println() } ``` ### Ranking by Similarity ```go // Query queryResult, err := ai.Embed(ctx, ai.EmbedOptions{ Model: model, Input: "neural networks", }) if err != nil { log.Fatal(err) } // Candidate documents candidates := []string{ "Deep learning uses neural networks", "The weather is sunny today", "Convolutional neural networks for images", "Recipe for chocolate cake", } // Batch embed candidates candidatesResult, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: model, Inputs: candidates, }) if err != nil { log.Fatal(err) } // Rank by similarity indices, scores, err := ai.RankBySimilarity( queryResult.Embedding, candidatesResult.Embeddings, ) if err != nil { log.Fatal(err) } fmt.Println("Ranked results:") for i, idx := range indices { fmt.Printf("%d. %s (score: %.4f)\n", i+1, candidates[idx], scores[i]) } ``` ### Chunked Processing for Large Datasets ```go // Process large dataset in chunks documents := loadDocuments() // Assume this loads many documents chunkSize := 100 var allEmbeddings [][]float64 for i := 0; i < len(documents); i += chunkSize { end := i + chunkSize if end > len(documents) { end = len(documents) } chunk := documents[i:end] result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: model, Inputs: chunk, }) if err != nil { log.Printf("Error embedding chunk %d: %v", i/chunkSize, err) continue } allEmbeddings = append(allEmbeddings, result.Embeddings...) fmt.Printf("Processed chunk %d/%d\n", (i/chunkSize)+1, (len(documents)+chunkSize-1)/chunkSize) } fmt.Printf("Total embeddings generated: %d\n", len(allEmbeddings)) ``` ### With Different Embedding Dimensions ```go // Different models have different dimensions providers := []struct { name string model provider.EmbeddingModel }{ {"OpenAI Small", openaiSmallModel}, {"OpenAI Large", openaiLargeModel}, {"Cohere", cohereModel}, } texts := []string{"test", "example", "sample"} for _, p := range providers { result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{ Model: p.model, Inputs: texts, }) if err != nil { log.Printf("Error with %s: %v", p.name, err) continue } fmt.Printf("%s: %d dimensions\n", p.name, len(result.Embeddings[0])) } ``` ## Error Handling ```go result, err := ai.EmbedMany(ctx, opts) if err != nil { switch { case strings.Contains(err.Error(), "model is required"): log.Println("Missing required model parameter") case strings.Contains(err.Error(), "at least one input is required"): log.Println("Inputs slice is empty") case strings.Contains(err.Error(), "batch embedding failed"): log.Println("Batch embedding error:", err) case errors.Is(err, context.DeadlineExceeded): log.Println("Request timed out") default: log.Println("Unknown error:", err) } return } // Verify we got embeddings for all inputs if len(result.Embeddings) != len(opts.Inputs) { log.Printf("Warning: expected %d embeddings, got %d", len(opts.Inputs), len(result.Embeddings)) } ``` ## See Also - [Embed](https://goaisdk.com/docs/reference/ai/embed.md) - Single text embedding - [CosineSimilarity](https://goaisdk.com/docs/reference/ai/cosine-similarity.md) - Calculate similarity - [Embedding Models Guide](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) - [RAG Guide](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) --- # Embed > API reference for Embed, the Go AI SDK function that generates a single vector embedding for one text input using an embedding model. Canonical URL: https://goaisdk.com/docs/reference/ai/embed Documentation index: https://goaisdk.com/llms.txt Generates a vector embedding for a single text input using an embedding model. ## Signature ```go func Embed(ctx context.Context, opts EmbedOptions) (*EmbedResult, error) ``` ## Parameters ### EmbedOptions | Field | Type | Required | Description | |-------|------|----------|-------------| | Model | provider.EmbeddingModel | Yes | Embedding model to use | | Input | string | Yes | Text to embed | ## Return Value ### EmbedResult | Field | Type | Description | |-------|------|-------------| | Embedding | []float64 | Vector embedding | | Usage | types.EmbeddingUsage | Token usage information | ## Examples ### Basic Embedding ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { provider := openai.New(openai.Config{ APIKey: "your-api-key", }) model, err := provider.EmbeddingModel("text-embedding-3-small") if err != nil { log.Fatal(err) } result, err := ai.Embed(context.Background(), ai.EmbedOptions{ Model: model, Input: "The quick brown fox jumps over the lazy dog", }) if err != nil { log.Fatal(err) } fmt.Printf("Embedding dimensions: %d\n", len(result.Embedding)) fmt.Printf("First 5 values: %v\n", result.Embedding[:5]) fmt.Printf("Tokens used: %d\n", result.Usage.TotalTokens) } ``` ### Semantic Search ```go // Embed query queryResult, err := ai.Embed(ctx, ai.EmbedOptions{ Model: model, Input: "machine learning algorithms", }) if err != nil { log.Fatal(err) } // Embed documents docs := []string{ "Neural networks are a type of machine learning model", "The weather today is sunny and warm", "Deep learning uses multiple layers of neural networks", } var docEmbeddings [][]float64 for _, doc := range docs { result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: model, Input: doc, }) if err != nil { log.Fatal(err) } docEmbeddings = append(docEmbeddings, result.Embedding) } // Find most similar document idx, similarity, err := ai.FindMostSimilar(queryResult.Embedding, docEmbeddings) if err != nil { log.Fatal(err) } fmt.Printf("Most similar document: %s\n", docs[idx]) fmt.Printf("Similarity score: %.4f\n", similarity) ``` ### Calculate Similarity ```go result1, err := ai.Embed(ctx, ai.EmbedOptions{ Model: model, Input: "artificial intelligence", }) if err != nil { log.Fatal(err) } result2, err := ai.Embed(ctx, ai.EmbedOptions{ Model: model, Input: "machine learning", }) if err != nil { log.Fatal(err) } similarity, err := ai.CosineSimilarity(result1.Embedding, result2.Embedding) if err != nil { log.Fatal(err) } fmt.Printf("Cosine similarity: %.4f\n", similarity) ``` ### With Different Providers ```go // OpenAI openaiProvider := openai.New(openai.Config{APIKey: "key"}) openaiModel, _ := openaiProvider.EmbeddingModel("text-embedding-3-small") result1, err := ai.Embed(ctx, ai.EmbedOptions{ Model: openaiModel, Input: "test input", }) // Cohere cohereProvider := cohere.New(cohere.Config{APIKey: "key"}) cohereModel, _ := cohereProvider.EmbeddingModel("embed-english-v3.0") result2, err := ai.Embed(ctx, ai.EmbedOptions{ Model: cohereModel, Input: "test input", }) fmt.Printf("OpenAI dimensions: %d\n", len(result1.Embedding)) fmt.Printf("Cohere dimensions: %d\n", len(result2.Embedding)) ``` ## Error Handling ```go result, err := ai.Embed(ctx, opts) if err != nil { switch { case strings.Contains(err.Error(), "model is required"): log.Println("Missing required model parameter") case strings.Contains(err.Error(), "input is required"): log.Println("Missing required input text") case strings.Contains(err.Error(), "embedding failed"): log.Println("Embedding generation error:", err) case errors.Is(err, context.DeadlineExceeded): log.Println("Request timed out") default: log.Println("Unknown error:", err) } return } ``` ## See Also - [EmbedMany](https://goaisdk.com/docs/reference/ai/embed-many.md) - Batch embedding generation - [CosineSimilarity](https://goaisdk.com/docs/reference/ai/cosine-similarity.md) - Calculate similarity between embeddings - [Embedding Models Guide](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) - [RAG Guide](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) --- # Files and skills > Reference for the Go AI SDK functions that upload, inspect, download and delete provider files and upload skills. Canonical URL: https://goaisdk.com/docs/reference/ai/files-and-skills Documentation index: https://goaisdk.com/llms.txt These functions call a provider's Files API or Skills API. The `API` field of each options struct takes a `provider.FilesAPI` (or `provider.SkillsAPI`) or a provider that exposes a `Files()` (or `Skills()`) method. ## Functions | Function | Description | | --- | --- | | `ai.UploadFile(ctx, UploadFileOptions) (*types.UploadFileResult, error)` | Uploads a file and returns a provider reference. | | `ai.GetFileMetadata(ctx, GetFileMetadataOptions) (*provider.FileMetadataResult, error)` | Returns metadata for an uploaded file. | | `ai.DownloadFile(ctx, DownloadFileOptions) (*provider.DownloadFileResult, error)` | Downloads an uploaded file. | | `ai.DeleteFile(ctx, DeleteFileOptions) (*provider.DeleteFileResult, error)` | Deletes an uploaded file. | | `ai.UploadSkill(ctx, UploadSkillOptions) (*types.UploadSkillResult, error)` | Uploads a skill. | ## UploadFileOptions {/* gen:fields ai.UploadFileOptions */} | Field | Type | Description | | --- | --- | --- | | `API` | `interface{}` | | | `Data` | `interface{}` | | | `MediaType` | `string` | | | `Filename` | `string` | | | `ProviderOptions` | `map[string]interface{}` | | | `Headers` | `map[string]string` | Headers are additional HTTP headers to send with the upload request. Only applicable for HTTP-based providers. | {/* /gen:fields */} ## GetFileMetadataOptions {/* gen:fields ai.GetFileMetadataOptions */} | Field | Type | Description | | --- | --- | --- | | `API` | `interface{}` | API is a provider.FilesAPI instance or a provider.FilesProvider (a Provider exposing a Files() method). | | `File` | `types.ProviderReference` | File is the provider reference of the file, as returned by UploadFile. | | `Headers` | `map[string]string` | | | `ProviderOptions` | `map[string]interface{}` | | {/* /gen:fields */} ## DownloadFileOptions {/* gen:fields ai.DownloadFileOptions */} | Field | Type | Description | | --- | --- | --- | | `API` | `interface{}` | API is a provider.FilesAPI instance or a provider.FilesProvider (a Provider exposing a Files() method). | | `File` | `types.ProviderReference` | File is the provider reference of the file, as returned by UploadFile. | | `Headers` | `map[string]string` | | | `ProviderOptions` | `map[string]interface{}` | | {/* /gen:fields */} ## DeleteFileOptions {/* gen:fields ai.DeleteFileOptions */} | Field | Type | Description | | --- | --- | --- | | `API` | `interface{}` | API is a provider.FilesAPI instance or a provider.FilesProvider (a Provider exposing a Files() method). | | `File` | `types.ProviderReference` | File is the provider reference of the file, as returned by UploadFile. | | `Headers` | `map[string]string` | | | `ProviderOptions` | `map[string]interface{}` | | {/* /gen:fields */} ## UploadSkillOptions {/* gen:fields ai.UploadSkillOptions */} | Field | Type | Description | | --- | --- | --- | | `API` | `interface{}` | | | `Files` | `[]UploadSkillFile` | | | `DisplayTitle` | `string` | | | `ProviderOptions` | `map[string]interface{}` | | {/* /gen:fields */} ## Errors | Variable | Returned when | | --- | --- | | `ai.ErrFilesAPINotSupported` | The provider does not expose a `Files()` method. | | `ai.ErrFileMetadataNotSupported` | The provider's Files API cannot return metadata. | | `ai.ErrFileDownloadNotSupported` | The provider's Files API cannot download files. | | `ai.ErrFileDeleteNotSupported` | The provider's Files API cannot delete files. | | `ai.ErrSkillsAPINotSupported` | The provider does not expose a `Skills()` method. | --- # GenerateID / CreateIDGenerator > API reference for GenerateID and CreateIDGenerator, Go AI SDK helpers that create unique identifier strings for tool calls and requests. Canonical URL: https://goaisdk.com/docs/reference/ai/generate-id Documentation index: https://goaisdk.com/llms.txt Generates a unique identifier string. Useful for creating tool call IDs or request IDs. ## Signature ```go func GenerateID() string ``` Not cryptographically secure (uses `math/rand`, matching the TypeScript AI SDK's `Math.random()`-based implementation). Equivalent to calling `CreateIDGenerator()()` with default options — a 16-character random string from the alphabet `0-9A-Za-z`. ## Return Value Returns a unique string identifier (16-character random alphanumeric string by default). ## CreateIDGenerator Builds a custom ID generator function, mirroring the TypeScript AI SDK's `createIdGenerator` / `generateId` from `@ai-sdk/provider-utils`. ```go func CreateIDGenerator(opts ...CreateIDGeneratorOptions) (IDGenerator, error) type CreateIDGeneratorOptions struct { Prefix string // prepended (with Separator) to every generated ID Separator string // between Prefix and the random part; default "-" Size int // length of the random part; default 16 Alphabet string // character set for the random part; default "0-9A-Za-z" } ``` `CreateIDGenerator` returns a `*errors.InvalidArgumentError` if `Prefix` is set and `Separator` appears inside `Alphabet` (this would make prefix detection ambiguous — matches the TS SDK's synchronous throw at construction time). ```go gen, err := ai.CreateIDGenerator(ai.CreateIDGeneratorOptions{ Prefix: "msg", Size: 24, }) if err != nil { log.Fatal(err) } id := gen() // e.g. "msg-aB3dE7fG9hJ2kL4mN6pQ8rS" ``` ## Examples ### Generate Request ID ```go package main import ( "fmt" "github.com/digitallysavvy/go-ai/pkg/ai" ) func main() { requestID := ai.GenerateID() fmt.Printf("Request ID: %s\n", requestID) } ``` ### Custom Tool Call ID ```go toolCall := types.ToolCall{ ID: ai.GenerateID(), ToolName: "my_tool", Arguments: map[string]interface{}{ "param": "value", }, } ``` ### Session Tracking ```go type Session struct { ID string CreatedAt time.Time } session := Session{ ID: ai.GenerateID(), CreatedAt: time.Now(), } fmt.Printf("Session: %s\n", session.ID) ``` ## See Also - [Tool](https://goaisdk.com/docs/reference/ai/tool.md) - Tool definition - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - Text generation --- # GenerateImage > API reference for GenerateImage, the Go AI SDK function that generates images from text prompts using a configured image model and provider. Canonical URL: https://goaisdk.com/docs/reference/ai/generate-image Documentation index: https://goaisdk.com/llms.txt Generates images from prompts using an image model. ## Signature ```go func GenerateImage(ctx context.Context, opts GenerateImageOptions) (*GenerateImageResult, error) ``` ## GenerateImageOptions | Field | Type | Required | Description | |---|---|---|---| | `Model` | `provider.ImageModel` | Yes | Image model implementation | | `Prompt` | `string` | Yes | Prompt text | | `N` | `*int` | No | Number of images requested | | `Size` | `string` | No | Provider image size string | | `AspectRatio` | `string` | No | Aspect ratio string | | `Seed` | `*int` | No | Deterministic seed | | `Quality` | `string` | No | Provider quality setting | | `Style` | `string` | No | Provider style setting | | `Files` | `[]provider.ImageFile` | No | Input files for edit/variation flows | | `Mask` | `*provider.ImageFile` | No | Optional mask file | | `Headers` | `map[string]string` | No | Extra request headers | | `MaxImagesPerCall` | `int` | No | Override the model's per-call image limit for automatic batching | | `MaxRetries` | `*int` | No | Maximum retries per image model call; defaults to 2, set to 0 to disable | | `ProviderOptions` | `map[string]interface{}` | No | Provider-specific options | ## GenerateImageResult | Field | Type | Description | |---|---|---| | `Image` | `types.GeneratedFile` | First generated image | | `Images` | `[]types.GeneratedFile` | All generated images | | `Warnings` | `[]types.Warning` | Provider warnings | | `Responses` | `[]*types.ResponseMetadata` | Provider response metadata entries | | `ProviderMetadata` | `map[string]interface{}` | Provider-specific metadata | | `Usage` | `types.ImageUsage` | Image generation usage metrics | `GeneratedFile.Data` contains binary image bytes. When a provider returns base64 image strings, `GenerateImage` decodes them into `GeneratedFile.Data` so JSON encoding of the result preserves the provider base64 value instead of double-encoding it. ## Example ```go result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: imageModel, Prompt: "A serene landscape at sunset", }) if err != nil { // handle } _ = os.WriteFile("output.png", result.Image.Data, 0644) ``` --- # GenerateObject > API reference for GenerateObject, the Go AI SDK function that generates structured JSON objects conforming to a schema using language models. Canonical URL: https://goaisdk.com/docs/reference/ai/generate-object Documentation index: https://goaisdk.com/llms.txt Generates structured JSON objects that conform to a specified schema using language models with structured output support. ## Signature ```go func GenerateObject(ctx context.Context, opts GenerateObjectOptions) (*GenerateObjectResult, error) ``` ## Parameters ### GenerateObjectOptions | Field | Type | Required | Description | |-------|------|----------|-------------| | Model | provider.LanguageModel | Yes | Language model to use (must support structured output) | | Prompt | string | No | Simple string prompt (alternative to Messages) | | Messages | []types.Message | No | List of conversation messages | | System | string | No | System instructions | | Schema | schema.Schema | No | Schema for output validation (required for object/array modes) | | OutputMode | ObjectOutputMode | No | Output mode: object, array, enum, or no-schema (default: object) | | EnumValues | []string | No | Enum values (required for enum mode) | | Temperature | *float64 | No | Sampling temperature (0.0 to 2.0) | | MaxTokens | *int | No | Maximum tokens to generate | | TopP | *float64 | No | Nucleus sampling parameter | | FrequencyPenalty | *float64 | No | Frequency penalty (-2.0 to 2.0) | | PresencePenalty | *float64 | No | Presence penalty (-2.0 to 2.0) | | Seed | *int | No | Random seed for reproducibility | | OnFinish | func | No | Called when generation completes | | ExperimentalContext | interface{} | No | User-defined context | ## Return Value ### GenerateObjectResult | Field | Type | Description | |-------|------|-------------| | Object | interface{} | Generated object (for object/no-schema modes) | | Array | []interface{} | Generated array (for array mode) | | EnumValue | string | Selected enum value (for enum mode) | | Text | string | Raw JSON text | | FinishReason | types.FinishReason | Why generation finished | | Usage | types.Usage | Token usage information | | Warnings | []types.Warning | Provider warnings | ## Examples ### Basic Object Generation ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/schema" ) type Person struct { Name string `json:"name"` Age int `json:"age"` City string `json:"city"` } func main() { provider := openai.New(openai.Config{ APIKey: "your-api-key", }) model, err := provider.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } // Define schema personSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "age": map[string]interface{}{"type": "integer"}, "city": map[string]interface{}{"type": "string"}, }, "required": []string{"name", "age", "city"}, }) // Generate object result, err := ai.GenerateObject(context.Background(), ai.GenerateObjectOptions{ Model: model, Prompt: "Generate information about a fictional character named Alice", Schema: personSchema, }) if err != nil { log.Fatal(err) } fmt.Printf("Result: %+v\n", result.Object) fmt.Printf("Tokens used: %d\n", result.Usage.GetTotalTokens()) } ``` ### Generate Into Struct ```go type Recipe struct { Name string `json:"name"` Ingredients []string `json:"ingredients"` Steps []string `json:"steps"` CookTime int `json:"cookTime"` } var recipe Recipe recipeSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "ingredients": map[string]interface{}{"type": "array", "items": map[string]interface{}{"type": "string"}}, "steps": map[string]interface{}{"type": "array", "items": map[string]interface{}{"type": "string"}}, "cookTime": map[string]interface{}{"type": "integer"}, }, "required": []string{"name", "ingredients", "steps", "cookTime"}, }) err := ai.GenerateObjectInto(ctx, ai.GenerateObjectOptions{ Model: model, Prompt: "Generate a recipe for chocolate chip cookies", Schema: recipeSchema, }, &recipe) if err != nil { log.Fatal(err) } fmt.Printf("Recipe: %s\n", recipe.Name) fmt.Printf("Ingredients: %v\n", recipe.Ingredients) ``` ### Array Mode ```go itemSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "price": map[string]interface{}{"type": "number"}, }, "required": []string{"name", "price"}, }) result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Prompt: "List 5 popular tech gadgets with prices", Schema: itemSchema, OutputMode: ai.ObjectModeArray, }) if err != nil { log.Fatal(err) } for i, item := range result.Array { fmt.Printf("%d. %+v\n", i+1, item) } ``` ### Enum Mode ```go result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Prompt: "Classify the sentiment: 'This product is amazing!'", OutputMode: ai.ObjectModeEnum, EnumValues: []string{"positive", "negative", "neutral"}, }) if err != nil { log.Fatal(err) } fmt.Printf("Sentiment: %s\n", result.EnumValue) ``` ### No-Schema Mode ```go result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Prompt: "Generate a JSON object with user profile data", OutputMode: ai.ObjectModeNoSchema, }) if err != nil { log.Fatal(err) } // Object is returned as map[string]interface{} fmt.Printf("Generated: %+v\n", result.Object) ``` ## Error Handling ```go result, err := ai.GenerateObject(ctx, opts) if err != nil { switch { case strings.Contains(err.Error(), "model is required"): log.Println("Missing required model parameter") case strings.Contains(err.Error(), "schema is required"): log.Println("Schema required for this output mode") case strings.Contains(err.Error(), "does not support structured output"): log.Println("Model doesn't support structured output") case strings.Contains(err.Error(), "validation failed"): log.Println("Output validation error:", err) case strings.Contains(err.Error(), "failed to parse JSON"): log.Println("Invalid JSON output:", err) default: log.Println("Unknown error:", err) } return } ``` ## Option fields (generated) This table is generated from the `ai.GenerateObjectOptions` source and lists every exported field. {/* gen:fields ai.GenerateObjectOptions */} | Field | Type | Description | | --- | --- | --- | | `Model` | `provider.LanguageModel` | Model to use for generation | | `Prompt` | `string` | Prompt can be a simple string or a list of messages | | `Messages` | `[]types.Message` | | | `System` | `string` | | | `Schema` | `schema.Schema` | Schema for the output object (not required for no-schema mode) | | `OutputMode` | `ObjectOutputMode` | Output mode - object, array, enum, or no-schema | | `EnumValues` | `[]string` | Enum values (required for enum mode) | | `Temperature` | `*float64` | Generation parameters | | `MaxTokens` | `*int` | | | `TopP` | `*float64` | | | `TopK` | `*int` | | | `FrequencyPenalty` | `*float64` | | | `PresencePenalty` | `*float64` | | | `Seed` | `*int` | | | `MaxRetries` | `int` | | | `RepairText` | `RepairTextFunc` | RepairText repairs invalid JSON or schema-invalid object output. Return nil when the output cannot be repaired. Takes precedence over ExperimentalRepairText when both are set (audit row 09a52cb). | | `ExperimentalRepairText` | `RepairTextFunc` | ExperimentalRepairText repairs invalid JSON or schema-invalid object output. Return nil when the output cannot be repaired. Deprecated: use RepairText. | | `Headers` | `map[string]string` | Additional HTTP headers sent with the request. | | `ProviderOptions` | `map[string]interface{}` | Additional provider-specific options. | | `SchemaName` | `string` | SchemaName is an optional name for the output schema. | | `SchemaDescription` | `string` | SchemaDescription is an optional description for the output schema. | | `Telemetry` | `*TelemetrySettings` | Telemetry configures observability for this operation. When both Telemetry and ExperimentalTelemetry are set, Telemetry wins. | | `ExperimentalTelemetry` | `*TelemetrySettings` | Telemetry configuration for observability. Deprecated: use Telemetry. | | `OnStart` | `func(ctx context.Context, e ObjectOnStartEvent)` | OnStart is called once before any LLM call is made. | | `ExperimentalOnStart` | `func(ctx context.Context, e ObjectOnStartEvent)` | ExperimentalOnStart is a deprecated alias for OnStart. Deprecated: use OnStart. | | `OnStepStart` | `func(ctx context.Context, e ObjectOnStepStartEvent)` | OnStepStart is called just before the provider is called. | | `ExperimentalOnStepStart` | `func(ctx context.Context, e ObjectOnStepStartEvent)` | ExperimentalOnStepStart is a deprecated alias for OnStepStart. Deprecated: use OnStepStart. | | `OnStepEnd` | `func(ctx context.Context, e ObjectOnStepFinishEvent)` | OnStepEnd is called after the provider returns, before JSON parsing. | | `OnStepFinish` | `func(ctx context.Context, e ObjectOnStepFinishEvent)` | OnStepFinish is called after the provider returns, before JSON parsing. Deprecated: use OnStepEnd. | | `OnEnd` | `func(ctx context.Context, e ObjectOnFinishEvent)` | OnEnd is called when the operation completes with a typed event. For GenerateObject, the event Error field is always nil. | | `OnFinishEvent` | `func(ctx context.Context, e ObjectOnFinishEvent)` | OnFinishEvent is a deprecated alias for OnEnd. Deprecated: use OnEnd. | | `OnFinish` | `func(ctx context.Context, result *GenerateObjectResult, userContext interface{})` | Legacy callback — kept for backward compatibility. Prefer OnEnd for structured access. | | `ExperimentalContext` | `interface{}` | ExperimentalContext allows passing custom context through generation lifecycle | {/* /gen:fields */} ## See Also - [StreamObject](https://goaisdk.com/docs/reference/ai/stream-object.md) - Streaming structured output - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - Text generation - [JSON Schema](https://goaisdk.com/docs/reference/schema/json-schema.md) - Schema definition - [Structured Output Guide](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) --- # GenerateSpeech > API reference for GenerateSpeech, the Go AI SDK function that generates speech audio from text using a configured speech model and provider. Canonical URL: https://goaisdk.com/docs/reference/ai/generate-speech Documentation index: https://goaisdk.com/llms.txt Generates speech audio from text. ## Signature ```go func GenerateSpeech(ctx context.Context, opts GenerateSpeechOptions) (*GenerateSpeechResult, error) ``` ## GenerateSpeechOptions | Field | Type | Required | Description | |---|---|---|---| | `Model` | `provider.SpeechModel` | Yes | Speech model | | `Text` | `string` | Yes | Input text | | `Voice` | `string` | No | Voice identifier | | `Speed` | `*float64` | No | Speech speed | | `OutputFormat` | `string` | No | Desired audio format, such as `mp3` or `wav` | | `Instructions` | `string` | No | Provider-dependent delivery instructions | | `Language` | `string` | No | Speech language code or provider-supported automatic language mode | | `ProviderOptions` | `map[string]interface{}` | No | Provider-specific options keyed by provider name | | `Headers` | `map[string]string` | No | Extra request headers | | `MaxRetries` | `*int` | No | Maximum retries per speech call. Defaults to 2; set to 0 to disable | | `Telemetry` | `*ai.TelemetrySettings` | No | Observability configuration for this call. `ExperimentalTelemetry` is a deprecated alias; `Telemetry` wins when both are set | ### Telemetry When a telemetry integration is registered (global via `telemetry.RegisterTelemetryIntegration`, or per-call via `Telemetry.Integrations`), `GenerateSpeech` emits an `ai.generateSpeech` root span covering the whole call (including retries), with `gen_ai.output.type: "speech"`, the input text (`ai.request.text`, input-gated), the generated audio's size/media type/format (`ai.response.audio.*`, output-gated), and the provider's reported usage flattened into numeric `ai.usage.*` / `gen_ai.usage.*` attributes (e.g. `ai.usage.characters` for a provider that reports `{"characters": N}`; see the [Telemetry guide](https://goaisdk.com/docs/ai-sdk-core/telemetry.md#speech-and-transcription) for the key-normalization rules). A failed call — including a provider response with no audio — ends the span with `codes.Error` status. ## GenerateSpeechResult | Field | Type | Description | |---|---|---| | `Audio` | `ai.GeneratedAudioFile` | Generated audio file payload | | `Warnings` | `[]types.Warning` | Provider warnings (if exposed) | | `Responses` | `[]ai.SpeechModelResponseMetadata` | Response metadata entries (`timestamp`, `modelId`, optional `headers`, optional `body`) | | `ProviderMetadata` | `map[string]interface{}` | Provider-specific metadata | ## GeneratedAudioFile ```go type GeneratedAudioFile struct { Data []byte `json:"data,omitempty"` URL string `json:"url,omitempty"` MediaType string `json:"mediaType"` Format string `json:"format"` } func (f GeneratedAudioFile) Base64() string func (f GeneratedAudioFile) Uint8Array() []byte ``` ## Deprecated Compatibility Aliases The TypeScript SDK still exports deprecated `experimental_generateSpeech` and `Experimental_SpeechResult` names. Go exposes equivalent migration aliases: ```go func ExperimentalGenerateSpeech(ctx context.Context, opts GenerateSpeechOptions) (*GenerateSpeechResult, error) type Experimental_SpeechResult = GenerateSpeechResult ``` ## Example ```go result, err := ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{ Model: speechModel, Text: "Hello from Go AI SDK", Voice: "alloy", }) if err != nil { // handle } _ = os.WriteFile("output.mp3", result.Audio.Data, 0644) ``` --- # GenerateText > API reference for GenerateText, the Go AI SDK function that generates text with support for tool calling, streaming, and multi-step conversations. Canonical URL: https://goaisdk.com/docs/reference/ai/generate-text Documentation index: https://goaisdk.com/llms.txt Generates text using a language model with support for tool calling, streaming, and multi-step conversations. ## Signature ```go func GenerateText(ctx context.Context, opts GenerateTextOptions) (*GenerateTextResult, error) ``` ## Parameters ### GenerateTextOptions `Model` is required, and so is one of `Prompt` or `Messages`. Every other field is optional. The table is generated from the source and lists every exported field. {/* gen:fields ai.GenerateTextOptions */} | Field | Type | Description | | --- | --- | --- | | `Model` | `provider.LanguageModel` | Model to use for generation | | `Prompt` | `string` | Prompt can be a simple string or a list of messages | | `Messages` | `[]types.Message` | | | `System` | `string` | | | `Instructions` | `*string` | Instructions is a TypeScript-compatible alias for System. When set, Instructions takes precedence over System. | | `InstructionMessages` | `[]types.Message` | InstructionMessages supplies the instructions as system messages (TS instructions: SystemModelMessage \| SystemModelMessage[]), preserving per-message ProviderOptions. When non-empty it takes precedence over Instructions and System. Every message must have the system role. | | `AllowSystemMessages` | `bool` | AllowSystemMessages permits system-role messages in Messages. Defaults to false; use System for system instructions unless you are intentionally passing provider-native system messages. | | `AllowSystemInMessages` | `bool` | AllowSystemInMessages is a TypeScript-compatible alias for AllowSystemMessages. Either field enables system-role messages. | | `Temperature` | `*float64` | Generation parameters | | `MaxTokens` | `*int` | | | `TopP` | `*float64` | | | `TopK` | `*int` | | | `FrequencyPenalty` | `*float64` | | | `PresencePenalty` | `*float64` | | | `StopSequences` | `[]string` | | | `Seed` | `*int` | | | `Headers` | `map[string]string` | | | `Tools` | `[]types.Tool` | Tools available for the model to call | | `ToolChoice` | `types.ToolChoice` | | | `ToolOrder` | `[]string` | ToolOrder controls the order tools are sent to providers. Listed tools are sent first in listed order; unlisted tools follow alphabetically. | | `ActiveTools` | `[]string` | ActiveTools restricts the tools available for this generation. | | `ExperimentalToolCallers` | `ExperimentalToolCallers` | ExperimentalToolCallers configures which tools may call which other tools (programmatic tool calling / code mode), and which tools stay hidden from the model until discovered via ai.ToolSearch. Mirrors the TypeScript SDK's experimental_toolCallers. | | `ToolApproval` | `types.ToolApprovalConfig` | ToolApproval configures approval handling for tool execution. | | `ExperimentalToolApprovalSecret` | `[]byte` | ExperimentalToolApprovalSecret signs server-issued approval requests so resumed approval responses can be verified before execution. | | `MaxSteps` | `*int` | MaxSteps is a convenience shorthand for StopWhen\{StepCountIs(N)\}. Deprecated: use StopWhen with StepCountIs instead. If StopWhen is set, MaxSteps is ignored. | | `StopWhen` | `[]StopCondition` | StopWhen defines conditions that terminate the tool-calling loop. Conditions are evaluated OR -- first non-empty string stops the loop. Evaluated after each step that produces tool results. | | `Timeout` | `*TimeoutConfig` | Timeout provides granular timeout controls Supports total timeout, per-step timeout, and per-chunk timeout (for streaming) | | `Output` | `interface{}` | Output specifies how to handle and parse model output Use TextOutput(), ObjectOutput(), ArrayOutput(), ChoiceOutput(), or JSONOutput() If nil, defaults to text output | | `ResponseFormat` | `*provider.ResponseFormat` | ResponseFormat (for structured output) - DEPRECATED: Use Output instead Kept for backward compatibility | | `RuntimeContext` | `interface{}` | RuntimeContext is user-defined context that flows through the conversation. This context is passed to tool execution functions, callbacks, and approval hooks. | | `SensitiveRuntimeContext` | `bool` | SensitiveRuntimeContext omits runtime context from telemetry payloads. | | `ToolsContext` | `map[string]interface{}` | ToolsContext is the per-tool context map passed to approval and execution hooks. | | `ExperimentalContext` | `interface{}` | ExperimentalContext is a deprecated alias for RuntimeContext. | | `ExperimentalRetention` | `*types.RetentionSettings` | ExperimentalRetention controls what data is retained from LLM requests/responses. Useful for reducing memory consumption with images or large contexts. Default (nil) retains everything for backwards compatibility. Example: retention := &types.RetentionSettings\{ RequestBody: types.BoolPtr(false), // Don't retain request ResponseBody: types.BoolPtr(false), // Don't retain response \} This can reduce memory consumption by 50-80% for image-heavy workloads. | | `Include` | `*IncludeOptions` | Include controls which large request/response details are retained in step results. Defaults match the TypeScript SDK: request/response bodies and request messages are excluded unless explicitly enabled. | | `ExperimentalInclude` | `*IncludeOptions` | ExperimentalInclude is a deprecated alias for Include. | | `Reasoning` | `*types.ReasoningLevel` | Reasoning controls how much thinking effort the model applies. nil means unset (use provider default). Set to types.ReasoningDefault to explicitly omit from the API request. Providers map this to their native reasoning APIs (Anthropic: thinking.budget_tokens, OpenAI: reasoning_effort, Google: thinkingConfig.thinkingBudget, Bedrock: reasoningConfig). | | `SendReasoning` | `*bool` | SendReasoning controls whether reasoning/thinking stream boundary chunks are exposed to callbacks. nil and false suppress reasoning-start/end, matching the TypeScript SDK default. | | `ProviderOptions` | `map[string]interface{}` | ProviderOptions allows passing provider-specific options Example: ProviderOptions: map[string]interface\{\}\{ "openai": map[string]interface\{\}\{ "promptCacheRetention": "24h", \}, \} | | `MaxRetries` | `*int` | MaxRetries controls transient provider call retries. For Gateway models, nil uses the TypeScript SDK default of 2 retries; set to 0 to disable. | | `ExperimentalSandbox` | `interface{}` | ExperimentalSandbox is passed through to tool execution. PrepareStep can override it for an individual step. | | `RepairToolCall` | `ToolCallRepairFunction` | RepairToolCall attempts to repair tool calls that fail to parse because the tool does not exist or its input is invalid. When it returns a call, the repaired call is parsed again; (nil, nil) keeps the call invalid. | | `ExperimentalRepairToolCall` | `ToolCallRepairFunction` | ExperimentalRepairToolCall is a deprecated alias for RepairToolCall. Deprecated: use RepairToolCall. | | `OnLanguageModelCallStart` | `OnLanguageModelCallStartCallback` | OnLanguageModelCallStart is called immediately before each provider model call begins. | | `ExperimentalOnLanguageModelCallStart` | `OnLanguageModelCallStartCallback` | ExperimentalOnLanguageModelCallStart is a deprecated alias for OnLanguageModelCallStart. Deprecated: use OnLanguageModelCallStart. | | `OnLanguageModelCallEnd` | `OnLanguageModelCallEndCallback` | OnLanguageModelCallEnd is called after each provider model response is normalized and parsed, before client-side tool execution. | | `ExperimentalOnLanguageModelCallEnd` | `OnLanguageModelCallEndCallback` | ExperimentalOnLanguageModelCallEnd is a deprecated alias for OnLanguageModelCallEnd. Deprecated: use OnLanguageModelCallEnd. | | `ExperimentalRefineToolInput` | `map[string]ToolInputRefiner` | ExperimentalRefineToolInput refines parsed tool inputs by tool name before approval, callbacks, telemetry, execution, and response messages. | | `ExperimentalDownload` | `DownloadFunction` | ExperimentalDownload downloads remote file URLs that need inline data before sending tool-result messages back to the model. Defaults to DefaultDownload, matching the TypeScript SDK's default download behavior. | | `Telemetry` | `*TelemetrySettings` | Telemetry configures observability for this operation. When both Telemetry and ExperimentalTelemetry are set, Telemetry wins. | | `ExperimentalTelemetry` | `*TelemetrySettings` | ExperimentalTelemetry enables OpenTelemetry tracing for this operation Deprecated: use Telemetry. When set, automatically records spans with prompts, responses, token usage, and latencies Example: import "github.com/digitallysavvy/go-ai/pkg/telemetry" result, err := ai.GenerateText(ctx, ai.GenerateTextOptions\{ Model: model, Prompt: "Hello", ExperimentalTelemetry: &telemetry.Options\{ IsEnabled: true, RecordInputs: true, RecordOutputs: true, \}, \}) For MLflow integration, see pkg/observability/mlflow | | `Internal` | `*InternalOptions` | Internal contains test-only ID generators matching the TypeScript _internal option. | | `PrepareStep` | `func(ctx context.Context, step PrepareStepOptions) PrepareStepOptions` | PrepareStep is called before each generation step Allows modification of options before the next step Receives the user context if ExperimentalContext is set | | `OnStepEnd` | `func(ctx context.Context, step types.StepResult, userContext interface{})` | OnStepEnd is called after each generation step completes. | | `OnStepFinish` | `func(ctx context.Context, step types.StepResult, userContext interface{})` | OnStepFinish is called after each generation step completes. Deprecated: use OnStepEnd. | | `OnEnd` | `func(ctx context.Context, result *GenerateTextResult, userContext interface{})` | OnEnd is called when generation completes. Receives the user context if RuntimeContext/ExperimentalContext is set. | | `OnFinish` | `func(ctx context.Context, result *GenerateTextResult, userContext interface{})` | OnFinish is called when generation completes. Receives the user context if RuntimeContext/ExperimentalContext is set. Deprecated: use OnEnd. | | `OnStart` | `func(ctx context.Context, e OnStartEvent)` | OnStart is called once, before the first LLM step begins. | | `OnStepStart` | `func(ctx context.Context, e OnStepStartEvent)` | OnStepStart is called at the beginning of each LLM step. | | `OnToolExecutionStart` | `func(ctx context.Context, e OnToolCallStartEvent)` | OnToolExecutionStart is called just before each tool's Execute function runs. | | `OnToolExecutionEnd` | `func(ctx context.Context, e OnToolCallFinishEvent)` | OnToolExecutionEnd is called after each tool's Execute function returns (whether the execution succeeded or failed). | | `OnToolCallStart` | `func(ctx context.Context, e OnToolCallStartEvent)` | Deprecated: use OnToolExecutionStart. | | `OnToolCallFinish` | `func(ctx context.Context, e OnToolCallFinishEvent)` | Deprecated: use OnToolExecutionEnd. | | `OnStepEndEvent` | `func(ctx context.Context, e OnStepFinishEvent)` | OnStepEndEvent is called at the end of each LLM step with a typed event. | | `OnStepFinishEvent` | `func(ctx context.Context, e OnStepFinishEvent)` | OnStepFinishEvent is called at the end of each LLM step with a typed event. Deprecated: use OnStepEndEvent. | | `OnEndEvent` | `func(ctx context.Context, e OnFinishEvent)` | OnEndEvent is called once when the entire generation completes with a typed event. | | `OnFinishEvent` | `func(ctx context.Context, e OnFinishEvent)` | OnFinishEvent is called once when the entire generation completes with a typed OnFinishEvent. Use OnEndEvent for new code. Deprecated: use OnEndEvent. | | `OnAbort` | `func(ctx context.Context, steps []types.StepResult)` | OnAbort is called when generation is aborted by context cancellation or deadline before normal completion. Deprecated: use OnAbortEvent, which also carries the call ID and abort reason. | | `OnAbortEvent` | `OnAbortCallback` | OnAbortEvent is called when generation is aborted by context cancellation or deadline before normal completion, with a GenerateTextAbortEvent carrying the call ID, completed steps, and abort reason. Takes precedence over the deprecated OnAbort. | {/* /gen:fields */} ## Return Value ### GenerateTextResult | Field | Type | Description | |-------|------|-------------| | Text | string | Generated text content | | ToolCalls | []types.ToolCall | Tool calls made during generation | | ToolResults | []types.ToolResult | Results from executed tools | | Steps | []types.StepResult | Steps taken during generation | | FinalStep | types.StepResult | Last step, including `Performance` statistics | | FinishReason | types.FinishReason | Why generation finished | | StopReason | string | Reason string from the `StopCondition` that stopped the loop; empty if no condition fired | | Usage | types.Usage | Token usage information | | ContextManagement | interface{} | Context management info (Anthropic) | | Warnings | []types.Warning | Provider warnings | | RawRequest | interface{} | Raw request for debugging | | RawResponse | interface{} | Raw response for debugging | Each `types.StepResult` includes `Performance` with `StepTimeMs`, `ResponseTimeMs`, `ToolExecutionMs`, `EffectiveOutputTokensPerSecond`, `EffectiveTotalTokensPerSecond`, and streaming-only `OutputTokensPerSecond`, `InputTokensPerSecond`, `TimeToFirstOutputMs`, and `TimeBetweenOutputChunksMs`. ## Examples ### Basic Text Generation ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { // Create provider and model provider := openai.New(openai.Config{ APIKey: "your-api-key", }) model, err := provider.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } // Generate text result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Write a haiku about Go programming", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) fmt.Printf("Tokens used: %d\n", result.Usage.GetTotalTokens()) } ``` ### With Generation Parameters ```go temperature := 0.7 maxTokens := 500 result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing", Temperature: &temperature, MaxTokens: &maxTokens, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` ### With Multi-turn Conversation ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What is the capital of France?"}, }, }, { Role: types.RoleAssistant, Content: []types.ContentPart{ types.TextContent{Text: "The capital of France is Paris."}, }, }, { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What is its population?"}, }, }, }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) ``` ### With Tool Calling ```go weatherTool := types.Tool{ Name: "get_weather", Description: "Get current weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "City name", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location := input["location"].(string) return fmt.Sprintf("Weather in %s: Sunny, 72°F", location), nil }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather like in San Francisco?", Tools: []types.Tool{weatherTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) fmt.Printf("Tool calls: %d\n", len(result.ToolCalls)) ``` ## Error Handling Common errors and how to handle them: ```go result, err := ai.GenerateText(ctx, opts) if err != nil { switch { case errors.Is(err, context.DeadlineExceeded): log.Println("Request timed out") case errors.Is(err, context.Canceled): log.Println("Request was canceled") case strings.Contains(err.Error(), "model is required"): log.Println("Missing required model parameter") case strings.Contains(err.Error(), "generation failed"): log.Println("Generation error:", err) default: log.Println("Unknown error:", err) } return } ``` ## Tool approval signing `ExperimentalToolApprovalSecret` is a byte slice. When it is set, the SDK signs every approval request it issues with HMAC-SHA256 and verifies the signature on the approval response before it runs the tool. A client that edits the tool name, the tool call ID or the tool input between the request and the response fails verification. When it is nil, requests are not signed and responses are not verified. The signing functions are `ai.SignToolApproval` and `ai.VerifyToolApprovalSignature`; see [Tool approval helpers](https://goaisdk.com/docs/reference/ai/tool-approvals.md). ## See Also - [StreamText](https://goaisdk.com/docs/reference/ai/stream-text.md) - Streaming text generation - [GenerateObject](https://goaisdk.com/docs/reference/ai/generate-object.md) - Structured object generation - [Tool Definition](https://goaisdk.com/docs/reference/ai/tool.md) - Defining custom tools - [Language Models Guide](https://goaisdk.com/docs/foundations/providers-and-models.md) --- # Harness Sandbox: Vercel > API reference for pkg/harness/sandbox/vercel, a Go port of the Vercel Sandbox harness provider Canonical URL: https://goaisdk.com/docs/reference/ai/harness-sandbox-vercel Documentation index: https://goaisdk.com/llms.txt `pkg/harness/sandbox/vercel` runs [`pkg/harness`](#harness-docs) sandboxes on [Vercel Sandbox](https://vercel.com/docs/vercel-sandbox), with network policies and template snapshots. It is a Go port of the TypeScript harness's Vercel Sandbox provider. ## Package ```go import "github.com/digitallysavvy/go-ai/pkg/harness/sandbox/vercel" ``` ## Auth options ```go type Credentials struct { Token string TeamID string ProjectID string } ``` `ResolveCredentials` mirrors the TypeScript provider's credential resolution: explicit `Token`/`TeamID`/`ProjectID` win when all three fields are set (setting only some of them is an error); otherwise `VERCEL_OIDC_TOKEN` is read from the environment and decoded as a JWT for its `owner_id`/`project_id` claims. ```go func ResolveCredentials(explicit Credentials) (Credentials, error) func HasConfiguredCredentials(explicit Credentials) bool ``` `ErrNoCredentials` is returned when neither explicit credentials nor `VERCEL_OIDC_TOKEN` are available. See [Known differences: Vercel Sandbox has no OIDC token refresh loop](https://goaisdk.com/docs/migration-guides/known-differences.md#vercel-sandbox-no-oidc-token-refresh-loop) — unlike the TypeScript provider, this port does not run a background `@vercel/oidc` refresh loop; `VERCEL_OIDC_TOKEN` is re-read from the environment on each call. ## Creating a session ```go func CreateNetworkSandboxSession(ctx context.Context, opts CreateSessionOptions) (harness.NetworkSandboxSession, error) type CreateSessionOptions struct { Credentials Credentials BaseURL string // overrides the Vercel API base URL (tests only) SandboxID string Name string Runtime string Image string Source *SnapshotSource TimeoutMs int64 Ports []int Persistent *bool NetworkPolicy *NetworkPolicy SnapshotExpiration *int64 Resources *ResourcesParams Env map[string]string Tags map[string]string Region string FailoverRegions []string KeepLastSnapshots *KeepLastSnapshotsParams Template *Template // one-time preparation recipe, cached by Template.Identity } ``` ```go session, err := vercel.CreateNetworkSandboxSession(ctx, vercel.CreateSessionOptions{ Runtime: "node22", TimeoutMs: 10 * 60 * 1000, Env: map[string]string{"NODE_ENV": "production"}, }) ``` When `Template` is set, `CreateNetworkSandboxSession` reuses (or creates and caches) a snapshot for `Template.Identity`, running `Template.Prepare` only the first time that identity is seen, then forks a live sandbox from the snapshot — this is the template-snapshot behavior mentioned in the v0.5.0 release notes. ## Resuming a session ```go func ResumeNetworkSandboxSession(ctx context.Context, opts ResumeSessionOptions) (harness.NetworkSandboxSession, error) type ResumeSessionOptions struct { Credentials Credentials BaseURL string SandboxID string } ``` ```go session, err := vercel.ResumeNetworkSandboxSession(ctx, vercel.ResumeSessionOptions{ SandboxID: existingSandboxID, }) ``` ## Wrapping an existing sandbox If you already created a `*vercel.Sandbox` through the lower-level `CreateSandbox`/`GetSandbox` client calls, wrap it instead of going through `CreateNetworkSandboxSession`: ```go func NetworkSessionFromNativeSandbox(sandbox *Sandbox) harness.NetworkSandboxSession func SessionFromNativeSandbox(sandbox *Sandbox) providerutils.SandboxSession ``` `SessionFromNativeSandbox` returns the restricted (file I/O + exec only) session interface; `NetworkSessionFromNativeSandbox` returns the full network-capable session, including `SetNetworkPolicy`, `SetRequestTransformations`, and port endpoint resolution. ## Harness docs {#harness-docs} There is no separate top-level reference page for `pkg/harness` itself yet. See the Harness section of [Migrating from the TypeScript AI SDK](https://goaisdk.com/docs/migration-guides/from-typescript-ai-sdk.md#harness) for the mapping from `@ai-sdk/harness` concepts (`HarnessV1Session`, `Agent`/`AgentSession`, adapter packages, sandbox providers) to their Go equivalents, and the [Migrating from v0.4.x to v0.5.0](https://goaisdk.com/docs/migration-guides/from-v0.4-to-v0.5.md) guide's "Provider updates to review" section for what's new this cycle. ## See Also - [Known differences from the TypeScript AI SDK](https://goaisdk.com/docs/migration-guides/known-differences.md) - [Migrating from the TypeScript AI SDK](https://goaisdk.com/docs/migration-guides/from-typescript-ai-sdk.md) --- # HasToolCall > Built-in StopCondition that stops the tool-calling loop when a specific tool is called Canonical URL: https://goaisdk.com/docs/reference/ai/has-tool-call Documentation index: https://goaisdk.com/llms.txt Returns a `StopCondition` that stops the tool-calling loop when the model calls a tool with one of the given names. Use this for **semantic completion signals** — for example, when the model is expected to call a `"finish"` or `"submit"` tool to indicate that it has completed its task. ## Signature ```go func HasToolCall(toolNames ...string) StopCondition ``` ## Parameters | Parameter | Type | Description | |-----------|------|-------------| | toolNames | ...string | Names of the tools whose invocation stops the loop | ## Return Value A `StopCondition` function. When the condition fires it returns the reason string `"tool 'toolName' was called"`; otherwise it returns `""`. ```go type StopCondition func(state StopConditionState) string ``` ## Behavior `HasToolCall` inspects the **most recent step's tool calls**. If any of the given `toolNames` appears in them the condition fires. It returns `""` when there are no completed steps yet (safe to evaluate on step zero). ## Examples ### Semantic completion signal Define a `finish` tool that the model calls when it is satisfied with its research. Use `HasToolCall("finish")` to stop the loop the moment it fires. Pair it with `IsStepCount` as a hard-limit fallback in case the model never calls the tool. ```go finishTool := types.Tool{ Name: "finish", Description: "Signal that the task is complete. Provide a summary of your findings.", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "summary": map[string]interface{}{ "type": "string", "description": "Final summary of the research", }, }, "required": []string{"summary"}, }, Execute: func(ctx context.Context, args map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{"status": "complete", "summary": args["summary"]}, nil }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: `Research Go concurrency patterns. When you are satisfied with your findings, call the finish tool.`, Tools: []types.Tool{searchTool, finishTool}, StopWhen: []ai.StopCondition{ ai.HasToolCall("finish"), // stop on semantic completion ai.IsStepCount(10), // safety ceiling }, }) if err != nil { log.Fatal(err) } fmt.Printf("StopReason: %s\n", result.StopReason) // → "StopReason: tool 'finish' was called" (if model finished early) // → "StopReason: maximum number of steps (10) reached" (if it hit the ceiling) ``` ### With ToolLoopAgent ```go myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: []types.Tool{searchTool, submitTool}, StopWhen: []ai.StopCondition{ ai.HasToolCall("submit"), ai.IsStepCount(20), }, }) result, err := myAgent.Execute(ctx, "Complete the report and submit it") fmt.Printf("StopReason: %s\n", result.StopReason) ``` ### Condition ordering and side effects `EvaluateStopConditions` runs **all** conditions before returning the first match. This means every condition in `StopWhen` always executes, regardless of order. Side-effectful conditions (e.g. logging, metric recording) can therefore be placed anywhere and are guaranteed to run. ```go // Both conditions always run; the first non-empty reason is returned. StopWhen: []ai.StopCondition{ ai.HasToolCall("finish"), // fires → "tool 'finish' was called" ai.IsStepCount(20), // also evaluated, but result ignored if finish fired first }, ``` ## See Also - [IsStepCount](https://goaisdk.com/docs/reference/ai/step-count-is.md) — stop after n steps (safety ceiling pattern) - [Loop Control](https://goaisdk.com/docs/agents/loop-control.md) — full guide with examples - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) — `GenerateTextOptions.StopWhen` field - [ToolLoopAgent](https://goaisdk.com/docs/reference/ai/tool-loop-agent.md) — `AgentConfig.StopWhen` field - [Stop-When Example](https://github.com/digitallysavvy/go-ai/blob/main/examples/agents/callbacks/stop-when/main.go) — runnable example --- # LlamaIndex Adapter > API reference for pkg/llamaindex, a Go port of @ai-sdk/llamaindex that adapts LlamaIndex chat-engine responses into AI SDK UI message stream chunks. Canonical URL: https://goaisdk.com/docs/reference/ai/llamaindex Documentation index: https://goaisdk.com/llms.txt `pkg/llamaindex` adapts a LlamaIndex chat-engine response stream into AI SDK UI message stream chunks. It is a Go port of `@ai-sdk/llamaindex`. The TypeScript package has no dependency on the real `llamaindex` npm package: its `EngineResponse` is a locally-defined `{ delta: string }` shape, and the adapter only ever reads `.delta` off of it. Go has no LlamaIndex SDK to bind against either — this package defines the same minimal input contract. You bridge whatever LlamaIndex client you use (HTTP, gRPC, a wrapped Python subprocess, etc.) into a `<-chan EngineResponse` of delta chunks; the package does not talk to LlamaIndex itself. ## Package ```go import "github.com/digitallysavvy/go-ai/pkg/llamaindex" ``` ## EngineResponse ```go type EngineResponse struct { Delta string } ``` The minimal Go analog of a LlamaIndex chat-engine streamed response chunk. ## StreamCallbacks ```go type StreamCallbacks struct { // OnStart is called once, before any chunk is processed. OnStart func() error // OnToken is called for every delta chunk, including empty ones // produced by leading-whitespace trimming. OnToken func(token string) error // OnText is called for every delta chunk, identically to OnToken. OnText func(text string) error // OnFinal is called once, with the full concatenation of every // (possibly trimmed) delta chunk, after the input stream ends and // before the "text-end" chunk is emitted. OnFinal func(completion string) error } ``` ## ToUIMessageStream ```go func ToUIMessageStream( ctx context.Context, stream <-chan EngineResponse, callbacks *StreamCallbacks, ) (<-chan ai.UIMessageChunk, <-chan error) ``` Converts a LlamaIndex chat-engine response stream into AI SDK UI message stream chunks: a single text part (id `"1"`) spanning `text-start` → `text-delta` (one or more) → `text-end`. Leading whitespace is trimmed from the stream exactly once: every delta is trimmed until the first delta that is non-empty *after* trimming is seen; after that, deltas pass through unmodified, even if empty. `callbacks` may be `nil`. The returned error channel receives at most one error: a callback error (from `OnStart`/`OnToken`/`OnText`/`OnFinal`) or `ctx.Err()` if `ctx` is cancelled mid-stream. Both channels are closed when the stream is fully drained. ### Example ```go package main import ( "context" "log" "github.com/digitallysavvy/go-ai/pkg/llamaindex" ) func consumeEngineStream(ctx context.Context, engineDeltas <-chan llamaindex.EngineResponse) { chunks, errs := llamaindex.ToUIMessageStream(ctx, engineDeltas, &llamaindex.StreamCallbacks{ OnFinal: func(completion string) error { log.Printf("final completion: %s", completion) return nil }, }) for chunk := range chunks { // chunk is an ai.UIMessageChunk (map[string]interface{}) with // "type": "text-start" | "text-delta" | "text-end". _ = chunk } if err := <-errs; err != nil { log.Printf("stream error: %v", err) } } ``` To serve these chunks over HTTP in the AI SDK's UI message stream protocol (so a TS `useChat` frontend can consume them directly), merge the returned channel into a stream built with `ai.CreateUIMessageStreamWithOptions` via `UIMessageStreamWriter.Merge`, the same integration point used for any non-`StreamText` chunk producer. See [UI message stream helpers](https://goaisdk.com/docs/reference/ai/stream-transport-helpers.md). ## See Also - [UI message stream helpers](https://goaisdk.com/docs/reference/ai/stream-transport-helpers.md) - [Migrating from v0.4.x to v0.5.0](https://goaisdk.com/docs/migration-guides/from-v0.4-to-v0.5.md) --- # Middleware Aliases > Reference for the middleware type aliases, wrapper functions, and built-in middleware constructors that pkg/ai re-exports for discoverability. Canonical URL: https://goaisdk.com/docs/reference/ai/middleware-aliases Documentation index: https://goaisdk.com/llms.txt The `pkg/ai` package re-exports middleware APIs for parity-oriented discoverability. ## Type aliases - `ai.LanguageModelMiddleware` - `ai.EmbeddingModelMiddleware` - `ai.ImageModelMiddleware` - `ai.ExtractReasoningOptions` - `ai.ExtractJSONOptions` - `ai.AddToolInputExamplesOptions` ## Wrapper functions ```go func WrapLanguageModel(model provider.LanguageModel, middleware []*ai.LanguageModelMiddleware, modelID, providerID *string) provider.LanguageModel func WrapEmbeddingModel(model provider.EmbeddingModel, middleware []*ai.EmbeddingModelMiddleware, modelID, providerID *string) provider.EmbeddingModel func WrapImageModel(model provider.ImageModel, middleware []*ai.ImageModelMiddleware, modelID, providerID *string) provider.ImageModel func WrapProvider(p provider.Provider, languageModelMiddleware []*ai.LanguageModelMiddleware, embeddingModelMiddleware []*ai.EmbeddingModelMiddleware) provider.Provider ``` `WrapProvider` accepts image model middleware through the variadic `WithImageModelMiddleware(...)` option (see `pkg/registry` for the same option on `NewRegistry`): ```go wrapped := ai.WrapProvider(baseProvider, languageModelMiddleware, embeddingModelMiddleware, ai.WithImageModelMiddleware([]*ai.ImageModelMiddleware{myImageMiddleware}), ) ``` ## Built-in middleware constructors ```go func SimulateStreamingMiddleware() *ai.LanguageModelMiddleware func DefaultSettingsMiddleware(settings *provider.GenerateOptions) *ai.LanguageModelMiddleware func ExtractReasoningMiddleware(options *ai.ExtractReasoningOptions) *ai.LanguageModelMiddleware func ExtractJSONMiddleware(options *ai.ExtractJSONOptions) *ai.LanguageModelMiddleware func AddToolInputExamplesMiddleware(options *ai.AddToolInputExamplesOptions) *ai.LanguageModelMiddleware ``` ### AddToolInputExamplesOptions.Remove `Remove` is `*bool`, not `bool`. `nil` means "remove the examples" (the TS default); pass an explicit pointer to keep them: ```go // Removes InputExamples from the tool schema (nil == remove, the default). ai.AddToolInputExamplesMiddleware(&ai.AddToolInputExamplesOptions{}) // Explicitly keeps InputExamples. keep := false ai.AddToolInputExamplesMiddleware(&ai.AddToolInputExamplesOptions{Remove: &keep}) ``` ### LanguageModelMiddleware.OverrideSupportedURLs `OverrideSupportedURLs func(model provider.LanguageModel) map[string][]string` lets middleware override which URL patterns (by media type) a wrapped model reports as natively supported. A wrapped model without this field set forwards the underlying model's `SupportedURLs()` unchanged. --- # Rerank > API reference for Rerank, the Go AI SDK function that reorders documents by relevance to a query using a specialized reranking model. Canonical URL: https://goaisdk.com/docs/reference/ai/rerank Documentation index: https://goaisdk.com/llms.txt Reranks documents according to their relevance to a query using a specialized reranking model. ## Signature ```go func Rerank(ctx context.Context, opts RerankOptions) (*RerankResult, error) ``` ## Parameters ### RerankOptions | Field | Type | Required | Description | |-------|------|----------|-------------| | Model | provider.RerankingModel | Yes | Reranking model to use | | Documents | interface{} | Yes | Documents to rerank ([]string, []map[string]interface{}, or []interface{}) | | Query | string | Yes | Query to rerank documents against | | TopN | *int | No | Number of top documents to return (nil returns all) | | OnFinish | func(*RerankResult) | No | Called when reranking finishes | ## Return Value ### RerankResult | Field | Type | Description | |-------|------|-------------| | OriginalDocuments | interface{} | Original documents in original order | | Ranking | []RerankItem | Indices and scores in relevance order | | RerankedDocuments | interface{} | Documents sorted by relevance | | Response | types.RerankResponse | Response metadata | | ProviderMetadata | interface{} | Provider-specific metadata | ### RerankItem | Field | Type | Description | |-------|------|-------------| | OriginalIndex | int | Index in original document list | | Score | float64 | Relevance score (higher is more relevant) | | Document | interface{} | The actual document | ## Examples ### Basic Reranking ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/cohere" ) func main() { provider := cohere.New(cohere.Config{ APIKey: "your-api-key", }) model, err := provider.RerankingModel("rerank-english-v3.0") if err != nil { log.Fatal(err) } documents := []string{ "The capital of France is Paris", "Python is a programming language", "Paris is known for the Eiffel Tower", "Machine learning is a subset of AI", } result, err := ai.Rerank(context.Background(), ai.RerankOptions{ Model: model, Documents: documents, Query: "What is the capital of France?", }) if err != nil { log.Fatal(err) } fmt.Println("Reranked documents:") for i, item := range result.Ranking { fmt.Printf("%d. %s (score: %.4f)\n", i+1, item.Document.(string), item.Score) } } ``` ### Top-N Results ```go topN := 3 result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: model, Documents: documents, Query: "machine learning algorithms", TopN: &topN, }) if err != nil { log.Fatal(err) } fmt.Printf("Top %d results:\n", topN) for i, item := range result.Ranking { fmt.Printf("%d. %s (score: %.4f)\n", i+1, item.Document.(string), item.Score) } ``` ### Reranking Search Results ```go // Initial search results from vector database searchResults := []string{ "Neural networks are computational models", "Deep learning uses multiple layers", "The weather forecast for tomorrow", "Convolutional neural networks for images", "Recurrent neural networks for sequences", "Today's news headlines", } // Rerank by relevance to query result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: model, Documents: searchResults, Query: "neural network architectures", }) if err != nil { log.Fatal(err) } // Use reranked results for _, item := range result.Ranking { if item.Score > 0.5 { // Threshold for relevance fmt.Printf("Relevant: %s (score: %.4f)\n", item.Document.(string), item.Score) } } ``` ### Reranking Structured Documents ```go type Document struct { ID string Title string Content string } documents := []map[string]interface{}{ { "id": "doc1", "title": "Introduction to AI", "content": "Artificial intelligence is...", }, { "id": "doc2", "title": "Machine Learning Basics", "content": "Machine learning is a method...", }, { "id": "doc3", "title": "Cooking Recipes", "content": "How to make pasta...", }, } result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: model, Documents: documents, Query: "artificial intelligence", }) if err != nil { log.Fatal(err) } for i, item := range result.Ranking { doc := item.Document.(map[string]interface{}) fmt.Printf("%d. %s (score: %.4f)\n", i+1, doc["title"], item.Score) } ``` ### Two-Stage Retrieval ```go // Stage 1: Fast vector search for candidates queryEmbed, err := ai.Embed(ctx, ai.EmbedOptions{ Model: embeddingModel, Input: "machine learning frameworks", }) if err != nil { log.Fatal(err) } candidates := vectorDB.Search(queryEmbed.Embedding, 100) // top 100 candidates // Stage 2: Precise reranking result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: rerankModel, Documents: candidates, Query: "machine learning frameworks", TopN: &topN, // Final top 10 }) if err != nil { log.Fatal(err) } // Use highly relevant results for _, item := range result.Ranking { fmt.Printf("Score: %.4f - %s\n", item.Score, item.Document) } ``` ### With Callback ```go result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: model, Documents: documents, Query: query, OnFinish: func(result *ai.RerankResult) { fmt.Printf("Reranking complete! Processed %d documents\n", len(result.Ranking)) // Log top result if len(result.Ranking) > 0 { top := result.Ranking[0] fmt.Printf("Top result score: %.4f\n", top.Score) } }, }) ``` ### Empty Documents Handling ```go // Rerank handles empty document list gracefully result, err := ai.Rerank(ctx, ai.RerankOptions{ Model: model, Documents: []string{}, Query: "test query", }) if err != nil { log.Fatal(err) } // Returns empty result (no error) fmt.Printf("Documents: %d\n", len(result.Ranking)) // 0 ``` ## Error Handling ```go result, err := ai.Rerank(ctx, opts) if err != nil { switch { case strings.Contains(err.Error(), "model is required"): log.Println("Missing required model parameter") case strings.Contains(err.Error(), "documents are required"): log.Println("Missing required documents") case strings.Contains(err.Error(), "query is required"): log.Println("Missing required query") case strings.Contains(err.Error(), "must be []string"): log.Println("Invalid document type") case strings.Contains(err.Error(), "reranking failed"): log.Println("Reranking error:", err) default: log.Println("Unknown error:", err) } return } // Validate results if len(result.Ranking) == 0 { log.Println("No results returned") } ``` ## See Also - [Embed](https://goaisdk.com/docs/reference/ai/embed.md) - Generate embeddings - [EmbedMany](https://goaisdk.com/docs/reference/ai/embed-many.md) - Batch embeddings - [Reranking Models Guide](https://goaisdk.com/docs/ai-sdk-core/reranking.md) - [RAG Guide](https://goaisdk.com/docs/ai-sdk-core/reranking.md) --- # IsStepCount > API reference for IsStepCount (formerly StepCountIs), a built-in StopCondition that stops the Go AI SDK's tool-calling loop once a given number of steps completes. Canonical URL: https://goaisdk.com/docs/reference/ai/step-count-is Documentation index: https://goaisdk.com/llms.txt Returns a `StopCondition` that stops the tool-calling loop once exactly `n` steps have completed (TS `isStepCount`). `StepCountIs` is the deprecated name for the same function and still works. Pass it to `StopWhen` in `GenerateTextOptions` or `AgentConfig`. ## Signature ```go func IsStepCount(n int) StopCondition // Deprecated: use IsStepCount. func StepCountIs(n int) StopCondition ``` ## Parameters | Parameter | Type | Description | |-----------|------|-------------| | n | int | Stop when `len(state.Steps) == n` | ## Return Value A `StopCondition` function. When the condition fires it returns the reason string `"maximum number of steps (n) reached"`; otherwise it returns `""`. ```go type StopCondition func(state StopConditionState) string ``` ## Behavior `IsStepCount(n)` checks `len(steps) == n`, the same exact match as TypeScript's `isStepCount(n)` (`steps.length === n`). Stop conditions are evaluated after every step and the step count grows by one each time, so the threshold is never skipped. ## Examples ### Basic ceiling ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Analyze the dataset and create a summary", Tools: tools, StopWhen: []ai.StopCondition{ ai.IsStepCount(5), }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) fmt.Printf("Stopped because: %s\n", result.StopReason) // → "Stopped because: maximum number of steps (5) reached" ``` ### Combined with HasToolCall (safety ceiling pattern) Place `HasToolCall` first so the semantic completion signal fires before the hard ceiling. Because `EvaluateStopConditions` runs **all** conditions before returning the first match, both conditions always execute regardless of order. ```go StopWhen: []ai.StopCondition{ ai.HasToolCall("finish"), // semantic completion ai.IsStepCount(10), // hard limit fallback }, ``` ### With ToolLoopAgent ```go myAgent := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: tools, StopWhen: []ai.StopCondition{ ai.IsStepCount(8), }, }) result, err := myAgent.Execute(ctx, "Complete the task") fmt.Printf("Steps: %d, StopReason: %s\n", len(result.Steps), result.StopReason) ``` ## Defaults | Scenario | Behavior | |----------|----------| | Neither `StopWhen` nor `MaxSteps` set | `GenerateText` / `StreamText` default to `IsStepCount(1)`; `agent.NewToolLoopAgent` defaults to `IsStepCount(20)` | | `MaxSteps: &n` set, `StopWhen` not set | Converts to `StopWhen{IsStepCount(n)}` | | Both set | `StopWhen` takes precedence | ## See Also - [HasToolCall](https://goaisdk.com/docs/reference/ai/has-tool-call.md) — stop when a specific tool is called - [Loop Control](https://goaisdk.com/docs/agents/loop-control.md) — full guide with examples - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) — `GenerateTextOptions.StopWhen` field - [ToolLoopAgent](https://goaisdk.com/docs/reference/ai/tool-loop-agent.md) — `AgentConfig.StopWhen` field --- # StopCondition > Type reference for StopCondition and StopConditionState in the Go AI SDK, covering the function type, evaluation order, and usage with StopWhen. Canonical URL: https://goaisdk.com/docs/reference/ai/stop-condition Documentation index: https://goaisdk.com/llms.txt A function type that determines whether the tool-calling loop should stop after a step. Pass one or more `StopCondition` values to the `StopWhen` field in `GenerateTextOptions` or `AgentConfig`. ## Type Definition ```go type StopCondition func(state StopConditionState) string ``` A `StopCondition` is called after each step. Return a non-empty reason string to stop the loop; return `""` to continue. ## StopConditionState `StopConditionState` is passed to every `StopCondition` after each step completes. ```go type StopConditionState struct { // Steps completed so far. The step that just finished is the last element. Steps []types.StepResult // Full message history including the latest tool-result messages. Messages []types.Message // Accumulated token usage across all completed steps. Usage types.Usage } ``` | Field | Type | Description | |-------|------|-------------| | Steps | `[]types.StepResult` | Completed steps; last element is the most recent step | | Messages | `[]types.Message` | Full conversation history at this point in the loop | | Usage | `types.Usage` | Token usage summed across all steps so far | ## EvaluateStopConditions ```go func EvaluateStopConditions(conditions []StopCondition, state StopConditionState) string ``` The loop engine calls `EvaluateStopConditions` internally after every step. It runs **all** conditions before inspecting any result, then returns the first non-empty reason string (or `""` if none fired). Key invariant: **every condition always executes**, regardless of order. Side-effectful conditions (logging, metrics) are therefore safe to place anywhere in the slice and are guaranteed to run even when an earlier condition fires first. ## Examples ### Custom closure Any function with the right signature is a valid `StopCondition`: ```go stopOnError := func(state ai.StopConditionState) string { if len(state.Steps) == 0 { return "" } last := state.Steps[len(state.Steps)-1] for _, result := range last.ToolResults { if result.Error != nil { return "tool returned an error" } } return "" } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Search for recent Go releases", Tools: []types.Tool{searchTool}, StopWhen: []ai.StopCondition{stopOnError, ai.IsStepCount(10)}, }) ``` ### Token budget Stop before the model runs out of context by inspecting `state.Usage`: ```go tokenBudget := func(maxTokens int) ai.StopCondition { return func(state ai.StopConditionState) string { if int(state.Usage.GetTotalTokens()) >= maxTokens { return fmt.Sprintf("token budget of %d reached", maxTokens) } return "" } } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Summarize this document", Tools: tools, StopWhen: []ai.StopCondition{ tokenBudget(50_000), ai.HasToolCall("finish"), ai.IsStepCount(20), }, }) if err != nil { log.Fatal(err) } fmt.Printf("StopReason: %s\n", result.StopReason) ``` ### Side-effectful condition (metrics) Because `EvaluateStopConditions` always runs all conditions, you can record metrics in a `StopCondition` without worrying about evaluation order: ```go recordStepMetric := func(state ai.StopConditionState) string { metrics.RecordStep(len(state.Steps), int(state.Usage.GetTotalTokens())) return "" // never stops the loop on its own } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Summarize this document", StopWhen: []ai.StopCondition{ recordStepMetric, // always runs; records every step ai.HasToolCall("done"), // may stop early ai.IsStepCount(15), // hard ceiling }, }) ``` ## Condition ordering guarantee `EvaluateStopConditions` collects all reason strings in a single pass, then returns the first non-empty one. This means: 1. **Every condition runs on every step** — no short-circuit evaluation. 2. The **first** condition with a non-empty reason wins (slice order matters for which reason is surfaced, not for which conditions execute). 3. The winning reason is returned as `GenerateTextResult.StopReason`. ```go // Both conditions always run; the first non-empty reason is returned. StopWhen: []ai.StopCondition{ ai.HasToolCall("finish"), // fires → "tool 'finish' was called" ai.IsStepCount(20), // also evaluated, but result ignored if finish fired first }, ``` ## See Also - [IsStepCount](https://goaisdk.com/docs/reference/ai/step-count-is.md) — built-in condition: stop after n steps - [HasToolCall](https://goaisdk.com/docs/reference/ai/has-tool-call.md) — built-in condition: stop when a tool is called - [Loop Control](https://goaisdk.com/docs/agents/loop-control.md) — full guide with examples - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) — `GenerateTextOptions.StopWhen` field - [ToolLoopAgent](https://goaisdk.com/docs/reference/ai/tool-loop-agent.md) — `AgentConfig.StopWhen` field --- # StreamObject > API reference for StreamObject, the Go AI SDK function that streams structured object generation, parsing and validating partial JSON as it arrives. Canonical URL: https://goaisdk.com/docs/reference/ai/stream-object Documentation index: https://goaisdk.com/llms.txt Performs streaming structured object generation where partial JSON objects are parsed and validated as they arrive. ## Signature ```go func StreamObject(ctx context.Context, opts StreamObjectOptions) (*GenerateObjectResult, error) ``` ## Parameters ### StreamObjectOptions | Field | Type | Required | Description | |-------|------|----------|-------------| | Model | provider.LanguageModel | Yes | Language model to use (must support structured output) | | Prompt | string | No | Simple string prompt (alternative to Messages) | | Messages | []types.Message | No | List of conversation messages | | System | string | No | System instructions | | Schema | schema.Schema | Yes | Schema for output validation | | Temperature | *float64 | No | Sampling temperature (0.0 to 2.0) | | MaxTokens | *int | No | Maximum tokens to generate | | TopP | *float64 | No | Nucleus sampling parameter | | FrequencyPenalty | *float64 | No | Frequency penalty (-2.0 to 2.0) | | PresencePenalty | *float64 | No | Presence penalty (-2.0 to 2.0) | | Seed | *int | No | Random seed for reproducibility | | OnChunk | func(interface{}) | No | Called for each partial object | | OnFinish | func | No | Called when generation completes | | ExperimentalContext | interface{} | No | User-defined context | ## Return Value ### GenerateObjectResult | Field | Type | Description | |-------|------|-------------| | Object | interface{} | Final generated object | | Text | string | Raw JSON text | | FinishReason | types.FinishReason | Why generation finished | | Usage | types.Usage | Token usage information | | Warnings | []types.Warning | Provider warnings | ## Examples ### Basic Streaming Object ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/schema" ) func main() { provider := openai.New(openai.Config{ APIKey: "your-api-key", }) model, err := provider.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } // Define schema storySchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "title": map[string]interface{}{"type": "string"}, "characters": map[string]interface{}{ "type": "array", "items": map[string]interface{}{"type": "string"}, }, "plot": map[string]interface{}{"type": "string"}, }, "required": []string{"title", "characters", "plot"}, }) // Stream object with partial updates result, err := ai.StreamObject(context.Background(), ai.StreamObjectOptions{ Model: model, Prompt: "Create a story about space exploration", Schema: storySchema, OnChunk: func(partial interface{}) { fmt.Printf("Partial object: %+v\n", partial) }, }) if err != nil { log.Fatal(err) } fmt.Printf("\nFinal object: %+v\n", result.Object) fmt.Printf("Tokens used: %d\n", result.Usage.GetTotalTokens()) } ``` ### Streaming with Progress Updates ```go type Article struct { Title string `json:"title"` Sections []string `json:"sections"` Summary string `json:"summary"` } articleSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "title": map[string]interface{}{"type": "string"}, "sections": map[string]interface{}{ "type": "array", "items": map[string]interface{}{"type": "string"}, }, "summary": map[string]interface{}{"type": "string"}, }, "required": []string{"title", "sections", "summary"}, }) result, err := ai.StreamObject(ctx, ai.StreamObjectOptions{ Model: model, Prompt: "Write an article about renewable energy", Schema: articleSchema, OnChunk: func(partial interface{}) { // Cast to map for inspection if obj, ok := partial.(map[string]interface{}); ok { if title, ok := obj["title"].(string); ok { fmt.Printf("Title: %s\n", title) } if sections, ok := obj["sections"].([]interface{}); ok { fmt.Printf("Sections so far: %d\n", len(sections)) } } }, OnFinish: func(ctx context.Context, result *ai.GenerateObjectResult, userCtx interface{}) { fmt.Println("Generation complete!") }, }) if err != nil { log.Fatal(err) } ``` ### Streaming Array of Objects ```go type Task struct { ID int `json:"id"` Description string `json:"description"` Priority string `json:"priority"` } taskSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "id": map[string]interface{}{"type": "integer"}, "description": map[string]interface{}{"type": "string"}, "priority": map[string]interface{}{"type": "string", "enum": []string{"low", "medium", "high"}}, }, "required": []string{"id", "description", "priority"}, }) completedTasks := 0 result, err := ai.StreamObject(ctx, ai.StreamObjectOptions{ Model: model, Prompt: "Generate 10 project tasks", Schema: taskSchema, OnChunk: func(partial interface{}) { // Count completed fields to show progress if obj, ok := partial.(map[string]interface{}); ok { if _, hasID := obj["id"]; hasID { if _, hasDesc := obj["description"]; hasDesc { if _, hasPrio := obj["priority"]; hasPrio { completedTasks++ fmt.Printf("✓ Task %d complete\n", completedTasks) } } } } }, }) if err != nil { log.Fatal(err) } ``` ### With Fallback to Non-Streaming ```go // StreamObject automatically falls back to non-streaming // if the model doesn't support streaming structured output result, err := ai.StreamObject(ctx, ai.StreamObjectOptions{ Model: model, Prompt: "Generate product information", Schema: productSchema, OnChunk: func(partial interface{}) { // This may not be called if fallback occurs fmt.Printf("Chunk: %+v\n", partial) }, }) if err != nil { log.Fatal(err) } // Final result is always available fmt.Printf("Final: %+v\n", result.Object) ``` ## Error Handling ```go result, err := ai.StreamObject(ctx, opts) if err != nil { switch { case strings.Contains(err.Error(), "model is required"): log.Println("Missing required model parameter") case strings.Contains(err.Error(), "schema is required"): log.Println("Schema is required for streaming objects") case strings.Contains(err.Error(), "stream error"): log.Println("Streaming failed:", err) case strings.Contains(err.Error(), "validation failed"): log.Println("Object validation error:", err) case strings.Contains(err.Error(), "failed to parse"): log.Println("JSON parsing error:", err) default: log.Println("Unknown error:", err) } return } ``` ## See Also - [GenerateObject](https://goaisdk.com/docs/reference/ai/generate-object.md) - Non-streaming object generation - [StreamText](https://goaisdk.com/docs/reference/ai/stream-text.md) - Streaming text generation - [JSON Schema](https://goaisdk.com/docs/reference/schema/json-schema.md) - Schema definition - [Structured Output Guide](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) --- # StreamText > API reference for StreamText, the Go AI SDK function that streams text generation, returning tokens incrementally as the model produces them. Canonical URL: https://goaisdk.com/docs/reference/ai/stream-text Documentation index: https://goaisdk.com/llms.txt Performs streaming text generation where tokens are returned incrementally as they are generated. ## Signature ```go func StreamText(ctx context.Context, opts StreamTextOptions) (*StreamTextResult, error) ``` ## Parameters ### StreamTextOptions `Model` is required, and so is one of `Prompt` or `Messages`. Every other field is optional. The table is generated from the source and lists every exported field. {/* gen:fields ai.StreamTextOptions */} | Field | Type | Description | | --- | --- | --- | | `Model` | `provider.LanguageModel` | Model to use for generation | | `Prompt` | `string` | Prompt can be a simple string or a list of messages | | `Messages` | `[]types.Message` | | | `System` | `string` | | | `Instructions` | `*string` | Instructions is a TypeScript-compatible alias for System. When set, Instructions takes precedence over System. | | `InstructionMessages` | `[]types.Message` | InstructionMessages supplies the instructions as system messages (TS instructions: SystemModelMessage \| SystemModelMessage[]), preserving per-message ProviderOptions. When non-empty it takes precedence over Instructions and System. Every message must have the system role. | | `AllowSystemMessages` | `bool` | AllowSystemMessages permits system-role messages in Messages. Defaults to false; use System for system instructions unless you are intentionally passing provider-native system messages. | | `AllowSystemInMessages` | `bool` | AllowSystemInMessages is a TypeScript-compatible alias for AllowSystemMessages. Either field enables system-role messages. | | `Temperature` | `*float64` | Generation parameters | | `MaxTokens` | `*int` | | | `TopP` | `*float64` | | | `TopK` | `*int` | | | `FrequencyPenalty` | `*float64` | | | `PresencePenalty` | `*float64` | | | `StopSequences` | `[]string` | | | `Seed` | `*int` | | | `Headers` | `map[string]string` | | | `Tools` | `[]types.Tool` | Tools available for the model to call | | `ToolChoice` | `types.ToolChoice` | | | `ToolOrder` | `[]string` | | | `ActiveTools` | `[]string` | | | `ExperimentalToolCallers` | `ExperimentalToolCallers` | ExperimentalToolCallers configures which tools may call which other tools (programmatic tool calling / code mode), and which tools stay hidden from the model until discovered via ai.ToolSearch. Mirrors the TypeScript SDK's experimental_toolCallers. | | `ToolApproval` | `types.ToolApprovalConfig` | ToolApproval configures approval handling for tool execution. | | `ExperimentalToolApprovalSecret` | `[]byte` | ExperimentalToolApprovalSecret signs server-issued approval requests so resumed approval responses can be verified before execution. | | `MaxSteps` | `*int` | MaxSteps is a convenience shorthand for StopWhen\{StepCountIs(N)\}. Deprecated: use StopWhen with StepCountIs instead. If StopWhen is set, MaxSteps is ignored. | | `StopWhen` | `[]StopCondition` | StopWhen defines conditions that terminate the tool-calling loop. Conditions are evaluated OR -- first non-empty string stops the loop. | | `ResponseFormat` | `*provider.ResponseFormat` | Response format (for structured output) Deprecated: Use Output instead. | | `Output` | `interface{}` | Output specifies how to handle and parse model output during streaming. Use TextOutput(), ObjectOutput(), ArrayOutput(), ChoiceOutput(), or JSONOutput(). When set, the model is called with the appropriate ResponseFormat and PartialOutput() is updated after each text chunk via ParsePartialOutput. If nil, defaults to plain text streaming. | | `Timeout` | `*TimeoutConfig` | Timeout provides granular timeout controls Supports total timeout, per-step timeout, and per-chunk timeout | | `ExperimentalRetention` | `*types.RetentionSettings` | ExperimentalRetention controls what data is retained from LLM requests/responses. Useful for reducing memory consumption with images or large contexts. Default (nil) retains everything for backwards compatibility. | | `Include` | `*IncludeOptions` | Include controls which large request/response details are retained in step results and whether raw provider chunks are requested. Defaults match the TypeScript SDK. | | `ExperimentalInclude` | `*IncludeOptions` | ExperimentalInclude is a deprecated alias for Include. | | `IncludeRawChunks` | `bool` | IncludeRawChunks is a deprecated alias for Include.RawChunks. | | `Reasoning` | `*types.ReasoningLevel` | Reasoning controls how much thinking effort the model applies. nil means unset (use provider default). Set to types.ReasoningDefault to explicitly omit from the API request. | | `SendReasoning` | `*bool` | SendReasoning controls whether reasoning/thinking boundary chunks are exposed to OnChunk. nil and false suppress reasoning-start/end, matching the TypeScript SDK default. | | `ProviderOptions` | `map[string]interface{}` | ProviderOptions allows passing provider-specific options | | `MaxRetries` | `*int` | MaxRetries controls transient provider call retries. For Gateway models, nil uses the TypeScript SDK default of 2 retries; set to 0 to disable. | | `ExperimentalSandbox` | `interface{}` | ExperimentalSandbox is passed through to tool execution. PrepareStep can override it for an individual step. | | `RepairToolCall` | `ToolCallRepairFunction` | RepairToolCall attempts to repair tool calls that fail to parse because the tool does not exist or its input is invalid. When it returns a call, the repaired call is parsed again; (nil, nil) keeps the call invalid. | | `ExperimentalRepairToolCall` | `ToolCallRepairFunction` | ExperimentalRepairToolCall is a deprecated alias for RepairToolCall. Deprecated: use RepairToolCall. | | `OnLanguageModelCallStart` | `OnLanguageModelCallStartCallback` | OnLanguageModelCallStart is called immediately before each provider model call begins. | | `ExperimentalOnLanguageModelCallStart` | `OnLanguageModelCallStartCallback` | ExperimentalOnLanguageModelCallStart is a deprecated alias for OnLanguageModelCallStart. Deprecated: use OnLanguageModelCallStart. | | `OnLanguageModelCallEnd` | `OnLanguageModelCallEndCallback` | OnLanguageModelCallEnd is called after each provider model response is normalized and parsed, before client-side tool execution. | | `ExperimentalOnLanguageModelCallEnd` | `OnLanguageModelCallEndCallback` | ExperimentalOnLanguageModelCallEnd is a deprecated alias for OnLanguageModelCallEnd. Deprecated: use OnLanguageModelCallEnd. | | `ExperimentalRefineToolInput` | `map[string]ToolInputRefiner` | ExperimentalRefineToolInput refines parsed tool inputs by tool name before approval, callbacks, telemetry, execution, and response messages. | | `ExperimentalDownload` | `DownloadFunction` | ExperimentalDownload downloads remote file URLs that need inline data before sending tool-result messages back to the model. Defaults to DefaultDownload, matching the TypeScript SDK's default download behavior. | | `RuntimeContext` | `interface{}` | RuntimeContext is user-defined context that flows through callbacks. It is passed as-is to all structured event callbacks. | | `SensitiveRuntimeContext` | `bool` | SensitiveRuntimeContext omits runtime context from telemetry payloads. | | `ToolsContext` | `map[string]interface{}` | ToolsContext is the per-tool context map passed to approval and execution hooks. | | `ExperimentalContext` | `interface{}` | ExperimentalContext is a deprecated alias for RuntimeContext. | | `PrepareStep` | `func(ctx context.Context, step PrepareStepOptions) PrepareStepOptions` | PrepareStep is called before each model step. | | `Telemetry` | `*TelemetrySettings` | Telemetry configures observability for this operation. When both Telemetry and ExperimentalTelemetry are set, Telemetry wins. | | `ExperimentalTelemetry` | `*TelemetrySettings` | Telemetry configuration for observability. Deprecated: use Telemetry. | | `Internal` | `*InternalOptions` | Internal contains test-only ID generators matching the TypeScript _internal option. | | `OnChunk` | `func(chunk provider.StreamChunk)` | Callbacks | | `InitialStreamChunks` | `[]provider.StreamChunk` | InitialStreamChunks are emitted before the provider stream. Higher-level adapters use this for work resolved before the next model step while preserving stream event ordering. | | `OnEnd` | `func(result *StreamTextResult)` | OnEnd is called when the stream is fully consumed, with the full \*StreamTextResult (Text(), Usage(), ToolCalls(), etc. give the same data TS's onEnd event carries). This is a Go-only convenience signature that predates OnEndEvent below; it is NOT the same field as GenerateTextOptions.OnEnd (whose signature is func(ctx, \*GenerateTextResult, userContext) — StreamText's per-step and per-call analogues of that tuple shape are OnStepEnd below and OnEndEvent/OnFinishEvent, since a StreamTextResult is a different Go type from a GenerateTextResult and can't share that field's type). For code that wants the same shared, TS-parity event struct that GenerateText/Agent use, prefer OnEndEvent. | | `OnFinish` | `func(result *StreamTextResult)` | OnFinish is called when the stream is fully consumed. Deprecated: use OnEnd. | | `OnStepEnd` | `func(ctx context.Context, step types.StepResult, userContext interface{})` | OnStepEnd is called after each step completes, mirroring GenerateTextOptions.OnStepEnd exactly (same func type, since types.StepResult — unlike \*GenerateTextResult/\*StreamTextResult — is shared between generate and stream). TS shares one onStepEnd callback type between generateText and streamText; this is the Go equivalent. | | `OnStepFinish` | `func(ctx context.Context, step types.StepResult, userContext interface{})` | OnStepFinish is called after each step completes. Deprecated: use OnStepEnd. | | `OnStart` | `func(ctx context.Context, e OnStartEvent)` | OnStart is called once, before streaming begins. | | `OnStepStart` | `func(ctx context.Context, e OnStepStartEvent)` | OnStepStart is called at the beginning of each LLM step. | | `OnToolExecutionStart` | `func(ctx context.Context, e OnToolCallStartEvent)` | OnToolExecutionStart is called just before each tool's Execute function runs. | | `OnToolExecutionEnd` | `func(ctx context.Context, e OnToolCallFinishEvent)` | OnToolExecutionEnd is called after each tool's Execute function returns. | | `OnToolCallStart` | `func(ctx context.Context, e OnToolCallStartEvent)` | Deprecated: use OnToolExecutionStart. | | `OnToolCallFinish` | `func(ctx context.Context, e OnToolCallFinishEvent)` | Deprecated: use OnToolExecutionEnd. | | `OnStepEndEvent` | `func(ctx context.Context, e OnStepFinishEvent)` | OnStepEndEvent is called at the end of each LLM step. | | `OnStepFinishEvent` | `func(ctx context.Context, e OnStepFinishEvent)` | OnStepFinishEvent is called at the end of each LLM step. Deprecated: use OnStepEndEvent. | | `OnEndEvent` | `func(ctx context.Context, e OnFinishEvent)` | OnEndEvent is called once when the stream fully completes. | | `OnFinishEvent` | `func(ctx context.Context, e OnFinishEvent)` | OnFinishEvent is called once when the stream fully completes. Deprecated: use OnEndEvent. | | `OnError` | `func(ctx context.Context, err error)` | OnError is called when an error chunk is received from the provider during streaming. Mirrors the TypeScript SDK's onError callback. Unlike a fatal stream error (which surfaces via TextStream.Err()), error chunks are non-fatal stream events that can be observed and logged without aborting the stream. If nil, error chunks are silently forwarded to OnChunk. | | `StreamRetries` | `*int` | StreamRetries is the maximum number of automatic retries for ANY provider error chunk received after streaming has already started (a mid-stream error chunk). Matching TS (stream-text.ts's `automaticStreamRetryCount < streamRetries`), this budget is consulted unconditionally — it does not check whether the normalized providererrors.StreamProviderError considers itself retryable; that classification is informational, for OnError/ OnErrorRetry to build their own heuristics on. nil (the option omitted entirely) disables ALL stream retry behavior, both automatic and OnErrorRetry-directed, and preserves pre-streamRetries incremental chunk delivery for existing OnError observers (TS: "Omit this option to disable all stream retry behavior and preserve incremental tool streaming for existing onError observers"). An explicit 0 disables automatic retries but still allows OnErrorRetry to request one retry (TS canRetryStreamViaOnError = streamRetries !== undefined && onErrorArg != null). A negative value is rejected synchronously by StreamText (TS "streamRetries must be >= 0"). A ToolChoiceViolationError is never retried, automatically or via OnErrorRetry (TS stream-text.ts isToolChoiceViolation check). Audit row 802af1e / WG8. | | `OnErrorRetry` | `func(ctx context.Context, event StreamTextOnErrorRetryEvent) bool` | OnErrorRetry is called for a mid-stream provider error (after OnError) and, if StreamRetries permits a retry, may request one by returning true. It is consulted only when StreamRetries was explicitly set (even to 0) and no automatic retry applies (StreamRetries exhausted), and is honored at most once per streamText call, matching TS's StreamTextOnErrorRetryCallback semantics. | | `OnAbort` | `func(ctx context.Context, steps []types.StepResult)` | OnAbort is called when streaming is aborted by context cancellation or deadline before normal completion. Deprecated: use OnAbortEvent, which also carries the call ID and abort reason. | | `OnAbortEvent` | `OnAbortCallback` | OnAbortEvent is called when streaming is aborted by context cancellation or deadline before normal completion, with a GenerateTextAbortEvent carrying the call ID, completed steps, and abort reason. Takes precedence over the deprecated OnAbort. | | `ExperimentalTransform` | `[]StreamTransformFunc` | ExperimentalTransform is an ordered list of transform functions applied to each stream chunk after provider emission but before forwarding to OnChunk. Each function receives a chunk and returns zero or more replacement chunks. Returning nil or an empty slice drops the chunk. Transforms are applied in order; each transform receives the output of the previous one. Mirrors the TypeScript SDK's experimental_transform option. | {/* /gen:fields */} ## Return Value ### StreamTextResult The result provides methods to consume the stream: | Method | Returns | Description | |--------|---------|-------------| | Stream() | provider.TextStream | Underlying text stream | | FullStream() | provider.TextStream | Every chunk type, including lifecycle chunks (see below) | | Text() | string | Accumulated text so far | | FinishReason() | types.FinishReason | Finish reason (after stream ends) | | Usage() | types.Usage | Usage info (after stream ends) | | ContextManagement() | interface{} | Context management info | | Err() | error | Error that occurred during streaming | | Close() | error | Close the stream | | ReadAll() | (string, error) | Read all chunks and return complete text | | Chunks() | `<-chan provider.StreamChunk` | Channel of chunks | | Steps() | []types.StepResult | Completed steps, including per-step `Performance` statistics | | FinalStep() | types.StepResult | Last completed step with `StepTimeMs`, `ResponseTimeMs`, `ToolExecutionMs`, `EffectiveOutputTokensPerSecond`, `EffectiveTotalTokensPerSecond`, and streaming throughput fields | | Content() | []types.ContentPart | Ordered generated content from all completed steps | ## Examples ### Basic Streaming ```go package main import ( "context" "fmt" "io" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { p := openai.New(openai.Config{ APIKey: "your-api-key", }) model, err := p.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } // Start streaming result, err := ai.StreamText(context.Background(), ai.StreamTextOptions{ Model: model, Prompt: "Write a story about a robot", }) if err != nil { log.Fatal(err) } defer result.Close() // Read chunks manually for { chunk, err := result.Stream().Next() if err == io.EOF { break } if err != nil { log.Fatal(err) } if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } } fmt.Printf("\n\nTotal tokens: %d\n", result.Usage().GetTotalTokens()) } ``` ### Using Callbacks ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Explain machine learning", OnChunk: func(chunk provider.StreamChunk) { if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } }, OnFinish: func(result *ai.StreamTextResult) { fmt.Printf("\n\nDone! Tokens: %d\n", result.Usage().GetTotalTokens()) }, }) if err != nil { log.Fatal(err) } ``` ### Using ReadAll ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a poem about nature", }) if err != nil { log.Fatal(err) } defer result.Close() // Read all at once text, err := result.ReadAll() if err != nil { log.Fatal(err) } fmt.Println(text) fmt.Printf("Tokens: %d\n", result.Usage().GetTotalTokens()) ``` ### Using Channel-based Consumption ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Describe quantum physics", }) if err != nil { log.Fatal(err) } defer result.Close() // Consume chunks from channel for chunk := range result.Chunks() { if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } } fmt.Printf("\n\nFinish reason: %s\n", result.FinishReason()) ``` ### With Per-Chunk Timeout ```go perChunk := 5 * time.Second timeout := &ai.TimeoutConfig{ PerChunk: &perChunk, } result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long essay", Timeout: timeout, }) if err != nil { log.Fatal(err) } defer result.Close() text, err := result.ReadAll() if err != nil { if strings.Contains(err.Error(), "chunk timeout") { log.Println("Stream timed out waiting for chunk") } else { log.Fatal(err) } } ``` ## Full Stream Lifecycle `FullStream()` (and `OnChunk`) receive every chunk the call produces, including lifecycle chunks that bracket the call and each step: - `provider.ChunkTypeStart` opens the stream — emitted exactly once, before the first step's provider chunks. - `provider.ChunkTypeStartStep` / `provider.ChunkTypeFinishStep` bracket each step. `ChunkTypeFinishStep` carries that step's request, response, usage, finish reason, and provider metadata. - `provider.ChunkTypeFinish` closes the call — emitted exactly once, after the last step, carrying total usage across all steps. This differs from earlier SDK versions, where `ChunkTypeFinish` was forwarded once per step. Code that used to detect step boundaries by watching for `ChunkTypeFinish` must switch to `ChunkTypeFinishStep`; only the very last `ChunkTypeFinish` in a multi-step run signals the whole call is done. ```go for chunk := range result.Chunks() { switch chunk.Type { case provider.ChunkTypeStart: fmt.Println("stream started") case provider.ChunkTypeStartStep: fmt.Println("step started") case provider.ChunkTypeText: fmt.Print(chunk.Text) case provider.ChunkTypeFinishStep: fmt.Printf("\nstep finished: %s\n", chunk.FinishReason) case provider.ChunkTypeFinish: fmt.Printf("call finished, total usage: %d tokens\n", chunk.Usage.GetTotalTokens()) } } ``` ## Error Handling ```go result, err := ai.StreamText(ctx, opts) if err != nil { log.Fatal("Failed to start stream:", err) } defer result.Close() for { chunk, err := result.Stream().Next() if err == io.EOF { break } if err != nil { if strings.Contains(err.Error(), "chunk timeout") { log.Println("Chunk timeout, continuing...") continue } log.Fatal("Stream error:", err) } // Process chunk fmt.Print(chunk.Text) } // Check for errors during streaming if result.Err() != nil { log.Println("Error during streaming:", result.Err()) } ``` ## Tool approval signing `ExperimentalToolApprovalSecret` is a byte slice. When it is set, the SDK signs every approval request it issues with HMAC-SHA256 and verifies the signature on the approval response before it runs the tool. A client that edits the tool name, the tool call ID or the tool input between the request and the response fails verification. When it is nil, requests are not signed and responses are not verified. The signing functions are `ai.SignToolApproval` and `ai.VerifyToolApprovalSignature`; see [Tool approval helpers](https://goaisdk.com/docs/reference/ai/tool-approvals.md). ## See Also - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - Non-streaming text generation - [StreamObject](https://goaisdk.com/docs/reference/ai/stream-object.md) - Streaming structured output - [Streaming Guide](https://goaisdk.com/docs/foundations/streaming.md) --- # Stream Transport Helpers > Reference for Go AI SDK stream transport helpers that convert stream results into HTTP-friendly payloads, including CreateUIMessageStreamResponse. Canonical URL: https://goaisdk.com/docs/reference/ai/stream-transport-helpers Documentation index: https://goaisdk.com/llms.txt The Go AI SDK provides transport-oriented helpers for converting stream results into HTTP-friendly payloads and reading them back. ## `CreateTextStreamResponse` ```go func CreateTextStreamResponse(ctx context.Context, result *ai.StreamTextResult) (*http.Response, error) ``` Returns an `*http.Response` with a plain text body (`text/plain; charset=utf-8`) built from text chunks. ## `PipeTextStreamToResponse` ```go func PipeTextStreamToResponse(ctx context.Context, result *ai.StreamTextResult, w io.Writer) error ``` Writes text chunks to `w` as plain UTF-8 text. On an `http.ResponseWriter` it first sets `Content-Type: text/plain; charset=utf-8` and writes status 200. `PipeTextStreamToResponseWithInit(ctx, result, w, init)` takes a `*ai.TextStreamResponseInit` to change the status or add headers. Mirrors TS `pipeTextStreamToResponse`. ## `CreateUIMessageStream` ```go func CreateUIMessageStream(ctx context.Context, result *ai.StreamTextResult, opts ...ai.UIMessageStreamResultOptions) (<-chan ai.UIMessageChunk, <-chan error) ``` Converts stream chunks to UI message chunks. ## `CreateUIMessageStreamResponse` ```go func CreateUIMessageStreamResponse(ctx context.Context, result *ai.StreamTextResult, opts ...ai.UIMessageStreamResultOptions) (*http.Response, error) ``` Returns an `*http.Response` with Server-Sent Events (`text/event-stream`) UI message chunks. ## `PipeUIMessageStreamToResponse` ```go func PipeUIMessageStreamToResponse(ctx context.Context, result *ai.StreamTextResult, w io.Writer, opts ...ai.UIMessageStreamResultOptions) error ``` Writes UI message chunks as Server-Sent Events. On an `http.ResponseWriter` it first sets the UI message stream headers (see `UIMessageStreamHeaders`) and writes status 200; `PipeUIMessageStreamToResponseWithInit` takes an init with a different status or extra headers. Other writers get only the body. Mirrors TS `pipeUIMessageStreamToResponse`. ## `PipeUIMessageChunksToResponse` ```go func PipeUIMessageChunksToResponse(chunks <-chan ai.UIMessageChunk, w io.Writer, init *ai.UIMessageStreamResponseInit) error ``` Writes any UI message chunk stream to `w` as Server-Sent Events, ending with `data: [DONE]`. Use it for a stream that isn't a single `StreamTextResult`, such as one built with `CreateUIMessageStreamWithOptions` that writes custom data parts and merges an agent stream. Mirrors TS `pipeUIMessageStreamToResponse({ response, stream })`. On an `http.ResponseWriter` it first sets the UI message stream headers (overridden by `init.Headers`) and writes the status (200 unless `init.Status` is set). `init` may be `nil`. When writing fails it returns without draining `chunks`, so create the stream with a context you cancel, such as the request's. ```go chunks, _ := ai.CreateUIMessageStreamWithOptions(r.Context(), ai.UIMessageStreamOptions{ Execute: func(writer ai.UIMessageStreamWriter) { writer.Write(ai.UIMessageChunk{"type": "data-status", "id": "status", "data": map[string]any{"step": "searching"}}) result, err := ai.StreamText(r.Context(), ai.StreamTextOptions{Model: model, Prompt: "Hello"}) if err != nil { writer.Write(ai.UIMessageChunk{"type": "error", "errorText": err.Error()}) return } uiChunks, _ := ai.CreateUIMessageStream(r.Context(), result) writer.Merge(uiChunks) }, }) if err := ai.PipeUIMessageChunksToResponse(chunks, w, nil); err != nil { log.Printf("stream: %v", err) } ``` ## `CreateUIMessageChunksResponse` ```go func CreateUIMessageChunksResponse(chunks <-chan ai.UIMessageChunk, init *ai.UIMessageStreamResponseInit) (*http.Response, error) ``` Returns an `*http.Response` that streams any UI message chunk stream, with the same status and headers as `CreateUIMessageStreamResponse`. Mirrors TS `createUIMessageStreamResponse({ stream })`. ## `UIMessageStreamHeaders` ```go func UIMessageStreamHeaders() http.Header ``` Returns the headers a UI message stream response carries: `Content-Type: text/event-stream`, `Cache-Control: no-cache`, `Connection: keep-alive`, `X-Vercel-AI-UI-Message-Stream: v1` and `X-Accel-Buffering: no`. The `Pipe*` helpers set them on an `http.ResponseWriter` for you; use this when you write the response another way. Mirrors TS `UI_MESSAGE_STREAM_HEADERS`. ### Keep-alive `CreateUIMessageStreamResponseWithInit` and `PipeUIMessageStreamToResponseWithInit` accept a `*ai.UIMessageStreamResponseInit` with a `KeepAliveMs *time.Duration` field. When set to a positive duration, the pipeline writes an opening `: stream-open\n\n` SSE comment immediately — so a reverse proxy sees the response start right away instead of buffering an empty connection — and then a `: keep-alive\n\n` comment whenever that duration elapses with no other write, resetting on every real chunk: ```go keepAlive := 15 * time.Second err := ai.PipeUIMessageStreamToResponseWithInit(ctx, result, w, &ai.UIMessageStreamResponseInit{ KeepAliveMs: &keepAlive, }) ``` `ReadUIMessageStream` and `ReadUIMessages` already skip SSE comment lines (any line starting with `:`), so a client reading a keep-alive-enabled stream needs no changes. ## `ReadUIMessageStream` ```go func ReadUIMessageStream(r io.Reader) ([]ai.UIMessageChunk, error) ``` Reads Server-Sent Events UI chunk payloads from a reader. ## `ConsumeStream` ```go func ConsumeStream(ctx context.Context, stream provider.TextStream) error ``` Consumes a stream until completion or context cancellation. ## `CreateTextStreamResponseWithInit` ```go func CreateTextStreamResponseWithInit(ctx context.Context, result *ai.StreamTextResult, init *ai.TextStreamResponseInit) (*http.Response, error) ``` Like `CreateTextStreamResponse`, with a custom status, status text and headers from `init`. {/* gen:fields ai.TextStreamResponseInit */} | Field | Type | Description | | --- | --- | --- | | `Status` | `int` | | | `StatusText` | `string` | | | `Headers` | `map[string]string` | | {/* /gen:fields */} ## `CreateTextStreamResponseFromStream` ```go func CreateTextStreamResponseFromStream(ctx context.Context, stream provider.TextStream, init *ai.TextStreamResponseInit) (*http.Response, error) ``` Builds the plain text response from a `provider.TextStream` instead of a `StreamTextResult`. ## `PipeTextStreamToWriter` ```go func PipeTextStreamToWriter(ctx context.Context, stream provider.TextStream, w io.Writer) error ``` Writes the text deltas of a provider stream to `w` as UTF-8. If a write fails or `ctx` ends first, it closes the stream before it returns, so a client that disconnects does not leave a provider connection open. ## `ToTextStream` ```go func ToTextStream(ctx context.Context, stream provider.TextStream) (<-chan string, <-chan error) ``` Returns the text deltas of a stream as a channel of strings, with a channel for the error. ## `CreateUIMessageStreamWithOptions` ```go func CreateUIMessageStreamWithOptions(ctx context.Context, options ai.UIMessageStreamOptions) (<-chan ai.UIMessageChunk, <-chan error) ``` Creates a UI message stream from an `Execute` function that writes chunks and merges other streams. Use it to add custom data parts or to combine several streams. The option fields are in [UI message chunks](https://goaisdk.com/docs/reference/ai/ui-message-chunks.md#uimessagestreamoptions). ## `ToUIMessageStream` and `ToUIMessageChunk` ```go func ToUIMessageStream(ctx context.Context, stream provider.TextStream, opts ...ai.UIMessageStreamResultOptions) (<-chan ai.UIMessageChunk, <-chan error) func ToUIMessageChunk(part provider.StreamChunk, opts ai.UIMessageStreamResultOptions) (ai.UIMessageChunk, bool) ``` `ToUIMessageStream` converts a provider stream to UI message chunks. `ToUIMessageChunk` converts one chunk and returns false when the part has no UI chunk. ## `CreateUIMessageStreamResponseWithInit` and `PipeUIMessageStreamToResponseWithInit` ```go func CreateUIMessageStreamResponseWithInit(ctx context.Context, result *ai.StreamTextResult, init *ai.UIMessageStreamResponseInit, opts ...ai.UIMessageStreamResultOptions) (*http.Response, error) func PipeUIMessageStreamToResponseWithInit(ctx context.Context, result *ai.StreamTextResult, w io.Writer, init *ai.UIMessageStreamResponseInit, opts ...ai.UIMessageStreamResultOptions) error ``` The `WithInit` forms take an `*ai.UIMessageStreamResponseInit` for the status, status text, headers, repeated headers, a stream consumer and keep-alive. The fields are in [UI message chunks](https://goaisdk.com/docs/reference/ai/ui-message-chunks.md#uimessagestreamresponseinit). ## `ReadUIMessages` ```go func ReadUIMessages(ctx context.Context, opts ai.ReadUIMessagesOptions) (<-chan ai.UIMessage, <-chan error) ``` Folds a chunk stream into progressively updated `ai.UIMessage` snapshots. See [UI messages](https://goaisdk.com/docs/reference/ai/ui-messages.md#readuimessages). ## `IsUIMessageStreamError` ```go func IsUIMessageStreamError(err error) bool ``` Reports whether `err` is an `*ai.UIMessageStreamError`, which the stream reports through `OnError` when a chunk refers to a part that does not exist. ## `SmoothStream` ```go func SmoothStream(opts ai.SmoothStreamOptions) (ai.StreamTransformFunc, error) ``` Returns a stream transform that re-chunks text deltas and delays them, so text appears at an even pace. Pass it as `ExperimentalTransform` in `ai.StreamTextOptions`. {/* gen:fields ai.SmoothStreamOptions */} | Field | Type | Description | | --- | --- | --- | | `DelayInMs` | `*int` | DelayInMs is the delay between emitted chunks. nil means unset and defaults to 10ms (TS default); a non-nil pointer to 0 disables the delay entirely (TS `delayInMs: null`). This mirrors the \*int "unset vs explicit zero" convention used elsewhere in this package (see GenerateTextOptions.MaxRetries), applied to TS's `number \| null`. | | `Chunking` | `interface{}` | Chunking controls how text is split for streaming. One of: - "word" (default): split on `\S+\s+` (word boundaries). - "line": split on `\n+`. - a \*regexp.Regexp: must not match an empty string. - a ChunkDetector function. | {/* /gen:fields */} ## `ElementStream` and `ElementStreamWithOutput` ```go func ElementStream[ELEMENT any](result *ai.StreamTextResult, opts ai.ElementStreamOptions[ELEMENT]) <-chan ai.ElementStreamResult[ELEMENT] func ElementStreamWithOutput[ELEMENT any](result *ai.StreamTextResult, output ai.Output[[]ELEMENT, []ELEMENT]) <-chan ai.ElementStreamResult[ELEMENT] ``` Emits each element of a streamed JSON array as soon as it is complete. See [Utilities](https://goaisdk.com/docs/reference/ai/utilities.md#elementstreamoptions) for the option and result fields. ## `ChatTransport` `ai.ChatTransport` is the client-side interface that sends a chat's message history and receives UI message chunks. `workflow.WorkflowChatTransport` and the LangSmith deployment transport implement it. ```go type ChatTransport interface { SendMessages(ctx context.Context, req ChatTransportSendMessagesRequest) (<-chan UIMessageChunk, <-chan error) ReconnectToStream(ctx context.Context, req ChatTransportReconnectToStreamRequest) (<-chan UIMessageChunk, <-chan error) } ``` `ReconnectToStream` closes both channels without values when no stream is active. {/* gen:fields ai.ChatTransportSendMessagesRequest */} | Field | Type | Description | | --- | --- | --- | | `ChatID` | `string` | ChatID identifies the conversation. Callers that run a single conversation per process (such as pkg/tui's AgentTUIRunner) typically generate this once per runner instance. | | `MessageID` | `string` | MessageID is the id of the message being responded to, when known. | | `Trigger` | `string` | Trigger describes why SendMessages was called (for example "submit-message" or "regenerate-message"), matching the TypeScript ChatTransport's `trigger` field. | | `Messages` | `[]types.Message` | Messages is the full message history sent to the transport. | {/* /gen:fields */} {/* gen:fields ai.ChatTransportReconnectToStreamRequest */} | Field | Type | Description | | --- | --- | --- | | `ChatID` | `string` | ChatID identifies the conversation whose in-progress stream should be resumed. | {/* /gen:fields */} --- # Streaming transcription and translation > Reference for the experimental Go AI SDK functions that stream audio to a model for live transcription and speech translation, with stream part types. Canonical URL: https://goaisdk.com/docs/reference/ai/streaming-audio Documentation index: https://goaisdk.com/llms.txt Two experimental functions send raw audio to a model as it arrives and stream results back. ```go func ExperimentalStreamTranscribe(ctx context.Context, opts StreamTranscribeOptions) (*StreamTranscriptionResult, error) func ExperimentalStreamTranslate(ctx context.Context, opts StreamTranslateOptions) (*StreamTranslationResult, error) ``` `FullStream()` on each result is single-consumer. Read it before any other result accessor. An accessor such as `ProviderMetadata()` consumes the stream, and `FullStream()` is unavailable afterwards. ## Audio sources | Function | Description | | --- | --- | | `ai.NewChannelAudioStream(chunks <-chan []byte, onCancel func(reason error)) provider.AudioStream` | Wraps a channel of raw audio chunks, for live input. The producer closes the channel at the end of the audio. `onCancel` runs at most once if the consumer cancels. | | `ai.NewReaderAudioStream(r io.Reader, chunkSize int) provider.AudioStream` | Reads audio from `r` in `chunkSize`-byte chunks. `chunkSize` defaults to 4096. Cancel closes `r` when it is an `io.Closer`. | ## StreamTranscribeOptions {/* gen:fields ai.StreamTranscribeOptions */} | Field | Type | Description | | --- | --- | --- | | `Model` | `provider.TranscriptionModel` | Model must implement provider.TranscriptionStreamer. | | `Audio` | `provider.AudioStream` | Audio is the source of raw audio chunks to transcribe. | | `InputAudioFormat` | `provider.AudioFormat` | InputAudioFormat is the input audio format for the raw audio chunks. | | `ProviderOptions` | `map[string]interface{}` | ProviderOptions contains provider-specific request options. | | `Headers` | `map[string]string` | Additional HTTP/WebSocket headers to send when supported by the provider. | | `IncludeRawChunks` | `bool` | IncludeRawChunks requests provider raw chunks in the stream. | | `Telemetry` | `*TelemetrySettings` | Telemetry configures observability for this operation. When both Telemetry and ExperimentalTelemetry are set, Telemetry wins. | | `ExperimentalTelemetry` | `*TelemetrySettings` | ExperimentalTelemetry configures observability for this operation. Deprecated: use Telemetry. | {/* /gen:fields */} ## StreamTranslateOptions {/* gen:fields ai.StreamTranslateOptions */} | Field | Type | Description | | --- | --- | --- | | `Model` | `provider.SpeechTranslationModel` | Model to use for speech-to-speech translation. | | `Audio` | `provider.AudioStream` | Audio is the source of raw audio chunks to translate. | | `InputAudioFormat` | `provider.AudioFormat` | InputAudioFormat is the input audio format for the raw audio chunks. | | `TargetLanguage` | `string` | TargetLanguage is the language to translate the audio into, as a BCP-47-style language tag (e.g. "en", "es", "fr-CA"). | | `SourceLanguage` | `string` | SourceLanguage is the language of the source audio. Auto-detected when empty. | | `OutputAudioFormat` | `*provider.AudioFormat` | OutputAudioFormat is the desired audio format for translated audio chunks. When nil, the provider default output format is used. | | `ProviderOptions` | `map[string]interface{}` | ProviderOptions contains provider-specific request options. | | `Headers` | `map[string]string` | Additional HTTP/WebSocket headers to send when supported by the provider. | | `IncludeRawChunks` | `bool` | IncludeRawChunks requests provider raw chunks in the stream. | {/* /gen:fields */} ## Stream types `ai.TranscriptionStream` and `ai.TranslationStream` have the same shape as `provider.TextStream`: `Next()`, `Err()` and `Close()`. ### TranscriptionStreamPart `Type` is one of `transcript-delta`, `transcript-partial`, `transcript-final`, `raw` or `error`. {/* gen:fields ai.TranscriptionStreamPart */} | Field | Type | Description | | --- | --- | --- | | `Type` | `string` | Type is one of "transcript-delta", "transcript-partial", "transcript-final", "raw", or "error". | | `ID` | `string` | | | `Delta` | `string` | | | `Text` | `string` | | | `StartSecond` | `*float64` | | | `EndSecond` | `*float64` | | | `DurationInSeconds` | `*float64` | | | `ChannelIndex` | `*int` | | | `ProviderMetadata` | `map[string]interface{}` | | | `RawValue` | `interface{}` | RawValue carries the provider's raw chunk when Type is "raw". | | `Err` | `interface{}` | Err carries the streamed error payload when Type is "error". | {/* /gen:fields */} ### TranslationStreamPart `Type` is one of `audio`, `output-text-delta`, `output-text-final`, `source-transcript-delta`, `source-transcript-partial`, `source-transcript-final`, `raw` or `error`. {/* gen:fields ai.TranslationStreamPart */} | Field | Type | Description | | --- | --- | --- | | `Type` | `string` | Type is one of "audio", "output-text-delta", "output-text-final", "source-transcript-delta", "source-transcript-partial", "source-transcript-final", "raw", or "error". | | `ID` | `string` | | | `AudioData` | `[]byte` | | | `Delta` | `string` | | | `Text` | `string` | | | `StartSecond` | `*float64` | | | `EndSecond` | `*float64` | | | `ChannelIndex` | `*int` | | | `ProviderMetadata` | `map[string]interface{}` | | | `RawValue` | `interface{}` | RawValue carries the provider's raw chunk when Type is "raw". | | `Err` | `interface{}` | Err carries the streamed error payload when Type is "error". | {/* /gen:fields */} ### SpeechTranslationModelResponseMetadata {/* gen:fields ai.SpeechTranslationModelResponseMetadata */} | Field | Type | Description | | --- | --- | --- | | `Timestamp` | `time.Time` | | | `ModelID` | `string` | | | `Headers` | `map[string]string` | | {/* /gen:fields */} ## Errors `ai.IsNoTranslationGeneratedError(err)` reports whether `err` is a `*ai.NoTranslationGeneratedError`. --- # Tool approval helpers > Reference for the Go AI SDK tool approval functions: signing, verification, collecting, validating and resuming approvals, with their option and result types and errors. Canonical URL: https://goaisdk.com/docs/reference/ai/tool-approvals Documentation index: https://goaisdk.com/llms.txt A tool with `ToolApproval` set pauses the run and sends a `tool-approval-request` chunk. The client answers, and the next request carries the response. These helpers sign requests, verify responses and resume the run. `GenerateText`, `StreamText` and `ToolLoopAgent` call them for you. Call them directly when you write your own generation loop. For the chunk shapes, see [UI message chunks](https://goaisdk.com/docs/reference/ai/ui-message-chunks.md). For the tool field, see [Tool](https://goaisdk.com/docs/reference/ai/tool.md). ## Signing Set `ExperimentalToolApprovalSecret` on `ai.GenerateTextOptions`, `ai.StreamTextOptions` or `agent.AgentConfig` to turn signing on. The SDK signs each approval request with HMAC-SHA256 over the approval ID, tool call ID, tool name and a digest of the tool input. On resume, it recomputes the signature for each approved call. A call with a missing or wrong signature is rejected with `*ai.InvalidToolApprovalSignatureError`. The signature format matches the TypeScript SDK, so a request signed by one SDK verifies in the other. ```go func SignToolApproval(secret []byte, approvalID, toolCallID, toolName string, input interface{}) (string, error) func VerifyToolApprovalSignature(secret []byte, signature, approvalID, toolCallID, toolName string, input interface{}) (bool, error) ``` `SignToolApproval` returns a base64url string. `VerifyToolApprovalSignature` also accepts the legacy newline-joined payload, but only when none of the IDs or the tool name contain a newline. ## Status values `ai.ToolApprovalStatus` is the result of an approval policy. The constants alias the values in `types`. | Constant | Meaning | | --- | --- | | `ai.ToolApprovalStatusNotApplicable` | No policy applies. The tool runs. | | `ai.ToolApprovalStatusApproved` | The policy approved the call. | | `ai.ToolApprovalStatusDenied` | The policy denied the call. | | `ai.ToolApprovalStatusUserApproval` | The user must decide. | `ai.ToolApprovalConfig` configures automatic approval on options structs. ## CollectToolApprovals ```go func CollectToolApprovals(messages []types.Message) (CollectToolApprovalsResult, error) ``` Reads the approval responses from the last message when it is a tool message. Approved responses whose tool call already has a result are skipped. Denied responses are skipped only when the existing result is not an execution-denied output, so a client denial can replace an earlier synthetic one. It returns `*ai.InvalidToolApprovalError` when a response names an unknown approval ID, and `*ai.ToolCallNotFoundForApprovalError` when a request names a tool call that is not in the history. {/* gen:fields ai.CollectToolApprovalsResult */} | Field | Type | Description | | --- | --- | --- | | `ApprovedToolApprovals` | `[]CollectedToolApproval` | | | `DeniedToolApprovals` | `[]CollectedToolApproval` | | {/* /gen:fields */} ### CollectedToolApproval {/* gen:fields ai.CollectedToolApproval */} | Field | Type | Description | | --- | --- | --- | | `ApprovalRequest` | `types.ToolApprovalRequestContent` | | | `ApprovalResponse` | `types.ToolApprovalResponseContent` | | | `ToolCall` | `types.ToolCall` | | | `ExistingToolResult` | `*types.ToolResultContent` | ExistingToolResult is the tool result for the call that is already present in the last tool message, if any. | {/* /gen:fields */} ## ValidateApprovedToolApprovals ```go func ValidateApprovedToolApprovals(ctx context.Context, opts ValidateApprovedToolApprovalsOptions) (ValidateApprovedToolApprovalsResult, error) ``` Re-validates approved approvals that were rebuilt from client-supplied history, before the tools run. It verifies the signature when `ToolApprovalSecret` is set, validates the input against the tool's current schema and input refinement, and resolves the server-side approval policy again. Approvals whose input no longer validates come back as `InvalidToolApproval` values and are reported to the model as tool errors. {/* gen:fields ai.ValidateApprovedToolApprovalsOptions */} | Field | Type | Description | | --- | --- | --- | | `ApprovedToolApprovals` | `[]CollectedToolApproval` | | | `Tools` | `[]types.Tool` | | | `ToolApproval` | `types.ToolApprovalConfig` | | | `Messages` | `[]types.Message` | | | `ToolsContext` | `map[string]interface{}` | | | `RuntimeContext` | `interface{}` | | | `ToolApprovalSecret` | `[]byte` | ToolApprovalSecret enables HMAC signature verification when non-nil. | | `RefineToolInput` | `map[string]ToolInputRefiner` | | {/* /gen:fields */} {/* gen:fields ai.ValidateApprovedToolApprovalsResult */} | Field | Type | Description | | --- | --- | --- | | `ApprovedToolApprovals` | `[]CollectedToolApproval` | | | `DeniedToolApprovals` | `[]CollectedToolApproval` | | | `InvalidToolApprovals` | `[]InvalidToolApproval` | | {/* /gen:fields */} ### InvalidToolApproval {/* gen:fields ai.InvalidToolApproval */} | Field | Type | Description | | --- | --- | --- | | (embedded) | `CollectedToolApproval` | Embedded. | | `Error` | `*InvalidToolInputError` | | {/* /gen:fields */} ## ResumeToolApprovals ```go func ResumeToolApprovals(ctx context.Context, opts ResumeToolApprovalsOptions) (ResumeToolApprovalsResult, error) ``` Applies the approvals at the end of `Messages`: it collects, validates and executes approved tools and records denials. `GenerateText` and `StreamText` run the same logic before the first model call. Use it when you implement your own loop. {/* gen:fields ai.ResumeToolApprovalsOptions */} | Field | Type | Description | | --- | --- | --- | | `Messages` | `[]types.Message` | Messages is the input history to scan for a trailing tool-approval-response message. | | `Tools` | `[]types.Tool` | Tools are the tools available to resolve approved calls against. | | `ToolApproval` | `types.ToolApprovalConfig` | ToolApproval configures automatic approval handling. | | `ToolsContext` | `map[string]interface{}` | ToolsContext contains per-tool execution context keyed by tool name. | | `RuntimeContext` | `interface{}` | RuntimeContext is user-defined context passed to callbacks and tools. | | `Secret` | `[]byte` | Secret verifies (and, when re-signing, produces) approval signatures. Nil disables signature verification. | | `RefineToolInput` | `map[string]ToolInputRefiner` | RefineToolInput refines parsed tool inputs by tool name before re-validation and execution. | | `CallID` | `string` | CallID correlates emitted tool execution events with the caller's call. | | `ModelProvider` | `string` | ModelProvider and ModelID identify the model for emitted events. | | `ModelID` | `string` | | | `OnToolExecutionStart` | `func(ctx context.Context, e OnToolCallStartEvent)` | OnToolExecutionStart/OnToolExecutionEnd are notified around each approved tool's execution. | | `OnToolExecutionEnd` | `func(ctx context.Context, e OnToolCallFinishEvent)` | | {/* /gen:fields */} {/* gen:fields ai.ResumeToolApprovalsResult */} | Field | Type | Description | | --- | --- | --- | | `ResponseMessages` | `[]types.Message` | ResponseMessages is the tool message (if any) produced by resuming approvals: executed results, invalid-input errors, and execution-denied outputs. Empty when the history had no pending or resolved approvals to resume. | | `Usage` | `types.Usage` | Usage is the token usage consumed by any tools that ran. | {/* /gen:fields */} ## Errors | Type | Test function | Returned when | | --- | --- | --- | | `*ai.InvalidToolApprovalError` | `ai.IsInvalidToolApprovalError(err)` | A response references an approval ID with no matching request. | | `*ai.ToolCallNotFoundForApprovalError` | `ai.IsToolCallNotFoundForApprovalError(err)` | A request references a tool call that is not in the history. | | `*ai.InvalidToolApprovalSignatureError` | `ai.IsInvalidToolApprovalSignatureError(err)` | The signature is missing or does not match. | --- # Tool call parsing and repair > Reference for the Go AI SDK functions that parse, repair and refine model tool calls, and the tool definition fingerprints that detect drift. Canonical URL: https://goaisdk.com/docs/reference/ai/tool-call-repair Documentation index: https://goaisdk.com/llms.txt The SDK parses each tool call the model returns against the tool's input schema. A call that names an unknown tool or has invalid input is marked invalid unless a repair function fixes it. `GenerateText`, `StreamText` and `ToolLoopAgent` run these steps for you. The functions on this page are the building blocks. ## RepairToolCall Set `RepairToolCall` on `ai.GenerateTextOptions` or `ai.StreamTextOptions` to repair calls that fail to parse. The function type is `ai.ToolCallRepairFunction`. ```go type ToolCallRepairFunction func(ctx context.Context, options ToolCallRepairOptions) (*types.ToolCall, error) ``` Return `(nil, nil)` when the call cannot be repaired. The returned call may carry the repaired input as `RawArguments` (JSON text) or `Arguments`. {/* gen:fields ai.ToolCallRepairOptions */} | Field | Type | Description | | --- | --- | --- | | `ToolCall` | `types.ToolCall` | ToolCall is the tool call that failed to parse. | | `Tools` | `[]types.Tool` | Tools are the tools available in the step. | | `InputSchema` | `func(toolName string) (map[string]interface{}, error)` | InputSchema returns the JSON schema of a tool's input. | | `Instructions` | `string` | Instructions is the text instructions (system prompt) of the step. | | `InstructionMessages` | `[]types.Message` | InstructionMessages are the system-message instructions of the step, when instructions were given as system messages. | | `System` | `string` | System is the step's system prompt. Deprecated: use Instructions. | | `Messages` | `[]types.Message` | Messages are the messages of the current generation step. | | `Error` | `error` | Error is the \*NoSuchToolError or \*InvalidToolInputError that occurred. | {/* /gen:fields */} `ai.RepairTextFunc` is a function type that repairs raw JSON text before parsing: `func(ctx context.Context, text string, parseErr error) (*string, error)`. Return nil to leave the text unchanged. ## ParseToolCall ```go func ParseToolCall(ctx context.Context, opts ParseToolCallOptions) (types.ToolCall, error) ``` Resolves a provider tool call against the available tools, parses and validates its input, and runs `RepairToolCall` for unknown-tool and invalid-input errors. A call that cannot be parsed or repaired comes back with `Invalid`, `Dynamic` and `Error` set. The only error it returns is a context error. {/* gen:fields ai.ParseToolCallOptions */} | Field | Type | Description | | --- | --- | --- | | `ToolCall` | `types.ToolCall` | | | `Tools` | `[]types.Tool` | | | `RepairToolCall` | `ToolCallRepairFunction` | | | `Instructions` | `string` | | | `InstructionMessages` | `[]types.Message` | | | `Messages` | `[]types.Message` | | {/* /gen:fields */} ## Basic repair helpers These helpers predate `ToolCallRepairFunction`. They take an `ai.ToolCallRepairFunc`, which sees the failing call and the error but not the tools or messages. | Name | Description | | --- | --- | | `ai.ToolCallRepairFunc` | `func(ctx context.Context, toolCall types.ToolCall, err error) (*types.ToolCall, error)`. | | `ai.DefaultToolCallRepair` | A `ToolCallRepairFunc` that fixes common argument problems. | | `ai.RepairOptions` | Options for `TryRepairToolCalls`: `MaxAttempts` and `RepairFunc`. | | `ai.DefaultRepairOptions()` | Returns the default `RepairOptions`. | | `ai.TryRepairToolCalls(ctx, toolCalls, opts)` | Repairs every call in a slice. | | `ai.IsToolCallRepairError(err)` | Reports whether `err` is an `*ai.ToolCallRepairError`. | ## Refining tool input A refiner adjusts parsed tool input before approval, callbacks, telemetry and execution. ```go type ToolInputRefiner func(ctx context.Context, opts ToolInputRefinementOptions) (map[string]interface{}, error) func RefineToolCalls(ctx context.Context, calls []types.ToolCall, tools []types.Tool, refiners map[string]ToolInputRefiner, runtimeContext interface{}, toolsContext map[string]interface{}) ([]types.ToolCall, error) ``` `refiners` is keyed by tool name. When a refiner changes the input, the original input is kept as `inputSchemaInput` on the approval request. {/* gen:fields ai.ToolInputRefinementOptions */} | Field | Type | Description | | --- | --- | --- | | `ToolCall` | `types.ToolCall` | | | `Tool` | `*types.Tool` | | | `RuntimeContext` | `interface{}` | | | `ToolsContext` | `map[string]interface{}` | | {/* /gen:fields */} ## Tool definition fingerprints A server-controlled tool definition, such as one fetched from an MCP server, can change after you reviewed it. Fingerprint the tools when you trust them and compare later. ```go func FingerprintTools(tools []types.Tool) (map[string]string, error) func DetectToolDrift(current, baseline map[string]string) ToolDrift ``` `FingerprintTools` hashes the description, the resolved input schema and the title of each tool with SHA-256 over canonical JSON. A `DescriptionFunc` is pinned by presence only, not by output. The digests match the TypeScript SDK's `fingerprintTools`. `DetectToolDrift` returns the tool names that were added, removed or changed. The slices are sorted. {/* gen:fields ai.ToolDrift */} | Field | Type | Description | | --- | --- | --- | | `Added` | `[]string` | Added lists tools present only in the current fingerprints. | | `Removed` | `[]string` | Removed lists tools present only in the baseline fingerprints. | | `Changed` | `[]string` | Changed lists tools whose pinned definition differs. | {/* /gen:fields */} --- # ToolLoopAgent > API reference for agent.ToolLoopAgent, the Go AI SDK type that creates an agent continuously using tools in a loop until the task completes. Canonical URL: https://goaisdk.com/docs/reference/ai/tool-loop-agent Documentation index: https://goaisdk.com/llms.txt `ToolLoopAgent` is an agent that loops through tool calls until the task completes. It lives in `github.com/digitallysavvy/go-ai/pkg/agent` (not `pkg/ai`) and implements the [`agent.Agent`](https://goaisdk.com/docs/reference/ai/agent.md) interface. ## Signature ```go func NewToolLoopAgent(config AgentConfig) *ToolLoopAgent ``` ## Parameters ### AgentConfig (selected fields) `AgentConfig` has many fields; these are the ones most commonly used. | Field | Type | Description | |-------|------|-------------| | Model | provider.LanguageModel | Language model for the agent | | System | string | System prompt for the agent | | Tools | []types.Tool | Tools available to the agent | | ActiveTools | []string | Restricts the available tools by name before model calls | | StopWhen | []ai.StopCondition | Conditions that terminate the tool-calling loop, e.g. `[]ai.StopCondition{ai.IsStepCount(20)}` (default) | | MaxSteps | int | **Deprecated** — use `StopWhen` with `ai.IsStepCount` instead | | Temperature | *float64 | Sampling temperature | | MaxTokens | *int | Maximum tokens per generation | | ToolChoice | types.ToolChoice | Controls how the model may call tools | | ProviderOptions | map[string]interface{} | Provider-specific options | | Skills | *SkillRegistry | Reusable agent behaviors, registered with `AddSkill` | | Subagents | *SubagentRegistry | Specialized agents that can be delegated to | | PrepareCall | func(ctx, PrepareCallConfig) PrepareCallConfig | Called before each generation step to adjust settings dynamically | | OnStepEnd | func(types.StepResult) | Called after each step completes | | OnFinish | func(*AgentResult) | Called when agent execution completes | | ToolApproval | types.ToolApprovalConfig | Approval policy for tool calls: a `types.GenericToolApprovalFunc` or a per-tool map. Tools can also set `Tool.ToolApproval` | | ExperimentalToolApprovalSecret | []byte | Signs approval requests and verifies approval responses when a client resumes a conversation | | ToolApprovalRequired, ToolApprover | bool, func | Legacy blocking callback for CLI-style programs. `ToolApprover` is called inline and returns true or false | See `go doc ./pkg/agent AgentConfig` for the complete field list, including telemetry, retry, timeout, and callback-event options. ## Methods ```go func (a *ToolLoopAgent) Execute(ctx context.Context, prompt string) (*AgentResult, error) func (a *ToolLoopAgent) ExecuteWithMessages(ctx context.Context, messages []types.Message) (*AgentResult, error) func (a *ToolLoopAgent) Generate(ctx context.Context, opts AgentGenerateOptions) (*ai.GenerateTextResult, error) func (a *ToolLoopAgent) GenerateAgent(ctx context.Context, opts AgentGenerateOptions) (*AgentResult, error) func (a *ToolLoopAgent) Stream(ctx context.Context, opts AgentStreamOptions) (*ai.StreamTextResult, error) func (a *ToolLoopAgent) AddTool(tool types.Tool) func (a *ToolLoopAgent) RemoveTool(toolName string) func (a *ToolLoopAgent) Tools() []types.Tool func (a *ToolLoopAgent) AddSkill(skill *Skill) error func (a *ToolLoopAgent) AddSubagent(name string, subagent Agent) error func (a *ToolLoopAgent) DelegateToSubagent(ctx context.Context, name string, prompt string) (*AgentResult, error) func (a *ToolLoopAgent) SetMaxSteps(maxSteps int) func (a *ToolLoopAgent) SetStopConditions(conditions []ai.StopCondition) func (a *ToolLoopAgent) SetSystem(system string) func (a *ToolLoopAgent) ID() string func (a *ToolLoopAgent) Version() string ``` ## Return Value ### AgentResult | Field | Type | Description | |-------|------|-------------| | Text | string | Agent's final text output | | Output | interface{} | Parsed/structured output when an output strategy is configured | | Steps | []types.StepResult | All steps in the loop | | ToolResults | []types.ToolResult | All tool results from execution | | FinishReason | types.FinishReason | Final finish reason | | StopReason | string | Why the loop stopped | | Usage | types.Usage | Total token usage | | Warnings | []types.Warning | Warnings from any step | ## Examples ### Basic Tool Loop Agent ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { p := openai.New(openai.Config{ APIKey: "your-api-key", }) model, _ := p.LanguageModel("gpt-6-astra") searchTool := types.Tool{Name: "search"} priceCompareTool := types.Tool{Name: "price_compare"} a := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: []types.Tool{searchTool, priceCompareTool}, }) result, err := a.Execute(context.Background(), "Find and compare prices of three popular smartphones") if err != nil { log.Fatal(err) } fmt.Printf("Answer: %s\n", result.Text) fmt.Printf("Tool calls: %d\n", len(result.ToolResults)) } ``` ### With a Custom Stop Condition ```go a := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: tools, StopWhen: []ai.StopCondition{ ai.HasToolCall("final_answer"), ai.IsStepCount(5), }, }) result, err := a.Execute(ctx, "Research topic and gather 3 key facts") ``` ### With Step and Tool Callbacks ```go a := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: tools, OnStepEnd: func(step types.StepResult) { fmt.Printf("Step finished, %d tool calls\n", len(step.ToolCalls)) }, OnToolResult: func(result types.ToolResult) { fmt.Printf("Tool result: %+v\n", result.Result) if result.Error != nil { fmt.Printf("Tool error: %v\n", result.Error) } }, }) result, err := a.Execute(ctx, "Complex multi-step research task") ``` ## Configuration fields (generated) These tables are generated from the `agent.AgentConfig` and `agent.PrepareCallConfig` source and list every exported field. ### AgentConfig `agent.DefaultAgentConfig()` returns an `AgentConfig` with default values. {/* gen:fields agent.AgentConfig */} | Field | Type | Description | | --- | --- | --- | | `ID` | `string` | ID is a unique identifier for the agent Useful for tracking and logging which agent performed actions | | `Version` | `string` | Version is the version of the agent (e.g., "agent-v1", "1.0.0") Useful for versioning agent behavior and configurations | | `Model` | `provider.LanguageModel` | Model to use for the agent | | `System` | `string` | System prompt for the agent | | `Prompt` | `string` | Prompt is a TypeScript-compatible alias for System/instructions. When System is empty, Prompt is used as the system prompt for each step. | | `Instructions` | `*string` | Instructions is the canonical TypeScript-compatible name for System. When set, Instructions takes precedence over System. | | `AllowSystemInMessages` | `bool` | AllowSystemInMessages mirrors TypeScript ToolLoopAgentSettings.allowSystemInMessages. | | `Tools` | `[]types.Tool` | Tools available to the agent | | `ActiveTools` | `[]string` | ActiveTools restricts the available tools by name before model calls. | | `ToolOrder` | `[]string` | ToolOrder controls the order tools are sent to providers. | | `ExperimentalToolCallers` | `ai.ExperimentalToolCallers` | ExperimentalToolCallers configures which tools may call which other tools, and which tools stay hidden until discovered via ai.ToolSearch. Forwarded to the underlying GenerateText/StreamText call. Mirrors the TypeScript SDK's ToolLoopAgentSettings.experimental_toolCallers. | | `Skills` | `*SkillRegistry` | Skills are reusable agent behaviors Skills can be registered and executed by the agent | | `Subagents` | `*SubagentRegistry` | Subagents are specialized agents that can be delegated to The main agent can delegate tasks to subagents for specialized processing | | `MaxSteps` | `int` | Maximum number of steps (iterations) the agent can take. If both MaxSteps and StopWhen are set, StopWhen takes precedence. Deprecated: use StopWhen with ai.IsStepCount instead. | | `StopWhen` | `[]ai.StopCondition` | StopWhen defines conditions that terminate the agent's tool-calling loop. Conditions are evaluated OR -- first non-empty string stops the loop. If neither StopWhen nor MaxSteps is set, defaults to []ai.StopCondition\{ai.IsStepCount(20)\}. | | `Temperature` | `*float64` | Temperature for generation | | `MaxTokens` | `*int` | MaxTokens per generation | | `TopP` | `*float64` | TopP controls nucleus sampling. | | `TopK` | `*int` | TopK controls top-k sampling for providers that support it. | | `FrequencyPenalty` | `*float64` | FrequencyPenalty reduces repetition. | | `PresencePenalty` | `*float64` | PresencePenalty encourages topic diversity. | | `StopSequences` | `[]string` | StopSequences halt generation when encountered. | | `ToolChoice` | `types.ToolChoice` | ToolChoice controls how the model may call tools. | | `Seed` | `*int` | Seed enables deterministic generation when supported by the provider. | | `Headers` | `map[string]string` | Headers are custom request headers forwarded to the provider. | | `Reasoning` | `*types.ReasoningLevel` | Reasoning configures provider reasoning effort. | | `SendReasoning` | `*bool` | SendReasoning controls whether reasoning stream/content parts are exposed. | | `ProviderOptions` | `map[string]interface{}` | ProviderOptions are provider-specific options keyed by provider name. | | `Output` | `interface{}` | Output specifies how to handle and parse model output during streaming. | | `Telemetry` | `*ai.TelemetrySettings` | Telemetry configures observability for this agent. | | `ExperimentalTelemetry` | `*ai.TelemetrySettings` | ExperimentalTelemetry is a deprecated alias for Telemetry, included for parity with TypeScript prepareCall. | | `SensitiveRuntimeContext` | `bool` | SensitiveRuntimeContext omits runtime context from telemetry payloads. | | `Include` | `*ai.IncludeOptions` | Include forwards stable include options to core stream/generate behavior. | | `Internal` | `*ai.InternalOptions` | Internal contains test-only ID generators matching TypeScript _internal. | | `Timeout` | `*ai.TimeoutConfig` | Timeout provides granular timeout controls Supports total timeout, per-step timeout, and per-chunk timeout | | `MaxRetries` | `*int` | MaxRetries controls transient provider call retries, forwarded to ai.GenerateTextOptions.MaxRetries / ai.StreamTextOptions.MaxRetries. nil uses the default retry count (2), matching TS WorkflowAgent's `mergedGenerationSettings.maxRetries ?? 2`. | | `PrepareCall` | `func(ctx context.Context, config PrepareCallConfig) PrepareCallConfig` | PrepareCall is called before each generation step Allows dynamic modification of the agent's configuration per step This is useful for: - Adjusting system prompts based on context - Changing temperature/parameters dynamically - Modifying available tools based on state - Adding contextual information Example: PrepareCall: func(ctx context.Context, config PrepareCallConfig) PrepareCallConfig \{ // Adjust system prompt based on step number if config.StepNumber > 5 \{ config.System = "Be more concise." \} return config \} | | `ExperimentalContext` | `interface{}` | ExperimentalContext is user-defined context that flows through all structured callback events (OnStartEvent, OnStepStartEvent, etc.). It is passed as-is to every event — useful for correlating events with application-level state (e.g. request IDs, session objects). | | `RuntimeContext` | `interface{}` | RuntimeContext is user-defined context that flows through callbacks, approval hooks, and tool execution. ExperimentalContext is a deprecated alias and is used only when RuntimeContext is nil. | | `ToolsContext` | `map[string]interface{}` | ToolsContext contains per-tool context keyed by tool name. Each value is validated against the matching tool's ContextSchema before approval callbacks and execution. | | `ExperimentalSandbox` | `interface{}` | ExperimentalSandbox is passed to tool execution and prepare-step hooks. | | `ExperimentalRefineToolInput` | `map[string]ai.ToolInputRefiner` | ExperimentalRefineToolInput refines parsed tool input before execution. | | `CallOptions` | `interface{}` | CallOptions contains agent-level call options passed through PrepareCall. When CallOptionsSchema is set, this value is validated before model calls. | | `CallOptionsSchema` | `schema.Schema` | CallOptionsSchema validates CallOptions before any generate/stream call. | | `FilterActiveTools` | `func(ctx context.Context, stepNumber int, tools []types.Tool) []types.Tool` | FilterActiveTools can restrict the available tools for each step. It mirrors the TypeScript filterActiveTools/activeTools behavior. | | `ExperimentalDownload` | `ai.DownloadFunction` | ExperimentalDownload customizes remote file URL downloads before model calls. When nil, the core default downloader is used for unsupported URLs. | | `InstructionMessages` | `[]types.Message` | InstructionMessages supplies the agent's instructions as system messages (ai.GenerateTextOptions.InstructionMessages / StreamTextOptions equivalent), preserving per-message ProviderOptions. Takes precedence over Instructions/System/Prompt when non-empty. | | `PrepareStep` | `func(ctx context.Context, step ai.PrepareStepOptions) ai.PrepareStepOptions` | PrepareStep lets you provide different settings for a step. Forwarded to ai.GenerateText/StreamText for Generate/Stream (TS ToolLoopAgentSettings.prepareStep). | | `RepairToolCall` | `ai.ToolCallRepairFunction` | RepairToolCall attempts to repair tool calls that fail to parse because the tool does not exist or its input is invalid. Forwarded to ai.GenerateText/StreamText for Generate/Stream. | | `ExperimentalRepairToolCall` | `ai.ToolCallRepairFunction` | ExperimentalRepairToolCall is a deprecated alias for RepairToolCall. Deprecated: use RepairToolCall. | | `OnLanguageModelCallStart` | `ai.OnLanguageModelCallStartCallback` | OnLanguageModelCallStart is called immediately before each provider model call begins. | | `ExperimentalOnLanguageModelCallStart` | `ai.OnLanguageModelCallStartCallback` | ExperimentalOnLanguageModelCallStart is a deprecated alias for OnLanguageModelCallStart. Deprecated: use OnLanguageModelCallStart. | | `OnLanguageModelCallEnd` | `ai.OnLanguageModelCallEndCallback` | OnLanguageModelCallEnd is called after each provider model response is normalized and parsed, before client-side tool execution. | | `ExperimentalOnLanguageModelCallEnd` | `ai.OnLanguageModelCallEndCallback` | ExperimentalOnLanguageModelCallEnd is a deprecated alias for OnLanguageModelCallEnd. Deprecated: use OnLanguageModelCallEnd. | | `OnStepStart` | `func(stepNum int)` | Legacy/Basic Callbacks | | `OnStepEnd` | `func(step types.StepResult)` | | | `OnStepFinish` | `func(step types.StepResult)` | Deprecated: use OnStepEnd. | | `OnToolCall` | `func(toolCall types.ToolCall)` | | | `OnToolResult` | `func(toolResult types.ToolResult)` | | | `OnFinish` | `func(result *AgentResult)` | | | `OnStart` | `func(ctx context.Context, e ai.OnStartEvent)` | OnStart is called once when agent execution begins. | | `OnStepStartEvent` | `func(ctx context.Context, e ai.OnStepStartEvent)` | OnStepStart is called at the beginning of each LLM step. | | `OnToolExecutionStart` | `func(ctx context.Context, e ai.OnToolCallStartEvent)` | OnToolExecutionStart is called just before each tool's Execute function runs. | | `OnToolExecutionEnd` | `func(ctx context.Context, e ai.OnToolCallFinishEvent)` | OnToolExecutionEnd is called after each tool's Execute function returns. | | `OnToolCallStart` | `func(ctx context.Context, e ai.OnToolCallStartEvent)` | OnToolCallStart is the previous Go name for OnToolExecutionStart. Deprecated: use OnToolExecutionStart. | | `OnToolCallFinish` | `func(ctx context.Context, e ai.OnToolCallFinishEvent)` | OnToolCallFinish is the previous Go name for OnToolExecutionEnd. Deprecated: use OnToolExecutionEnd. | | `OnStepEndEvent` | `func(ctx context.Context, e ai.OnStepFinishEvent)` | OnStepEndEvent is called at the end of each LLM step. | | `OnStepFinishEvent` | `func(ctx context.Context, e ai.OnStepFinishEvent)` | OnStepFinishEvent is called at the end of each LLM step. Deprecated: use OnStepEndEvent. | | `OnEndEvent` | `func(ctx context.Context, e ai.OnFinishEvent)` | OnEndEvent is called once when agent execution completes. | | `OnFinishEvent` | `func(ctx context.Context, e ai.OnFinishEvent)` | OnFinishEvent is a deprecated alias for OnEndEvent. Deprecated: use OnEndEvent. | | `OnChainStart` | `func(input string, messages []types.Message)` | OnChainStart is called when the agent begins execution Use this to initialize resources, start timers, or log the beginning of a chain | | `OnChainEnd` | `func(result *AgentResult)` | OnChainEnd is called when the agent completes successfully Use this to clean up resources, log completion, or process final results | | `OnChainError` | `func(err error)` | OnChainError is called when the agent encounters an error during execution Use this for error logging, cleanup, or triggering fallback mechanisms | | `OnAgentAction` | `func(action AgentAction)` | OnAgentAction is called when the agent decides to take an action (e.g., call a tool) This is invoked before tool execution begins Use this to log agent decisions or implement custom action approval logic | | `OnAgentFinish` | `func(finish AgentFinish)` | OnAgentFinish is called when the agent reaches a final answer (no more tool calls) This differs from OnChainEnd as it's called when the agent decides it's done, even if more processing remains | | `OnToolStart` | `func(toolCall types.ToolCall)` | OnToolStart is called immediately before a tool begins execution Use this to log tool invocations or start performance timers | | `OnToolEnd` | `func(toolResult types.ToolResult)` | OnToolEnd is called after a tool successfully completes execution Use this to log successful tool completions or process tool outputs | | `OnToolError` | `func(toolCall types.ToolCall, err error)` | OnToolError is called when a tool execution fails Use this for error logging, retry logic, or fallback mechanisms | | `ToolApprovalRequired` | `bool` | ToolApprovalRequired determines if tools require approval before execution | | `ToolApprover` | `func(toolCall types.ToolCall) bool` | ToolApprover is called when a tool needs approval (if ToolApprovalRequired is true) Should return true to approve, false to reject | | `ToolApproval` | `types.ToolApprovalConfig` | ToolApproval configures automatic approval handling. Supported values are types.ToolApprovalFunc and maps keyed by tool name. Nil or zero-valued results are treated as not-applicable. | | `ExperimentalToolApprovalSecret` | `[]byte` | ExperimentalToolApprovalSecret signs issued tool approval requests and verifies resumed approval responses before tools execute (TS experimental_toolApprovalSecret). Nil disables signing/verification. | {/* /gen:fields */} ### PrepareCallConfig `PrepareCall` returns a `PrepareCallConfig` to override settings for one call. {/* gen:fields agent.PrepareCallConfig */} | Field | Type | Description | | --- | --- | --- | | `StepNumber` | `int` | StepNumber is the current step number | | `Model` | `provider.LanguageModel` | Model is the language model for this call. | | `System` | `string` | System prompt for this call | | `AllowSystemInMessages` | `bool` | AllowSystemInMessages controls whether system-role messages are accepted in this call's message list. | | `Prompt` | `string` | Prompt is the per-call user prompt when the caller used Prompt instead of Messages. | | `Messages` | `[]types.Message` | Messages for this call | | `Tools` | `[]types.Tool` | Tools available for this call | | `ActiveTools` | `[]string` | ActiveTools restricts Tools by name before calling the model. | | `ToolChoice` | `types.ToolChoice` | ToolChoice controls how the model may call tools for this call. | | `ToolOrder` | `[]string` | ToolOrder controls the order tools are sent to providers for this call. | | `ExperimentalToolCallers` | `ai.ExperimentalToolCallers` | ExperimentalToolCallers configures which tools may call which other tools for this call. | | `ToolApproval` | `types.ToolApprovalConfig` | ToolApproval configures automatic approval handling for this call. | | `ExperimentalToolApprovalSecret` | `[]byte` | ExperimentalToolApprovalSecret overrides the approval signing secret for this call. | | `SensitiveRuntimeContext` | `bool` | SensitiveRuntimeContext omits runtime context from telemetry payloads. | | `CallOptions` | `interface{}` | CallOptions are validated and passed through from AgentConfig. | | `Temperature` | `*float64` | Temperature for this call | | `MaxTokens` | `*int` | MaxTokens for this call | | `TopP` | `*float64` | TopP controls nucleus sampling. | | `TopK` | `*int` | TopK controls top-k sampling for providers that support it. | | `FrequencyPenalty` | `*float64` | FrequencyPenalty reduces repetition. | | `PresencePenalty` | `*float64` | PresencePenalty encourages topic diversity. | | `StopSequences` | `[]string` | StopSequences halt generation when encountered. | | `Seed` | `*int` | Seed enables deterministic generation when supported by the provider. | | `Headers` | `map[string]string` | Headers are custom request headers forwarded to the provider. | | `Reasoning` | `*types.ReasoningLevel` | Reasoning configures provider reasoning effort. | | `SendReasoning` | `*bool` | SendReasoning controls whether reasoning stream/content parts are exposed. | | `ProviderOptions` | `map[string]interface{}` | ProviderOptions are provider-specific options keyed by provider name. | | `StopWhen` | `[]ai.StopCondition` | StopWhen defines conditions that can terminate the tool loop. | | `Output` | `interface{}` | Output specifies how to handle and parse model output. | | `Telemetry` | `*ai.TelemetrySettings` | Telemetry configures observability for this call. | | `ExperimentalTelemetry` | `*ai.TelemetrySettings` | ExperimentalTelemetry is a deprecated alias for Telemetry. | | `Include` | `*ai.IncludeOptions` | Include controls retained request/response details. | | `Internal` | `*ai.InternalOptions` | Internal contains test-only ID generators matching TypeScript _internal. | | `ExperimentalSandbox` | `interface{}` | ExperimentalSandbox is the sandbox environment for this step. | | `ExperimentalRefineToolInput` | `map[string]ai.ToolInputRefiner` | ExperimentalRefineToolInput refines parsed tool input before execution. | | `ExperimentalDownload` | `ai.DownloadFunction` | ExperimentalDownload customizes prompt URL download handling. | | `InstructionMessages` | `[]types.Message` | InstructionMessages supplies this step's instructions as system messages. See AgentConfig.InstructionMessages. | | `PrepareStep` | `func(ctx context.Context, step ai.PrepareStepOptions) ai.PrepareStepOptions` | PrepareStep lets you provide different settings for a step. See AgentConfig.PrepareStep. | | `RepairToolCall` | `ai.ToolCallRepairFunction` | RepairToolCall attempts to repair tool calls that fail to parse for this step. | | `OnLanguageModelCallStart` | `ai.OnLanguageModelCallStartCallback` | OnLanguageModelCallStart is called immediately before this step's provider model call begins. | | `OnLanguageModelCallEnd` | `ai.OnLanguageModelCallEndCallback` | OnLanguageModelCallEnd is called after this step's provider model response is normalized and parsed, before client-side tool execution. | | `RuntimeContext` | `interface{}` | RuntimeContext is user-defined runtime data for this call. | | `ToolsContext` | `map[string]interface{}` | ToolsContext contains per-tool execution context keyed by tool name. | | `PreviousSteps` | `[]types.StepResult` | PreviousSteps contains completed step results before this call. | | `AccumulatedUsage` | `types.Usage` | AccumulatedUsage is the total usage so far | | `CustomData` | `interface{}` | CustomData allows passing custom data between PrepareCall invocations | | `MaxRetries` | `*int` | MaxRetries controls transient provider call retries for this call. | | `Timeout` | `*ai.TimeoutConfig` | Timeout provides granular timeout controls for this call, the Go stand-in for TS's per-call abortSignal. | {/* /gen:fields */} ### Tool approval signing `ExperimentalToolApprovalSecret` is a byte slice. When it is set, the SDK signs every approval request it issues with HMAC-SHA256 and verifies the signature on the approval response before it runs the tool. A client that edits the tool name, the tool call ID or the tool input between the request and the response fails verification. When it is nil, requests are not signed and responses are not verified. The signing functions are `ai.SignToolApproval` and `ai.VerifyToolApprovalSignature`; see [Tool approval helpers](https://goaisdk.com/docs/reference/ai/tool-approvals.md). ## See Also - [Agent](https://goaisdk.com/docs/reference/ai/agent.md) - The `agent.Agent` interface - [Tool](https://goaisdk.com/docs/reference/ai/tool.md) - Tool definition - [Agents Guide](https://goaisdk.com/docs/agents/overview.md) --- # Tool search and tool callers > Reference for ai.ToolSearch, deferred tool loading and the experimental tool caller routing helpers in the Go AI SDK. Canonical URL: https://goaisdk.com/docs/reference/ai/tool-search-and-callers Documentation index: https://goaisdk.com/llms.txt Two features change which tools the model sees on each step. Tool search hides tools marked `DeferLoading` until the model finds them. Tool callers route a tool so that only another tool, such as a code execution tool, can call it. Both are experimental. ## ToolSearch ```go func ToolSearch(config ...ToolSearchConfig) types.Tool ``` Returns the native tool-search tool. Pair it with tools that set `DeferLoading: true`. The generation loop binds a search registry, so a match becomes available on the next model step. `ToolSearch` panics with `*providererrors.InvalidArgumentError` when `MaxResults` is negative. {/* gen:fields ai.ToolSearchConfig */} | Field | Type | Description | | --- | --- | --- | | `Name` | `string` | Name overrides the tool's registered name (default: ToolSearchDefaultName). Set it to register more than one tool-search instance in the same request (e.g. one per tool-caller boundary), or to avoid a collision with an existing tool name. Whatever name is used here must match the corresponding entry in ExperimentalToolCallers and the tools list passed to GenerateText/StreamText. | | `MaxResults` | `int` | MaxResults caps the number of matching tools returned per search. Zero (the default) uses ToolSearchDefaultMaxResults (5). A negative value panics with an InvalidArgumentError, mirroring the TypeScript SDK's toolSearch(\{ maxResults \}) synchronous validation throw. | | `Search` | `ToolSearchRankFunc` | Search optionally selects and ranks eligible deferred tools, instead of the built-in keyword scoring over tool names and descriptions. Mirrors the TypeScript SDK's toolSearch(\{ search \}) callback. | {/* /gen:fields */} | Constant | Value | | --- | --- | | `ai.ToolSearchDefaultName` | `"toolSearch"` | | `ai.ToolSearchDefaultMaxResults` | `5` | `ai.ToolSearchRankFunc` is the type of `Search`. It is an alias for `types.ToolSearchRankFunc`. ### ToolSearchState For custom loops, `ai.NewToolSearchState(tools, toolCallers)` creates the discovery state for one generation. Do not share it across generations. Call `Apply(activeTools, toolsContext, sandbox)` once per step. It hides deferred tools that were not discovered, and rebinds any search tool to the deferred tools reachable through its callers. When no tool uses `DeferLoading` and no search tool is present, `Apply` returns its input unchanged. ## Tool callers `ai.ExperimentalToolCallers` is a `map[string][]string`. Each key is a tool name, and each value lists the tool names allowed to call it. Include `ai.DirectToolCall` (`"AI_SDK_DIRECT_TOOL_CALL"`) to let the model call the tool directly as well. A tool that is absent from the map is not routed. Set it as `ExperimentalToolCallers` on `ai.GenerateTextOptions`, `ai.StreamTextOptions` or `agent.AgentConfig`. Mark a caller tool with `types.Tool.ExperimentalToolCaller`. | Function | Description | | --- | --- | | `ai.ResolveToolCallerConfiguration(tools, toolCallers) (ResolvedToolCallers, error)` | Validates a configuration against the available tools. | | `ai.PrepareToolsForToolCallers(tools, toolCallers) (executionTools, modelTools, toolCallerMessages)` | Splits tools into the set used for execution and the set shown to the model. Collects the messages that announce local caller catalogs. | | `ai.AppendToolCallerMessages(messages, toolCallerMessages) []types.Message` | Appends the announcement messages and skips duplicates. Returns a copy. | `ai.ResolvedToolCallers` is the validated form, also a `map[string][]string`. --- # Tool > API reference for the Tool type in the Go AI SDK, defining tools that language models can call to perform actions or retrieve information. Canonical URL: https://goaisdk.com/docs/reference/ai/tool Documentation index: https://goaisdk.com/llms.txt Defines a tool that language models can call to perform actions or retrieve information. ## Type Definition ```go type Tool struct { Name string Description string Title string Parameters interface{} Execute ToolExecutor ToModelOutput ToModelOutputFunc InputExamples []ToolInputExample Strict *bool ToolApproval interface{} ProviderExecuted bool ProviderOptions interface{} OnInputStart OnInputStartFunc OnInputDelta OnInputDeltaFunc OnInputAvailable OnInputAvailableFunc } ``` ## Fields `Name` and `Parameters` are required, and `Execute` is required unless the provider runs the tool. The table is generated from the source and lists every exported field. {/* gen:fields types.Tool */} | Field | Type | Description | | --- | --- | --- | | `Name` | `string` | Name of the tool (must be unique) | | `Description` | `string` | Description of what the tool does (helps the model decide when to use it) | | `DescriptionFunc` | `func(ctx context.Context, options ToolDescriptionOptions) string` | DescriptionFunc resolves the tool description dynamically at call-prep time using the active per-tool context and sandbox. | | `Title` | `string` | Title is a short, human-readable title for the tool (optional) | | `Parameters` | `interface{}` | Parameters schema for the tool input | | `OutputSchema` | `interface{}` | OutputSchema is the schema for provider-executed tool output. It mirrors TypeScript provider tool outputSchema and is not sent as part of function tool definitions. | | `Execute` | `ToolExecutor` | Execute function that runs the tool This is not serialized to JSON | | `ToModelOutput` | `ToModelOutputFunc` | ToModelOutput allows custom conversion of tool results to model-readable output This can be used to format results in a specific way for the model If nil, the raw result will be used | | `InputExamples` | `[]ToolInputExample` | InputExamples provides example inputs to help guide the LLM These examples can improve the model's ability to use the tool correctly | | `Strict` | `*bool` | Strict enables strict schema enforcement for tool parameters. When true, the model must follow the schema exactly. A \*bool (rather than bool) lets callers distinguish "unset" from "explicitly false", so a provider can forward/warn about an explicit strict: false the same way TS does (TS checks `strict != null`), instead of only ever seeing the zero value (hand-off: "Tool.Strict bool -> \*bool"). | | `ContextSchema` | `schema.Schema` | ContextSchema optionally validates the tool-specific context passed to the tool execution and approval callbacks. | | `ToolApproval` | `interface{}` | ToolApproval indicates whether tool execution requires user approval. Can be a boolean, ToolApprovalStatus string, ToolNeedsApprovalFunc, or deprecated NeedsApprovalFunc. | | `NeedsApproval` | `interface{}` | NeedsApproval is a deprecated alias for ToolApproval. Can be a boolean or a function that determines approval based on input. Deprecated: use ToolApproval. | | `ProviderExecuted` | `bool` | ProviderExecuted indicates whether this tool is executed by the provider (not locally) When true, the tool is executed by the LLM provider (e.g., Anthropic tool-search, xAI file-search) When false or unset, the tool is executed locally by the client using the Execute function This affects error handling and validation behavior | | `SupportsDeferredResults` | `bool` | SupportsDeferredResults indicates that this provider-executed tool may not return its result in the same response that contains the tool call. When true, the step loop continues even if FinishReason is not ToolCalls, waiting for the provider to deliver the result in a later response. Only meaningful when ProviderExecuted is also true. | | `ProviderOptions` | `interface{}` | ProviderOptions contains provider-specific options for the tool This allows passing provider-specific configuration without polluting the main Tool struct Example: Anthropic cache_control, OpenAI response_format, etc. The value should be provider-specific types (e.g., anthropic.ToolOptions) | | `ProviderMetadata` | `map[string]interface{}` | ProviderMetadata carries provider-specific metadata for the tool and is propagated onto tool calls and results produced for this tool. | | `Metadata` | `map[string]interface{}` | Metadata carries tool-specific metadata that is not sent to models and is propagated via ToolCall.ToolMetadata and ToolResult.ToolMetadata. | | `Meta` | `map[string]interface{}` | Meta preserves provider/tool-definition metadata fields such as MCP `_meta`. It is exposed for callers that need the raw provider metadata but is not sent to models. | | `ProviderName` | `string` | ProviderName identifies the provider that owns a provider-defined tool. | | `Type` | `string` | Type is "function" (default), "dynamic", or "provider" for provider-defined native tools. When Type is "provider", ProviderID and ProviderArgs specify the native tool. | | `ProviderID` | `string` | ProviderID is the identifier for provider-defined tools. Example: "google.google_search", "google.url_context", "google.code_execution" | | `ProviderArgs` | `map[string]interface{}` | ProviderArgs are optional arguments for provider-defined tools. The structure depends on the specific provider tool. | | `OnInputStart` | `OnInputStartFunc` | OnInputStart is called when tool input streaming starts | | `OnInputDelta` | `OnInputDeltaFunc` | OnInputDelta is called for each delta during tool input streaming | | `OnInputAvailable` | `OnInputAvailableFunc` | OnInputAvailable is called when complete tool input is available | | `DeferLoading` | `bool` | DeferLoading defers exposing this tool to the model until it is discovered by a toolSearch tool (see ai.ToolSearch). Deferred tools still support direct calls and local callers that announce tools in conversation messages. Discovered tools become available on the next model step. | | `ExperimentalToolCaller` | `*ToolCallerDefinition` | ExperimentalToolCaller marks this tool as usable as a tool caller for other tools (TS experimental_toolCaller / ToolCallerTool). Nil means this tool is not a caller. | | `IsToolSearch` | `bool` | IsToolSearch marks this tool as the native tool search tool created by ai.ToolSearch(). It mirrors the TypeScript SDK's internal vercel.ai.toolSearch symbol tag and must not be set directly. | | `ToolSearchMaxResults` | `int` | ToolSearchMaxResults caps the number of matching tools a toolSearch tool (ai.ToolSearch()) returns per search. Zero means unset, in which case ai.ToolSearch applies its own default (5). Set only by ai.ToolSearch(); must not be set directly. | | `ToolSearchRank` | `ToolSearchRankFunc` | ToolSearchRank optionally selects and ranks eligible deferred tools for a toolSearch tool (ai.ToolSearch()) instead of its built-in keyword scoring. Mirrors the TypeScript SDK's toolSearch(\{ search \}) callback. Set only by ai.ToolSearch(); must not be set directly. | {/* /gen:fields */} `NeedsApproval` still compiles as a deprecated alias for `ToolApproval`. New code should set `ToolApproval`. Call-level approval policies (`AgentConfig.ToolApproval`, `GenerateTextOptions.ToolApproval`) take precedence and accept a function or a per-tool map. When a function tool is serialized for provider requests, nil parameters or a schema with `properties` but no `type` are normalized to an explicit object schema. For example, an empty tool schema is sent as `{"type":"object","properties":{}}`, matching the TypeScript SDK request shape. ## ToolExecutor Signature ```go type ToolExecutor func( ctx context.Context, input map[string]interface{}, options ToolExecutionOptions, ) (interface{}, error) ``` ## Examples ### Basic Tool ```go package main import ( "context" "fmt" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func main() { weatherTool := types.Tool{ Name: "get_weather", Description: "Get current weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", "description": "City name or coordinates", }, "units": map[string]interface{}{ "type": "string", "enum": []string{"celsius", "fahrenheit"}, "description": "Temperature units", }, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location := input["location"].(string) units := "celsius" if u, ok := input["units"].(string); ok { units = u } // Implement weather API call return map[string]interface{}{ "location": location, "temperature": 22, "units": units, "condition": "sunny", }, nil }, } fmt.Printf("Tool: %s\n", weatherTool.Name) } ``` ### Tool with Input Examples ```go calculatorTool := types.Tool{ Name: "calculator", Description: "Perform mathematical calculations", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "expression": map[string]interface{}{ "type": "string", "description": "Mathematical expression to evaluate", }, }, "required": []string{"expression"}, }, InputExamples: []types.ToolInputExample{ { Input: map[string]interface{}{ "expression": "2 + 2", }, Description: "Simple addition", }, { Input: map[string]interface{}{ "expression": "sqrt(16) * 3", }, Description: "Expression with function", }, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { expr := input["expression"].(string) // Implement calculator logic return fmt.Sprintf("Result: %s", expr), nil }, } ``` ### Tool with Strict Schema ```go strictTool := types.Tool{ Name: "create_user", Description: "Create a new user account", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "username": map[string]interface{}{ "type": "string", "minLength": 3, "maxLength": 20, }, "email": map[string]interface{}{ "type": "string", "format": "email", }, "age": map[string]interface{}{ "type": "integer", "minimum": 13, }, }, "required": []string{"username", "email"}, "additionalProperties": false, }, Strict: types.BoolPtr(true), // Enforce strict schema validation Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Create user return map[string]interface{}{ "success": true, "userId": "user123", }, nil }, } ``` ### Tool with Approval Required ```go deleteFileTool := types.Tool{ Name: "delete_file", Description: "Delete a file from the system", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "filepath": map[string]interface{}{ "type": "string", }, }, "required": []string{"filepath"}, }, ToolApproval: true, // Require approval before execution Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { filepath := input["filepath"].(string) // Delete file logic return fmt.Sprintf("Deleted: %s", filepath), nil }, } ``` ### Tool with Conditional Approval ```go sendEmailTool := types.Tool{ Name: "send_email", Description: "Send an email", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "to": map[string]interface{}{"type": "string"}, "subject": map[string]interface{}{"type": "string"}, "body": map[string]interface{}{"type": "string"}, }, "required": []string{"to", "subject", "body"}, }, ToolApproval: types.ToolNeedsApprovalFunc(func(ctx context.Context, input map[string]interface{}, options types.ToolNeedsApprovalOptions) bool { // Only require approval for external emails to := input["to"].(string) return !strings.HasSuffix(to, "@company.com") }), Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Send email logic return "Email sent", nil }, } ``` ### Tool with Custom Output Formatting ```go searchTool := types.Tool{ Name: "web_search", Description: "Search the web", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{"type": "string"}, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Return structured search results return []map[string]interface{}{ {"title": "Result 1", "url": "https://example.com/1"}, {"title": "Result 2", "url": "https://example.com/2"}, }, nil }, ToModelOutput: func(ctx context.Context, opts types.ToModelOutputOptions) (*types.ToolResultOutput, error) { // Format results for model consumption results := opts.Output.([]map[string]interface{}) var formatted string for i, r := range results { formatted += fmt.Sprintf("%d. %s (%s)\n", i+1, r["title"], r["url"]) } return &types.ToolResultOutput{ Type: types.ToolResultOutputText, Value: formatted, }, nil }, } ``` ### Database Query Tool ```go queryTool := types.Tool{ Name: "query_database", Description: "Query the database for information", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "table": map[string]interface{}{ "type": "string", "enum": []string{"users", "orders", "products"}, }, "filters": map[string]interface{}{ "type": "object", }, }, "required": []string{"table"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { table := input["table"].(string) filters := input["filters"] // Query database results := []map[string]interface{}{ {"id": 1, "name": "Item 1"}, {"id": 2, "name": "Item 2"}, } return results, nil }, } ``` ## See Also - [Dynamic Tools](https://goaisdk.com/docs/reference/ai/dynamic-tool.md) - Registering tools at runtime - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - Using tools with generation - [Tool Calling Guide](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) --- # Transcribe > API reference for Transcribe, the Go AI SDK function that converts audio to text using a transcription model, including URL and base64 input. Canonical URL: https://goaisdk.com/docs/reference/ai/transcribe Documentation index: https://goaisdk.com/llms.txt Converts audio to text using a transcription model. ## Signature ```go func Transcribe(ctx context.Context, opts TranscribeOptions) (*TranscribeResult, error) ``` ## TranscribeOptions | Field | Type | Required | Description | |---|---|---|---| | `Model` | `provider.TranscriptionModel` | Yes | Transcription model | | `Audio` | `[]byte` | Yes* | Audio bytes | | `AudioBase64` | `string` | Yes* | Base64 or base64url audio data, matching TypeScript `DataContent` string input | | `AudioURL` | `string` | Yes* | Remote audio URL to download before transcription | | `MimeType` | `string` | No | Audio MIME type. When omitted, `Transcribe` detects it from the audio bytes and falls back to `audio/wav` | | `ProviderOptions` | `map[string]interface{}` | No | Provider-specific options keyed by provider name | | `Headers` | `map[string]string` | No | Extra request headers | | `Download` | `ai.URLDownloadFunction` | No | Custom single-URL downloader for `AudioURL`; defaults to `CreateURLDownload(nil)` | | `MaxRetries` | `*int` | No | Maximum retries per transcription call. Defaults to 2; set to 0 to disable | | `Telemetry` | `*ai.TelemetrySettings` | No | Observability configuration for this call. `ExperimentalTelemetry` is a deprecated alias; `Telemetry` wins when both are set | *Provide one of `Audio`, `AudioBase64`, or `AudioURL`. ### Telemetry When a telemetry integration is registered (global via `telemetry.RegisterTelemetryIntegration`, or per-call via `Telemetry.Integrations`), `Transcribe` emits an `ai.transcribe` root span covering the whole call (including retries), with the input audio's size/media type (`ai.request.audio.*`, input-gated), the transcript text (`ai.response.text`, output-gated), and the provider's reported usage flattened into numeric `ai.usage.*` / `gen_ai.usage.*` attributes (e.g. `ai.usage.seconds` for a provider that reports `{"seconds": N}`; see the [Telemetry guide](https://goaisdk.com/docs/ai-sdk-core/telemetry.md#speech-and-transcription) for the key-normalization rules). A failed call — including a provider response with no transcript — ends the span with `codes.Error` status. `ExperimentalStreamTranscribe` (below) emits the analogous `ai.streamTranscribe` span. ## TranscribeResult | Field | Type | Description | |---|---|---| | `Text` | `string` | Transcribed text | | `Segments` | `[]TranscriptionSegment` | Segment-level timing entries | | `Language` | `string` | Detected language when returned by the provider | | `DurationInSeconds` | `*float64` | Optional duration (provider-dependent) | | `Warnings` | `[]types.Warning` | Provider warnings (if exposed) | | `Responses` | `[]ai.TranscriptionModelResponseMetadata` | Response metadata entries (`timestamp`, `modelId`, optional `headers`, optional `body`) | | `ProviderMetadata` | `map[string]interface{}` | Provider-specific metadata | ## Deprecated Compatibility Aliases The TypeScript SDK still exports deprecated `experimental_transcribe` and `Experimental_TranscriptionResult` names. Go exposes equivalent migration aliases: ```go func ExperimentalTranscribe(ctx context.Context, opts TranscribeOptions) (*TranscribeResult, error) type Experimental_TranscriptionResult = TranscribeResult ``` ## Example ```go audioData, _ := os.ReadFile("audio.mp3") result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, Audio: audioData, MimeType: "audio/mpeg", }) if err != nil { // handle } fmt.Println(result.Text) ``` ## URL and Base64 Input ```go result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, AudioURL: "https://example.com/audio.wav", }) ``` ```go result, err := ai.Transcribe(ctx, ai.TranscribeOptions{ Model: transcriptionModel, AudioBase64: "UklGRi...", }) ``` ## Errors When the provider response contains no transcript text, `Transcribe` returns a typed `NoTranscriptGeneratedError` with response metadata: ```go result, err := ai.Transcribe(ctx, opts) if ai.IsNoTranscriptGeneratedError(err) { // inspect err.(*ai.NoTranscriptGeneratedError).Responses if needed } ``` ## ExperimentalStreamTranscribe Streams transcripts using a transcription model that implements `provider.TranscriptionStreamer`. ### Signature ```go func ExperimentalStreamTranscribe(ctx context.Context, opts StreamTranscribeOptions) (*StreamTranscriptionResult, error) ``` ### StreamTranscribeOptions | Field | Type | Required | Description | |---|---|---|---| | `Model` | `provider.TranscriptionModel` | Yes | Must also implement `provider.TranscriptionStreamer` | | `Audio` | `provider.AudioStream` | Yes | Source of raw audio chunks to transcribe | | `InputAudioFormat` | `provider.AudioFormat` | No | Input audio format for the raw chunks | | `ProviderOptions` | `map[string]interface{}` | No | Provider-specific options keyed by provider name | | `Headers` | `map[string]string` | No | Extra HTTP/WebSocket headers when supported by the provider | | `IncludeRawChunks` | `bool` | No | Request provider raw chunks in the stream | | `Telemetry` | `*ai.TelemetrySettings` | No | Observability configuration for this call. `ExperimentalTelemetry` is a deprecated alias; `Telemetry` wins when both are set | ### Telemetry When a telemetry integration is registered, `ExperimentalStreamTranscribe` emits an `ai.streamTranscribe` root span spanning the whole stream (from before `DoStream` is called until the final `finish` part), with `gen_ai.request.stream: true`, the input audio's accumulated byte length/media type, the final transcript text, and the provider's reported usage flattened into `ai.usage.*` / `gen_ai.usage.*` attributes — the same shape as `Transcribe`'s `ai.transcribe` span (see the [Telemetry guide](https://goaisdk.com/docs/ai-sdk-core/telemetry.md#speech-and-transcription)). The span ends with `codes.Error` status, and `OnError` fires, if the stream fails or is cancelled before a `finish` part arrives. ### Example ```go sampleRate := 16000 result, err := ai.ExperimentalStreamTranscribe(ctx, ai.StreamTranscribeOptions{ Model: streamingTranscriptionModel, Audio: audioStream, InputAudioFormat: provider.AudioFormat{Type: "audio/pcm", Rate: &sampleRate}, Telemetry: &ai.TelemetrySettings{ IsEnabled: telemetry.Bool(true), }, }) if err != nil { // handle } text, err := result.Text() ``` --- # UI message chunks > Reference for every chunk type the Go AI SDK UI message stream emits, with fields, matching the AI SDK UIMessageChunk union. Canonical URL: https://goaisdk.com/docs/reference/ai/ui-message-chunks Documentation index: https://goaisdk.com/llms.txt The UI message stream is a sequence of JSON chunks sent as Server-Sent Events. Each chunk is an `ai.UIMessageChunk`, a `map[string]interface{}` with a `type` key. The chunk set matches the TypeScript AI SDK `UIMessageChunk` union, so `useChat` from `@ai-sdk/react` reads the stream without an adapter. Functions that produce the stream are in [Stream transport helpers](https://goaisdk.com/docs/reference/ai/stream-transport-helpers.md). Functions that fold chunks back into messages are in [UI messages](https://goaisdk.com/docs/reference/ai/ui-messages.md). All field names are camelCase. A field marked optional is omitted when it has no value. Several chunks accept `providerMetadata` (a `map[string]interface{}` keyed by provider name); it appears on a chunk when the provider returned metadata for it. ## Wire format Every chunk is one SSE event, `data: \n\n`. After the last chunk the stream writes `data: [DONE]\n\n`. The response carries the headers from `ai.UIMessageStreamHeaders()`: `Content-Type: text/event-stream`, `Cache-Control: no-cache`, `Connection: keep-alive`, `X-Vercel-AI-UI-Message-Stream: v1` and `X-Accel-Buffering: no`. ## Message lifecycle A stream for one assistant message looks like this: ``` start start-step text-start, text-delta..., text-end tool-input-start, tool-input-delta..., tool-input-available tool-output-available finish-step finish ``` A tool-calling run repeats the `start-step` to `finish-step` block for each model step. ### `start` Begins the message. Sent once, when `SendStart` is not false. | Field | Type | Description | | --- | --- | --- | | `messageId` | `string` | Optional. ID of the assistant message. Set from `ResponseMessageID` or `GenerateMessageID`. | | `messageMetadata` | `object` | Optional. Result of the `MessageMetadata` callback for the `start` part. | ### `finish` Ends the message. Sent once, when `SendFinish` is not false. | Field | Type | Description | | --- | --- | --- | | `finishReason` | `string` | Why generation ended: `stop`, `length`, `content-filter`, `tool-calls`, `user-approval`, `error` or `other`. | | `messageMetadata` | `object` | Optional. Result of the `MessageMetadata` callback for the `finish` part. | ### `abort` The stream was aborted. | Field | Type | Description | | --- | --- | --- | | `reason` | `string` | Optional. Why the stream was aborted. | ### `message-metadata` Updates the message metadata mid-stream. Emitted when the `MessageMetadata` callback returns a value for a part other than `start` and `finish`. A client merges the value into the message metadata. | Field | Type | Description | | --- | --- | --- | | `messageMetadata` | `object` | The metadata to merge. | ### `start-step` Begins one model call. Carries no fields. A client adds a `step-start` part to the message. ### `finish-step` Ends one model call. Carries no fields. ### `reset-step` Discards the parts added since the last `start-step` because the step is being retried. A client truncates the message to the last `step-start` part and clears in-flight text, reasoning and tool state. The Go UI stream does not emit this chunk itself. `ai.ReadUIMessages` and the workflow chat transport client handle it, so a stream from a retrying workflow agent reads correctly. ### `error` A stream error. | Field | Type | Description | | --- | --- | --- | | `errorText` | `string` | Error message. `UIMessageStreamOptions.OnError` and `UIMessageStreamResultOptions.OnError` map the Go error to this text. The default text is "An error occurred." | ## Text Text is streamed as a start, one or more deltas and an end, all sharing an `id`. | Type | Fields | | --- | --- | | `text-start` | `id` (string), `providerMetadata` (optional) | | `text-delta` | `id` (string), `delta` (string), `providerMetadata` (optional) | | `text-end` | `id` (string), `providerMetadata` (optional) | ## Reasoning Reasoning chunks have the same shape as text chunks. They are sent only when `SendReasoning` is not false (the default is true). | Type | Fields | | --- | --- | | `reasoning-start` | `id` (string), `providerMetadata` (optional) | | `reasoning-delta` | `id` (string), `delta` (string), `providerMetadata` (optional) | | `reasoning-end` | `id` (string), `providerMetadata` (optional) | ## Sources, files and custom content ### `source-url` Sent only when `SendSources` is true (the default is false). | Field | Type | Description | | --- | --- | --- | | `sourceId` | `string` | Source ID. | | `url` | `string` | Source URL. | | `title` | `string` | Optional. Source title. | | `providerMetadata` | `object` | Optional. | ### `source-document` Sent only when `SendSources` is true. | Field | Type | Description | | --- | --- | --- | | `sourceId` | `string` | Source ID. | | `mediaType` | `string` | IANA media type of the document. | | `title` | `string` | Optional. Document title. | | `filename` | `string` | Optional. File name. | | `providerMetadata` | `object` | Optional. | ### `file` A model-generated file. | Field | Type | Description | | --- | --- | --- | | `mediaType` | `string` | IANA media type. | | `url` | `string` | The file URL, or a `data:` URL when the provider returned inline bytes. | | `providerMetadata` | `object` | Optional. | ### `reasoning-file` A file generated as part of reasoning. Same fields as `file`. Sent only when `SendReasoning` is not false. ### `custom` Provider-specific content that has no standard part type. | Field | Type | Description | | --- | --- | --- | | `kind` | `string` | Provider-defined content kind. | | `providerMetadata` | `object` | Optional. | ## Tool calls ### `tool-input-start` The model started writing a tool call. Sent when the provider streams tool input. | Field | Type | Description | | --- | --- | --- | | `toolCallId` | `string` | Tool call ID. | | `toolName` | `string` | Tool name. | | `providerExecuted` | `bool` | Optional. True when the provider runs the tool. | | `dynamic` | `bool` | Optional. True for dynamic tools. | | `title` | `string` | Optional. Tool title. | | `toolMetadata` | `object` | Optional. Tool metadata from `types.Tool.Metadata`. | | `providerMetadata` | `object` | Optional. | ### `tool-input-delta` A piece of the raw JSON tool input. | Field | Type | Description | | --- | --- | --- | | `toolCallId` | `string` | Tool call ID. | | `inputTextDelta` | `string` | The next piece of input text. | ### `tool-input-available` The complete, parsed tool input. This chunk is sent even when the provider did not stream the input. | Field | Type | Description | | --- | --- | --- | | `toolCallId` | `string` | Tool call ID. | | `toolName` | `string` | Tool name. | | `input` | `any` | Parsed tool input. | | `providerExecuted` | `bool` | Optional. | | `dynamic` | `bool` | Optional. | | `title` | `string` | Optional. | | `toolMetadata` | `object` | Optional. | | `providerMetadata` | `object` | Optional. | ### `tool-input-error` The model produced a tool call that is invalid (unknown tool or input that fails validation). | Field | Type | Description | | --- | --- | --- | | `toolCallId` | `string` | Tool call ID. | | `toolName` | `string` | Tool name. | | `input` | `any` | The raw input. | | `errorText` | `string` | Error message, mapped by `OnError`. | | `providerExecuted` | `bool` | Optional. | | `dynamic` | `bool` | Optional. | | `title` | `string` | Optional. | | `toolMetadata` | `object` | Optional. | | `providerMetadata` | `object` | Optional. | ## Tool results ### `tool-output-available` The tool returned a result. | Field | Type | Description | | --- | --- | --- | | `toolCallId` | `string` | Tool call ID. | | `output` | `any` | Tool output. | | `providerExecuted` | `bool` | Optional. | | `dynamic` | `bool` | Optional. | | `preliminary` | `bool` | Optional. True for a preliminary result from a streaming tool. A later chunk replaces it. | | `toolMetadata` | `object` | Optional. | | `providerMetadata` | `object` | Optional. | ### `tool-output-error` The tool failed. | Field | Type | Description | | --- | --- | --- | | `toolCallId` | `string` | Tool call ID. | | `errorText` | `string` | Error message, mapped by `OnError`. Provider-executed tools report the raw error text. | | `providerExecuted` | `bool` | Optional. | | `dynamic` | `bool` | Optional. | | `toolMetadata` | `object` | Optional. | | `providerMetadata` | `object` | Optional. | ### `tool-output-denied` The user denied the tool call, so the tool did not run. | Field | Type | Description | | --- | --- | --- | | `toolCallId` | `string` | Tool call ID. | ## Tool approval See [Tool approval helpers](https://goaisdk.com/docs/reference/ai/tool-approvals.md) for the server-side flow. ### `tool-approval-request` The tool needs approval before it runs. The client shows a prompt and answers with a tool approval response. | Field | Type | Description | | --- | --- | --- | | `approvalId` | `string` | Approval ID. | | `toolCallId` | `string` | ID of the tool call that needs approval. | | `isAutomatic` | `bool` | Optional. True when a `ToolApproval` function decided without asking the user. | | `signature` | `string` | Optional. HMAC signature, present when `ExperimentalToolApprovalSecret` is set. | | `inputSchemaInput` | `any` | Optional. The tool input before schema parsing and input refinement, when it differs from the parsed input. | | `reason` | `string` | Optional. Why approval is requested. | ### `tool-approval-response` The approval decision. The stream emits this chunk for automatic decisions. Clients send the same shape back on the next request, as a tool part approval. | Field | Type | Description | | --- | --- | --- | | `approvalId` | `string` | Approval ID. | | `approved` | `bool` | Whether the call is approved. | | `reason` | `string` | Optional. Reason for the decision. | | `providerExecuted` | `bool` | Optional. | ## Data chunks ### `data-` A custom data part. The type is `data-` followed by a name you choose, for example `data-weather`. Write one with `UIMessageStreamWriter.Write`. | Field | Type | Description | | --- | --- | --- | | `id` | `string` | Optional. When set, a later chunk with the same type and ID replaces the data of the earlier part instead of adding a new part. | | `data` | `any` | The payload. | | `transient` | `bool` | Optional. When true, the part is sent to the client but not stored in the message. It does not reach `OnEnd` callbacks. | ## Options ### UIMessageStreamOptions Options for `ai.CreateUIMessageStreamWithOptions`. {/* gen:fields ai.UIMessageStreamOptions */} | Field | Type | Description | | --- | --- | --- | | `Execute` | `func(writer UIMessageStreamWriter)` | | | `OnError` | `func(error) string` | | | `OriginalMessages` | `[]UIMessageChunk` | | | `OnStepEnd` | `UIMessageStreamOnStepEndCallback` | | | `OnStepFinish` | `UIMessageStreamOnStepFinishCallback` | Deprecated: use OnStepEnd. | | `OnEnd` | `UIMessageStreamOnEndCallback` | | | `OnFinish` | `UIMessageStreamOnFinishCallback` | Deprecated: use OnEnd. | | `GenerateMessageID` | `IDGenerator` | | {/* /gen:fields */} ### UIMessageStreamResultOptions Options for `ai.CreateUIMessageStream`, `ai.ToUIMessageStream` and the response helpers. {/* gen:fields ai.UIMessageStreamResultOptions */} | Field | Type | Description | | --- | --- | --- | | `OriginalMessages` | `[]UIMessageChunk` | | | `GenerateMessageID` | `IDGenerator` | | | `ResponseMessageID` | `string` | | | `Tools` | `[]types.Tool` | | | `MessageMetadata` | `func(part map[string]interface{}) map[string]interface{}` | | | `SendReasoning` | `*bool` | | | `SendSources` | `*bool` | | | `SendStart` | `*bool` | | | `SendFinish` | `*bool` | | | `OnStepEnd` | `UIMessageStreamOnStepEndCallback` | | | `OnStepFinish` | `UIMessageStreamOnStepFinishCallback` | Deprecated: use OnStepEnd. | | `OnEnd` | `UIMessageStreamOnEndCallback` | | | `OnFinish` | `UIMessageStreamOnFinishCallback` | Deprecated: use OnEnd. | | `OnError` | `func(error) string` | | {/* /gen:fields */} ### UIMessageStreamResponseInit Options for the response helpers that take an init value. {/* gen:fields ai.UIMessageStreamResponseInit */} | Field | Type | Description | | --- | --- | --- | | `Status` | `int` | | | `StatusText` | `string` | | | `Headers` | `map[string]string` | Headers sets a single value per header key. Use Header instead when a header (e.g. Set-Cookie) needs multiple values. | | `Header` | `http.Header` | Header carries repeated header values (e.g. multiple Set-Cookie entries), which map[string]string cannot represent. Values here are added in addition to Headers, so both can be set together. | | `ConsumeSSEStream` | `func(io.Reader) error` | | | `KeepAliveMs` | `*time.Duration` | KeepAliveMs, when set, makes the SSE pipeline flush an opening ": stream-open\n\n" comment immediately (so headers reach the client right away even if the first real chunk is slow) and send a ": keep-alive\n\n" comment whenever this duration elapses without any other write, resetting on every chunk. Must be a positive duration; zero or negative disables it the same as a nil value. Mirrors TS createSseStreamWithKeepAlive's `keepAliveMs` option (TS #21672). | {/* /gen:fields */} ### UIMessageStreamOutcome The operation-level result of a stream. Declare one with `UIMessageStreamWriter.SetOutcome`. {/* gen:fields ai.UIMessageStreamOutcome */} | Field | Type | Description | | --- | --- | --- | | `Status` | `UIMessageStreamOutcomeStatus` | | | `Error` | `error` | Error is set when Status is UIMessageStreamOutcomeFailed. | {/* /gen:fields */} {/* gen:consts ai.UIMessageStreamOutcomeStatus */} | Constant | Value | Description | | --- | --- | --- | | `UIMessageStreamOutcomeCompleted` | `"completed"` | | | `UIMessageStreamOutcomeFailed` | `"failed"` | | | `UIMessageStreamOutcomeAborted` | `"aborted"` | | | `UIMessageStreamOutcomeUnknown` | `"unknown"` | | {/* /gen:consts */} ## Writer `ai.UIMessageStreamWriter` is passed to `UIMessageStreamOptions.Execute`. | Method | Description | | --- | --- | | `Write(part UIMessageChunk)` | Appends a chunk. | | `Merge(stream <-chan UIMessageChunk)` | Forwards every chunk from another stream. | | `SetOutcome(outcome UIMessageStreamOutcome)` | Declares the outcome. The first outcome that is not `unknown` wins. | ## Conversion functions | Function | Description | | --- | --- | | `ai.ToUIMessageChunk(part provider.StreamChunk, opts UIMessageStreamResultOptions) (UIMessageChunk, bool)` | Converts one provider stream chunk. The boolean is false when the part produces no UI chunk. | | `ai.ToUIMessageStream(ctx, stream provider.TextStream, opts ...UIMessageStreamResultOptions)` | Converts a whole provider stream to a chunk channel. | | `ai.IsUIMessageStreamError(err error) bool` | Reports whether `err` is an `*ai.UIMessageStreamError`. | `ai.UIMessageStreamError` is the error type reported through `OnError` when a chunk refers to a part that does not exist, for example a `text-delta` with no matching `text-start`. --- # UI messages > Reference for the Go AI SDK UI message types, part types, type guards, validation, conversion to model messages and stream reading. Canonical URL: https://goaisdk.com/docs/reference/ai/ui-messages Documentation index: https://goaisdk.com/llms.txt `ai.UIMessage` is the message shape `useChat` stores and sends. The Go AI SDK models it as a typed struct and as a map form (`ai.UIMessageChunk`) that matches the JSON on the wire. For the chunks that build a message while it streams, see [UI message chunks](https://goaisdk.com/docs/reference/ai/ui-message-chunks.md). ## UIMessage {/* gen:fields ai.UIMessage */} | Field | Type | Description | | --- | --- | --- | | `ID` | `string` | ID is a unique identifier for the message. | | `Role` | `UIMessageRole` | Role is system, user or assistant. | | `Metadata` | `interface{}` | Metadata is optional application metadata (omitted when nil). | | `Parts` | `[]UIMessagePart` | Parts are the message parts. Parts decoded from JSON are pointers to the concrete part types (\*TextUIPart, \*ToolUIPart, ...). | {/* /gen:fields */} {/* gen:consts ai.UIMessageRole */} | Constant | Value | Description | | --- | --- | --- | | `UIMessageRoleSystem` | `"system"` | | | `UIMessageRoleUser` | `"user"` | | | `UIMessageRoleAssistant` | `"assistant"` | | {/* /gen:consts */} ## Parts `ai.UIMessagePart` is the interface every part implements. `UIPartType()` returns the `type` discriminator. Parts decoded from JSON are pointers to the concrete types. ### TextUIPart {/* gen:fields ai.TextUIPart */} | Field | Type | Description | | --- | --- | --- | | `Text` | `string` | | | `State` | `UIPartState` | | | `ProviderMetadata` | `map[string]interface{}` | | {/* /gen:fields */} ### ReasoningUIPart {/* gen:fields ai.ReasoningUIPart */} | Field | Type | Description | | --- | --- | --- | | `ID` | `string` | | | `Text` | `string` | | | `State` | `UIPartState` | | | `ProviderMetadata` | `map[string]interface{}` | | {/* /gen:fields */} {/* gen:consts ai.UIPartState */} | Constant | Value | Description | | --- | --- | --- | | `UIPartStateStreaming` | `"streaming"` | | | `UIPartStateDone` | `"done"` | | {/* /gen:consts */} ### ToolUIPart A static tool part has type `tool-`. A dynamic tool part has type `dynamic-tool` and sets `ToolName`. `ai.DynamicToolUIPart` and `ai.ToolOutputErrorUIPart` are aliases for `ToolUIPart`. {/* gen:fields ai.ToolUIPart */} | Field | Type | Description | | --- | --- | --- | | `Type` | `string` | Type is `tool-` for static tools or `dynamic-tool`. | | `ToolName` | `string` | ToolName is set (and serialized) for dynamic tools only. | | `ToolCallID` | `string` | | | `State` | `ToolUIPartState` | | | `Title` | `string` | | | `ToolMetadata` | `map[string]interface{}` | | | `ProviderExecuted` | `*bool` | | | `Input` | `interface{}` | Input is the tool input. For input-streaming and output-error nil means absent; for other states nil is serialized as null. | | `Output` | `interface{}` | Output is the tool output (output-available). | | `RawInput` | `interface{}` | RawInput serves two purposes depending on State: - input-streaming: the accumulated raw (partial-JSON) tool input text received so far. Used to continue input streaming when a message is persisted and later resumed (see seedPartialToolCalls). - output-error: the deprecated raw input of an output-error part. Deprecated: for output-error, use Input instead. | | `ErrorText` | `string` | | | `Preliminary` | `*bool` | | | `CallProviderMetadata` | `map[string]interface{}` | | | `ResultProviderMetadata` | `map[string]interface{}` | | | `Approval` | `*ToolUIPartApproval` | | {/* /gen:fields */} {/* gen:consts ai.ToolUIPartState */} | Constant | Value | Description | | --- | --- | --- | | `ToolStateInputStreaming` | `"input-streaming"` | | | `ToolStateInputAvailable` | `"input-available"` | | | `ToolStateApprovalRequested` | `"approval-requested"` | | | `ToolStateApprovalResponded` | `"approval-responded"` | | | `ToolStateOutputAvailable` | `"output-available"` | | | `ToolStateOutputError` | `"output-error"` | | | `ToolStateOutputDenied` | `"output-denied"` | | {/* /gen:consts */} ### ToolUIPartApproval {/* gen:fields ai.ToolUIPartApproval */} | Field | Type | Description | | --- | --- | --- | | `ID` | `string` | | | `Approved` | `*bool` | Approved is nil while approval is requested. | | `Descriptor` | `interface{}` | Descriptor is an application-defined approval descriptor. | | `RequestReason` | `string` | RequestReason is the reason the approval was requested (user-approval status with a reason). | | `Reason` | `string` | Reason is the approval response reason. | | `IsAutomatic` | `*bool` | | | `Signature` | `string` | | | `InputSchemaInput` | `interface{}` | InputSchemaInput is the tool input before schema parsing and input refinement (present only when it differs from Input). | {/* /gen:fields */} ### SourceURLUIPart {/* gen:fields ai.SourceURLUIPart */} | Field | Type | Description | | --- | --- | --- | | `SourceID` | `string` | | | `URL` | `string` | | | `Title` | `string` | | | `ProviderMetadata` | `map[string]interface{}` | | {/* /gen:fields */} ### SourceDocumentUIPart {/* gen:fields ai.SourceDocumentUIPart */} | Field | Type | Description | | --- | --- | --- | | `SourceID` | `string` | | | `MediaType` | `string` | | | `Title` | `string` | | | `Filename` | `string` | | | `ProviderMetadata` | `map[string]interface{}` | | {/* /gen:fields */} ### FileUIPart {/* gen:fields ai.FileUIPart */} | Field | Type | Description | | --- | --- | --- | | `MediaType` | `string` | | | `Filename` | `string` | | | `URL` | `string` | | | `ProviderReference` | `map[string]string` | | | `ProviderMetadata` | `map[string]interface{}` | | {/* /gen:fields */} ### ReasoningFileUIPart {/* gen:fields ai.ReasoningFileUIPart */} | Field | Type | Description | | --- | --- | --- | | `MediaType` | `string` | | | `URL` | `string` | | | `ProviderMetadata` | `map[string]interface{}` | | {/* /gen:fields */} ### CustomContentUIPart {/* gen:fields ai.CustomContentUIPart */} | Field | Type | Description | | --- | --- | --- | | `Kind` | `string` | | | `ProviderMetadata` | `map[string]interface{}` | | {/* /gen:fields */} ### DataUIPart `Type` is `data-`. `DataName()` returns the name without the prefix. {/* gen:fields ai.DataUIPart */} | Field | Type | Description | | --- | --- | --- | | `Type` | `string` | | | `ID` | `string` | | | `Data` | `interface{}` | | {/* /gen:fields */} ### StepStartUIPart An empty struct with type `step-start`. It marks the start of a model step. ## Type guards and accessors | Function | Description | | --- | --- | | `ai.IsTextUIPart(part)` | True for `text` parts. | | `ai.IsReasoningUIPart(part)` | True for `reasoning` parts. | | `ai.IsReasoningFileUIPart(part)` | True for `reasoning-file` parts. | | `ai.IsFileUIPart(part)` | True for `file` parts. | | `ai.IsCustomContentUIPart(part)` | True for `custom` parts. | | `ai.IsDataUIPart(part)` | True for `data-*` parts. | | `ai.IsToolUIPart(part)` | True for static and dynamic tool parts. | | `ai.IsStaticToolUIPart(part)` | True for `tool-` parts. | | `ai.IsDynamicToolUIPart(part)` | True for `dynamic-tool` parts. | | `ai.IsToolOutputErrorUIPart(part)` | True for tool parts in the `output-error` state. | | `ai.AsToolUIPart(part) (*ToolUIPart, bool)` | Returns the tool part behind a value or pointer. | | `ai.GetStaticToolName(part *ToolUIPart) string` | The tool name of a static tool part. | | `ai.GetToolName(part *ToolUIPart) string` | The tool name of a static or dynamic tool part. | | `ai.GetToolOrDynamicToolName(part *ToolUIPart) string` | Alias for `GetToolName`. | ## JSON and map conversion | Function | Description | | --- | --- | | `ai.UnmarshalUIMessagePart(data []byte) (UIMessagePart, error)` | Decodes one part into its concrete pointer type. | | `ai.UIMessageFromChunk(message UIMessageChunk) (UIMessage, error)` | Converts a map-shaped message, such as `responseMessage` in an end callback, into a `UIMessage`. | | `ai.UIMessagesFromChunks(messages []UIMessageChunk) ([]UIMessage, error)` | Converts a slice of map-shaped messages. | | `ai.UIMessagesToChunks(messages []UIMessage) ([]UIMessageChunk, error)` | Converts typed messages to map-shaped messages. | | `UIMessage.ToChunk() (UIMessageChunk, error)` | Converts one typed message to a map. | ## ReadUIMessages ```go func ReadUIMessages(ctx context.Context, opts ReadUIMessagesOptions) (<-chan UIMessage, <-chan error) ``` Folds a chunk stream into a channel of progressively updated messages. Each value is the assistant message so far. {/* gen:fields ai.ReadUIMessagesOptions */} | Field | Type | Description | | --- | --- | --- | | `Message` | `*UIMessage` | Message is the last assistant message to continue when a conversation is resumed. Optional. | | `Stream` | `<-chan UIMessageChunk` | Stream is the UI message chunk stream to read. | | `OnError` | `func(error)` | OnError is called for processing errors (e.g. a delta for a missing part). Optional. | | `TerminateOnError` | `bool` | TerminateOnError stops reading and reports the first error on the error channel. | {/* /gen:fields */} ## Validation ```go func ValidateUIMessages(ctx context.Context, opts ValidateUIMessagesOptions) ([]UIMessage, error) func SafeValidateUIMessages(ctx context.Context, opts ValidateUIMessagesOptions) SafeValidateUIMessagesResult func ValidateUIMessagesForAgent(ctx context.Context, opts ValidateUIMessagesOptions) ([]UIMessage, error) ``` `SafeValidateUIMessages` returns a result instead of an error. `ValidateUIMessagesForAgent` converts terminal tool parts whose tools are not registered, for example tools from a disconnected MCP server, to dynamic tool parts instead of failing. ### ValidateUIMessagesOptions {/* gen:fields ai.ValidateUIMessagesOptions */} | Field | Type | Description | | --- | --- | --- | | `Messages` | `interface{}` | Messages are the messages to validate. Accepts []UIMessage, []UIMessageChunk, decoded JSON ([]interface\{\}), json.RawMessage, []byte or string JSON, or any value that marshals to a UI message array. nil means "not provided". | | `MetadataSchema` | `schema.Schema` | MetadataSchema validates each message's metadata. The validated value (with schema defaults applied) replaces the metadata. | | `DataSchemas` | `map[string]schema.Schema` | DataSchemas validate data parts by name (without the `data-` prefix). When set, a data part without a schema is an error. Validated values (with schema defaults applied) replace the part data. | | `Tools` | `[]types.Tool` | Tools validate static tool parts against the current tool input (Parameters) and output (OutputSchema) schemas. A nil slice means no tools were provided (tool parts are not validated); a non-nil empty slice validates against an empty tool set. | | `ExperimentalRefineToolInput` | `map[string]ToolInputRefiner` | ExperimentalRefineToolInput reconstructs the approved input from an approval's inputSchemaInput before comparing it with the part input. | {/* /gen:fields */} ### SafeValidateUIMessagesResult {/* gen:fields ai.SafeValidateUIMessagesResult */} | Field | Type | Description | | --- | --- | --- | | `Success` | `bool` | | | `Data` | `[]UIMessage` | | | `Error` | `error` | | {/* /gen:fields */} ## Conversion to model messages ```go func ConvertToModelMessages(ctx context.Context, messages []UIMessage, opts ...ConvertToModelMessagesOptions) ([]types.Message, error) ``` Converts UI messages to the `types.Message` slice a model call takes. A failure returns a `*ai.MessageConversionError`; `ai.IsMessageConversionError(err)` tests for it. {/* gen:fields ai.ConvertToModelMessagesOptions */} | Field | Type | Description | | --- | --- | --- | | `Tools` | `[]types.Tool` | Tools are used to convert tool outputs with Tool.ToModelOutput. | | `IgnoreIncompleteToolCalls` | `bool` | IgnoreIncompleteToolCalls drops tool parts that are not in a completed state (approval-responded, final output-available, output-error, output-denied). | | `ConvertDataPart` | `func(part DataUIPart) types.ContentPart` | ConvertDataPart converts data parts to text or file model message parts (types.TextContent / types.FileContent). Returning nil skips the part. Without it, data parts are ignored. | {/* /gen:fields */} ## Auto-resubmit predicates | Function | Description | | --- | --- | | `ai.LastAssistantMessageIsCompleteWithToolCalls(messages []UIMessage) bool` | True when the last message is an assistant message whose last step has client-executed tool calls and every one has reached a final output or an error, with no text still streaming. Mirrors the TypeScript helper of the same name. | | `ai.LastAssistantMessageIsCompleteWithApprovalResponses(messages []UIMessage) bool` | True when the last message is an assistant message whose last step has at least one tool call in the `approval-responded` state and every tool call in that step has reached a terminal state. Provider-executed calls count. Mirrors the TypeScript helper of the same name. | --- # Utilities > Reference for Go AI SDK utility types and functions: message pruning, timeouts, stream adapters, partial output parsing, evaluation, realtime sessions, sandboxes and warnings. Canonical URL: https://goaisdk.com/docs/reference/ai/utilities Documentation index: https://goaisdk.com/llms.txt Smaller exported types and functions in `pkg/ai` that the other reference pages do not cover. ## Message pruning Prune a conversation to fit a token budget. ```go func PruneMessages(ctx context.Context, messages []types.Message, opts PruneOptions) ([]types.Message, error) func PruneToFitContext(ctx context.Context, messages []types.Message, contextWindow int, reserveTokens int) ([]types.Message, error) func DefaultMessagePrune(ctx context.Context, messages []types.Message, maxTokens int) ([]types.Message, error) func DefaultPruneOptions(maxTokens int) PruneOptions ``` `ai.MessagePruneFunc` is the function type of `DefaultMessagePrune`: `func(ctx context.Context, messages []types.Message, maxTokens int) ([]types.Message, error)`. `DefaultMessagePrune` keeps the system message and the last messages and drops older ones. {/* gen:fields ai.PruneOptions */} | Field | Type | Description | | --- | --- | --- | | `MaxTokens` | `int` | MaxTokens is the maximum number of tokens allowed | | `PreserveSystemMessage` | `bool` | PreserveSystemMessage keeps the system message even if pruning is needed | | `PreserveLastN` | `int` | PreserveLastN keeps the last N messages | | `PruneFunc` | `MessagePruneFunc` | PruneFunc is the custom pruning function | {/* /gen:fields */} ## Timeouts `ai.TimeoutConfig` sets total, per-step and per-chunk timeouts on `GenerateTextOptions` and `StreamTextOptions`. `ai.NewTimeoutConfig(total, perStep, perChunk *time.Duration)` builds one. Pass nil for a timeout you do not want. `ai.TimeoutReason` names the boundary that aborted an operation. {/* gen:consts ai.TimeoutReason */} | Constant | Value | Description | | --- | --- | --- | | `TimeoutReasonTotal` | `"total"` | | | `TimeoutReasonStep` | `"step"` | | | `TimeoutReasonChunk` | `"chunk"` | | | `TimeoutReasonFirstChunk` | `"firstChunk"` | | | `TimeoutReasonTool` | `"tool"` | | {/* /gen:consts */} ## Stream status `ai.StreamStatus` reports the state of a stream for UI code. {/* gen:consts ai.StreamStatus */} | Constant | Value | Description | | --- | --- | --- | | `StreamStatusSubmitted` | `"submitted"` | StreamStatusSubmitted indicates the request has been submitted and the stream is actively receiving data from the model. | | `StreamStatusStreaming` | `"streaming"` | StreamStatusStreaming indicates at least one chunk has been received. | | `StreamStatusDone` | `"done"` | StreamStatusDone indicates the stream has completed successfully. | {/* /gen:consts */} ## Building stream results ### NewStreamTextResultFromParts ```go func NewStreamTextResultFromParts(ctx context.Context, src provider.TextStream, opts ExternalStreamOptions) *StreamTextResult ``` Builds a `*StreamTextResult` from a stream you already have, instead of calling a model. The stream carries every step's chunks in order. Use it to put the SDK's result accessors and UI helpers over a custom source. {/* gen:fields ai.ExternalStreamOptions */} | Field | Type | Description | | --- | --- | --- | | `CallID` | `string` | CallID correlates this result with the caller's own callback/telemetry events. A new one is generated when empty. | | `Provider` | `string` | Provider and ModelID label steps and the response metadata when a chunk does not carry its own provider.ResponseMetadata. | | `ModelID` | `string` | | | `OnChunk` | `func(chunk provider.StreamChunk)` | OnChunk is invoked for every chunk as it is consumed from src, mirroring StreamTextOptions.OnChunk. | | `OnFinish` | `func(result *StreamTextResult)` | OnFinish/OnEnd are invoked once, after src is fully consumed (whether it ended in success or with an error), mirroring StreamTextOptions' callbacks of the same name. OnEnd takes precedence when both are set. | | `OnEnd` | `func(result *StreamTextResult)` | | | `Output` | `interface{}` | Output, when it satisfies the internal outputProcessor interface (any value returned by TextOutput/ObjectOutput/ArrayOutput/ChoiceOutput/ JSONOutput), enables structured-output parsing on the returned result: Output()/OutputErr() resolve once src is fully consumed by parsing the final step's accumulated text, and PartialOutput() updates after every text chunk of that step, deduplicated exactly like StreamTextOptions. Output. A value that does not satisfy outputProcessor is ignored (mirrors StreamTextOptions.Output's `interface{}` + type-assertion contract; see stream.go's own opts.Output handling for the pattern this mirrors). | {/* /gen:fields */} ### SimulateReadableStream ```go func SimulateReadableStream(chunks []provider.StreamChunk, perChunkDelay time.Duration) provider.TextStream ``` Creates a stream from in-memory chunks, with a delay between chunks. Use it in tests. ### Stream transform emitters A `StreamTransformFunc` can emit chunks before it returns, which keeps output flowing for transforms that wait. | Name | Description | | --- | --- | | `ai.StreamTransformEmitter` | `func(provider.StreamChunk)`. Emits one chunk immediately. | | `ai.WithStreamTransformEmitter(ctx, emit)` | Returns a context that carries an emitter. The stream loop installs one around every transform call. | | `ai.StreamTransformEmitterFromContext(ctx)` | Returns the emitter and true inside a transform call. Outside one it returns false, and the transform should return its chunks as a batch. | ### ChunkDetector `ai.ChunkDetector` is one of the values `ai.SmoothStreamOptions.Chunking` accepts, next to `"word"`, `"line"` and a `*regexp.Regexp`. It takes the buffered text and returns the first complete chunk and true, or an empty string and false when the buffer holds no complete chunk. It mirrors the TypeScript `ChunkDetector`. ### Text stream response init `ai.CreateTextStreamResponseWithInit(ctx, result, init *TextStreamResponseInit)` is `CreateTextStreamResponse` with a custom status, status text and headers. `ai.PipeTextStreamToWriter(ctx, stream provider.TextStream, w io.Writer) error` writes text deltas from a provider stream to `w` as UTF-8. If a write fails or `ctx` ends first, it closes the stream before it returns, so upstream resources are released. ## Parsing partial output | Name | Description | | --- | --- | | `ai.ParsePartialJSON(text string) jsonparser.ParseResult` | Parses JSON that may be incomplete. | | `ai.ParsePartialOutputOptions` | Input for an output specification's partial parse: `Text`. | | `ai.ParseCompleteOutputOptions` | Input for the final parse: `Text`, `Response`, `Usage` and `FinishReason`. | | `ai.GetTextFromDataURL(dataURL string) (string, error)` | Decodes a base64 `text/*` data URL, using its declared charset (UTF-8 by default). | ### Object output modes {/* gen:consts ai.ObjectOutputMode */} | Constant | Value | Description | | --- | --- | --- | | `ObjectModeObject` | `"object"` | ObjectModeObject returns a single object (default) | | `ObjectModeArray` | `"array"` | ObjectModeArray returns an array of objects (streaming) | | `ObjectModeEnum` | `"enum"` | ObjectModeEnum forces selection from enum values | | `ObjectModeNoSchema` | `"no-schema"` | ObjectModeNoSchema returns raw JSON without validation | {/* /gen:consts */} ### ElementStreamOptions Options for streaming array elements as they complete. {/* gen:fields ai.ElementStreamOptions */} | Field | Type | Description | | --- | --- | --- | | `ElementSchema` | `schema.Schema` | ElementSchema defines the structure of each array element | | `OnElement` | `func(element ElementStreamResult[ELEMENT])` | OnElement is called when a new element is parsed | | `OnError` | `func(err error)` | OnError is called when an error occurs during parsing | | `OnComplete` | `func()` | OnComplete is called when the stream completes | {/* /gen:fields */} {/* gen:fields ai.ElementStreamResult */} | Field | Type | Description | | --- | --- | --- | | `Element` | `ELEMENT` | The parsed element | | `Index` | `int` | Index of this element in the array | | `IsFinal` | `bool` | Whether this is the final element in the array | {/* /gen:fields */} ## Evaluation `ai.ExperimentalEvaluate` asks an evaluation model typed questions and returns one typed answer per question. `ai.NewEvaluationLanguageModel` returns an `*ai.EvaluationLanguageModel`, which adapts any language model into an evaluation model by prompting it with a JSON-schema-constrained request. Reasoning defaults to `none` unless you override it with provider options. {/* gen:fields ai.EvaluationLanguageModelOptions */} | Field | Type | Description | | --- | --- | --- | | `Model` | `provider.LanguageModel` | | | `Provider` | `string` | Provider overrides the reported provider ID. Defaults to "<model.Provider()>.evaluation". | {/* /gen:fields */} ### EvaluateResult {/* gen:fields ai.EvaluateResult */} | Field | Type | Description | | --- | --- | --- | | `Answers` | `map[string]provider.EvaluationAnswer` | Answers has exactly one typed answer per question, keyed by question ID. | | `Usage` | `EvaluateUsage` | | | `Warnings` | `[]types.Warning` | | | `Rounding` | `*provider.EvaluationRounding` | | | `ProviderMetadata` | `map[string]interface{}` | | | `Response` | `EvaluateResponse` | | {/* /gen:fields */} {/* gen:fields ai.EvaluateUsage */} | Field | Type | Description | | --- | --- | --- | | `InputTokens` | `*int` | | | `OutputTokens` | `*int` | | | `TotalTokens` | `*int` | | {/* /gen:fields */} {/* gen:fields ai.EvaluateResponse */} | Field | Type | Description | | --- | --- | --- | | `ID` | `string` | | | `Timestamp` | `time.Time` | | | `ModelID` | `string` | | | `Headers` | `map[string]string` | | | `Body` | `interface{}` | | {/* /gen:fields */} ## Per-call diagnostics `GenerateImageResult.Calls` holds one `ai.GenerateImageCall` per provider call. Each step result from text generation carries an `ai.GenerateStepRequest` and an `ai.GenerateStepResponse`. {/* gen:fields ai.GenerateImageCall */} | Field | Type | Description | | --- | --- | --- | | `Images` | `[]types.GeneratedFile` | | | `ProviderMetadata` | `map[string]interface{}` | | | `Response` | `*types.ResponseMetadata` | | | `Warnings` | `[]types.Warning` | | | `Usage` | `types.ImageUsage` | | {/* /gen:fields */} {/* gen:fields ai.GenerateStepRequest */} | Field | Type | Description | | --- | --- | --- | | `Body` | `interface{}` | Body is the raw request body sent to the provider API (for debugging). | {/* /gen:fields */} {/* gen:fields ai.GenerateStepResponse */} | Field | Type | Description | | --- | --- | --- | | `ID` | `string` | ID is the provider-assigned response identifier when available. For streaming paths, this may be a generated fallback value. | | `Timestamp` | `time.Time` | Timestamp is when the provider started generating the response. Zero value means it was not available. | | `ModelID` | `string` | ModelID is the model that handled the request, when available. | | `Headers` | `map[string]string` | Headers are the raw HTTP response headers from the provider. | | `Messages` | `[]types.Message` | Messages are the response messages generated in this step (assistant message + any tool messages). | | `Body` | `interface{}` | Body is the raw response body from the provider (for debugging). | {/* /gen:fields */} ## Realtime sessions `ai.ConnectRealtime(ctx, model, opts RealtimeSessionOptions)` opens a `*ai.RealtimeSession` on an experimental realtime model. `Send(ctx, event)` writes a client event. `Read(ctx)` returns the next batch of server events. `Close()` ends the session. | Name | Description | | --- | --- | | `ai.RealtimeDialer` | Opens a connection: `Dial(ctx, provider.WebSocketConfig) (RealtimeWebSocketConn, error)`. | | `ai.WebSocketRealtimeDialer` | The default dialer, built on `golang.org/x/net/websocket`. The `Origin` field is unused. | | `ai.RealtimeWebSocketConn` | A connection with `Send`, `Receive` and `Close`. | | `ai.RealtimeBinaryConn` | Optional capability for a connection that can send and receive binary frames (`SendBinary`, `ReceiveFrame`). Without it, every frame is text. | | `ai.GetRealtimeToolDefinitions(tools) ([]provider.RealtimeToolDefinition, error)` | Converts tools to realtime tool definitions. | | `ai.GetRealtimeToolDefinitionsWithOptions(opts RealtimeToolDefinitionsOptions)` | Same, with a per-tool context. | {/* gen:fields ai.RealtimeToolDefinitionsOptions */} | Field | Type | Description | | --- | --- | --- | | `Tools` | `[]types.Tool` | | | `ToolsContext` | `map[string]interface{}` | | {/* /gen:fields */} ## Sandboxes `ai.ShellSandbox` runs commands through the local shell with context cancellation. `ai.NewShellSandbox(opts ...ShellSandboxOption)` creates one, and `ai.WithShellSandboxDescription(description)` sets its description. The sandbox types `ai.SandboxProcess`, `ai.SandboxProcessResult`, `ai.SandboxReadTextFileOptions` and `ai.SandboxWriteTextFileOptions` re-export the types in `pkg/providerutils`. ## Warnings and telemetry | Name | Description | | --- | --- | | `ai.LogWarningsFunction` | `func(options LogWarningsOptions)`. Install one with `ai.SetLogWarnings` to replace the default stderr logger. | | `ai.DisableLogWarnings()` | Turns warning logging off. The `AI_SDK_LOG_WARNINGS=false` environment variable does the same. | | `ai.TelemetryOptions` | Alias for `telemetry.Options`. | ## URL support checks | Function | Description | | --- | --- | | `ai.SupportedURLCheckerForModel(model)` | Returns a `func(mediaType, rawURL string) bool` for a model that exposes `SupportedURLs`. A model without it supports no direct URLs. | | `ai.SupportedURLCheckerFromPatterns(patternsByMediaType)` | Compiles a `SupportedURLs`-style map of regular expressions, keyed by media type, into the same checker. | --- # Video operations > Reference for the experimental Go AI SDK functions that start a video generation and check its status later, with option and result types. Canonical URL: https://goaisdk.com/docs/reference/ai/video-operations Documentation index: https://goaisdk.com/llms.txt `ai.ExperimentalStartVideo` starts a video generation and returns without waiting. `ai.ExperimentalGetVideoStatus` checks the operation later, from any process. Use them for long-running generations. The model must implement `provider.VideoModelStarter`. ```go func ExperimentalStartVideo(ctx context.Context, opts StartVideoOptions) (*StartVideoResult, error) func ExperimentalGetVideoStatus(ctx context.Context, model provider.VideoModelV3, opts GetVideoStatusOptions) (*GetVideoStatusResult, error) ``` ## StartVideoOptions {/* gen:fields ai.StartVideoOptions */} | Field | Type | Description | | --- | --- | --- | | `Model` | `provider.VideoModelV3` | Model to use. Must implement provider.VideoModelStarter. | | `Prompt` | `VideoPrompt` | Prompt can be text-only or image+text for image-to-video. | | `N` | `int` | Number of videos to generate (default: 1). Must not exceed the model's MaxVideosPerCall — fan out with multiple StartVideo calls. | | `MaxVideosPerCall` | `*int` | Maximum videos per API call (provider-specific). If not set, uses the model's MaxVideosPerCall() value for the limit check. | | `AspectRatio` | `string` | Aspect ratio in format "width:height", or "adaptive". | | `Resolution` | `string` | Resolution in format "widthxheight". | | `Duration` | `*float64` | Duration in seconds. | | `FPS` | `*int` | Frames per second. | | `Seed` | `*int` | Seed for reproducible generation. | | `FrameImages` | `[]VideoFrameImageInput` | FrameImages are role-tagged image inputs for image-to-video and first-last-frame generation. | | `InputReferences` | `[]VideoReferenceInput` | InputReferences are reference image or video inputs for reference-to-video generation. | | `GenerateAudio` | `*bool` | GenerateAudio requests that the model generate audio alongside the video, when supported. | | `ProviderOptions` | `map[string]interface{}` | Provider-specific options. | | `MaxRetries` | `*int` | Maximum retries for the start call (default: 2). Set to 0 to disable. | | `Headers` | `map[string]string` | Additional HTTP headers. | | `WebhookURL` | `string` | WebhookURL, when set, asks the provider to notify this URL when the generation reaches a terminal state. | {/* /gen:fields */} ## StartVideoResult `Operation` is an opaque JSON value. Persist it and pass it to `ExperimentalGetVideoStatus`. {/* gen:fields ai.StartVideoResult */} | Field | Type | Description | | --- | --- | --- | | `Operation` | `json.RawMessage` | Operation is a JSON-serializable opaque reference to the started generation. Persist it and pass it to ExperimentalGetVideoStatus to retrieve the status and result later, from any process. | | `Warnings` | `[]types.Warning` | Warnings for the call, e.g. unsupported settings. | | `ProviderMetadata` | `map[string]interface{}` | ProviderMetadata is passed through from the provider. Carries the provider's own job identifiers (e.g. the AI Gateway's providerMetadata.gateway.asyncJob.jobId). | | `Response` | `VideoModelResponseMetadata` | Response is response metadata from the provider. | {/* /gen:fields */} ## GetVideoStatusOptions {/* gen:fields ai.GetVideoStatusOptions */} | Field | Type | Description | | --- | --- | --- | | `Operation` | `json.RawMessage` | Operation is the opaque reference returned by ExperimentalStartVideo. | | `Headers` | `map[string]string` | Additional HTTP headers. | | `MaxRetries` | `*int` | Maximum retries for the status call (default: 2). Set to 0 to disable. | {/* /gen:fields */} ## GetVideoStatusResult `Status` is one of the `provider.VideoOperationStatus*` constants. `Videos` is set when the status is completed. `Error` is set when the status is error. {/* gen:fields ai.GetVideoStatusResult */} | Field | Type | Description | | --- | --- | --- | | `Status` | `string` | | | `Videos` | `[]provider.VideoModelV3VideoData` | Videos is set when Status is provider.VideoOperationStatusCompleted. | | `Error` | `string` | Error is set when Status is provider.VideoOperationStatusError. | | `Warnings` | `[]types.Warning` | | | `ProviderMetadata` | `map[string]interface{}` | | | `Response` | `VideoModelResponseMetadata` | | {/* /gen:fields */} ## VideoFrameImageInput A role-tagged image for image-to-video and first-and-last-frame generation. {/* gen:fields ai.VideoFrameImageInput */} | Field | Type | Description | | --- | --- | --- | | `Image` | `VideoPromptImage` | Image is the file used for this frame. | | `FrameType` | `string` | FrameType is provider.VideoFrameTypeFirstFrame or provider.VideoFrameTypeLastFrame. | {/* /gen:fields */} ## Errors `ai.IsNoVideoGeneratedError(err)` reports whether `err` is a `*ai.NoVideoGeneratedError`. --- # Harness adapters > Reference for the Harness adapter interface, sessions and prompt controls in the Go AI SDK, and for the Claude Code, Codex and other built-in adapters. Canonical URL: https://goaisdk.com/docs/reference/harness/adapters Documentation index: https://goaisdk.com/llms.txt An adapter wraps one coding-agent runtime. It implements `harness.Harness` and returns a `harness.Session`. `harness.Agent` is the only consumer of that interface, so a test double is a first-class substitute for a real adapter. Application code usually builds an adapter with its constructor and hands it to `harness.NewAgent`. See [Harness agent](https://goaisdk.com/docs/reference/harness/agent.md). ## Harness ```go type Harness interface { SpecificationVersion() string // always "harness-v1" HarnessID() string BuiltinTools() map[string]BuiltinTool DoStart(ctx context.Context, opts StartOptions) (Session, error) } ``` `harness.SpecificationVersion` is the version string. `BuiltinTools` is keyed by what the runtime reports on tool-call events: the common name when it has one, otherwise the native name. Optional capabilities are separate interfaces that an adapter may also implement: | Interface | Purpose | | --- | --- | | `harness.BootstrapProvider` | Declare a [bootstrap recipe](https://goaisdk.com/docs/reference/harness/sandbox.md#bootstrap-recipes). | | `harness.BuiltinToolApprovalSupport` | Emit approval requests for built-in tools. | | `harness.BuiltinToolFilteringSupport` | Hide inactive built-in tools from the runtime. | | `harness.LifecycleStateValidator` | Validate the adapter-defined `data` in lifecycle state. | ### StartOptions {/* gen:fields harness.StartOptions */} | Field | Type | Description | | --- | --- | --- | | `Headers` | `map[string]string` | Headers are additional normalized HTTP headers to send with model requests. | | `SessionID` | `string` | SessionID is the stable identifier for this harness session. | | `ResumeFrom` | `*ResumeSessionState` | ResumeFrom is the resume payload from a prior lifecycle method. | | `ContinueFrom` | `*ContinueTurnState` | ContinueFrom is the continuation payload from DoSuspendTurn, or nested in ResumeFrom. | | `PermissionMode` | `PermissionMode` | PermissionMode is the approval policy for built-in adapter-native tools. | | `BuiltinToolFiltering` | `*BuiltinToolFiltering` | BuiltinToolFiltering lists adapter-native built-in tools available for this session. | | `Observability` | `*Observability` | Observability is diagnostics wiring; nil when diagnostics are disabled. | | `SandboxSession` | `providerutils.SandboxSession` | SandboxSession is the sandbox the adapter operates against. It is either a NetworkSandboxSession or a caller-provided plain providerutils.SandboxSession (filesystem + process only). Adapters must not stop or destroy the sandbox themselves. | | `SessionWorkDir` | `string` | SessionWorkDir is the absolute path the adapter runs the agent in. | {/* /gen:fields */} `StartOptions.SandboxSession` is a `harness.NetworkSandboxSession` or a plain `providerutils.SandboxSession` that you supply. An adapter must not stop or destroy the sandbox itself. ## Session ```go type Session interface { SessionID() string IsResume() bool DoPromptTurn(ctx context.Context, opts PromptTurnOptions) (PromptControl, error) DoCompact(ctx context.Context, customInstructions string) error DoContinueTurn(ctx context.Context, opts ContinueTurnOptions) (PromptControl, error) DoSuspendTurn(ctx context.Context) (*ContinueTurnState, error) DoDetach(ctx context.Context) (*ResumeSessionState, error) DoStop(ctx context.Context) (*ResumeSessionState, error) DoDestroy(ctx context.Context) error } ``` An adapter whose runtime cannot compact returns `*harness.CapabilityUnsupportedError` from `DoCompact`. A session may also implement `harness.HistoryReader` (`DoReadHistory(ctx, since string) (*ReadHistoryResult, error)`). `since` is a previous result's cursor, which is opaque to the host. An adapter that cannot reach its runtime's store returns `*harness.HistoryUnavailableError`. ### PromptTurnOptions and ContinueTurnOptions Both embed `harness.TurnSettings`, the framework-owned settings captured when a turn begins: `Model`, `Skills`, `Instructions` and `Tools`. {/* gen:fields harness.PromptTurnOptions */} | Field | Type | Description | | --- | --- | --- | | (embedded) | `TurnSettings` | Embedded. | | `Prompt` | `Prompt` | Prompt is the fresh input for this turn. | | `ResponseFormat` | `*ResponseFormat` | ResponseFormat requested for this turn. Adapters that cannot honor a JSON response format must return a CapabilityUnsupportedError. | | `Emit` | `EmitFunc` | Emit is invoked once for each event the adapter produces. | {/* /gen:fields */} {/* gen:fields harness.ContinueTurnOptions */} | Field | Type | Description | | --- | --- | --- | | (embedded) | `TurnSettings` | Embedded. | | `ResponseFormat` | `*ResponseFormat` | ResponseFormat of the in-flight turn. | | `Emit` | `EmitFunc` | Emit is invoked once for each event the adapter produces. | {/* /gen:fields */} {/* gen:fields harness.TurnSettings */} | Field | Type | Description | | --- | --- | --- | | `Model` | `string` | | | `Skills` | `[]Skill` | | | `Instructions` | `string` | | | `Tools` | `[]ToolSpec` | | {/* /gen:fields */} `harness.Prompt` is plain text or one user message. Build one with `harness.TextPrompt(text)`. {/* gen:fields harness.Prompt */} | Field | Type | Description | | --- | --- | --- | | `Text` | `string` | | | `Message` | `*types.Message` | | {/* /gen:fields */} `harness.ResponseFormat` is the response format for a turn. `Type` is `harness.ResponseFormatText` or `harness.ResponseFormatJSON`. Adapters that cannot produce JSON return `*harness.CapabilityUnsupportedError`. {/* gen:fields harness.ResponseFormat */} | Field | Type | Description | | --- | --- | --- | | `Type` | `ResponseFormatType` | | | `Schema` | `map[string]any` | | | `Name` | `string` | | | `Description` | `string` | | {/* /gen:fields */} ### Skills `harness.Skill` is a self-contained instruction bundle the runtime can load. `harness.SkillFile` is an extra file, with a skill-relative POSIX `Path` and its `Content`. `harness.AgentSkill` is the consumer-facing alias. {/* gen:fields harness.Skill */} | Field | Type | Description | | --- | --- | --- | | `Name` | `string` | | | `Description` | `string` | | | `Content` | `string` | | | `Files` | `[]SkillFile` | | {/* /gen:fields */} ## PromptControl `DoPromptTurn` and `DoContinueTurn` return a `harness.PromptControl`, the control surface for the turn. ```go type PromptControl interface { SubmitToolResult(ctx context.Context, result ToolResultSubmission) error Done() <-chan struct{} // closed when the turn ends Err() error // the turn error, after Done is closed } ``` {/* gen:fields harness.ToolResultSubmission */} | Field | Type | Description | | --- | --- | --- | | `ToolCallID` | `string` | | | `Output` | `any` | | | `IsError` | `bool` | | | `ToolResult` | `any` | ToolResult optionally carries the full tool-result part (TS `ToolResultPart`) for adapters that forward it verbatim. | {/* /gen:fields */} A control can also implement: | Interface | Method | | --- | --- | | `harness.ToolApprovalSubmitter` | `SubmitToolApproval(ctx, ToolApprovalSubmission) error` | | `harness.UserMessageSubmitter` | Accepts a user message mid-turn, which is what `ExperimentalSteer` uses. | | `harness.CheckpointPinner` | `PinCheckpoint() (release func())`. Pins the replay checkpoint so already-delivered bridge events are kept until you call `release`. | {/* gen:fields harness.ToolApprovalSubmission */} | Field | Type | Description | | --- | --- | --- | | `ApprovalID` | `string` | | | `Approved` | `bool` | | | `Reason` | `string` | | {/* /gen:fields */} ## Consumer-facing aliases The consumer-facing `harness.Agent*` names are aliases of the specification types, matching the TypeScript `HarnessAgent*` names. | Alias | Type | | --- | --- | | `harness.AgentAdapter` | `harness.Harness` | | `harness.AgentAdapterSession` | `harness.Session` | | `harness.AgentPrompt` | `harness.Prompt` | | `harness.AgentPromptControl` | `harness.PromptControl` | | `harness.AgentPromptTurnOptions` | `harness.PromptTurnOptions` | | `harness.AgentStartOptions` | `harness.StartOptions` | ## Built-in adapters Each adapter is its own package under `pkg/harness`. Every adapter has a `HarnessID` and a `CredentialEnvironmentVariables` list, and most have a bootstrap recipe under a `.harness-bootstrap/` directory. | Package | Constructor | Runtime | | --- | --- | --- | | `harness/claudecode` | `claudecode.New(Settings) (*Harness, error)` | Claude Code. `HarnessID` is `claude-code`. | | `harness/codex` | `codex.New(Settings) *Harness` | OpenAI Codex. `HarnessID` is `codex`. | | `harness/opencode` | `opencode.CreateOpenCode(...Settings) (harness.Harness, error)` | OpenCode. | | `harness/deepagents` | `deepagents.CreateDeepAgents(...Settings) harness.Harness` | Deep Agents. | | `harness/cursor` | `cursor.CreateCursor(...Settings) (harness.Harness, error)` | Cursor, through ACP. | | `harness/githubcopilot` | `githubcopilot.CreateGitHubCopilot(...Settings) (harness.Harness, error)` | GitHub Copilot, through ACP. | | `harness/grokbuild` | `grokbuild.CreateGrokBuild(...Settings) (harness.Harness, error)` | Grok Build, through ACP. | | `harness/acp` | `acp.CreateACP(Settings) (harness.Harness, error)` | Any runtime that speaks the Agent Client Protocol. The Cursor, GitHub Copilot and Grok Build adapters are configurations of it. | ### Claude Code `claudecode.New` starts the Claude Agent SDK in a bridge process inside the sandbox and talks to it over a port the sandbox exposes. {/* gen:fields harness/claudecode.Settings */} | Field | Type | Description | | --- | --- | --- | | `Auth` | `harness.Authentication` | Auth selects the authentication route: unset (auto-detect), "direct", "ai-gateway", or an isolated environment via harness.AuthEnvironment. | | `CredentialForwarding` | `harness.CredentialForwarding` | CredentialForwarding customizes each credential value immediately before it is forwarded into the sandbox. | | `MCPServers` | `map[string]any` | MCPServers are additional MCP server definitions, keyed by name, in the Claude Agent SDK's native configuration format. The name "harness-tools" is reserved. | | `MaxTurns` | `int` | MaxTurns caps how many internal turns the CLI can take before yielding back to the caller. Zero means the CLI's default. | | `AgentProgressSummaries` | `bool` | AgentProgressSummaries enables periodic AI-generated progress summaries for running subagents. The summaries are forwarded in raw `task_progress` stream parts. | | `ForwardSubagentText` | `bool` | ForwardSubagentText forwards subagent text and thinking messages in addition to tool activity. Subagent messages are exposed as raw stream parts. | | `Env` | `map[string]string` | Env are additional environment variables for the Claude Code process, merged over the resolved authentication environment. | | `Thinking` | `*ThinkingConfig` | Thinking controls extended-thinking behavior. Defaults to \{Type: "adaptive", Display: "summarized"\}. | | `Effort` | `string` | Effort controls adaptive-thinking effort: "low", "medium", "high", "xhigh" or "max". Empty uses the Claude Agent SDK default. | | `Port` | `int` | Port overrides the port the bridge binds inside the sandbox. Defaults to the first port the sandbox exposes. | | `PortEndpoint` | `*harness.PortEndpoint` | PortEndpoint overrides the host endpoint used to reach the bridge. Required together with Port for a sandbox that is not a harness.NetworkSandboxSession. | | `StartupTimeout` | `time.Duration` | StartupTimeout bounds how long to wait for the bridge to announce its port. Defaults to 120s. | | `Reconnect` | `bridge.ReconnectOptions` | Reconnect tunes reconnection after an established bridge connection drops. | | `MintBridgeToken` | `harness.MintBridgeTokenCallback` | MintBridgeToken creates the bridge channel token. Defaults to a random 32-byte hex token. Requires the sandbox to expose an ID. | {/* /gen:fields */} {/* gen:fields harness/claudecode.ThinkingConfig */} | Field | Type | Description | | --- | --- | --- | | `Type` | `string` | Type is "adaptive" (default), "enabled" or "disabled". | | `Display` | `string` | Display is "summarized" (default) or "omitted"; ignored when Type is "disabled". | {/* /gen:fields */} `claudecode.GetBootstrap(ctx)` returns the bootstrap recipe. `claudecode.BootstrapDir` is `.harness-bootstrap/claude-code`. `claudecode.AuthModeDirect` and its siblings are the resolved authentication modes. ### Codex {/* gen:fields harness/codex.Settings */} | Field | Type | Description | | --- | --- | --- | | `Auth` | `harness.Authentication` | Auth selects the authentication route: unset (auto-detect), "direct", "ai-gateway", or an isolated environment via harness.AuthEnvironment. | | `CredentialForwarding` | `harness.CredentialForwarding` | CredentialForwarding customizes each credential value immediately before it is forwarded into the sandbox. | | `CodexConfig` | `map[string]any` | CodexConfig is additional configuration passed through to Codex as-is (snake_case config.toml keys). Values managed by this adapter take precedence over conflicting entries. | | `MCPServers` | `map[string]any` | MCPServers are additional MCP server definitions, keyed by name, in Codex's native configuration format. | | `ReasoningEffort` | `string` | ReasoningEffort for reasoning-capable models: "low", "medium", "high", "xhigh" or "max". Empty defers to the CLI's default. | | `WebSearch` | `*bool` | WebSearch, when true, allows the runtime to use live web search. | | `Port` | `int` | Port overrides the port the bridge binds inside the sandbox. Defaults to the first port the sandbox exposes. | | `PortEndpoint` | `*harness.PortEndpoint` | PortEndpoint overrides the host endpoint used to reach the bridge. | | `StartupTimeout` | `time.Duration` | StartupTimeout bounds how long to wait for the bridge to announce its port. Defaults to 120s. | | `Reconnect` | `bridge.ReconnectOptions` | Reconnect tunes reconnection after an established bridge connection drops. | | `MintBridgeToken` | `harness.MintBridgeTokenCallback` | MintBridgeToken creates the bridge channel token. Defaults to a random 32-byte hex token. Requires the sandbox to expose an ID. | {/* /gen:fields */} `codex.DefaultModel` is the model used when none is set. `codex.DefaultOpenAIBaseURL` is `https://api.openai.com/v1`. ### Authentication Every adapter takes an `Auth` setting of type `harness.Authentication`. Leave it unset to detect the route, or choose `harness.AuthModeDirect`, `harness.AuthModeAIGateway` or an isolated environment with `harness.AuthEnvironment`. Adapters forward credentials into the sandbox through request transformations where the provider supports it, so the sandbox does not hold the real secret. See [Request transformations](https://goaisdk.com/docs/reference/harness/sandbox.md#request-transformations). --- # Harness agent > Reference for harness.Agent, AgentSettings and AgentSession in the Go AI SDK: running coding-agent runtimes such as Claude Code and Codex as an agent, with sessions, approvals and lifecycle state. Canonical URL: https://goaisdk.com/docs/reference/harness/agent Documentation index: https://goaisdk.com/llms.txt Package `github.com/digitallysavvy/go-ai/pkg/harness` runs a third-party coding-agent runtime, such as Claude Code or Codex, as an `agent.Agent`. A harness adapter wraps the runtime. `harness.Agent` merges the adapter's built-in tools with your tools, creates a sandbox session, runs prompt turns and streams the result in the same form as `StreamText`. This page covers the consumer API. The adapter interface is in [Harness adapters](https://goaisdk.com/docs/reference/harness/adapters.md), sandboxes in [Harness sandboxes](https://goaisdk.com/docs/reference/harness/sandbox.md) and the stream part types in [Harness stream parts](https://goaisdk.com/docs/reference/harness/stream-parts.md). ## NewAgent ```go func NewAgent(settings AgentSettings) (*Agent, error) ``` Validates the settings and returns an agent. `*harness.Agent` implements `agent.Agent`, so you can use it wherever an agent is expected, including the [agent UI stream helpers](https://goaisdk.com/docs/reference/ai/agent-ui-stream-helpers.md). ```go import ( "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/harness" "github.com/digitallysavvy/go-ai/pkg/harness/claudecode" "github.com/digitallysavvy/go-ai/pkg/harness/sandbox/local" ) h, err := claudecode.New(claudecode.Settings{}) if err != nil { return err } coder, err := harness.NewAgent(harness.AgentSettings{ Harness: h, Sandbox: local.NewProvider(local.Options{}), }) if err != nil { return err } session, err := coder.CreateSession(ctx, harness.CreateSessionOptions{}) if err != nil { return err } defer session.Destroy(ctx) result, err := coder.Generate(ctx, agent.AgentGenerateOptions{ Prompt: "Add a README to this project.", HarnessSession: session, }) ``` ## AgentSettings {/* gen:fields harness.AgentSettings */} | Field | Type | Description | | --- | --- | --- | | `Harness` | `Harness` | Harness is the adapter driving the underlying agent runtime. Its builtin tools are merged with UserTools and exposed via Agent.Tools(). | | `ID` | `string` | ID is exposed via HarnessAgent.ID(). | | `Model` | `string` | Model is the model identifier used by the harness adapter. PrepareCall can replace it between completed turns. | | `UserTools` | `map[string]types.Tool` | UserTools are tools available to the runtime in addition to the harness's own builtins. User tools take precedence over harness builtins on key collision. | | `ToolsContext` | `map[string]interface{}` | ToolsContext is per-tool context passed to host-executed tools. | | `RuntimeContext` | `interface{}` | RuntimeContext is user-defined context passed to lifecycle callbacks and telemetry (subject to Telemetry.IncludeRuntimeContext). A per-call agent.AgentGenerateOptions.RuntimeContext takes precedence when set. Mirrors TS `HarnessAgentSettings.runtimeContext`. | | `Skills` | `[]Skill` | Skills made available to the underlying runtime. | | `Instructions` | `interface{}` | Instructions for the underlying agent runtime. Adapters append these to a native system/developer prompt when supported. Accepts either a plain string or a \*types.Message (a system message; only its Content text is forwarded to the harness adapter) for parity with ToolLoopAgent's Instructions field, mirroring TS `HarnessAgentSettings.instructions: string \| SystemModelMessage` (TS 4d1bf28). PrepareCall can replace it (with the same two shapes) between completed turns. | | `Headers` | `map[string]string` | Headers are additional HTTP headers sent with every model request. "authorization", "x-api-key", "user-agent" and "x-client-app" are managed by the harness and must not be set here. | | `PrepareCall` | `func(ctx context.Context, opts PrepareCallOptions) (PrepareCallResult, error)` | PrepareCall derives the prompt and the settings that may vary between completed turns. See PrepareCallOptions/PrepareCallResult. | | `Output` | `interface{}` | Output is an optional specification for generating typed output (e.g. ai.ObjectOutput/ai.ArrayOutput/ai.ChoiceOutput/ai.JSONOutput/ ai.TextOutput), active for every turn this agent runs. Mirrors TS `HarnessAgentSettings.output`. The same value type StreamText's StreamTextOptions.Output/GenerateTextOptions.Output accept — it is resolved against pkg/ai's internal outputProcessor interface, so any value that does not implement it is silently ignored, exactly like those options. See Agent.HasOutput and startTurn's ResponseFormat derivation. | | `StopWhen` | `[]ai.StopCondition` | StopWhen are the conditions that stop the current result after a completed harness tool step that could continue into another model step. The underlying turn remains unfinished and can be suspended and continued. Nil means the harness runs until the turn naturally finishes or pauses. | | `Callbacks` | `Callbacks` | | | `PermissionMode` | `PermissionMode` | PermissionMode is the baseline permission mode for adapter-native built-in tools. Defaults to PermissionModeAllowAll. | | `ToolApproval` | `AgentToolApprovalConfiguration` | ToolApproval is the per-tool static approval configuration for host-executed tools. | | `ActiveTools` | `[]string` | ActiveTools / InactiveTools limit the tools available to the harness without changing the tool call/result types in the result. Mutually exclusive; leave both nil to allow every tool. | | `InactiveTools` | `[]string` | | | `Sandbox` | `SandboxProvider` | Sandbox is the provider used to create or resume network sandbox sessions. When nil, every CreateSession call must provide an existing sandbox session. | | `SandboxConfig` | `AgentSandboxConfig` | SandboxConfig is the sandbox working-directory and lifecycle-hook configuration. | | `Telemetry` | `*telemetry.Options` | Telemetry configures OpenTelemetry span/attribute reporting for every turn this agent runs, via pkg/telemetry's dispatch pattern (see run_prompt.go's telemetry.go): a turn span nests step spans, which nest model-call and tool-execution spans. Nil disables it. Mirrors TS `HarnessAgentSettings.telemetry`. | {/* /gen:fields */} ### Callbacks The callbacks reuse the event types from `pkg/ai`, so a harness turn appears in callback-driven tooling like a `StreamText` call. See [Callback events](https://goaisdk.com/docs/reference/ai/callback-events.md). {/* gen:fields harness.Callbacks */} | Field | Type | Description | | --- | --- | --- | | `OnStart` | `func(ctx context.Context, e ai.OnStartEvent)` | | | `OnStepStart` | `func(ctx context.Context, e ai.OnStepStartEvent)` | | | `OnLanguageModelCallStart` | `ai.OnLanguageModelCallStartCallback` | | | `OnLanguageModelCallEnd` | `ai.OnLanguageModelCallEndCallback` | | | `OnToolExecutionStart` | `func(ctx context.Context, e ai.OnToolCallStartEvent)` | | | `OnToolExecutionEnd` | `func(ctx context.Context, e ai.OnToolCallFinishEvent)` | | | `OnStepEnd` | `func(ctx context.Context, e ai.OnStepFinishEvent)` | | | `OnEnd` | `func(ctx context.Context, e ai.OnFinishEvent)` | | {/* /gen:fields */} ### PrepareCall `PrepareCall` runs before each turn. It can replace the model, skills, instructions and tools between completed turns, and it can rewrite the prompt. {/* gen:fields harness.PrepareCallOptions */} | Field | Type | Description | | --- | --- | --- | | `CallOptions` | `interface{}` | | | `Prompt` | `Prompt` | | | `Model` | `string` | | | `Skills` | `[]Skill` | | | `Instructions` | `interface{}` | Instructions is either a string or a \*types.Message, exactly as configured on AgentSettings.Instructions (or overridden by the per-call Instructions option). See PrepareCallResult.Instructions. | | `Tools` | `map[string]types.Tool` | | | `ToolsContext` | `map[string]interface{}` | | | `RuntimeContext` | `interface{}` | RuntimeContext is the turn's runtime context before PrepareCall runs: AgentSettings.RuntimeContext, or a per-call agent.AgentGenerateOptions.RuntimeContext override when the caller supplied one. Mirrors TS `prepareCall`'s `runtimeContext` input field. | {/* /gen:fields */} {/* gen:fields harness.PrepareCallResult */} | Field | Type | Description | | --- | --- | --- | | `Model` | `string` | | | `Skills` | `[]Skill` | | | `Instructions` | `interface{}` | Instructions is either a string or a \*types.Message (a system message; only its Content text is used), mirroring TS `HarnessAgentSettings`'s `instructions?: string \| SystemModelMessage` (TS 4d1bf28). Extracted to a plain string via instructionsText before being handed to the harness adapter. | | `HasInstructions` | `bool` | HasInstructions distinguishes "clear the instructions" (Instructions == "", HasInstructions == true) from "leave them as configured" (HasInstructions == false). Go has no `undefined`, so this flag plays its role for prepareCall's rest-spread removal pattern. | | `Tools` | `map[string]types.Tool` | | | `ToolsContext` | `map[string]interface{}` | | | `RuntimeContext` | `interface{}` | RuntimeContext replaces the turn's runtime context when HasRuntimeContext is true (including clearing it, by leaving RuntimeContext nil) — mirrors TS `prepareCall`'s result spreading its own `runtimeContext` key (even `undefined`) over the configured default. See HasInstructions for why Go needs this flag where TS uses `undefined`. | | `HasRuntimeContext` | `bool` | | | `Prompt` | `Prompt` | | {/* /gen:fields */} ## Agent methods | Method | Description | | --- | --- | | `Generate(ctx, agent.AgentGenerateOptions) (*ai.GenerateTextResult, error)` | Runs a fresh prompt turn on `HarnessSession` and waits for the result. | | `Stream(ctx, agent.AgentStreamOptions) (*ai.StreamTextResult, error)` | Runs a fresh prompt turn on `HarnessSession` and streams the result. | | `ContinueGenerate(ctx, opts, toolApprovalContinuations, toolResultContinuations)` | Resumes a paused turn and waits for the result. | | `ContinueStream(ctx, opts, toolApprovalContinuations, toolResultContinuations)` | Resumes a paused turn and streams the result. | | `Execute(ctx, prompt string) (*agent.AgentResult, error)`, `ExecuteWithMessages(ctx, messages)` | The `agent.Agent` entry points. | | `CreateSession(ctx, CreateSessionOptions) (*AgentSession, error)` | Creates or resumes a session. | | `ExperimentalSteer(ctx, session, text string) error` | Sends an extra user message to the session's running turn, if the adapter supports it. | | `GetSandboxTemplate(ctx) (*HarnessSandboxTemplate, error)` | Returns the reusable sandbox template for this agent. Returns `nil, nil` when there is nothing to prepare. | | `Tools() []types.Tool` | The merged tool set. User tools win over built-in tools on a name collision. | | `ID() string`, `HarnessID() string`, `Version() string`, `HasOutput() bool` | Identity and capability accessors. | `harness.AgentVersion` is the agent interface specification version. `Generate`, `Stream`, `ContinueGenerate` and `ContinueStream` need a session. Pass it as `HarnessSession` on `agent.AgentGenerateOptions` (or on the embedded options of `agent.AgentStreamOptions`). `Execute` and `ExecuteWithMessages` create a session, run one turn and destroy the session. To keep a runtime across turns, create a session with `CreateSession`. ## Sessions `CreateSession` returns an `*harness.AgentSession`, the handle for one runtime session. One turn can be in flight at a time. `Generate` and `Stream` need the turn state to be `idle`. `ContinueGenerate` and `ContinueStream` need `awaiting-approval`, `awaiting-tool-result` or `suspended`. ### CreateSessionOptions {/* gen:fields harness.CreateSessionOptions */} | Field | Type | Description | | --- | --- | --- | | `SessionID` | `string` | SessionID is the stable identifier for the underlying sandbox/session. Generated when empty. | | `ResumeFrom` | `*ResumeSessionState` | ResumeFrom is the payload from a prior Detach/Stop. Mutually exclusive with ContinueFrom. | | `ContinueFrom` | `*ContinueTurnState` | ContinueFrom is the payload from a prior SuspendTurn. Mutually exclusive with ResumeFrom. | | `ToolsContext` | `map[string]interface{}` | ToolsContext rebinds host-only tool context for an unfinished turn resumed with ContinueFrom (directly or nested in ResumeFrom). Only valid together with one of those. | | `RuntimeContext` | `interface{}` | RuntimeContext rebinds host-only runtime context for an unfinished turn resumed with ContinueFrom (directly or nested in ResumeFrom). Runtime context is not serialized into lifecycle state. Mirrors TS `HarnessAgent.createSession`'s `options.runtimeContext`. | | `SandboxSession` | `providerutils.SandboxSession` | SandboxSession is a caller-owned sandbox session. When set, the caller retains ownership of its lifecycle; Agent.Sandbox is not consulted. | {/* /gen:fields */} ### AgentSession methods | Method | Description | | --- | --- | | `SessionID() string` | The session ID. | | `GetSessionWorkDir() string` | The working directory for this session in the sandbox. | | `GetSandboxSession() providerutils.SandboxSession` | The underlying sandbox session. | | `HasUnfinishedTurn() bool` | Whether a turn is in flight or paused. | | `Compact(ctx, customInstructions string) error` | Asks the runtime to compact its context. | | `ReadHistory(ctx, since string) (*ReadHistoryResult, error)` | Reads the conversation history the runtime persisted. Returns `*HistoryUnavailableError` when the adapter supports it but cannot reach the history. | | `ExperimentalSteerTurn(ctx, text string) error` | Sends an extra user message to the running turn. | | `SuspendTurn(ctx) (*ContinueTurnState, error)` | Freezes the active turn and returns the state to persist. The handle is detached afterward, so create a new session from the state to continue. | | `Detach(ctx) (*ResumeSessionState, error)` | Detaches without tearing down the runtime. Pass the state to `CreateSessionOptions.ResumeFrom` later. | | `Stop(ctx) (*ResumeSessionState, error)` | Persists enough state to resume, then stops the runtime. | | `Destroy(ctx) error` | Stops the runtime and returns no state. | {/* gen:consts harness.SessionState */} | Constant | Value | Description | | --- | --- | --- | | `SessionStateActive` | `"active"` | | | `SessionStateDetached` | `"detached"` | | | `SessionStateStopped` | `"stopped"` | | | `SessionStateDestroyed` | `"destroyed"` | | {/* /gen:consts */} {/* gen:consts harness.TurnState */} | Constant | Value | Description | | --- | --- | --- | | `TurnStateIdle` | `"idle"` | | | `TurnStateRunning` | `"running"` | | | `TurnStateAwaitingApproval` | `"awaiting-approval"` | | | `TurnStateAwaitingResult` | `"awaiting-tool-result"` | | | `TurnStateSuspended` | `"suspended"` | | {/* /gen:consts */} ### History {/* gen:fields harness.ReadHistoryResult */} | Field | Type | Description | | --- | --- | --- | | `Messages` | `[]HistoryMessage` | | | `Cursor` | `string` | Cursor is an opaque position; pass it back as the next read's Since to read only what follows. | {/* /gen:fields */} {/* gen:fields harness.HistoryMessage */} | Field | Type | Description | | --- | --- | --- | | (embedded) | `types.Message` | Embedded. | | `At` | `string` | At is the message's timestamp, when the adapter's runtime records one. | | `HarnessMetadata` | `Metadata` | HarnessMetadata is adapter-namespaced opaque data for this message, e.g. the runtime's raw record. | {/* /gen:fields */} ## Lifecycle state `Detach`, `Stop` and `SuspendTurn` return state you can persist as JSON and pass back to `CreateSession`. State written by the TypeScript SDK decodes in Go and the reverse. | Name | Description | | --- | --- | | `harness.LifecycleState` | Either a `*ResumeSessionState` or a `*ContinueTurnState`. | | `harness.ResumeSessionState` | State between turns. Accepted as `StartOptions.ResumeFrom`. | | `harness.ContinueTurnState` | State of a suspended turn. Accepted as `StartOptions.ContinueFrom`. | | `harness.NewResumeSessionState(...)`, `harness.NewContinueTurnState(...)` | Build a state. The adapter-defined `data` is marshaled to JSON. | | `harness.DecodeLifecycleState(data []byte) (LifecycleState, error)` | Parses and validates persisted JSON. | | `harness.LifecycleStateResumeSession`, `harness.LifecycleStateContinueTurn` | The `type` discriminators. | | `harness.LifecycleStateValidator` | Optional adapter interface that validates the adapter-defined `data`. | | `harness.AgentLifecycleState`, `harness.AgentResumeSessionState`, `harness.AgentContinueTurnState`, `harness.AgentContinueTurnOptions` | Consumer-facing aliases of the types above. | {/* gen:fields harness.ResumeSessionState */} | Field | Type | Description | | --- | --- | --- | | `Type` | `string` | | | `HarnessID` | `string` | | | `SpecificationVersion` | `string` | | | `Data` | `json.RawMessage` | | | `ContinueFrom` | `*ContinueTurnState` | ContinueFrom is optional unfinished-turn state. | {/* /gen:fields */} {/* gen:fields harness.ContinueTurnState */} | Field | Type | Description | | --- | --- | --- | | `Type` | `string` | | | `HarnessID` | `string` | | | `SpecificationVersion` | `string` | | | `Data` | `json.RawMessage` | | | `PendingToolApprovals` | `[]PendingToolApproval` | | | `PendingToolResults` | `[]PendingToolResult` | | | `TurnSettings` | `*TurnSettings` | | {/* /gen:fields */} ## Permissions and approvals `PermissionMode` sets the baseline for the adapter's native built-in tools. The default is `harness.DefaultPermissionMode` (`allow-all`). {/* gen:consts harness.PermissionMode */} | Constant | Value | Description | | --- | --- | --- | | `PermissionModeAllowReads` | `"allow-reads"` | | | `PermissionModeAllowEdits` | `"allow-edits"` | | | `PermissionModeAllowAll` | `"allow-all"` | | {/* /gen:consts */} | Name | Description | | --- | --- | | `harness.ResolvePermissionMode(mode)` | Returns `mode`, or the default when empty. | | `PermissionMode.Valid()` | Reports whether the mode is one of the three values. | | `harness.PermissionModeNeedsBuiltinSupport(mode)` | Reports whether the mode needs the adapter to support built-in tool approvals. | | `harness.SupportsBuiltinToolApprovals(h Harness) bool` | Reports whether the adapter implements `BuiltinToolApprovalSupport`. | | `harness.AgentPermissionMode` | Consumer-facing alias of `PermissionMode`. | `AgentSettings.ToolApproval` is an `harness.AgentToolApprovalConfiguration`, a `harness.ToolApprovalConfiguration`: a map of host-tool names to a static approval status. `harness.ResolveCustomToolApproval(...)` maps one entry to a `harness.CustomToolApprovalDecision` for one host tool call. {/* gen:fields harness.CustomToolApprovalDecision */} | Field | Type | Description | | --- | --- | --- | | `Type` | `CustomToolApprovalDecisionType` | | | `Reason` | `string` | | {/* /gen:fields */} `harness.CustomToolApprovalDecisionType` discriminates the decision: `harness.CustomToolApprovalAllow`, `harness.CustomToolApprovalDeny` or `harness.CustomToolApprovalRequest`. ### Pausing for approval or results A turn pauses when the runtime wants approval for a tool, or when it calls a client-executed tool. The paused state is recorded as a `harness.PendingToolApproval` or `harness.PendingToolResult`. Their kinds are `harness.PendingToolApprovalBuiltin` and `harness.PendingToolApprovalCustom`. A `harness.CompletedToolResult` is a client tool result executed while suspending, submitted on resume without running the tool again. {/* gen:fields harness.PendingToolApproval */} | Field | Type | Description | | --- | --- | --- | | `ApprovalID` | `string` | | | `ToolCallID` | `string` | | | `ToolName` | `string` | | | `Input` | `string` | | | `Kind` | `string` | "builtin" \| "custom" | | `ProviderExecuted` | `*bool` | | | `NativeName` | `string` | | {/* /gen:fields */} {/* gen:fields harness.PendingToolResult */} | Field | Type | Description | | --- | --- | --- | | `ToolCallID` | `string` | | | `ToolName` | `string` | | | `Input` | `string` | | | `ProviderOptions` | `map[string]map[string]any` | | | `CompletedResult` | `*CompletedToolResult` | | {/* /gen:fields */} {/* gen:fields harness.CompletedToolResult */} | Field | Type | Description | | --- | --- | --- | | `Output` | `any` | | | `IsError` | `*bool` | | | `ToolResult` | `any` | | {/* /gen:fields */} To resume, take the client's trailing tool message and pass the continuations to `ContinueGenerate` or `ContinueStream`: | Function | Description | | --- | --- | | `harness.CollectToolApprovalContinuations(messages []types.Message) ([]types.ToolApprovalResponseContent, error)` | Extracts approval decisions from a trailing `tool` message. Responses that already have a tool result are ignored. | | `harness.CollectToolResultContinuations(messages []types.Message) []types.ToolResultContent` | Extracts client-provided tool results from the trailing `tool` message. | `harness.AgentPendingToolApproval` and `harness.AgentPendingToolResult` are the consumer-facing aliases. ## Tools | Name | Description | | --- | --- | | `harness.BuiltinTool` | A tool the adapter's runtime exposes natively: a `types.Tool` plus harness metadata. | | `harness.CommonTool(name, CommonToolOptions)` | Declares a built-in tool that maps to a cross-harness common name. | | `harness.StandardBuiltinTools()` | The cross-harness vocabulary of common built-in tools with baseline input schemas. | | `harness.BuiltinToolNames` | The common built-in tool names, in declaration order: `harness.BuiltinToolRead`, `harness.BuiltinToolWrite`, `harness.BuiltinToolEdit`, `harness.BuiltinToolBash`, `harness.BuiltinToolGrep`, `harness.BuiltinToolGlob`, `harness.BuiltinToolWebSearch` and `harness.BuiltinToolAskUserQuestions`. | | `harness.BuiltinToolName`, `harness.BuiltinToolUseKind` | A common name, and a classification for permission handling: `harness.BuiltinToolUseKindReadonly`, `harness.BuiltinToolUseKindEdit` or `harness.BuiltinToolUseKindBash`. | | `harness.AgentBuiltinTool`, `harness.AgentBuiltinToolName`, `harness.AgentBuiltinToolUseKind`, `harness.AgentToolSpec` | Consumer-facing aliases. | | `harness.ToolSpec` | Describes a host-defined tool made available to the runtime. | | `harness.BuiltinToolFiltering` | Selects the adapter-native built-in tools for a session. `Mode` is `harness.BuiltinToolFilteringAllow` or `harness.BuiltinToolFilteringDeny`. | | `harness.BuiltinToolFilteringMode` | The type of `Mode`. | | `harness.ResolveToolFiltering(ResolveToolFilteringOptions) ResolvedToolFiltering` | Computes the active user tools and the built-in filtering to send to the adapter from `ActiveTools` and `InactiveTools`. | | `harness.IsBuiltinToolIncluded(...)` | Reports whether a built-in tool is allowed by a filtering. | | `harness.BuiltinToolFilteringDenialReason(...)` | The reason a built-in tool is denied. | | `harness.SupportsBuiltinToolFiltering(h Harness) bool` | Reports whether the adapter implements `BuiltinToolFilteringSupport`. | {/* gen:fields harness.ResolveToolFilteringOptions */} | Field | Type | Description | | --- | --- | --- | | `Harness` | `Harness` | | | `UserTools` | `map[string]types.Tool` | | | `AllTools` | `map[string]types.Tool` | | | `ActiveTools` | `[]string` | nil when unset | | `InactiveTools` | `[]string` | nil when unset | {/* /gen:fields */} {/* gen:fields harness.ResolvedToolFiltering */} | Field | Type | Description | | --- | --- | --- | | `ActiveUserTools` | `map[string]types.Tool` | ActiveUserTools is the (possibly narrowed) set of host-executed tools. | | `BuiltinToolFiltering` | `*BuiltinToolFiltering` | BuiltinToolFiltering is nil when every built-in tool stays available. | {/* /gen:fields */} ### Questions tool The `askUserQuestions` tool lets the runtime ask the user a question and wait for the answer. | Name | Description | | --- | --- | | `harness.QuestionsToolDescription` | The tool description. | | `harness.QuestionsToolInput`, `harness.QuestionsToolOutput` | Input and output of the tool. | | `harness.Question`, `harness.QuestionOption`, `harness.QuestionAnswer` | One question, one choice and one answer. | | `harness.AllowFreeForm` | `boolean` or `{ secret: boolean }`. Controls free-form answers. | | `harness.ParseQuestionsToolInput(...)`, `harness.ParseQuestionsToolOutput(...)` | Decode and validate the tool input and output. | | `harness.QuestionsToolInputJSONSchema()`, `harness.QuestionsToolOutputJSONSchema()` | The JSON Schemas. | | `harness.QuestionsActionAnswered`, `harness.QuestionsActionPartiallyAnswered`, `harness.QuestionsActionDeclined`, `harness.QuestionsActionCancelled` | The output `action` values. | ## Authentication `harness.Authentication` is an adapter `auth` setting: a mode string or an isolated environment. | Name | Description | | --- | --- | | `harness.AuthMode(mode string) Authentication` | Selects a mode. | | `harness.AuthEnvironment(env map[string]string) Authentication` | Uses an isolated environment. | | `harness.AuthModeAuto`, `harness.AuthModeAIGateway`, `harness.AuthModeDirect` | The shared modes. | | `harness.ErrInvalidAuthentication` | Returned for an authentication record that is not flat. | | `harness.CredentialForwarding` | Callback that customizes a credential value just before an adapter forwards it into the sandbox. | | `harness.CredentialForwardingOptions` | Input of that callback. | | `harness.MintBridgeTokenCallback` | Creates the bridge channel token for one session. | ## Telemetry and diagnostics `AgentSettings.Telemetry` takes a `*telemetry.Options`. A turn span contains step spans, which contain model-call and tool-execution spans. | Name | Description | | --- | --- | | `harness.DebugConfig` | Per-session diagnostics configuration. | | `harness.DebugLevel` | `harness.DebugLevelError`, `harness.DebugLevelWarn`, `harness.DebugLevelInfo`, `harness.DebugLevelDebug` or `harness.DebugLevelTrace`, ordered most to least severe. | | `harness.Diagnostic`, `harness.DiagnosticError` | A diagnostic and its error payload. | | `harness.Observability` | Diagnostics wiring handed to `DoStart`. | | `harness.CallWarning`, `harness.CallWarningType` | A non-fatal adapter warning. The types are `harness.CallWarningUnsupportedSetting`, `harness.CallWarningUnsupportedTool` and `harness.CallWarningOther`. | ## Errors | Type | Test function | Description | | --- | --- | --- | | `*harness.HarnessError` | `harness.IsHarnessError(err)` | Base error. `harness.NewHarnessError` creates one. | | `*harness.CapabilityUnsupportedError` | `harness.IsCapabilityUnsupportedError(err)` | The adapter or sandbox cannot do what you asked. `harness.NewCapabilityUnsupportedError` creates one. | | `*harness.SandboxAuthenticationError` | `harness.IsSandboxAuthenticationError(err)` | A sandbox provider cannot authenticate. `harness.NewSandboxAuthenticationError` creates one. | | `*harness.HistoryUnavailableError` | `harness.IsHistoryUnavailableError(err)` | `ReadHistory` cannot reach the runtime's history. `harness.NewHistoryUnavailableError` creates one. | `harness.GetHarnessErrorMessage(err)` returns a message that is safe to show a client. The error name constants are `harness.HarnessErrorName`, `harness.CapabilityUnsupportedErrorName`, `harness.SandboxAuthenticationErrorName` and `harness.HistoryUnavailableErrorName`. --- # Harness sandboxes > Reference for harness sandbox configuration, the SandboxProvider interface, bootstrap recipes and templates, network policy, request transformations and the local and Vercel providers. Canonical URL: https://goaisdk.com/docs/reference/harness/sandbox Documentation index: https://goaisdk.com/llms.txt A harness runtime runs inside a sandbox. A sandbox provider creates a network sandbox session, which gives the runtime file I/O, process execution, exposed ports and a network policy. This page covers the sandbox types in `pkg/harness`, bootstrap and templates, and the two built-in providers. The agent API is in [Harness agent](https://goaisdk.com/docs/reference/harness/agent.md). ## SandboxConfig `AgentSettings.SandboxConfig` configures the working directory and lifecycle hooks. Its type is `harness.AgentSandboxConfig`, an alias of `harness.SandboxConfig`. {/* gen:fields harness.SandboxConfig */} | Field | Type | Description | | --- | --- | --- | | `WorkDir` | `string` | WorkDir is an optional fixed working directory for all sessions, relative to the sandbox default working directory. | | `BootstrapHash` | `string` | BootstrapHash is the caller-controlled identity for OnBootstrap. | | `OnBootstrap` | `func(ctx context.Context, bc SandboxBootstrapContext) error` | OnBootstrap runs during sandbox template creation after the adapter's own bootstrap, before snapshot-capable providers publish a snapshot (a83a367). Must be provided together with BootstrapHash. | | `OnSession` | `func(ctx context.Context, sc SandboxSessionContext) error` | OnSession runs after each sandbox session is acquired and the session work directory exists, before the adapter starts. | {/* /gen:fields */} `OnBootstrap` runs once per sandbox template, after the adapter's own bootstrap and before a snapshot-capable provider takes a snapshot. `BootstrapHash` identifies that work, so changing it invalidates the snapshot. Set the two together. `OnSession` runs for every session, after the session work directory exists and before the adapter starts. {/* gen:fields harness.SandboxBootstrapContext */} | Field | Type | Description | | --- | --- | --- | | `Session` | `providerutils.SandboxSession` | | | `WorkDir` | `string` | | {/* /gen:fields */} {/* gen:fields harness.SandboxSessionContext */} | Field | Type | Description | | --- | --- | --- | | `Session` | `providerutils.SandboxSession` | | | `SessionWorkDir` | `string` | | {/* /gen:fields */} `harness.ValidateSandboxBootstrapSettings(cfg SandboxConfig) error` checks that `OnBootstrap` and `BootstrapHash` are set together. ## SandboxProvider ```go type SandboxProvider interface { SpecificationVersion() string // "harness-sandbox-v1" ProviderID() string CreateSession(ctx context.Context, opts CreateSandboxSessionOptions) (NetworkSandboxSession, error) } ``` A provider returns `*harness.SandboxAuthenticationError` when credentials are missing or invalid. A provider that can reattach by ID also implements `harness.SandboxSessionResumer` (`ResumeSession(ctx, sessionID) (NetworkSandboxSession, error)`). `harness.SandboxSpecificationVersion` is the specification version. {/* gen:fields harness.CreateSandboxSessionOptions */} | Field | Type | Description | | --- | --- | --- | | `SessionID` | `string` | SessionID names the underlying resource deterministically so a future ResumeSession can find it. Empty for prewarm paths. | | `Identity` | `string` | Identity is the stable identity for snapshot-based reuse. | | `OnFirstCreate` | `OnFirstCreateFunc` | OnFirstCreate is called exactly once per identity, on fresh creation. | {/* /gen:fields */} `harness.OnFirstCreateFunc` runs once per identity on fresh sandbox creation. It lets a snapshot-capable provider run the template's prepare step before it snapshots. ## NetworkSandboxSession `harness.NetworkSandboxSession` extends `providerutils.SandboxSession` (file I/O, exec and spawn) with the infrastructure surface. | Method | Description | | --- | --- | | `ID() string` | Stable ID of the underlying sandbox resource. | | `DefaultWorkingDirectory() string` | Absolute path that relative commands resolve against. | | `Ports() []int` | Ports the sandbox exposes. | | `GetPortEndpoint(ctx, PortEndpointOptions) (PortEndpoint, error)` | Connection details for an exposed port. | | `GetPortURL(ctx, PortEndpointOptions) (string, error)` | Deprecated. Use `GetPortEndpoint`. | | `Stop(ctx) error` | Stops the sandbox. Idempotent. | | `Destroy(ctx) error` | Stops the sandbox, then cleans up. Works on running and stopped sandboxes. | | `Restricted() providerutils.SandboxSession` | A reduced view with file I/O and process APIs only. | `harness.AsNetworkSandboxSession(s)` returns the network view of a session when it has one. `harness.GetRestrictedSandboxSession(s)` returns the restricted view, or the session itself. Optional interfaces extend a session. A provider implements the ones it supports. | Interface | Method | | --- | --- | | `harness.NetworkPolicySetter` | `SetNetworkPolicy(ctx, NetworkPolicy) error` | | `harness.RequestTransformationSetter` | `SetRequestTransformations(ctx, []RequestTransformation) error`. Replaces the full set. | | `harness.RequestTransformationAdder` | Adds rules without replacing existing ones. | | `harness.PortsSetter` | `SetPorts(ctx, []int) error`. Replaces the set of exposed ports. | ### Ports {/* gen:fields harness.PortEndpoint */} | Field | Type | Description | | --- | --- | --- | | `URL` | `string` | | | `Headers` | `map[string]string` | | {/* /gen:fields */} {/* gen:fields harness.PortEndpointOptions */} | Field | Type | Description | | --- | --- | --- | | `Port` | `int` | | | `Protocol` | `PortProtocol` | | {/* /gen:fields */} `harness.PortProtocol` selects the URL scheme: `harness.PortProtocolHTTP`, `harness.PortProtocolHTTPS` or `harness.PortProtocolWS`. ## Network policy {/* gen:fields harness.NetworkPolicy */} | Field | Type | Description | | --- | --- | --- | | `Mode` | `string` | | | `AllowedHosts` | `[]string` | | | `AllowedCIDRs` | `[]string` | | | `DeniedCIDRs` | `[]string` | | {/* /gen:fields */} `Mode` is `harness.NetworkPolicyAllowAll`, `harness.NetworkPolicyDenyAll` or `harness.NetworkPolicyCustom`. A custom policy needs at least one of `AllowedHosts` or `AllowedCIDRs`. `DeniedCIDRs` takes precedence over both. ## Request transformations A request transformation rewrites an outbound HTTPS request outside the sandbox security boundary. Adapters use them for credential brokering: the sandbox holds a placeholder, and the transformation adds the real credential on the way out, so the credential never enters the sandbox. {/* gen:fields harness.RequestTransformation */} | Field | Type | Description | | --- | --- | --- | | `Match` | `RequestTransformationMatch` | | | `Transform` | `RequestTransformationTransform` | | {/* /gen:fields */} {/* gen:fields harness.RequestTransformationMatch */} | Field | Type | Description | | --- | --- | --- | | `Host` | `string` | | | `Path` | `*StringMatcher` | | | `Method` | `[]string` | | | `QueryString` | `[]KeyValueMatcher` | | | `Headers` | `[]KeyValueMatcher` | | {/* /gen:fields */} {/* gen:fields harness.RequestTransformationTransform */} | Field | Type | Description | | --- | --- | --- | | `Headers` | `map[string]string` | | {/* /gen:fields */} `harness.StringMatcher` matches a string exactly, by prefix or by regular expression. `harness.KeyValueMatcher` matches a header or a query parameter. `harness.RequestTransformationSources` holds the inputs adapters use to build credential transformations. ## Bootstrap recipes An adapter can declare a recipe that installs its runtime in the sandbox. The recipe is a set of files and commands. It runs once per sandbox, and a marker file records completion, so repeated calls are cheap. | Name | Description | | --- | --- | | `harness.Bootstrap` | The recipe: `HarnessID`, `BootstrapDir`, `Files` and `Commands`. | | `harness.BootstrapFile`, `harness.BootstrapCommand` | One file written, and one command run from the bootstrap directory after all files are written. | | `harness.BootstrapProvider` | Interface an adapter implements to declare a recipe. | | `harness.GetBootstrap(ctx, h Harness) (*Bootstrap, error)` | Returns the adapter's recipe, or nil. | | `harness.BootstrapSchemaVersion` | The version of the recipe shape. | | `harness.HashHarnessBootstrap(recipe Bootstrap) string` | A deterministic 16-character hex identity of a recipe. Identical to the TypeScript hash for identical input. | | `harness.ApplyBootstrapRecipe(ctx, session, recipe, identity, stateDirectory) error` | Applies a recipe idempotently. | | `harness.BootstrapMarkerPath(recipe, identity, stateDirectory) string` | The path of the marker written after a recipe is applied. | | `harness.StateDirectoryName`, `harness.StateDirectoryPath(sandboxHomeDir) string` | The fixed directory under the sandbox `HOME` that holds all harness state. | | `harness.SessionDataDirectoryPath(stateDirectory, sessionID) string` | The per-session state path under `.agent-runs`. | ### OnBootstrap markers | Function | Description | | --- | --- | | `harness.OnBootstrapMarkerPath(ctx, session, bootstrapHash) (string, error)` | The marker path for a caller's `OnBootstrap` hook, keyed by the SHA-256 of the hash. | | `harness.HasOnBootstrapMarker(ctx, session, bootstrapHash) (bool, error)` | Reports whether `OnBootstrap` already finished for this hash on this sandbox. | | `harness.WriteOnBootstrapMarker(ctx, session, bootstrapHash) error` | Records completion. | ## Templates A template packages the bootstrap work so a snapshot-capable provider can reuse it. | Function | Description | | --- | --- | | `harness.CreateSandboxBootstrapPlan(recipe *Bootstrap, cfg SandboxConfig) (SandboxBootstrapPlan, error)` | Computes the recipe identity and the combined sandbox identity. | | `harness.CreateHarnessSandboxTemplate(ctx, CreateHarnessSandboxTemplateOptions) (*HarnessSandboxTemplate, error)` | Builds a template for one or more harnesses. Returns `nil, nil` when there is nothing to prepare. | | `harness.PrepareHarnessSandboxTemplate(ctx, PrepareHarnessSandboxTemplateOptions) error` | Prepares a template without running an agent. Idempotent. | | `harness.PrepareSandboxForHarness(ctx, PrepareSandboxForHarnessOptions) (*PrepareSandboxForHarnessResult, error)` | Applies recipes to a sandbox you own and returns an identity you can use as a snapshot key. | | `harness.RunSandboxBootstrap(ctx, RunSandboxBootstrapOptions) error` | Applies the adapter recipe, then runs `OnBootstrap`. | | `harness.PrewarmHarness(ctx, opts)` | Deprecated alias of `PrepareHarnessSandboxTemplate`. | {/* gen:fields harness.HarnessSandboxTemplate */} | Field | Type | Description | | --- | --- | --- | | `Identity` | `string` | | | `Prepare` | `func(ctx context.Context, session providerutils.SandboxSession) error` | | {/* /gen:fields */} {/* gen:fields harness.SandboxBootstrapPlan */} | Field | Type | Description | | --- | --- | --- | | `Recipe` | `*Bootstrap` | | | `RecipeIdentity` | `string` | | | `Identity` | `string` | Identity is the sandbox identity passed to SandboxProvider.CreateSession. | | `WorkDir` | `string` | | | `OnFirstCreate` | `OnFirstCreateFunc` | | {/* /gen:fields */} {/* gen:fields harness.CreateHarnessSandboxTemplateOptions */} | Field | Type | Description | | --- | --- | --- | | `Harnesses` | `[]Harness` | | | `SandboxConfig` | `*SandboxConfig` | SandboxConfig is optional; OnSession is ignored. | {/* /gen:fields */} {/* gen:fields harness.PrepareHarnessSandboxTemplateOptions */} | Field | Type | Description | | --- | --- | --- | | `Harness` | `Harness` | | | `SandboxProvider` | `SandboxProvider` | | | `SandboxConfig` | `*SandboxConfig` | SandboxConfig is optional; OnSession is ignored. | {/* /gen:fields */} {/* gen:fields harness.PrepareSandboxForHarnessOptions */} | Field | Type | Description | | --- | --- | --- | | `Session` | `providerutils.SandboxSession` | | | `Harnesses` | `[]Harness` | | | `SandboxConfig` | `*SandboxConfig` | SandboxConfig is optional; OnSession is ignored. | {/* /gen:fields */} {/* gen:fields harness.PrepareSandboxForHarnessResult */} | Field | Type | Description | | --- | --- | --- | | `Identity` | `string` | Identity is empty when no recipe was applied and no bootstrapHash set. | | `RecipeIdentities` | `map[string]string` | | | `SkippedHarnessIDs` | `[]string` | | {/* /gen:fields */} {/* gen:fields harness.RunSandboxBootstrapOptions */} | Field | Type | Description | | --- | --- | --- | | `Session` | `providerutils.SandboxSession` | | | `Recipe` | `*Bootstrap` | | | `RecipeIdentity` | `string` | | | `WorkDir` | `string` | | | `OnBootstrap` | `func(ctx context.Context, bc SandboxBootstrapContext) error` | | | `BootstrapHash` | `string` | BootstrapHash is the caller-controlled identity OnBootstrap is keyed by. Required for SkipOnBootstrapIfMarked and for the completion marker this writes after a successful OnBootstrap run. | | `SkipOnBootstrapIfMarked` | `bool` | SkipOnBootstrapIfMarked skips OnBootstrap entirely when a marker for BootstrapHash already exists on this sandbox (from a prior call, e.g. a snapshot-provider's OnFirstCreate having already run it). Mirrors TS `runSandboxBootstrap`'s `skipOnBootstrapIfMarked` (31742b9a1b). | | `DefaultWorkingDirectory` | `string` | DefaultWorkingDirectory skips the `pwd` lookup when known. | {/* /gen:fields */} ## Working directory helpers | Function | Description | | --- | --- | | `harness.NormalizeSandboxWorkDir(workDir string) (string, error)` | Validates and normalizes a relative work directory. | | `harness.ResolveSessionWorkDir(defaultWorkingDirectory, harnessID, sessionID, workDir string) string` | Composes `/-`, or `/` when `workDir` is set. | | `harness.EncodePathSegment(value string) string` | Encodes a value as one safe path segment. | | `harness.EnsureSandboxDirectory(ctx, session, workDir) error` | Runs `mkdir -p` for the directory. | | `harness.ResolveSandboxHomeDir(ctx, sandbox) (string, error)` | Resolves the sandbox `HOME`. | | `harness.ResolveSandboxDefaultWorkingDirectory(ctx, sandbox) (string, error)` | Returns the default working directory. | | `harness.StripWorkDir(...)`, `harness.NewToolInputWorkDirStripper(...)` | Remove the session work directory prefix from path fields of stream parts before they reach the consumer. | ## Local provider Package `github.com/digitallysavvy/go-ai/pkg/harness/sandbox/local` runs each session in a directory on the host. It has no isolation. Use it for development and tests. `local.ProviderID` is `"local"`. ```go func NewProvider(opts Options) *Provider ``` {/* gen:fields harness/sandbox/local.Options */} | Field | Type | Description | | --- | --- | --- | | `RootDir` | `string` | RootDir holds one directory per session. Defaults to <os.TempDir()>/ai-sdk-harness-local. | | `HomeDir` | `string` | HomeDir, when set, is used as HOME for every session (for example the real user home to reuse caches and native subscriptions). When empty, each session gets an isolated HOME under its session directory. | | `Shell` | `string` | Shell runs commands. Defaults to /bin/sh. | | `Env` | `map[string]string` | Env is added to every command's environment (after the host environment and HOME, before per-command env). | | `Ports` | `[]int` | Ports lists the ports reported by Ports(); all localhost ports are reachable regardless. | | `Host` | `string` | Host used in port endpoints. Defaults to 127.0.0.1. | | `KeepFiles` | `bool` | KeepFiles keeps the session directory on Destroy. | {/* /gen:fields */} ## Vercel provider Package `github.com/digitallysavvy/go-ai/pkg/harness/sandbox/vercel` runs sandboxes on Vercel Sandbox, with network policies and template snapshots. See [Harness sandbox: Vercel](https://goaisdk.com/docs/reference/ai/harness-sandbox-vercel.md) for the full API. --- # Harness stream parts > Reference for the stream part types a harness adapter emits during a prompt turn, how they map to AI SDK stream chunks, and the JSON helpers. Canonical URL: https://goaisdk.com/docs/reference/harness/stream-parts Documentation index: https://goaisdk.com/llms.txt An adapter emits `harness.StreamPart` values while a turn runs. `harness.Agent` translates them into `provider.StreamChunk` values for `ai.NewStreamTextResultFromParts`, so a harness turn streams like a `StreamText` call. Each part marshals to the JSON shape of the TypeScript `HarnessV1StreamPart`, so parts persisted by one SDK decode in the other. The adapter interface is in [Harness adapters](https://goaisdk.com/docs/reference/harness/adapters.md). ## Part types `StreamPart.PartType()` returns the `type` discriminator. The `harness.PartType*` constants hold the values. | Constant | `type` | Go type | Description | | --- | --- | --- | --- | | `harness.PartTypeStreamStart` | `stream-start` | `*harness.StreamStartPart` | Start of a turn. | | `harness.PartTypeTextStart` | `text-start` | `*harness.TextStartPart` | A text block opens. | | `harness.PartTypeTextDelta` | `text-delta` | `*harness.TextDeltaPart` | A piece of text. | | `harness.PartTypeTextEnd` | `text-end` | `*harness.TextEndPart` | A text block closes. | | `harness.PartTypeReasoningStart` | `reasoning-start` | `*harness.ReasoningStartPart` | A reasoning block opens. | | `harness.PartTypeReasoningDelta` | `reasoning-delta` | `*harness.ReasoningDeltaPart` | A piece of reasoning. | | `harness.PartTypeReasoningEnd` | `reasoning-end` | `*harness.ReasoningEndPart` | A reasoning block closes. | | `harness.PartTypeToolInputStart` | `tool-input-start` | `*harness.ToolInputStartPart` | Tool input starts streaming. | | `harness.PartTypeToolInputDelta` | `tool-input-delta` | `*harness.ToolInputDeltaPart` | A piece of tool input. | | `harness.PartTypeToolInputEnd` | `tool-input-end` | `*harness.ToolInputEndPart` | Tool input is complete. | | `harness.PartTypeToolCall` | `tool-call` | `*harness.ToolCallPart` | A tool call. Carries `nativeName` and `stepToolCallCount` in addition to the model tool call fields. | | `harness.PartTypeToolApprovalRequest` | `tool-approval-request` | `*harness.ToolApprovalRequestPart` | The runtime asks for approval. | | `harness.PartTypeToolResult` | `tool-result` | `*harness.ToolResultPart` | A tool result. | | `harness.PartTypeFinishStep` | `finish-step` | `*harness.FinishStepPart` | End of a step inside a turn. | | `harness.PartTypeFinish` | `finish` | `*harness.FinishPart` | End of the turn. | | `harness.PartTypeFileChange` | `file-change` | `*harness.FileChangePart` | A workspace change made through an opaque mechanism. | | `harness.PartTypeCompaction` | `compaction` | `*harness.CompactionPart` | The runtime compacted its context. | | `harness.PartTypeError` | `error` | `*harness.ErrorPart` | An error. | | `harness.PartTypeRaw` | `raw` | `*harness.RawPart` | Adapter-specific passthrough. | ### Fields {/* gen:fields harness.StreamStartPart */} | Field | Type | Description | | --- | --- | --- | | `Warnings` | `[]CallWarning` | | | `ModelID` | `string` | ModelID is the model the runtime resolved to for this turn, when known. | {/* /gen:fields */} {/* gen:fields harness.TextDeltaPart */} | Field | Type | Description | | --- | --- | --- | | `ID` | `string` | | | `Delta` | `string` | | | `HarnessMetadata` | `Metadata` | | {/* /gen:fields */} {/* gen:fields harness.ToolCallPart */} | Field | Type | Description | | --- | --- | --- | | `ToolCallID` | `string` | | | `ToolName` | `string` | | | `Input` | `string` | Input is the stringified JSON tool input. | | `ProviderExecuted` | `bool` | ProviderExecuted is true for runtime-executed builtins. | | `Dynamic` | `bool` | | | `ProviderMetadata` | `ProviderMetadata` | | | `NativeName` | `string` | NativeName is the runtime's native name when it differs from ToolName. | | `StepToolCallCount` | `*int` | StepToolCallCount is the total tool calls in the current model step, when known up front. Must be a positive integer when set. | {/* /gen:fields */} {/* gen:fields harness.ToolApprovalRequestPart */} | Field | Type | Description | | --- | --- | --- | | `ApprovalID` | `string` | | | `ToolCallID` | `string` | | | `ProviderMetadata` | `ProviderMetadata` | | {/* /gen:fields */} {/* gen:fields harness.ToolResultPart */} | Field | Type | Description | | --- | --- | --- | | `ToolCallID` | `string` | | | `ToolName` | `string` | | | `Result` | `any` | | | `IsError` | `bool` | | | `Preliminary` | `bool` | | | `Dynamic` | `bool` | | | `ProviderMetadata` | `ProviderMetadata` | | {/* /gen:fields */} {/* gen:fields harness.FinishStepPart */} | Field | Type | Description | | --- | --- | --- | | `FinishReason` | `FinishReason` | | | `Usage` | `Usage` | | | `HarnessMetadata` | `Metadata` | | {/* /gen:fields */} {/* gen:fields harness.FinishPart */} | Field | Type | Description | | --- | --- | --- | | `FinishReason` | `FinishReason` | | | `TotalUsage` | `Usage` | | | `HarnessMetadata` | `Metadata` | | {/* /gen:fields */} {/* gen:fields harness.FileChangePart */} | Field | Type | Description | | --- | --- | --- | | `Event` | `string` | | | `Path` | `string` | | | `HarnessMetadata` | `Metadata` | | {/* /gen:fields */} {/* gen:fields harness.CompactionPart */} | Field | Type | Description | | --- | --- | --- | | `Trigger` | `string` | | | `Summary` | `string` | | | `TokensBefore` | `*float64` | | | `TokensAfter` | `*float64` | | | `HarnessMetadata` | `Metadata` | | {/* /gen:fields */} {/* gen:fields harness.ErrorPart */} | Field | Type | Description | | --- | --- | --- | | `Error` | `any` | | {/* /gen:fields */} {/* gen:fields harness.RawPart */} | Field | Type | Description | | --- | --- | --- | | `RawValue` | `any` | | {/* /gen:fields */} The block parts all carry an `id`. Text and reasoning parts also carry `harnessMetadata`, and the delta parts add a `delta` string. The tool input parts carry `providerMetadata`. `harness.ToolInputStartPart` adds `toolName`, `providerExecuted`, `dynamic` and `title`. These parts are `harness.TextStartPart`, `harness.TextEndPart`, `harness.ReasoningStartPart`, `harness.ReasoningDeltaPart`, `harness.ReasoningEndPart`, `harness.ToolInputStartPart`, `harness.ToolInputDeltaPart` and `harness.ToolInputEndPart`. ### Constants | Group | Values | | --- | --- | | File change kinds | `harness.FileChangeCreate`, `harness.FileChangeModify`, `harness.FileChangeDelete` | | Compaction triggers | `harness.CompactionTriggerManual`, `harness.CompactionTriggerAuto` | | Finish reasons | `harness.FinishReasonStop`, `harness.FinishReasonLength`, `harness.FinishReasonContentFilter`, `harness.FinishReasonToolCalls`, `harness.FinishReasonError`, `harness.FinishReasonOther` | ### Usage `harness.FinishReason` is the harness wire encoding of the model finish reason. `harness.Usage` is the wire encoding of model usage, with `harness.InputTokenUsage` and `harness.OutputTokenUsage` for the token breakdowns. {/* gen:fields harness.Usage */} | Field | Type | Description | | --- | --- | --- | | `InputTokens` | `InputTokenUsage` | | | `OutputTokens` | `OutputTokenUsage` | | | `Raw` | `map[string]any` | | {/* /gen:fields */} {/* gen:fields harness.InputTokenUsage */} | Field | Type | Description | | --- | --- | --- | | `Total` | `*int` | | | `NoCache` | `*int` | | | `CacheRead` | `*int` | | | `CacheWrite` | `*int` | | {/* /gen:fields */} {/* gen:fields harness.OutputTokenUsage */} | Field | Type | Description | | --- | --- | --- | | `Total` | `*int` | | | `Text` | `*int` | | | `Reasoning` | `*int` | | {/* /gen:fields */} `harness.Metadata` is adapter-namespaced opaque data attached to events, keyed by harness ID. `harness.ProviderMetadata` is the provider metadata shape on tool parts. ## Translation ```go func TranslatePart(part StreamPart, opts TranslateOptions) []provider.StreamChunk ``` Converts one part to zero or more chunks. Most parts map one to one. These are the exceptions: - `tool-call` is not translated here. `harness.Agent` validates it against the merged tool set first. - A failed `tool-result` from a provider-executed tool becomes a tool result with an error. - `file-change` and `compaction` have no chunk type of their own. Each becomes a dynamic, provider-executed tool call and tool result pair named `fileChange` or `compaction`, so the event stays visible. - `stream-start`, `finish-step` and `finish` return nil. `harness.Agent` owns step and turn boundaries. {/* gen:fields harness.TranslateOptions */} | Field | Type | Description | | --- | --- | --- | | `IsProviderExecuted` | `func(toolCallID string) bool` | IsProviderExecuted reports whether the tool call that produced a tool-result event ran inside the harness runtime. Host results are echoed back as tool-result events too, so the event alone cannot say who ran the tool — only the originating tool-call can, and correlating the two is the caller's job (run_prompt.go tracks it). A nil func treats every failure as provider-executed, matching TS's `?? true` default. | {/* /gen:fields */} ## JSON helpers | Function | Description | | --- | --- | | `harness.DecodeStreamPart(data []byte) (StreamPart, error)` | Parses and validates one JSON part. Returns `harness.ErrUnknownPartType` for an unknown `type`. | | `harness.MarshalStreamPart(part StreamPart) ([]byte, error)` | Encodes a part with `type` as the first key. | | `harness.IsStreamPartType(typ string) bool` | Reports whether `typ` is a part discriminator. | | `harness.MarshalTagged(typ string, v any) ([]byte, error)` | Encodes a struct as a JSON object with `"type"` first. | | `harness.ReadTagged(data []byte) (string, map[string]json.RawMessage, error)` | Returns the `type` and the raw top-level fields of a JSON object. | | `harness.RequireKeys(typ string, fields map[string]json.RawMessage, required []string) error` | Checks that the required keys are present. | | `harness.StripWorkDir(part StreamPart, sessionWorkDir string) StreamPart` | Returns a copy of the part with the session work directory removed from path fields. | | `harness.EmitFunc` | `func(part StreamPart)`. Receives every part an adapter produces during a turn. | | `harness.AgentStreamPart` | Consumer-facing alias of `StreamPart`. | --- # API Reference > Complete API reference index for the Go AI SDK, linking core functions, tool and helper APIs, type definitions, provider interfaces, and middleware. Canonical URL: https://goaisdk.com/docs/reference/ Documentation index: https://goaisdk.com/llms.txt Complete reference documentation for all Go AI SDK APIs, types, and interfaces. ## Core AI Functions High-level functions for working with AI models: - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - Generate text completions - [StreamText](https://goaisdk.com/docs/reference/ai/stream-text.md) - Stream text completions in real-time - [GenerateObject](https://goaisdk.com/docs/reference/ai/generate-object.md) - Generate structured data - [StreamObject](https://goaisdk.com/docs/reference/ai/stream-object.md) - Stream structured data - [Embed](https://goaisdk.com/docs/reference/ai/embed.md) - Generate embeddings for text - [EmbedMany](https://goaisdk.com/docs/reference/ai/embed-many.md) - Generate embeddings for multiple texts - [Rerank](https://goaisdk.com/docs/reference/ai/rerank.md) - Rerank documents by relevance - [GenerateImage](https://goaisdk.com/docs/reference/ai/generate-image.md) - Generate images from text - [Transcribe](https://goaisdk.com/docs/reference/ai/transcribe.md) - Transcribe audio to text - [GenerateSpeech](https://goaisdk.com/docs/reference/ai/generate-speech.md) - Generate speech from text - [Agent](https://goaisdk.com/docs/reference/ai/agent.md) - The `agent.Agent` interface (in `pkg/agent`) - [ToolLoopAgent](https://goaisdk.com/docs/reference/ai/tool-loop-agent.md) - `agent.ToolLoopAgent`, the SDK's tool-using agent (in `pkg/agent`) ## Tool & Helper APIs Functions for working with tools and utilities: - [Tool](https://goaisdk.com/docs/reference/ai/tool.md) - Define tools for AI models - [Dynamic Tools](https://goaisdk.com/docs/reference/ai/dynamic-tool.md) - Register tools at runtime - [CosineSimilarity](https://goaisdk.com/docs/reference/ai/cosine-similarity.md) - Calculate similarity between vectors - [IsStepCount](https://goaisdk.com/docs/reference/ai/step-count-is.md) - Stop condition: stop after n steps - [HasToolCall](https://goaisdk.com/docs/reference/ai/has-tool-call.md) - Stop condition: stop when a tool is called - [StopCondition](https://goaisdk.com/docs/reference/ai/stop-condition.md) - The StopCondition type and StopWhen evaluation - [GenerateID](https://goaisdk.com/docs/reference/ai/generate-id.md) - Generate unique IDs - [Callback Events](https://goaisdk.com/docs/reference/ai/callback-events.md) - Structured lifecycle events for observability - [Stream Transport Helpers](https://goaisdk.com/docs/reference/ai/stream-transport-helpers.md) - Convert stream results to HTTP-friendly payloads - [Agent UI Stream Helpers](https://goaisdk.com/docs/reference/ai/agent-ui-stream-helpers.md) - Bridge ToolLoopAgent streams to UI transport - [Middleware Aliases](https://goaisdk.com/docs/reference/ai/middleware-aliases.md) - Middleware APIs re-exported from pkg/ai ## Type Definitions Core types used throughout the SDK: - [Messages](https://goaisdk.com/docs/reference/types/messages.md) - Message types and structures - [Tools](https://goaisdk.com/docs/reference/types/tools.md) - Tool definition types - [Usage](https://goaisdk.com/docs/reference/types/usage.md) - Token usage tracking types - [Errors](https://goaisdk.com/docs/reference/types/errors.md) - Error types and handling ## Provider Interfaces Interfaces for implementing custom providers: - [LanguageModel](https://goaisdk.com/docs/reference/providers/language-model.md) - Language model interface - [EmbeddingModel](https://goaisdk.com/docs/reference/providers/embedding-model.md) - Embedding model interface - [ImageModel](https://goaisdk.com/docs/reference/providers/image-model.md) - Image generation interface - [SpeechModel](https://goaisdk.com/docs/reference/providers/speech-model.md) - Speech synthesis interface - [TranscriptionModel](https://goaisdk.com/docs/reference/providers/transcription-model.md) - Audio transcription interface - [CustomProvider](https://goaisdk.com/docs/reference/providers/custom-provider.md) - Create custom providers ## Middleware Middleware system for modifying model behavior: - [WrapLanguageModel](https://goaisdk.com/docs/reference/middleware/wrap-language-model.md) - Wrap models with middleware - [MiddlewareInterface](https://goaisdk.com/docs/reference/middleware/middleware-interface.md) - Middleware interface specification - [BuiltInMiddleware](https://goaisdk.com/docs/reference/middleware/built-in-middleware.md) - Built-in middleware functions ## Schema & Registry Schema validation and provider registry: - [JSONSchema](https://goaisdk.com/docs/reference/schema/json-schema.md) - JSON schema definitions - [ProviderRegistry](https://goaisdk.com/docs/reference/registry/provider-registry.md) - Provider registration system ## Error Types Comprehensive error reference: - [Error Types Overview](https://goaisdk.com/docs/reference/types/errors.md) - All error types and handling ## Integrations (experimental) Adapters and experimental runtimes added in v0.5.0: - [Code Mode](https://goaisdk.com/docs/reference/ai/code-mode.md) - Run model-written JavaScript that calls host tools, in a QuickJS-on-WebAssembly sandbox - [Harness Sandbox: Vercel](https://goaisdk.com/docs/reference/ai/harness-sandbox-vercel.md) - Run harness sandboxes on Vercel Sandbox - [LlamaIndex Adapter](https://goaisdk.com/docs/reference/ai/llamaindex.md) - Convert a LlamaIndex chat-engine stream to UI message stream chunks ## See Also - [Getting Started Guide](https://goaisdk.com/docs/getting-started.md) - [AI SDK Core](https://goaisdk.com/docs/ai-sdk-core/overview.md) - [Provider Documentation](https://goaisdk.com/docs/providers.md) --- # MCP client > Reference for the Go AI SDK MCP client in pkg/mcp: client creation, transports, tool conversion, resources, prompts, completions, elicitation, MCP Apps and errors. Canonical URL: https://goaisdk.com/docs/reference/mcp/client Documentation index: https://goaisdk.com/llms.txt Package `github.com/digitallysavvy/go-ai/pkg/mcp` is a Model Context Protocol client. It connects to an MCP server, lists its tools, resources and prompts, and converts server tools into `types.Tool` values for `GenerateText`, `StreamText` and agents. OAuth helpers are in [MCP OAuth](https://goaisdk.com/docs/reference/mcp/oauth.md). Message and capability types are in [MCP protocol types](https://goaisdk.com/docs/reference/mcp/protocol-types.md). ## Creating a client | Function | Description | | --- | --- | | `mcp.CreateStdioMCPClient(command string, args []string) (*MCPClient, error)` | Starts a local server as a child process and talks to it over stdin and stdout. | | `mcp.CreateHTTPMCPClient(url string, oauth *OAuthConfig) (*MCPClient, error)` | Connects to a remote server over streamable HTTP. | | `mcp.CreateSSEMCPClient(url string, oauth *OAuthConfig) (*MCPClient, error)` | Connects over the legacy SSE transport. | | `mcp.CreateMCPClient(config MCPClientConfig, transport Transport) (*MCPClient, error)` | Creates a client over any transport with explicit configuration. | | `mcp.NewMCPClient(transport Transport, config MCPClientConfig) *MCPClient` | Creates a client without connecting. Call `Connect` yourself. | ### MCPClientConfig {/* gen:fields mcp.MCPClientConfig */} | Field | Type | Description | | --- | --- | --- | | `ClientName` | `string` | ClientName is the name of the client | | `Name` | `string` | Name is a deprecated alias for ClientName. Deprecated: use ClientName. | | `ClientVersion` | `string` | ClientVersion is the version of the client | | `Capabilities` | `ClientCapabilities` | Capabilities are optional client capabilities advertised during initialize. | | `RequestTimeoutMS` | `int` | RequestTimeoutMS is the timeout for individual requests in milliseconds Default: 30000 (30 seconds) | | `EnableLogging` | `bool` | EnableLogging enables client-level logging | | `MaxRetries` | `int` | MaxRetries is the maximum number of times a failed "tools/call" request is retried with exponential backoff, matching TypeScript's MCPClient maxRetries option (TS prepareMaxRetries). Default: 0 (no retries). A negative value makes Connect return an error (TS throws MCPClientError "maxRetries must be >= 0" from the constructor; Go has no fallible constructor, so the check runs at Connect instead). Only transient failures are retried: HTTP status 408/409/429/>=500 and transport-level connection errors (refused/reset/timeout/broken pipe/ closed). JSON-RPC application errors (a non-zero MCPClientError.Code) and successful results with IsError=true are never retried. | | `ProtocolVersionDiscovery` | `*bool` | ProtocolVersionDiscovery controls whether the client probes the 2026-07-28 `server/discover` method before falling back to the legacy `initialize` handshake (hash e6a9927). Default true. Only takes effect when the transport implements ProtocolVersionDiscoveryTransport and reports support (only HTTPTransport does today). | | `OnError` | `func(error)` | OnError, when set, receives non-fatal diagnostics the client would otherwise drop silently: a tool skipped because its x-mcp-header annotation is invalid, a tool call whose header binding failed (hash 0c60a40), or a registered OnElicitationRequest handler that returned an error or an invalid ElicitResult (matching TS DefaultMCPClient.onRequestMessage's this.onError(error) call). | {/* /gen:fields */} ## MCPClient methods | Method | Description | | --- | --- | | `Connect(ctx) error` | Connects the transport and runs the protocol handshake. | | `Close() error` | Closes the connection. | | `ListTools(ctx) ([]MCPTool, error)` | Lists the first page of tools. | | `ListToolsWithCursor(ctx, cursor string) (*ListToolsResult, error)` | Lists one page of tools from a cursor. | | `ListAllTools(ctx) ([]MCPTool, error)` | Lists every tool across pages. | | `GetSerializableTools(ctx) (*ListToolsResult, error)` | Returns the tool definitions in a form you can store, for example for [fingerprinting](https://goaisdk.com/docs/reference/ai/tool-call-repair.md#tool-definition-fingerprints). | | `CallTool(ctx, name string, arguments map[string]interface{}) (*CallToolResult, error)` | Calls a tool. | | `ListResources(ctx) ([]MCPResource, error)` | Lists resources. | | `ListResourceTemplates(ctx) (*ListResourceTemplatesResult, error)` | Lists resource templates. | | `ReadResource(ctx, uri string) (*ReadResourceResult, error)` | Reads a resource. | | `ListPrompts(ctx) ([]MCPPrompt, error)` | Lists prompts. | | `GetPrompt(ctx, name string, arguments map[string]interface{}) (*GetPromptResult, error)` | Gets a prompt with arguments filled in. | | `Complete(ctx, params CompleteRequestParams) (*CompleteResult, error)` | Requests argument completions for a prompt or resource template. | | `OnElicitationRequest(handler func(context.Context, ElicitationRequest) (ElicitResult, error))` | Registers the handler for server-initiated elicitation requests. | | `ServerCapabilities() ServerCapabilities` | Capabilities the server reported. | | `ServerInfo() ServerInfo` | Server name and version. | | `ServerInstructions() string` | Instructions the server sent. | | `InitializeResult() InitializeResult` | The raw initialize result. | `MaxRetries` retries only `tools/call` requests that fail with a transient error: HTTP 408, 409, 429 or 5xx, or a connection-level failure. JSON-RPC application errors and results with `IsError` set are never retried. ## Transports All transports implement `mcp.Transport`: ```go type Transport interface { Connect(ctx context.Context) error Close() error Send(ctx context.Context, message *MCPMessage) error Receive(ctx context.Context) (*MCPMessage, error) // io.EOF when closed IsConnected() bool } ``` | Constructor | Transport | | --- | --- | | `mcp.NewStdioTransport(StdioTransportConfig) *StdioTransport` | Child process over stdin and stdout. | | `mcp.NewHTTPTransport(HTTPTransportConfig) *HTTPTransport` | Streamable HTTP, with session resumption and an inbound SSE stream. | | `mcp.NewSSETransport(SSETransportConfig) *SSETransport` | Legacy SSE. | ### TransportConfig Common settings that every transport config embeds as `Config`. {/* gen:fields mcp.TransportConfig */} | Field | Type | Description | | --- | --- | --- | | `TimeoutMS` | `int` | Timeout for operations (optional) | | `Headers` | `map[string]string` | Headers for HTTP-based transports | | `EnableLogging` | `bool` | EnableLogging enables transport-level logging | | `Redirect` | `MCPRedirectMode` | Redirect controls how HTTP redirects are handled. Default (zero value) is MCPRedirectError — redirects cause an error. Set to MCPRedirectFollow to allow the transport to follow redirects. | {/* /gen:fields */} `mcp.MCPRedirectMode` controls HTTP redirects. The default, `mcp.MCPRedirectError`, fails on a redirect, so a server cannot silently send the client somewhere else. `mcp.MCPRedirectFollow` follows them. {/* gen:consts mcp.MCPRedirectMode */} | Constant | Value | Description | | --- | --- | --- | | `MCPRedirectError` | `"error"` | MCPRedirectError causes the transport to return an error when the server responds with an HTTP redirect (3xx). This is the default. | | `MCPRedirectFollow` | `"follow"` | MCPRedirectFollow allows the transport to follow HTTP redirects automatically. Use this only when you trust the MCP server and expect redirects (e.g. behind a reverse proxy that normalises URLs). | {/* /gen:consts */} ### StdioTransportConfig The child process does not inherit the whole parent environment. It gets `Env` plus a short allowlist of safe variables such as `PATH` and `HOME`. {/* gen:fields mcp.StdioTransportConfig */} | Field | Type | Description | | --- | --- | --- | | `Command` | `string` | Command is the command to execute | | `Args` | `[]string` | Args are the arguments to pass to the command | | `Env` | `[]string` | Env are additional environment variables to set, in "KEY=VALUE" form (the os.Environ()/exec.Cmd.Env convention). The child process does not inherit the full parent environment: only Env plus a small safe allowlist of inherited vars (PATH, HOME, etc. — see getEnvironment) are passed through, mirroring TS mcp-stdio/get-environment.ts. | | `WorkingDir` | `string` | WorkingDir is the working directory for the command | | `Config` | `TransportConfig` | Config is the base transport configuration | {/* /gen:fields */} ### HTTPTransportConfig {/* gen:fields mcp.HTTPTransportConfig */} | Field | Type | Description | | --- | --- | --- | | `URL` | `string` | URL is the URL of the MCP server | | `TimeoutMS` | `int` | Timeout is the HTTP request timeout | | `OAuth` | `*OAuthConfig` | OAuth configuration (optional) | | `Config` | `TransportConfig` | Config is the base transport configuration | | `HTTPClient` | `*http.Client` | HTTPClient is an optional custom HTTP client used for all MCP HTTP requests. Useful for TLS customization, proxy settings, and custom dialers. | | `SSEClient` | `SSEClient` | SSEClient is an optional custom SSE-capable client used instead of HTTPClient when supplied. It lets callers provide custom event-stream transports, TLS settings, proxy behavior, or dialers. | | `InitialSessionID` | `string` | InitialSessionID resumes a previously established streamable HTTP session (hash 241a8c5), sent as `mcp-session-id` on every request until the server issues a new one. | | `OnSessionIDChange` | `func(sessionID string)` | OnSessionIDChange is called whenever the server assigns or changes the session id (from an `mcp-session-id` response header). | | `OnSessionExpired` | `func(sessionID string)` | OnSessionExpired is called when the server responds 404 to a request carrying a session id, indicating that session is no longer valid. | | `TerminateSessionOnClose` | `*bool` | TerminateSessionOnClose sends a best-effort DELETE with the session id when Close is called. Default true. | | `OnError` | `func(error)` | OnError, when set, receives non-fatal diagnostics from the background inbound SSE listener (GET reconnect failures, malformed messages, etc.), the send()/POST path (non-2xx responses, fetch failures, OAuth refresh failures, session expiry), and the POST-response SSE reader (stream read failures, malformed messages), matching TS HttpMCPTransport's `onerror` callback. Optional. | {/* /gen:fields */} ### SSETransportConfig {/* gen:fields mcp.SSETransportConfig */} | Field | Type | Description | | --- | --- | --- | | `URL` | `string` | URL is the SSE endpoint URL of the MCP server. | | `TimeoutMS` | `int` | TimeoutMS is the HTTP request timeout. | | `OAuth` | `*OAuthConfig` | OAuth configuration (optional). | | `Config` | `TransportConfig` | Config is the base transport configuration. | | `HTTPClient` | `*http.Client` | HTTPClient is an optional custom HTTP client used for GET and POST requests. | | `SSEClient` | `SSEClient` | SSEClient is an optional custom SSE-capable client used instead of HTTPClient when supplied. | {/* /gen:fields */} ### OAuthConfig A static OAuth configuration for `HTTPTransportConfig.OAuth` and `SSETransportConfig.OAuth`. For the full authorization-code flow, use [`mcp.Auth`](https://goaisdk.com/docs/reference/mcp/oauth.md). {/* gen:fields mcp.OAuthConfig */} | Field | Type | Description | | --- | --- | --- | | `TokenURL` | `string` | TokenURL is the OAuth token endpoint | | `ClientID` | `string` | ClientID is the OAuth client ID | | `ClientSecret` | `string` | ClientSecret is the OAuth client secret | | `Scopes` | `[]string` | Scopes are the OAuth scopes to request | | `AccessToken` | `string` | AccessToken is the current access token | | `RefreshToken` | `string` | RefreshToken is the refresh token | | `ExpiresAt` | `time.Time` | ExpiresAt is when the access token expires | | `RefreshTokenFunc` | `func(ctx context.Context, cfg *OAuthConfig) (accessToken string, expiresIn time.Duration, err error)` | RefreshTokenFunc refreshes the OAuth token. When nil, refresh attempts return an error and callers should provide a fresh access token manually. | {/* /gen:fields */} ### Optional transport capabilities A transport can implement these interfaces. The client checks for them with a type assertion. | Interface | Purpose | | --- | --- | | `mcp.HeaderedSendTransport` | Attach extra HTTP headers to one request. | | `mcp.MCPToolParameterHeadersTransport` | Support binding tool arguments to HTTP headers (see below). | | `mcp.ProtocolVersionTransport` | Receive the protocol version negotiated during initialize. | | `mcp.ProtocolVersionDiscoveryTransport` | Support probing `server/discover` before falling back to `initialize`. Only `HTTPTransport` does. | | `mcp.SendOptionsTransport` | Accept `mcp.MCPTransportSendOptions` on each send. | | `mcp.SSEClient` | A custom SSE-capable client for the HTTP and SSE transports. | {/* gen:fields mcp.MCPTransportSendOptions */} | Field | Type | Description | | --- | --- | --- | | `RelatedRequestID` | `interface{}` | RelatedRequestID associates this outgoing message with an incoming request (TS: `relatedRequestId?: string \| number`). | | `ResumptionToken` | `string` | ResumptionToken resumes a previously interrupted request. | | `OnResumptionToken` | `func(token string)` | OnResumptionToken receives updated resumption tokens from transports that support them. | {/* /gen:fields */} ### Tool parameter headers A tool input schema can mark a property with `x-mcp-header`. The client then sends that argument as an `Mcp-Param-` header on the `tools/call` request. | Name | Description | | --- | --- | | `mcp.GetMCPToolHeaderBindings(inputSchema) ([]MCPToolHeaderBinding, error)` | Finds annotated properties reachable through `properties` only. | | `mcp.CreateMCPToolHeaders(bindings, args) (map[string]string, error)` | Builds the headers for one call. | | `mcp.EncodeMCPHeaderValue(value string) string` | Encodes a value for an HTTP header. Plain printable ASCII passes through. Other values become `=?base64??=`. | | `mcp.MCPHeaderValueType` | The JSON Schema `type` an annotated property must declare: `mcp.MCPHeaderValueString`, `mcp.MCPHeaderValueInteger` or `mcp.MCPHeaderValueBoolean`. | | `mcp.MCPToolHeaderBinding` | One binding of a schema property to a header. | ## Protocol versions | Name | Description | | --- | --- | | `mcp.LatestProtocolVersion` | `"2026-07-28"`. Negotiated with `server/discover`. | | `mcp.LatestLegacyProtocolVersion` | `"2025-11-25"`. Negotiated with `initialize`. | | `mcp.ProtocolVersion` | The legacy version the client sends in `initialize`. | | `mcp.SupportedProtocolVersions` | Every version the client understands, in order of preference. | | `mcp.DefaultProtocolDiscoveryTimeoutMS` | How long the client waits for a `server/discover` response before it falls back to `initialize`. | | `mcp.ModernProtocolErrorCodes` | JSON-RPC error codes a server uses for a hard failure during modern-era discovery. | `mcp.DiscoverResult` is the result of `server/discover`. `mcp.DiscoverResultServerInfo(result)` extracts the server info from its metadata. ## Converting tools ```go tools, err := mcp.GetMCPToolsForAgent(ctx, client) ``` `mcp.GetMCPToolsForAgent(ctx, client) ([]types.Tool, error)` fetches and converts every tool. For finer control, create an `mcp.MCPToolConverter` with `mcp.NewMCPToolConverter(client)`: | Method | Description | | --- | --- | | `ConvertToGoAITools(ctx) ([]types.Tool, error)` | Fetches and converts every tool. | | `ConvertToGoAIToolsWithSchemas(ctx, schemas map[string]MCPToolSchema) ([]types.Tool, error)` | Converts only the tools named in the map. A schema replaces the discovered input schema, and `OutputSchema` validates structured output. | | `ToolsFromDefinitions(mcpTools []MCPTool, schemas map[string]MCPToolSchema) ([]types.Tool, error)` | Converts tool definitions you already hold, for example from `GetSerializableTools`. | `mcp.MCPToolSchema` supplies the Go equivalent of the TypeScript `schemas` map. {/* gen:fields mcp.MCPToolSchema */} | Field | Type | Description | | --- | --- | --- | | `InputSchema` | `interface{}` | | | `OutputSchema` | `interface{}` | | {/* /gen:fields */} Converted tools carry `mcp.McpProviderMetadata` under the `mcp` provider metadata key. `mcp.McpToolAnnotations` carries the behavior hints a server reports (`readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint`). Treat these hints as untrusted unless you trust the server. {/* gen:fields mcp.McpProviderMetadata */} | Field | Type | Description | | --- | --- | --- | | `ClientName` | `string` | | | `Title` | `string` | | | `ToolName` | `string` | | | `Annotations` | `*McpToolAnnotations` | | | `App` | `map[string]interface{}` | | {/* /gen:fields */} {/* gen:fields mcp.McpToolAnnotations */} | Field | Type | Description | | --- | --- | --- | | `Title` | `*string` | | | `ReadOnlyHint` | `*bool` | | | `DestructiveHint` | `*bool` | | | `IdempotentHint` | `*bool` | | | `OpenWorldHint` | `*bool` | | {/* /gen:fields */} ### Converting results `mcp.ConvertMCPContentToAISDK(content []ToolResultContent) ([]types.ContentPart, error)` converts MCP tool result content to SDK content parts. Image content becomes an image part with decoded bytes, not text, so a large base64 payload does not count as prompt tokens. It accepts base64 data, `data:` URLs and `http` or `https` URLs. `mcp.ToolResultContent` is one content item. Its `type` is `text`, `image`, `resource` or `resource_link`. {/* gen:fields mcp.ToolResultContent */} | Field | Type | Description | | --- | --- | --- | | `Type` | `string` | "text", "image", "resource", "resource_link" | | `Text` | `string` | | | `Data` | `string` | base64 for image | | `MimeType` | `string` | | | `URI` | `string` | for resource/resource_link | | `Name` | `string` | | | `Title` | `string` | | | `Description` | `string` | | | `Resource` | `*ResourceContent` | | | `Metadata` | `interface{}` | | | `Meta` | `map[string]interface{}` | | {/* /gen:fields */} ### Detecting definition drift A server can change a tool definition after you reviewed it. Use `ai.FingerprintTools` and `ai.DetectToolDrift` on the converted tools. See [Tool call parsing and repair](https://goaisdk.com/docs/reference/ai/tool-call-repair.md#tool-definition-fingerprints). ## Resources `ListResources`, `ListResourceTemplates` and `ReadResource` return these types. | Type | Description | | --- | --- | | `mcp.MCPResource` | A resource the server exposes. | | `mcp.MCPResourceTemplate` | A parameterized resource URI. | | `mcp.ListResourcesParams`, `mcp.ListResourcesResult` | Parameters and result of `resources/list`. | | `mcp.ListResourceTemplatesResult` | Result of `resources/templates/list`. | | `mcp.ReadResourceParams`, `mcp.ReadResourceResult` | Parameters and result of `resources/read`. | | `mcp.ResourceContent` | One content item of a resource, with `text` or a base64 `blob`. | | `mcp.ResourceReference` | Identifies a resource template argument in a completion request. | {/* gen:fields mcp.MCPResource */} | Field | Type | Description | | --- | --- | --- | | `URI` | `string` | | | `Name` | `string` | | | `Title` | `string` | | | `Description` | `string` | | | `MimeType` | `string` | | | `Size` | `*int64` | | | `Metadata` | `interface{}` | | | `Meta` | `map[string]interface{}` | | {/* /gen:fields */} {/* gen:fields mcp.ResourceContent */} | Field | Type | Description | | --- | --- | --- | | `URI` | `string` | | | `Name` | `string` | | | `Title` | `string` | | | `MimeType` | `string` | | | `Text` | `string` | | | `Blob` | `string` | base64 encoded | | `Meta` | `map[string]interface{}` | | {/* /gen:fields */} ## Prompts | Type | Description | | --- | --- | | `mcp.MCPPrompt` | A prompt template the server exposes. | | `mcp.MCPPromptArgument` | An argument of a prompt template. | | `mcp.ListPromptsParams`, `mcp.ListPromptsResult` | Parameters and result of `prompts/list`. | | `mcp.GetPromptParams`, `mcp.GetPromptResult` | Parameters and result of `prompts/get`. | | `mcp.PromptMessage` | One message in a prompt result. | | `mcp.PromptContent` | The content of a prompt message. | | `mcp.PromptReference` | Identifies a prompt argument in a completion request. | {/* gen:fields mcp.MCPPrompt */} | Field | Type | Description | | --- | --- | --- | | `Name` | `string` | | | `Title` | `string` | | | `Description` | `string` | | | `Arguments` | `[]MCPPromptArgument` | | | `Template` | `string` | | | `Metadata` | `map[string]interface{}` | | {/* /gen:fields */} {/* gen:fields mcp.GetPromptResult */} | Field | Type | Description | | --- | --- | --- | | `Description` | `string` | | | `Messages` | `[]PromptMessage` | | | `Metadata` | `map[string]interface{}` | | {/* /gen:fields */} ## Completions `Complete` asks the server for suggestions for a prompt argument or a resource template argument. {/* gen:fields mcp.CompleteRequestParams */} | Field | Type | Description | | --- | --- | --- | | `PromptRef` | `*PromptReference` | | | `ResourceRef` | `*ResourceReference` | | | `Argument` | `CompletionArgument` | | | `Context` | `*CompletionContext` | | {/* /gen:fields */} {/* gen:fields mcp.CompleteResult */} | Field | Type | Description | | --- | --- | --- | | `Completion` | `CompleteResultCompletion` | | | `ResultType` | `string` | | {/* /gen:fields */} `mcp.CompleteResultCompletion` is the payload of `CompleteResult`. `mcp.CompletionArgument` names the argument being completed. `mcp.CompletionContext` carries argument values already resolved. ## Elicitation A server can ask the user for more input during a tool call. Register a handler with `OnElicitationRequest`. The handler receives an `mcp.ElicitationRequest` (a message and a JSON Schema for the answer) and returns an `mcp.ElicitResult`. {/* gen:fields mcp.ElicitationRequest */} | Field | Type | Description | | --- | --- | --- | | `Message` | `string` | | | `RequestedSchema` | `interface{}` | | | `Meta` | `map[string]interface{}` | | {/* /gen:fields */} {/* gen:fields mcp.ElicitResult */} | Field | Type | Description | | --- | --- | --- | | `Action` | `string` | | | `Content` | `map[string]interface{}` | | | `Meta` | `map[string]interface{}` | | {/* /gen:fields */} ## MCP Apps MCP Apps let a tool attach an interactive HTML resource, served from a `ui://` URI, for the host to render. | Name | Description | | --- | --- | | `mcp.MCPAppExtensionName` | The extension name, `io.modelcontextprotocol/ui`. | | `mcp.MCPAppMimeType` | The MIME type of an app resource. | | `mcp.MCPAppLegacyResourceURIMetaKey` | The legacy metadata key that holds the resource URI. | | `mcp.MCPAppClientCapabilities() ClientCapabilities` | The capabilities a host that supports MCP Apps advertises. | | `mcp.IsMCPAppTool(tool MCPTool) (bool, error)` | Reports whether a tool has an app resource. | | `mcp.GetMCPAppToolMeta(tool MCPTool) (*MCPAppToolMeta, error)` | Reads and validates the tool's app metadata. | | `mcp.GetMCPAppResourceURI(tool MCPTool) (string, error)` | The tool's `ui://` resource URI. | | `mcp.GetMCPAppResourceURIs(definitions ListToolsResult) ([]string, error)` | The unique resource URIs across tools. | | `mcp.SplitMCPAppTools(definitions ListToolsResult) (modelVisible, appVisible ListToolsResult, err error)` | Splits tools into the set the model sees and the set only the app can call. | | `mcp.ReadMCPAppResource(ctx, client, uri) (MCPAppResource, error)` | Reads and normalizes an app resource. | | `mcp.GetMCPAppResourceFromReadResult(uri, ReadResourceResult) (MCPAppResource, error)` | Extracts the app HTML from a read result you already have. | | `mcp.FingerprintMCPAppResource(resource MCPAppResource) string` | A stable SHA-256 digest of a resource. | | `mcp.DetectMCPAppResourceDrift(current, baseline string) bool` | True when two fingerprints differ. | {/* gen:fields mcp.MCPAppResource */} | Field | Type | Description | | --- | --- | --- | | `URI` | `string` | | | `MimeType` | `string` | | | `HTML` | `string` | | | `Meta` | `map[string]interface{}` | | {/* /gen:fields */} {/* gen:fields mcp.MCPAppToolMeta */} | Field | Type | Description | | --- | --- | --- | | `ResourceURI` | `string` | | | `Visibility` | `[]string` | | | `Extra` | `map[string]interface{}` | | {/* /gen:fields */} ## Errors | Type or value | Description | | --- | --- | | `*mcp.MCPClientError` | An error from the client. `Code` is the JSON-RPC error code. `StatusCode` is the HTTP status for transport failures. | | `mcp.NewMCPClientError(code, message, data, opts...)` | Creates one. | | `mcp.MCPClientErrorOption` | Option for `NewMCPClientError`: `mcp.WithMCPHTTPResponse`, `mcp.WithMCPHTTPResponseBody`, `mcp.WithMCPHTTPStatus`, `mcp.WithMCPHTTPURL`. | | `*mcp.MCPError` | A protocol error payload. | | `*mcp.TimeoutError`, `mcp.NewTimeoutError(operation)` | A request timed out. | | `*mcp.TransportError`, `mcp.NewTransportError(message, cause)` | A transport-level failure. | | `mcp.ErrParseError`, `mcp.ErrInvalidRequest`, `mcp.ErrMethodNotFound`, `mcp.ErrInvalidParams`, `mcp.ErrInternalError`, `mcp.ErrToolNotFound`, `mcp.ErrToolExecutionFail`, `mcp.ErrResourceNotFound`, `mcp.ErrUnauthorized` | Sentinel `*MCPClientError` values. | The matching codes are `mcp.ErrorCodeParseError` (-32700), `mcp.ErrorCodeInvalidRequest`, `mcp.ErrorCodeMethodNotFound`, `mcp.ErrorCodeInvalidParams`, `mcp.ErrorCodeInternalError`, and the MCP-specific `mcp.ErrorCodeToolNotFound` (-32000), `mcp.ErrorCodeToolExecutionFail`, `mcp.ErrorCodeResourceNotFound` and `mcp.ErrorCodeUnauthorized`. OAuth errors are in [MCP OAuth](https://goaisdk.com/docs/reference/mcp/oauth.md#errors). --- # MCP OAuth > Reference for the Go AI SDK MCP OAuth helpers: the Auth flow, client provider interfaces, discovery, registration, token exchange, PKCE and OAuth errors. Canonical URL: https://goaisdk.com/docs/reference/mcp/oauth Documentation index: https://goaisdk.com/llms.txt `mcp.Auth` runs the OAuth 2.1 authorization-code flow with PKCE for an MCP server. It matches the TypeScript SDK's `auth()`. The lower-level functions on this page are the steps it composes. For a static token, use `mcp.OAuthConfig` instead (see [MCP client](https://goaisdk.com/docs/reference/mcp/client.md#oauthconfig)). ## Auth ```go func Auth(ctx context.Context, provider OAuthClientProvider, options AuthOptions) (AuthResult, error) ``` One call advances the flow one step. It refreshes tokens when it can, registers a client when it has none, and starts an authorization when it needs the user. It returns: {/* gen:consts mcp.AuthResult */} | Constant | Value | Description | | --- | --- | --- | | `AuthResultAuthorized` | `"AUTHORIZED"` | | | `AuthResultRedirect` | `"REDIRECT"` | | {/* /gen:consts */} On `REDIRECT`, send the user to the authorization URL (the provider's `RedirectToAuthorization` runs), then call `Auth` again with the authorization code from the callback. When the authorization server answers `invalid_client` or `unauthorized_client`, `Auth` invalidates credentials and retries once, but only when the client information came from Dynamic Client Registration. A client you registered yourself is never replaced silently. On `invalid_grant`, it invalidates the stored tokens and retries once. ### AuthOptions Set `HasAuthorizationCode` (with `AuthorizationCode`, `CallbackState` and `CallbackIssuer`) when you complete an authorization-code callback. Leave it false to check for or refresh tokens, or to start a new authorization. {/* gen:fields mcp.AuthOptions */} | Field | Type | Description | | --- | --- | --- | | `ServerURL` | `string` | | | `HasAuthorizationCode` | `bool` | | | `AuthorizationCode` | `string` | | | `CallbackState` | `string` | | | `CallbackIssuer` | `string` | CallbackIssuer is the `iss` parameter from the authorization response, validated against the pinned issuer when present (hash 1f29230). | | `Scope` | `string` | Scope is an explicit scope override, typically extracted from a 401 challenge via ExtractWWWAuthenticateParams (hash 1011e33). | | `ResourceMetadataURL` | `*url.URL` | | | `HTTPClient` | `*http.Client` | | {/* /gen:fields */} ## OAuthClientProvider You implement this interface. It holds the storage and policy hooks `Auth` needs. Back it with a keychain, a database row or a struct in memory. | Method | Description | | --- | --- | | `Tokens(ctx) (*OAuthTokens, error)` | Returns stored tokens, or nil. | | `SaveTokens(ctx, tokens OAuthTokens) error` | Persists tokens from an exchange or a refresh. | | `RedirectToAuthorization(ctx, authorizationURL string) error` | Sends the user to the authorization URL. | | `SaveCodeVerifier(ctx, codeVerifier string) error` | Persists the PKCE verifier. | | `CodeVerifier(ctx) (string, error)` | Returns the saved verifier. | | `RedirectURL() string` | The registered redirect URI. | | `ClientMetadata() OAuthClientMetadata` | Metadata for Dynamic Client Registration and scope selection. | | `ClientInformation(ctx) (*OAuthClientInformation, error)` | Returns stored client credentials, or nil. | ### Optional provider interfaces A provider can also implement these. `Auth` checks for them with a type assertion. | Interface | Method | Purpose | | --- | --- | --- | | `mcp.OAuthClientInformationSaver` | `SaveClientInformation` | Persist newly registered client information. | | `mcp.OAuthDynamicRegistrationReporter` | `IsClientInformationDynamicallyRegistered` | Report whether the client came from Dynamic Client Registration. | | `mcp.OAuthCredentialInvalidator` | `InvalidateCredentials` | Clear stored credentials after the server rejects them. | | `mcp.OAuthTokenInvalidator` | `InvalidateCredentialsForTokens` | Receive the specific tokens being invalidated. | | `mcp.OAuthClientAuthenticator` | `AddClientAuthentication` | Customize how client credentials are added to token requests. | | `mcp.OAuthStateProvider` | `State`, `SaveState`, `StoredState` | Generate and verify the CSRF `state` parameter. | | `mcp.OAuthResourceURLValidator` | `ValidateResourceURL` | Override the RFC 8707 `resource` parameter. | | `mcp.OAuthAuthorizationServerURLValidator` | `ValidateAuthorizationServerURL` | Limit which authorization servers `Auth` fetches metadata from. | | `mcp.OAuthAuthorizationServerInformationStore` | `AuthorizationServerInformation`, `SaveAuthorizationServerInformation` | Persist the authorization server pin. | `mcp.OAuthCredentialInvalidationScope` selects what `InvalidateCredentials` discards. {/* gen:consts mcp.OAuthCredentialInvalidationScope */} | Constant | Value | Description | | --- | --- | --- | | `OAuthInvalidateAll` | `"all"` | | | `OAuthInvalidateClient` | `"client"` | | | `OAuthInvalidateTokens` | `"tokens"` | | | `OAuthInvalidateVerifier` | `"verifier"` | | {/* /gen:consts */} ### Authorization server pinning `mcp.OAuthAuthorizationServerInformation` pins the authorization server, token endpoint and issuer that issued stored credentials, so credentials are not reused with a different server. `mcp.CreateOAuthAuthorizationServerInformation(authorizationServerURL, metadata)` creates a pin. `mcp.AssertOAuthAuthorizationServerInformationMatches(stored, current)` returns an error when rediscovered metadata does not match. The error satisfies `mcp.IsAuthorizationServerMismatchError`. ## Data types {/* gen:fields mcp.OAuthTokens */} | Field | Type | Description | | --- | --- | --- | | `AccessToken` | `string` | | | `IDToken` | `string` | | | `TokenType` | `string` | | | `ExpiresIn` | `*int` | | | `Scope` | `string` | | | `RefreshToken` | `string` | | | `Issuer` | `string` | | | `AuthorizationServer` | `string` | | | `TokenEndpoint` | `string` | | {/* /gen:fields */} {/* gen:fields mcp.OAuthClientMetadata */} | Field | Type | Description | | --- | --- | --- | | `RedirectURIs` | `[]string` | | | `ApplicationType` | `string` | "native" \| "web" | | `TokenEndpointAuthMethod` | `string` | | | `GrantTypes` | `[]string` | | | `ResponseTypes` | `[]string` | | | `ClientName` | `string` | | | `ClientURI` | `string` | | | `LogoURI` | `string` | | | `Scope` | `string` | | | `Contacts` | `[]string` | | | `TosURI` | `string` | | | `PolicyURI` | `string` | | | `JWKSURI` | `string` | | | `SoftwareID` | `string` | | | `SoftwareVersion` | `string` | | | `SoftwareStatement` | `string` | | {/* /gen:fields */} {/* gen:fields mcp.OAuthClientInformation */} | Field | Type | Description | | --- | --- | --- | | `ClientID` | `string` | | | `ClientSecret` | `string` | | | `ClientIDIssuedAt` | `*int64` | | | `ClientSecretExpiresAt` | `*int64` | | | `Issuer` | `string` | | | `AuthorizationServer` | `string` | | | `TokenEndpoint` | `string` | | {/* /gen:fields */} `mcp.OAuthClientInformationFull` is the full result of Dynamic Client Registration: the registered metadata plus the client credentials. ## Discovery | Function | Description | | --- | --- | | `mcp.DiscoverOAuthProtectedResourceMetadata(ctx, serverURL, opts) (OAuthProtectedResourceMetadata, error)` | Fetches the server's RFC 9728 protected resource metadata. | | `mcp.SelectOAuthAuthorizationServerURL(ctx, serverURL, opts) (string, *OAuthProtectedResourceMetadata, error)` | Returns the first authorization server in the resource metadata, or the MCP server's own origin. | | `mcp.BuildAuthorizationServerDiscoveryURLs(authorizationServerURL) ([]OAuthDiscoveryURL, error)` | The OAuth and OIDC metadata URLs to try, in priority order. | | `mcp.DiscoverAuthorizationServerMetadata(ctx, authorizationServerURL, opts) (*OAuthAuthorizationServerMetadata, error)` | Fetches authorization server metadata and checks the issuer. | | `mcp.ExtractWWWAuthenticateParams(resp) WWWAuthenticateParams` | Parses a `Bearer` challenge for `resource_metadata` and `scope`. | | `mcp.ExtractResourceMetadataURL(resp) (*url.URL, bool)` | Extracts the RFC 9728 `resource_metadata` URL from a challenge. | | `mcp.AssertOAuthResourceMetadataURLSameOrigin(serverURL, resourceMetadataURL) error` | Rejects a `resource_metadata` URL on a different origin from the server. | | `mcp.SelectOAuthResourceURL(ctx, serverURL, provider, ...) (*url.URL, error)` | Selects the RFC 8707 `resource` value for the authorization server. | | `mcp.SelectOAuthScope(challengeScope, resourceMetadata, ...) string` | Picks the scope: the challenge scope first, then metadata, then client metadata. | | `mcp.TrustedOAuthAuthorizationServerOrigin(serverURL, authorizationServerURL) string` | Returns the origin that may skip the SSRF guard, only when it is a loopback address of the same server. | Discovery refuses unsafe endpoints. It applies an SSRF guard to every hop of metadata discovery and rejects redirects on token requests. {/* gen:fields mcp.OAuthDiscoveryOptions */} | Field | Type | Description | | --- | --- | --- | | `ProtocolVersion` | `string` | | | `ResourceMetadataURL` | `string` | | | `HTTPClient` | `*http.Client` | | | `ValidateAuthorizationServerURL` | `func(serverURL string, authorizationServerURL string) error` | | | `TrustedOrigin` | `string` | TrustedOrigin is a developer-configured origin whose discovery hops skip the SSRF guard in DiscoverAuthorizationServerMetadata (TS trustedOrigin). It must never be derived from response data. See TrustedOAuthAuthorizationServerOrigin. | {/* /gen:fields */} `mcp.OAuthProtectedResourceMetadata`, `mcp.OAuthAuthorizationServerMetadata`, `mcp.OAuthDiscoveryURL` and `mcp.WWWAuthenticateParams` are the result types. ## Flow steps | Function | Description | | --- | --- | | `mcp.RegisterOAuthClient(ctx, authorizationServerURL, ...) (*OAuthClientInformationFull, error)` | Dynamic Client Registration (RFC 7591). | | `mcp.StartOAuthAuthorization(authorizationServerURL, params) (authorizationURL, codeVerifier string, err error)` | Builds the authorization URL and a fresh PKCE verifier. | | `mcp.ExchangeOAuthAuthorization(ctx, authorizationServerURL, ...) (*OAuthTokens, error)` | Exchanges an authorization code for tokens. | | `mcp.RefreshOAuthAuthorization(ctx, authorizationServerURL, ...) (*OAuthTokens, error)` | Exchanges a refresh token for new tokens. | | `mcp.GenerateOAuthPKCE() (codeVerifier, codeChallenge string, err error)` | Generates a PKCE (RFC 7636) verifier and its S256 challenge. | | `mcp.ValidateOAuthState(returnedState, sentState string) error` | Compares the callback `state` with the one you sent. | {/* gen:fields mcp.StartOAuthAuthorizationParams */} | Field | Type | Description | | --- | --- | --- | | `Metadata` | `*OAuthAuthorizationServerMetadata` | | | `ClientInformation` | `OAuthClientInformation` | | | `RedirectURL` | `string` | | | `Scope` | `string` | | | `State` | `string` | | | `Resource` | `*url.URL` | | {/* /gen:fields */} {/* gen:fields mcp.ExchangeOAuthAuthorizationParams */} | Field | Type | Description | | --- | --- | --- | | `Metadata` | `*OAuthAuthorizationServerMetadata` | | | `ClientInformation` | `OAuthClientInformation` | | | `AuthorizationCode` | `string` | | | `CodeVerifier` | `string` | | | `RedirectURI` | `string` | | | `Resource` | `*url.URL` | | | `AddClientAuthentication` | `func(ctx context.Context, headers http.Header, params url.Values, tokenURL string, metadata *OAuthAuthorizationServerMetadata) error` | | | `HTTPClient` | `*http.Client` | | {/* /gen:fields */} {/* gen:fields mcp.RefreshOAuthAuthorizationParams */} | Field | Type | Description | | --- | --- | --- | | `Metadata` | `*OAuthAuthorizationServerMetadata` | | | `ClientInformation` | `OAuthClientInformation` | | | `RefreshToken` | `string` | | | `Resource` | `*url.URL` | | | `AddClientAuthentication` | `func(ctx context.Context, headers http.Header, params url.Values, tokenURL string, metadata *OAuthAuthorizationServerMetadata) error` | | | `HTTPClient` | `*http.Client` | | {/* /gen:fields */} ## Errors | Name | Description | | --- | --- | | `*mcp.MCPClientOAuthError` | An error from the OAuth flow. `Code` is the OAuth error code when known. | | `mcp.ParseOAuthErrorResponse(resp) *MCPClientOAuthError` | Parses an OAuth error response. | | `mcp.IsInvalidClientError(err)` | True for `invalid_client`. | | `mcp.IsInvalidGrantError(err)` | True for `invalid_grant`. | | `mcp.IsUnauthorizedClientError(err)` | True for `unauthorized_client`. | | `mcp.IsAuthorizationServerMismatchError(err)` | True when rediscovered authorization server metadata does not match the pin. | The codes are `mcp.OAuthErrorCodeServerError`, `mcp.OAuthErrorCodeInvalidClient`, `mcp.OAuthErrorCodeInvalidGrant`, `mcp.OAuthErrorCodeUnauthorizedClient` and `mcp.OAuthErrorCodeAuthorizationServerMismatch`. --- # MCP protocol types > Reference for the Go AI SDK MCP protocol types: JSON-RPC messages, capabilities, initialize and discover, tool, resource and prompt payloads, and logging. Canonical URL: https://goaisdk.com/docs/reference/mcp/protocol-types Documentation index: https://goaisdk.com/llms.txt Types in `pkg/mcp` that model the protocol. Most code uses the [client](https://goaisdk.com/docs/reference/mcp/client.md) and never builds these directly. They matter when you write a transport or a server-side bridge. ## JSON-RPC messages `mcp.MCPMessage` is a generic JSON-RPC 2.0 message. {/* gen:fields mcp.MCPMessage */} | Field | Type | Description | | --- | --- | --- | | `JSONRpc` | `string` | | | `ID` | `interface{}` | | | `Method` | `string` | | | `Params` | `json.RawMessage` | | | `Result` | `json.RawMessage` | | | `Error` | `*MCPError` | | {/* /gen:fields */} | Function | Description | | --- | --- | | `mcp.CreateRequest(id, method, params) (*MCPMessage, error)` | Builds a request. | | `mcp.CreateNotification(method, params) (*MCPMessage, error)` | Builds a notification (no ID). | | `mcp.CreateResponse(id, result) (*MCPMessage, error)` | Builds a success response. | | `mcp.CreateErrorResponse(id, code, message, data) *MCPMessage` | Builds an error response. | | `mcp.ValidateJSONRPCMessage(raw []byte) (*MCPMessage, error)` | Checks that `raw` is a well-formed request, notification or response. | | `mcp.IsRequest(msg)`, `mcp.IsNotification(msg)`, `mcp.IsResponse(msg)`, `mcp.IsError(msg)` | Classify a message. | | `mcp.GetError(msg) error` | Returns the error from an error response. | | `mcp.ParseParams(msg, target)`, `mcp.ParseResult(msg, target)` | Decode `params` or `result` into `target`. | | `mcp.NewIDGenerator() *IDGenerator` | Creates a generator. `Next()` returns a new request ID. | ## Initialize and discover | Type | Description | | --- | --- | | `mcp.InitializeParams` | Parameters of the `initialize` request. | | `mcp.InitializeResult` | Result of `initialize`. | | `mcp.ClientInfo`, `mcp.ServerInfo` | Name and version of each side. | | `mcp.ClientCapabilities` | What the client supports. | | `mcp.ServerCapabilities` | What the server supports. | | `mcp.DiscoverResult` | Result of `server/discover`. | {/* gen:fields mcp.ClientCapabilities */} | Field | Type | Description | | --- | --- | --- | | `Experimental` | `map[string]interface{}` | | | `Extensions` | `map[string]interface{}` | | | `Roots` | `*RootsCapability` | | | `Sampling` | `*SamplingCapability` | | | `Elicitation` | `map[string]interface{}` | | {/* /gen:fields */} {/* gen:fields mcp.ServerCapabilities */} | Field | Type | Description | | --- | --- | --- | | `Experimental` | `map[string]interface{}` | | | `Logging` | `*LoggingCapability` | | | `Completions` | `*CompletionsCapability` | Completions advertises support for the completion/complete method (hash 68a739a). See (\*MCPClient).Complete. | | `Prompts` | `*PromptsCapability` | | | `Resources` | `*ResourcesCapability` | | | `Tools` | `*ToolsCapability` | | {/* /gen:fields */} The capability types are `mcp.RootsCapability`, `mcp.SamplingCapability`, `mcp.ToolsCapability`, `mcp.ResourcesCapability`, `mcp.PromptsCapability`, `mcp.LoggingCapability` and `mcp.CompletionsCapability`. ## Tools {/* gen:fields mcp.MCPTool */} | Field | Type | Description | | --- | --- | --- | | `Name` | `string` | | | `Title` | `string` | | | `Description` | `string` | | | `InputSchema` | `map[string]interface{}` | | | `OutputSchema` | `map[string]interface{}` | | | `Annotations` | `map[string]interface{}` | | | `Meta` | `map[string]interface{}` | | {/* /gen:fields */} | Type | Description | | --- | --- | | `mcp.ListToolsParams`, `mcp.ListToolsResult` | Parameters and result of `tools/list`, including pagination. | | `mcp.CallToolParams`, `mcp.CallToolResult` | Parameters and result of `tools/call`. | | `mcp.ToolResultContent` | One content item in a tool result. | {/* gen:fields mcp.CallToolResult */} | Field | Type | Description | | --- | --- | --- | | `Content` | `[]ToolResultContent` | | | `StructuredContent` | `interface{}` | | | `ToolResult` | `interface{}` | | | `IsError` | `bool` | | | `Metadata` | `map[string]interface{}` | | {/* /gen:fields */} ## Resources, prompts and completions See [MCP client](https://goaisdk.com/docs/reference/mcp/client.md#resources) for `MCPResource`, `MCPResourceTemplate`, `ReadResourceParams`, `ReadResourceResult`, `ResourceContent`, `MCPPrompt`, `PromptMessage`, `CompleteRequestParams` and related types. ## Logging A client can set the server's log level and receive log messages. {/* gen:consts mcp.LoggingLevel */} | Constant | Value | Description | | --- | --- | --- | | `LoggingLevelDebug` | `"debug"` | | | `LoggingLevelInfo` | `"info"` | | | `LoggingLevelNotice` | `"notice"` | | | `LoggingLevelWarning` | `"warning"` | | | `LoggingLevelError` | `"error"` | | | `LoggingLevelCritical` | `"critical"` | | | `LoggingLevelAlert` | `"alert"` | | | `LoggingLevelEmergency` | `"emergency"` | | {/* /gen:consts */} {/* gen:fields mcp.SetLevelParams */} | Field | Type | Description | | --- | --- | --- | | `Level` | `LoggingLevel` | | {/* /gen:fields */} {/* gen:fields mcp.LoggingMessageNotification */} | Field | Type | Description | | --- | --- | --- | | `Level` | `LoggingLevel` | | | `Logger` | `string` | | | `Data` | `interface{}` | | {/* /gen:fields */} --- # Built-in Middleware > Catalog of built-in Go AI SDK middleware, including DefaultSettingsMiddleware, DefaultInstructionsMiddleware, ExtractJSONMiddleware, ExtractReasoningMiddleware, SimulateStreamingMiddleware, and AddToolInputExamplesMiddleware. Canonical URL: https://goaisdk.com/docs/reference/middleware/built-in-middleware Documentation index: https://goaisdk.com/llms.txt Pre-built middleware functions available in the Go AI SDK's `pkg/middleware` package. All of them return `*middleware.LanguageModelMiddleware` for use with [`WrapLanguageModel`](https://goaisdk.com/docs/reference/middleware/wrap-language-model.md) or [`WrapProvider`](#wrapprovider). ## DefaultSettingsMiddleware Applies default settings to generation requests when the caller didn't set them. ```go func DefaultSettingsMiddleware(settings *provider.GenerateOptions) *LanguageModelMiddleware ``` ### Example ```go temperature := 0.7 maxTokens := 500 defaultsMw := middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &temperature, MaxTokens: &maxTokens, }) wrapped := middleware.WrapLanguageModel(baseModel, []*middleware.LanguageModelMiddleware{defaultsMw}, nil, nil) ``` ## DefaultInstructionsMiddleware Applies default system instructions when the call doesn't provide any. ```go func DefaultInstructionsMiddleware(opts DefaultInstructionsOptions) *LanguageModelMiddleware type DefaultInstructionsOptions struct { // Instructions are plain-string default instructions, applied as // Prompt.System when the call has none. Instructions string // InstructionMessages supplies the default instructions as system // messages instead, preserving each message's ProviderOptions. Takes // precedence over Instructions when non-empty. InstructionMessages []types.Message } ``` ### Example ```go instructionsMw := middleware.DefaultInstructionsMiddleware(middleware.DefaultInstructionsOptions{ Instructions: "You are a concise, helpful assistant.", }) wrapped := middleware.WrapLanguageModel(baseModel, []*middleware.LanguageModelMiddleware{instructionsMw}, nil, nil) ``` ## ExtractJSONMiddleware Extracts JSON from model output, stripping markdown code fences by default. ```go func ExtractJSONMiddleware(options *ExtractJSONOptions) *LanguageModelMiddleware type ExtractJSONOptions struct { // Transform is a custom function to extract JSON from text. // If nil, the default transform strips markdown code fences. Transform func(text string) string } ``` ### Example ```go jsonMw := middleware.ExtractJSONMiddleware(nil) // use the default fence-stripping transform wrapped := middleware.WrapLanguageModel(baseModel, []*middleware.LanguageModelMiddleware{jsonMw}, nil, nil) ``` ## ExtractReasoningMiddleware Extracts reasoning content wrapped in an XML-style tag (for example ``) out of the model's text output. ```go func ExtractReasoningMiddleware(options *ExtractReasoningOptions) *LanguageModelMiddleware type ExtractReasoningOptions struct { // TagName is the XML tag name to extract reasoning from // (e.g., "think" for Anthropic, "reasoning" for OpenAI) TagName string // Separator is the separator to use between reasoning and text sections // Default: "\n" Separator string // StartWithReasoning indicates whether reasoning tokens appear at the beginning // Default: false StartWithReasoning bool } ``` ### Example ```go reasoningMw := middleware.ExtractReasoningMiddleware(&middleware.ExtractReasoningOptions{ TagName: "think", }) wrapped := middleware.WrapLanguageModel(baseModel, []*middleware.LanguageModelMiddleware{reasoningMw}, nil, nil) ``` ## SimulateStreamingMiddleware Simulates a streaming response from a model that only supports non-streaming generation, by generating the full result and then replaying it as stream chunks. ```go func SimulateStreamingMiddleware() *LanguageModelMiddleware ``` ### Example ```go simulateMw := middleware.SimulateStreamingMiddleware() wrapped := middleware.WrapLanguageModel(baseModel, []*middleware.LanguageModelMiddleware{simulateMw}, nil, nil) ``` ## AddToolInputExamplesMiddleware Appends each tool's `InputExamples` to its description so the model sees example inputs, then (by default) removes `InputExamples` from the tool definition sent to the provider. ```go func AddToolInputExamplesMiddleware(options *AddToolInputExamplesOptions) *LanguageModelMiddleware type AddToolInputExamplesOptions struct { // Prefix is the text to prepend before examples // Default: "Input Examples:" Prefix string // Format is a custom formatter for each example // If nil, uses JSON.stringify on the example input Format func(example types.ToolInputExample, index int) string // Remove indicates whether to remove the InputExamples property // after adding them to the description. Default: true. Remove *bool } ``` ### Example ```go examplesMw := middleware.AddToolInputExamplesMiddleware(&middleware.AddToolInputExamplesOptions{ Prefix: "Examples:", }) wrapped := middleware.WrapLanguageModel(baseModel, []*middleware.LanguageModelMiddleware{examplesMw}, nil, nil) ``` ## Combining Middleware ```go temperature := 0.7 wrapped := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &temperature, }), middleware.ExtractReasoningMiddleware(&middleware.ExtractReasoningOptions{ TagName: "think", }), middleware.ExtractJSONMiddleware(nil), }, nil, nil, ) ``` ## Image and embedding model wrappers There are no built-in `*ImageModelMiddleware` or `*EmbeddingModelMiddleware` constructors analogous to the ones above, but the same wrapping mechanism is available for image and embedding models if you write your own middleware struct (see [Middleware Interface](https://goaisdk.com/docs/reference/middleware/middleware-interface.md)): ```go func WrapImageModel(model provider.ImageModel, middleware []*ImageModelMiddleware, modelID, providerID *string) provider.ImageModel func WrapEmbeddingModel(model provider.EmbeddingModel, middleware []*EmbeddingModelMiddleware, modelID, providerID *string) provider.EmbeddingModel ``` ## WrapProvider Applies language model and embedding model middleware to every model returned by a provider, instead of wrapping one model at a time. ```go func WrapProvider( p provider.Provider, languageModelMiddleware []*LanguageModelMiddleware, embeddingModelMiddleware []*EmbeddingModelMiddleware, opts ...ProviderMiddlewareOption, ) provider.Provider ``` Pass `middleware.WithImageModelMiddleware([]*middleware.ImageModelMiddleware{...})` as a trailing option to also apply image model middleware. ### Example ```go temperature := 0.7 p := openai.New(openai.Config{APIKey: "your-api-key"}) wrappedProvider := middleware.WrapProvider( p, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &temperature, }), }, nil, ) model, _ := wrappedProvider.LanguageModel("gpt-6-astra") ``` ## See Also - [Middleware Interface](https://goaisdk.com/docs/reference/middleware/middleware-interface.md) - Create custom middleware - [WrapLanguageModel](https://goaisdk.com/docs/reference/middleware/wrap-language-model.md) - Apply middleware to models --- # Middleware Interface > API reference for LanguageModelMiddleware, EmbeddingModelMiddleware, and ImageModelMiddleware, the Go AI SDK structs for implementing custom model-behavior middleware. Canonical URL: https://goaisdk.com/docs/reference/middleware/middleware-interface Documentation index: https://goaisdk.com/llms.txt Structs for implementing custom middleware to intercept and modify model behavior. Each field is an optional hook; a middleware only needs to set the hooks it uses, and leave the rest nil. ## LanguageModelMiddleware ```go type LanguageModelMiddleware struct { // SpecificationVersion should be "v3" for the current version SpecificationVersion string // OverrideProvider allows overriding the provider name OverrideProvider func(model provider.LanguageModel) string // OverrideModelID allows overriding the model ID OverrideModelID func(model provider.LanguageModel) string // OverrideSupportedURLs allows overriding the URL patterns (regular // expressions keyed by media type) the wrapped model reports supporting. OverrideSupportedURLs func(model provider.LanguageModel) map[string][]string // TransformParams transforms the parameters before they are passed to the language model TransformParams func(ctx context.Context, callType string, params *provider.GenerateOptions, model provider.LanguageModel) (*provider.GenerateOptions, error) // WrapGenerate wraps the generate operation of the language model WrapGenerate func(ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel) (*types.GenerateResult, error) // WrapStream wraps the stream operation of the language model WrapStream func(ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel) (provider.TextStream, error) } ``` Middleware for language models that can intercept both generation and streaming calls. `WrapGenerate` and `WrapStream` each receive both `doGenerate` and `doStream` closures — a `WrapGenerate` hook calls `doGenerate()` to invoke the wrapped model (or the next middleware in the chain), and a `WrapStream` hook calls `doStream()`. ## EmbeddingModelMiddleware ```go type EmbeddingModelMiddleware struct { SpecificationVersion string OverrideProvider func(model provider.EmbeddingModel) string OverrideModelID func(model provider.EmbeddingModel) string OverrideMaxEmbeddingsPerCall func(model provider.EmbeddingModel) int OverrideSupportsParallelCalls func(model provider.EmbeddingModel) bool // TransformInput transforms the input before it is passed to the embedding model TransformInput func(ctx context.Context, input string, model provider.EmbeddingModel) (string, error) // WrapEmbed wraps a single embed call WrapEmbed func(ctx context.Context, doEmbed func() (*types.EmbeddingResult, error), input string, model provider.EmbeddingModel) (*types.EmbeddingResult, error) // WrapEmbedMany wraps a batch embed call WrapEmbedMany func(ctx context.Context, doEmbedMany func() (*types.EmbeddingsResult, error), inputs []string, model provider.EmbeddingModel) (*types.EmbeddingsResult, error) } ``` Middleware for embedding models. ## ImageModelMiddleware ```go type ImageModelMiddleware struct { // SpecificationVersion should be "v3" or "v4" for the current version SpecificationVersion string OverrideProvider func(model provider.ImageModel) string OverrideModelID func(model provider.ImageModel) string OverrideMaxImagesPerCall func(model provider.ImageModel) int // TransformParams transforms the parameters before they are passed to the image model TransformParams func(ctx context.Context, params *provider.ImageGenerateOptions, model provider.ImageModel) (*provider.ImageGenerateOptions, error) // WrapGenerate wraps the generate operation of the image model WrapGenerate func(ctx context.Context, doGenerate func() (*types.ImageResult, error), params *provider.ImageGenerateOptions, model provider.ImageModel) (*types.ImageResult, error) } ``` Middleware for image models. ## Examples ### Logging Middleware ```go package main import ( "context" "log" "time" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func LoggingMiddleware() *middleware.LanguageModelMiddleware { return &middleware.LanguageModelMiddleware{ WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { start := time.Now() log.Printf("Starting generation...") result, err := doGenerate() duration := time.Since(start) if err != nil { log.Printf("Generation failed after %v: %v", duration, err) } else { log.Printf("Generation completed in %v (%d tokens)", duration, result.Usage.GetTotalTokens()) } return result, err }, } } ``` ### Retry Middleware ```go func RetryMiddleware(maxRetries int) *middleware.LanguageModelMiddleware { return &middleware.LanguageModelMiddleware{ WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { var result *types.GenerateResult var err error for attempt := 0; attempt < maxRetries; attempt++ { result, err = doGenerate() if err == nil { return result, nil } log.Printf("Retry attempt %d/%d", attempt+1, maxRetries) time.Sleep(time.Duration(attempt+1) * time.Second) } return nil, fmt.Errorf("max retries exceeded: %w", err) }, } } ``` ### Token Counter Middleware ```go type TokenCounter struct { totalTokens int mu sync.Mutex } func (tc *TokenCounter) Middleware() *middleware.LanguageModelMiddleware { return &middleware.LanguageModelMiddleware{ WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { result, err := doGenerate() if err != nil { return nil, err } tc.mu.Lock() tc.totalTokens += int(result.Usage.GetTotalTokens()) tc.mu.Unlock() return result, nil }, } } func (tc *TokenCounter) Total() int { tc.mu.Lock() defer tc.mu.Unlock() return tc.totalTokens } ``` ## See Also - [WrapLanguageModel](https://goaisdk.com/docs/reference/middleware/wrap-language-model.md) - Apply middleware to models - [Built-in Middleware](https://goaisdk.com/docs/reference/middleware/built-in-middleware.md) - Available middleware catalog --- # WrapLanguageModel > API reference for WrapLanguageModel, the Go AI SDK function that wraps a language model with middleware to add default settings, logging, or other behavior. Canonical URL: https://goaisdk.com/docs/reference/middleware/wrap-language-model Documentation index: https://goaisdk.com/llms.txt Wraps a language model with middleware to add additional behavior like default settings, logging, or extraction. ## Signature ```go func WrapLanguageModel( model provider.LanguageModel, middleware []*middleware.LanguageModelMiddleware, modelID *string, providerID *string, ) provider.LanguageModel ``` When multiple middleware are provided, the first middleware transforms the input first, and the last middleware is wrapped directly around the model. ## Parameters | Parameter | Type | Description | |-----------|------|-------------| | model | provider.LanguageModel | Base language model to wrap | | middleware | []*middleware.LanguageModelMiddleware | Middleware to apply, in order | | modelID | *string | Optional override for the reported model ID (nil keeps the wrapped model's ID) | | providerID | *string | Optional override for the reported provider name (nil keeps the wrapped model's provider) | ## Examples ### Basic Wrapping ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/middleware" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { p := openai.New(openai.Config{ APIKey: "your-api-key", }) baseModel, _ := p.LanguageModel("gpt-6-astra") temperature := 0.7 wrapped := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &temperature, }), }, nil, nil, ) result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: wrapped, Prompt: "Hello!", // Temperature will default to 0.7 }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Multiple Middleware ```go temperature := 0.7 maxTokens := 500 wrapped := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{ middleware.DefaultSettingsMiddleware(&provider.GenerateOptions{ Temperature: &temperature, MaxTokens: &maxTokens, }), middleware.ExtractReasoningMiddleware(&middleware.ExtractReasoningOptions{ TagName: "think", }), }, nil, nil, ) ``` ### Overriding the Reported Model ID ```go modelID := "custom-model-alias" wrapped := middleware.WrapLanguageModel( baseModel, nil, &modelID, nil, ) ``` ### Custom Middleware ```go customMiddleware := &middleware.LanguageModelMiddleware{ WrapGenerate: func( ctx context.Context, doGenerate func() (*types.GenerateResult, error), doStream func() (provider.TextStream, error), params *provider.GenerateOptions, model provider.LanguageModel, ) (*types.GenerateResult, error) { start := time.Now() result, err := doGenerate() log.Printf("Generation took %v", time.Since(start)) return result, err }, } wrapped := middleware.WrapLanguageModel( baseModel, []*middleware.LanguageModelMiddleware{customMiddleware}, nil, nil, ) ``` ## See Also - [Middleware Interface](https://goaisdk.com/docs/reference/middleware/middleware-interface.md) - Middleware struct definitions - [Built-in Middleware](https://goaisdk.com/docs/reference/middleware/built-in-middleware.md) - Available middleware catalog --- # Anthropic configuration > Reference for the Anthropic provider Config struct and model option types, generated from the Go source. Canonical URL: https://goaisdk.com/docs/reference/providers/anthropic-config Documentation index: https://goaisdk.com/llms.txt Package `github.com/digitallysavvy/go-ai/pkg/providers/anthropic`. Create a provider with `anthropic.New(anthropic.Config{...})`. When `BaseURL` is empty, the provider reads `ANTHROPIC_BASE_URL`. The tables on this page are generated from the Go source. For usage, see the [Anthropic provider page](https://goaisdk.com/docs/providers/anthropic.md). ## Config {/* gen:fields providers/anthropic.Config */} | Field | Type | Description | | --- | --- | --- | | `APIKey` | `string` | APIKey is the Anthropic API key | | `Name` | `string` | Name overrides the provider name returned by Provider.Name() and LanguageModel.Provider(). Defaults to "anthropic". | | `BaseURL` | `string` | BaseURL is the URL prefix for API calls, including the version path (default: https://api.anthropic.com/v1, or ANTHROPIC_BASE_URL). The bare host https://api.anthropic.com is normalized to .../v1. A base URL that is set but empty after trimming whitespace makes New panic, matching the TS validateBaseURL error ("baseURL must be a non-empty string."). | | `APIVersion` | `string` | APIVersion is the Anthropic API version (default: 2023-06-01) | | `OmitAPIVersionHeader` | `bool` | OmitAPIVersionHeader disables the anthropic-version HTTP header. This is used by Vertex Anthropic, which sends anthropic_version in the JSON body. | | `HTTPClient` | `*stdhttp.Client` | HTTPClient overrides the HTTP client used for all requests. | | `MessagesPath` | `func(modelID string, stream bool) string` | MessagesPath builds the request path for messages API calls. Defaults to "/messages" (relative to BaseURL). | | `TransformRequestBody` | `func(body map[string]interface{}, stream bool) map[string]interface{}` | TransformRequestBody can rewrite the Anthropic messages request body before it is sent. It is used by Vertex Anthropic to remove the model field and inject anthropic_version in the JSON body. | | `TransformRequestBodyWithBetas` | `func(body map[string]interface{}, betas []string, stream bool) map[string]interface{}` | TransformRequestBodyWithBetas rewrites the request body with access to the request's anthropic-beta flags (TS transformRequestBody(args, betas)). It runs before TransformRequestBody. Bedrock-Anthropic uses it. | | `TransformStreamBody` | `func(body io.ReadCloser, header stdhttp.Header) io.ReadCloser` | TransformStreamBody wraps the streaming response body before SSE parsing (for example to convert an AWS event stream into SSE). | | `TransformErrorBody` | `func(body []byte) []byte` | TransformErrorBody rewrites a non-2xx response body into the Anthropic error shape before it is parsed. | | `SupportsNativeStructuredOutput` | `*bool` | SupportsNativeStructuredOutput gates native structured output. A nil value means true; the model capability must also allow it. | | `SupportsImageInput` | `*bool` | SupportsImageInput overrides model capability detection. A nil value preserves the default Anthropic model-based behavior. | | `SupportsStrictTools` | `*bool` | SupportsStrictTools controls whether strict mode on function tools is sent to Anthropic. A nil value means true; the model capability must also allow it. | | `SupportedURLs` | `func(modelID string) map[string][]string` | SupportedURLs overrides the URL patterns the model accepts directly without downloading first (TS languageModelConfig.supportedUrls). A nil value falls back to DefaultSupportedURLs (https image/\* and application/pdf), matching the direct Anthropic and anthropic-aws providers. Vertex-Anthropic and Bedrock-Anthropic set this to a function that returns an empty map to force base64 conversion, matching TS. | | `Headers` | `map[string]string` | Headers are custom HTTP headers to include in requests. | | `UserAgentName` | `string` | UserAgentName selects the `ai-sdk-/VERSION` User-Agent tag this provider construction adds (version.ProviderUserAgent). Defaults to "anthropic". TS's anthropic-aws and minimax packages are each their own npm package with their own tag ("ai-sdk-anthropic-aws", "ai-sdk-minimax"); Go's anthropicaws and minimax packages implement this by reusing anthropic.New as their Messages-API transport, so they set this field to avoid inheriting the wrong "ai-sdk-anthropic" tag. | | `NoUserAgentTag` | `bool` | NoUserAgentTag disables the `ai-sdk-/VERSION` tag entirely (UserAgentName is ignored when this is true). TS's google-vertex-anthropic-provider.ts builds its AnthropicLanguageModel directly instead of going through createAnthropic (the only place TS's own `ai-sdk-anthropic/VERSION` tag is added), so Vertex-Anthropic requests carry no anthropic-package tag at all -- only the runtime tag the shared HTTP client appends downstream. pkg/providers/googlevertex/anthropic sets this to match. | {/* /gen:fields */} ## Model factories | Method | Returns | | --- | --- | | `LanguageModel(id)` | A language model with default options. | | `LanguageModelWithOptions(id, *ModelOptions)` | A language model configured with `ModelOptions`. | | `EmbeddingModel(id)`, `ImageModel(id)`, `SpeechModel(id)`, `TranscriptionModel(id)`, `RerankingModel(id)` | Return an error. Anthropic does not offer these model types. | | `Files()` | The Files API client. | | `Skills()` | The Skills API client. | ## ModelOptions Options for `LanguageModelWithOptions`. {/* gen:fields providers/anthropic.ModelOptions */} | Field | Type | Description | | --- | --- | --- | | `ContextManagement` | `*ContextManagement` | ContextManagement enables automatic cleanup of conversation history to prevent context window overflow in long conversations. This is a beta feature that requires Claude 4.5+ models. When enabled, Anthropic will automatically remove or truncate old content based on the configured strategies. Example: options := anthropic.ModelOptions\{ ContextManagement: &anthropic.ContextManagement\{ Strategies: []string\{anthropic.StrategyClearToolUses\}, \}, \} See ContextManagement for available strategies. | | `Compaction` | `*CompactionOption` | Compaction requests an on-demand summary of the supplied conversation. Mutually exclusive with ContextManagement (setting both returns an error). Adds the "compact-2026-09-04" beta header automatically. Example: options := anthropic.ModelOptions\{ Compaction: &anthropic.CompactionOption\{Type: "summarize"\}, \} | | `Thinking` | `*ThinkingConfig` | Thinking configures Claude's extended thinking capabilities. When enabled, responses include thinking content blocks showing Claude's reasoning process before the final answer. For Opus 4.6 and newer models, use ThinkingTypeAdaptive: options := anthropic.ModelOptions\{ Thinking: &anthropic.ThinkingConfig\{ Type: anthropic.ThinkingTypeAdaptive, \}, \} For models before Opus 4.6, use ThinkingTypeEnabled with optional budget: budget := 5000 options := anthropic.ModelOptions\{ Thinking: &anthropic.ThinkingConfig\{ Type: anthropic.ThinkingTypeEnabled, BudgetTokens: &budget, \}, \} | | `Speed` | `Speed` | Speed configures the inference speed mode. Fast mode provides 2.5x faster output token speeds but is only supported with claude-opus-4-6. Example: options := anthropic.ModelOptions\{ Speed: anthropic.SpeedFast, \} | | `AutomaticCaching` | `bool` | AutomaticCaching enables Anthropic's automatic prompt caching feature. When enabled, the API automatically identifies and caches reusable prompt segments without requiring explicit cache_control markers in individual messages. This requires the "prompt-caching-2024-07-31" beta header. Example: options := anthropic.ModelOptions\{ AutomaticCaching: true, \} See https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching for details. | | `CacheControl` | `*CacheControlOption` | CacheControl configures explicit ephemeral prompt caching. Mutually exclusive with AutomaticCaching; CacheControl takes precedence if both are set. Example: options := anthropic.ModelOptions\{ CacheControl: &anthropic.CacheControlOption\{Type: "ephemeral", TTL: "5m"\}, \} | | `Effort` | `Effort` | Effort controls the model's reasoning effort level. Supported values: EffortLow, EffortMedium, EffortHigh, EffortXHigh, EffortMax. Sent as output_config.effort (no beta header is required). Example: options := anthropic.ModelOptions\{ Effort: anthropic.EffortHigh, \} | | `TaskBudget` | `*TaskBudget` | TaskBudget informs the model of the total token budget available for the current task. This is advisory only; it does not enforce a hard limit. Requires the "task-budgets-2026-03-13" beta header (injected automatically). | | `InferenceGeo` | `string` | InferenceGeo controls where Anthropic inference may run for this request. Supported values match the TypeScript SDK: "us" or "global". | | `Fallbacks` | `[]FallbackConfig` | Fallbacks configures Anthropic server-side fallback attempts. | | `FallbacksDefault` | `bool` | FallbacksDefault sends fallbacks: "default" to use Anthropic's default server-side fallback chain (beta server-side-fallback-2026-07-01). It takes precedence over Fallbacks. | | `ServiceTier` | `string` | ServiceTier selects the Anthropic service tier: "auto" or "standard_only". Serialized as service_tier. | | `AnthropicBeta` | `[]string` | AnthropicBeta lists additional anthropic-beta flags to send. | | `Safeguards` | `[]Safeguard` | Safeguards configures safeguard classifiers (for example dangerous_tool_use). Classifier verdicts are returned in providerMetadata.anthropic.safeguardResults. | | `ToolStreaming` | `*bool` | ToolStreaming controls whether fine-grained tool streaming is enabled. When nil or true, streaming requests set eager_input_streaming: true on function tools that do not set ToolOptions.EagerInputStreaming. Set it to false to disable that default. Example (disable): disabled := false options := anthropic.ModelOptions\{ToolStreaming: &disabled\} | | `DisableParallelToolUse` | `*bool` | DisableParallelToolUse prevents the model from calling multiple tools in a single response. When true, adds \{disable_parallel_tool_use: true\} to the tool_choice object sent to the API. An explicit false is ignored (with a warning) when the JSON response tool is used for structured output. Example: disable := true options := anthropic.ModelOptions\{ DisableParallelToolUse: &disable, \} | | `MCPServers` | `[]MCPServerConfig` | MCPServers configures remote MCP servers for native server-side tool invocation. The Anthropic API connects to these MCP servers directly, exposing their tools to the model without the caller having to proxy individual tool calls. Requires the "mcp-client-2025-04-04" beta header (injected automatically). Example: options := anthropic.ModelOptions\{ MCPServers: []anthropic.MCPServerConfig\{ \{Type: "url", Name: "my-server", URL: "https://mcp.example.com/sse"\}, \}, \} | | `Container` | `*ContainerConfig` | Container configures an Anthropic agent container for code execution and skills. When Skills are provided, the code-execution-2025-08-25, skills-2025-10-02, and files-api-2025-04-14 beta headers are automatically injected. Example (full config with skills): options := anthropic.ModelOptions\{ Container: &anthropic.ContainerConfig\{ Skills: []anthropic.ContainerSkill\{ \{Type: "anthropic", SkillID: "web_search"\}, \}, \}, \} | | `ContainerID` | `string` | ContainerID is a shorthand for specifying a plain container ID string. When set, the container body field is sent as a plain string value. Mutually exclusive with Container; ContainerID takes precedence if both are set. Example: options := anthropic.ModelOptions\{ ContainerID: "container-abc123", \} | | `StructuredOutputMode` | `StructuredOutputMode` | StructuredOutputMode controls how ResponseFormat is sent to the API. Default (empty/"auto"): uses output_config.format for models that support it (claude-\*-4-6, claude-\*-4-5), falls back to a synthetic json tool for models that don't (e.g. claude-sonnet-4-20250514, claude-3-7-sonnet). Example: options := anthropic.ModelOptions\{ StructuredOutputMode: anthropic.StructuredOutputJSONTool, \} | | `SendReasoning` | `*bool` | SendReasoning controls whether ReasoningContent (thinking) blocks from message history are included when sending messages to the Anthropic API. Default (nil or true): reasoning blocks with a valid Signature are included in outgoing requests — required when the receiving model supports extended thinking and the blocks were originally produced by that model. Set to false when routing a conversation to a model that does not support thinking input (e.g. an older Claude model or one without thinking enabled). This strips all ReasoningContent parts from outgoing message history before the API call, preventing errors caused by unexpected thinking content. Example (disable when switching to non-thinking model): disabled := false options := anthropic.ModelOptions\{SendReasoning: &disabled\} | | `Metadata` | `*Metadata` | Metadata to include with the request (TS anthropicLanguageModelOptions.metadata). Example: options := anthropic.ModelOptions\{ Metadata: &anthropic.Metadata\{UserID: "user-123"\}, \} | {/* /gen:fields */} ## ThinkingConfig {/* gen:fields providers/anthropic.ThinkingConfig */} | Field | Type | Description | | --- | --- | --- | | `Type` | `ThinkingType` | Type specifies the thinking mode | | `BudgetTokens` | `*int` | BudgetTokens specifies the maximum tokens for thinking (only for "enabled" type) Requires a minimum of 1,024 tokens and counts towards the max_tokens limit. Optional for "enabled" type, not used for "adaptive" type. | | `Display` | `ThinkingDisplay` | Display controls how thinking is returned for "adaptive" thinking: ThinkingDisplayOmitted, ThinkingDisplaySummarized or ThinkingDisplayUpdates (adds the thinking-display-updates beta). Ignored for other thinking types. | | `BlockBinding` | `*ThinkingBlockBinding` | BlockBinding configures preserved-thinking block binding (thinking.block_binding). It may be set with Type "adaptive" or alone (empty Type) for binding-only recovery requests. Adds the thinking-binding-controls beta. | {/* /gen:fields */} ## ToolOptions Per-tool options, set on `types.Tool.ProviderOptions`. {/* gen:fields providers/anthropic.ToolOptions */} | Field | Type | Description | | --- | --- | --- | | `CacheControl` | `*CacheControl` | CacheControl enables prompt caching for this tool definition. When set, the tool definition will be cached for reuse across requests. | | `EagerInputStreaming` | `*bool` | EagerInputStreaming enables eager (larger-chunk) streaming of tool input deltas for this custom function tool. When true, Anthropic streams tool input in larger batches rather than byte-by-byte, improving streaming responsiveness. Only applies to custom function tools. Do NOT set on provider tools (web_search_20260209, web_fetch_20260209, etc.). When enabled, the stream emits tool-input-start, tool-input-delta, and tool-input-end chunks alongside the final tool-call chunk. | | `DeferLoading` | `*bool` | DeferLoading marks this tool for deferred loading when used alongside a tool search tool. When true, Claude does not load this tool's full definition upfront; instead it discovers and loads the tool on demand via tool_search. Use with AnthropicTools.ToolSearchBm2520251119 or AnthropicTools.ToolSearchRegex20251119. Serialized as defer_loading in the API request. | | `AllowedCallers` | `[]string` | AllowedCallers restricts which tool types can invoke this tool programmatically. Valid values: "direct", "code_execution_20250825", "code_execution_20260120" When set, the anthropic-beta: advanced-tool-use-2025-11-20 header is automatically injected. Serialized as allowed_callers in the API request. | {/* /gen:fields */} ## ContainerConfig {/* gen:fields providers/anthropic.ContainerConfig */} | Field | Type | Description | | --- | --- | --- | | `ID` | `string` | ID is an optional existing container ID to reuse | | `Skills` | `[]ContainerSkill` | Skills are the capability bundles to load into the container | {/* /gen:fields */} ## MCPServerConfig {/* gen:fields providers/anthropic.MCPServerConfig */} | Field | Type | Description | | --- | --- | --- | | `Type` | `string` | Type must be "url" | | `Name` | `string` | Name is a unique identifier for this server within the request | | `URL` | `string` | URL is the HTTP(S) endpoint of the MCP server | | `AuthorizationToken` | `string` | AuthorizationToken is an optional bearer token for authentication | | `ToolConfiguration` | `*MCPToolConfiguration` | ToolConfiguration optionally restricts which tools from this server are available | {/* /gen:fields */} --- # Custom Provider Guide > Guide for implementing custom AI providers for the Go AI SDK, covering the provider interface, a custom language model, and streaming support. Canonical URL: https://goaisdk.com/docs/reference/providers/custom-provider Documentation index: https://goaisdk.com/llms.txt Learn how to implement custom AI providers for the Go AI SDK. ## Provider Interface ```go type Provider interface { Name() string LanguageModel(modelID string) (LanguageModel, error) EmbeddingModel(modelID string) (EmbeddingModel, error) ImageModel(modelID string) (ImageModel, error) SpeechModel(modelID string) (SpeechModel, error) TranscriptionModel(modelID string) (TranscriptionModel, error) RerankingModel(modelID string) (RerankingModel, error) } ``` ## Implementing a Language Model ```go package main import ( "context" "fmt" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) type CustomLanguageModel struct { modelID string apiKey string } func (m *CustomLanguageModel) SpecificationVersion() string { return "v3" } func (m *CustomLanguageModel) Provider() string { return "custom" } func (m *CustomLanguageModel) ModelID() string { return m.modelID } func (m *CustomLanguageModel) SupportsTools() bool { return false } func (m *CustomLanguageModel) SupportsStructuredOutput() bool { return false } func (m *CustomLanguageModel) SupportsImageInput() bool { return false } func int64Ptr(v int64) *int64 { return &v } func (m *CustomLanguageModel) DoGenerate(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { // Implement your API call here return &types.GenerateResult{ Text: "Generated text", FinishReason: types.FinishReasonStop, Usage: types.Usage{ InputTokens: int64Ptr(10), OutputTokens: int64Ptr(20), TotalTokens: int64Ptr(30), }, }, nil } func (m *CustomLanguageModel) DoStream(ctx context.Context, opts *provider.GenerateOptions) (provider.TextStream, error) { return nil, fmt.Errorf("streaming not implemented") } ``` ## Implementing a Provider ```go type CustomProvider struct { apiKey string } func NewCustomProvider(apiKey string) *CustomProvider { return &CustomProvider{apiKey: apiKey} } func (p *CustomProvider) Name() string { return "custom" } func (p *CustomProvider) LanguageModel(modelID string) (provider.LanguageModel, error) { return &CustomLanguageModel{ modelID: modelID, apiKey: p.apiKey, }, nil } func (p *CustomProvider) EmbeddingModel(modelID string) (provider.EmbeddingModel, error) { return nil, fmt.Errorf("embedding models not supported") } func (p *CustomProvider) ImageModel(modelID string) (provider.ImageModel, error) { return nil, fmt.Errorf("image models not supported") } func (p *CustomProvider) SpeechModel(modelID string) (provider.SpeechModel, error) { return nil, fmt.Errorf("speech models not supported") } func (p *CustomProvider) TranscriptionModel(modelID string) (provider.TranscriptionModel, error) { return nil, fmt.Errorf("transcription models not supported") } func (p *CustomProvider) RerankingModel(modelID string) (provider.RerankingModel, error) { return nil, fmt.Errorf("reranking models not supported") } ``` ## Using Your Custom Provider ```go func main() { // Create provider provider := NewCustomProvider("your-api-key") // Get model model, err := provider.LanguageModel("custom-model") if err != nil { log.Fatal(err) } // Use with ai package result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Implementing Streaming `provider.TextStream` only has `Next`, `Err`, and `Close` — it does **not** implement `io.Reader`. Don't add a `Read` method to a stream wrapper; consume chunks with `Next()` in a loop, as shown in the [`provider.TextStream` doc comment](https://pkg.go.dev/github.com/digitallysavvy/go-ai/pkg/provider#TextStream). ```go type CustomTextStream struct { chunks chan *provider.StreamChunk err error } func NewCustomTextStream() *CustomTextStream { return &CustomTextStream{ chunks: make(chan *provider.StreamChunk, 10), } } func (s *CustomTextStream) Next() (*provider.StreamChunk, error) { chunk, ok := <-s.chunks if !ok { return nil, io.EOF } return chunk, nil } func (s *CustomTextStream) Close() error { close(s.chunks) return nil } func (s *CustomTextStream) Err() error { return s.err } ``` ## See Also - [LanguageModel](https://goaisdk.com/docs/reference/providers/language-model.md) - Language model interface - [EmbeddingModel](https://goaisdk.com/docs/reference/providers/embedding-model.md) - Embedding model interface - [Provider Registry](https://goaisdk.com/docs/reference/registry/provider-registry.md) - Registering providers --- # EmbeddingModel Interface > API reference for the EmbeddingModel interface in the Go AI SDK, defining the methods every embedding model implementation must satisfy. Canonical URL: https://goaisdk.com/docs/reference/providers/embedding-model Documentation index: https://goaisdk.com/llms.txt Interface that all embedding model implementations must satisfy. ## Interface Definition ```go type EmbeddingModel interface { // Metadata SpecificationVersion() string Provider() string ModelID() string // Capability methods MaxEmbeddingsPerCall() int SupportsParallelCalls() bool // Embedding methods DoEmbed(ctx context.Context, input string, opts *provider.EmbedModelOptions) (*types.EmbeddingResult, error) DoEmbedMany(ctx context.Context, inputs []string, opts *provider.EmbedModelOptions) (*types.EmbeddingsResult, error) } ``` `opts` may be `nil`. Use `provider.EmbedModelOptions` when forwarding provider-specific options or custom headers. ```go type EmbedModelOptions struct { ProviderOptions map[string]interface{} Headers map[string]string } ``` ## Methods ### Metadata Methods | Method | Returns | Description | |--------|---------|-------------| | SpecificationVersion() | string | Returns specification version | | Provider() | string | Provider name (e.g., "openai", "cohere") | | ModelID() | string | Model ID (e.g., "text-embedding-3-small") | ### Capability Methods | Method | Returns | Description | |--------|---------|-------------| | MaxEmbeddingsPerCall() | int | Maximum embeddings per API call (0 for unlimited) | | SupportsParallelCalls() | bool | Whether model supports parallel calls for batching | ### Embedding Methods | Method | Parameters | Returns | Description | |--------|------------|---------|-------------| | DoEmbed() | ctx, string, opts | *types.EmbeddingResult, error | Generate single embedding | | DoEmbedMany() | ctx, []string, opts | *types.EmbeddingsResult, error | Generate multiple embeddings | ## Examples ### Using an Embedding Model ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { // Get an embedding model provider := openai.New(openai.Config{ APIKey: "your-api-key", }) model, err := provider.EmbeddingModel("text-embedding-3-small") if err != nil { log.Fatal(err) } // Check capabilities fmt.Printf("Model: %s\n", model.ModelID()) fmt.Printf("Provider: %s\n", model.Provider()) fmt.Printf("Max embeddings per call: %d\n", model.MaxEmbeddingsPerCall()) fmt.Printf("Supports parallel calls: %v\n", model.SupportsParallelCalls()) // Generate embedding result, err := model.DoEmbed(context.Background(), "Hello, world!", nil) if err != nil { log.Fatal(err) } fmt.Printf("Embedding dimensions: %d\n", len(result.Embedding)) } ``` ### Batch Embeddings ```go texts := []string{ "First text", "Second text", "Third text", } result, err := model.DoEmbedMany(ctx, texts, nil) if err != nil { log.Fatal(err) } fmt.Printf("Generated %d embeddings\n", len(result.Embeddings)) for i, emb := range result.Embeddings { fmt.Printf("Text %d: %d dimensions\n", i+1, len(emb)) } ``` ### Provider Options Provider-specific embedding options use the same provider-keyed shape as generation APIs: ```go dimensions := 768 result, err := model.DoEmbed(ctx, "document text", &provider.EmbedModelOptions{ ProviderOptions: map[string]interface{}{ "google": google.GoogleEmbeddingProviderOptions{ TaskType: "RETRIEVAL_DOCUMENT", OutputDimensionality: &dimensions, }, }, }) ``` ### Respecting Max Embeddings Limit ```go func embedLargeDataset(model provider.EmbeddingModel, texts []string) ([][]float64, error) { maxPerCall := model.MaxEmbeddingsPerCall() if maxPerCall <= 0 { maxPerCall = len(texts) // Unlimited } var allEmbeddings [][]float64 for i := 0; i < len(texts); i += maxPerCall { end := i + maxPerCall if end > len(texts) { end = len(texts) } batch := texts[i:end] result, err := model.DoEmbedMany(ctx, batch, nil) if err != nil { return nil, err } allEmbeddings = append(allEmbeddings, result.Embeddings...) } return allEmbeddings, nil } ``` ## See Also - [LanguageModel](https://goaisdk.com/docs/reference/providers/language-model.md) - Language model interface - [Embed](https://goaisdk.com/docs/reference/ai/embed.md) - High-level embedding function - [EmbedMany](https://goaisdk.com/docs/reference/ai/embed-many.md) - Batch embedding function - [Custom Provider](https://goaisdk.com/docs/reference/providers/custom-provider.md) - Implementing custom providers --- # Google configuration > Reference for the Google Generative AI provider Config struct and model option types, generated from the Go source. Canonical URL: https://goaisdk.com/docs/reference/providers/google-config Documentation index: https://goaisdk.com/llms.txt Package `github.com/digitallysavvy/go-ai/pkg/providers/google`. Create a provider with `google.New(google.Config{...})`. When `APIKey` is empty, the provider reads `GOOGLE_GENERATIVE_AI_API_KEY`. The tables on this page are generated from the Go source. For usage, see the [Google provider page](https://goaisdk.com/docs/providers/google.md). ## Config {/* gen:fields providers/google.Config */} | Field | Type | Description | | --- | --- | --- | | `APIKey` | `string` | APIKey is the Google API key | | `BaseURL` | `string` | BaseURL is the base URL for the Google API (default: https://generativelanguage.googleapis.com) | | `Headers` | `map[string]string` | Headers are custom HTTP headers to include in requests. | | `Name` | `string` | Name overrides the provider name returned by Provider.Name(). Defaults to "google.generative-ai". | | `HTTPClient` | `*http.Client` | HTTPClient overrides the HTTP client used for requests. | {/* /gen:fields */} ## Model factories | Method | Returns | | --- | --- | | `LanguageModel(id)` | A Gemini language model. | | `Interactions(id)`, `InteractionsAgent(agent)`, `InteractionsManagedAgent(id)` | Language models that use the Interactions API. | | `EmbeddingModel(id)` | An embedding model. | | `ImageModel(id)` | An image generation model. | | `SpeechModel(id)` | A speech synthesis model. | | `TranscriptionModel(id)` | A transcription model. | | `VideoModel(id)` | A video generation model. | | `RerankingModel(id)` | A reranking model. | | `Files()` | The Files API client. | ## GoogleEmbeddingProviderOptions {/* gen:fields providers/google.GoogleEmbeddingProviderOptions */} | Field | Type | Description | | --- | --- | --- | | `Parts` | `[]EmbeddingPart` | Parts provides additional multimodal content parts alongside the text input. Each element corresponds to a single embedding value and is appended to that value's content.parts in the API request. | | `Content` | `[][]EmbeddingPart` | Content provides TypeScript-compatible per-value multimodal content. Each entry corresponds to the embedding value at the same index. A nil entry is text-only; a non-nil entry must contain at least one part. | | `OutputDimensionality` | `*int` | OutputDimensionality optionally truncates the embedding vector length. | | `TaskType` | `string` | TaskType maps to the Google embedding taskType request field. | {/* /gen:fields */} ## GoogleSpeechModelOptions {/* gen:fields providers/google.GoogleSpeechModelOptions */} | Field | Type | Description | | --- | --- | --- | | `MultiSpeakerVoiceConfig` | `map[string]interface{}` | MultiSpeakerVoiceConfig configures Gemini multi-speaker TTS. When set it takes precedence over the top-level voice. | {/* /gen:fields */} ## GoogleVideoModelOptions {/* gen:fields providers/google.GoogleVideoModelOptions */} | Field | Type | Description | | --- | --- | --- | | `PersonGeneration` | `*string` | | | `NegativePrompt` | `*string` | | | `ReferenceImages` | `[]googleVideoReferenceImageOpt` | | {/* /gen:fields */} ## GoogleInteractionsProviderOptions {/* gen:fields providers/google.GoogleInteractionsProviderOptions */} | Field | Type | Description | | --- | --- | --- | | `PreviousInteractionID` | `string` | | | `Store` | `*bool` | | | `MediaResolution` | `string` | | | `SystemInstruction` | `string` | | | `ResponseModalities` | `[]string` | | | `ServiceTier` | `string` | | | `ThinkingLevel` | `string` | | | `ThinkingSummaries` | `string` | | | `ResponseFormat` | `[]map[string]interface{}` | ResponseFormat holds raw (camelCase) response_format entries from providerOptions.google.responseFormat, e.g. \{"type":"video","aspectRatio":"16:9",...\}. Converted to snake_case wire entries in buildArgs. | | `ImageConfig` | `map[string]interface{}` | | | `Agent` | `string` | | | `AgentConfig` | `map[string]interface{}` | | | `Environment` | `interface{}` | | | `PollingTimeoutMs` | `int` | | | `Background` | `*bool` | | {/* /gen:fields */} --- # ImageModel Interface > API reference for the ImageModel interface in the Go AI SDK, defining the methods and ImageGenerateOptions every image model implementation must satisfy. Canonical URL: https://goaisdk.com/docs/reference/providers/image-model Documentation index: https://goaisdk.com/llms.txt Interface that all image generation model implementations must satisfy. ## Interface Definition ```go type ImageModel interface { // Metadata SpecificationVersion() string Provider() string ModelID() string // Image generation DoGenerate(ctx context.Context, opts *ImageGenerateOptions) (*types.ImageResult, error) } ``` ## Methods | Method | Parameters | Returns | Description | |--------|------------|---------|-------------| | SpecificationVersion() | - | string | Specification version | | Provider() | - | string | Provider name | | ModelID() | - | string | Model identifier | | DoGenerate() | ctx, *ImageGenerateOptions | *types.ImageResult, error | Generate images | ## Optional Capability Methods A model may optionally implement these methods to advertise whether it consumes `ImageGenerateOptions.Files` / `.Mask` for image editing. They are not part of the `ImageModel` interface itself, so implementing them is optional; callers read them through `provider.ImageModelSupportsFileInputs` / `provider.ImageModelSupportsMaskInputs`, which return `nil` (unknown) for a model that doesn't implement the method: ```go type ImageEditSupport interface { SupportsFileInputs() *bool // nil means unknown; only route file-input edits when this is non-nil true SupportsMaskInputs() *bool // nil means unknown; advertised separately since some models accept files without masks } ``` `middleware.WrapImageModel` preserves these from the wrapped model and lets an `ImageModelMiddleware` override either one via `OverrideSupportsFileInputs` / `OverrideSupportsMaskInputs`. ## ImageGenerateOptions ```go type ImageGenerateOptions struct { Prompt string N *int Size string Quality string Style string } ``` ## Examples ### Using an Image Model ```go package main import ( "context" "log" "os" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { p := openai.New(openai.Config{ APIKey: "your-api-key", }) model, err := p.ImageModel("dall-e-3") if err != nil { log.Fatal(err) } result, err := model.DoGenerate(context.Background(), &provider.ImageGenerateOptions{ Prompt: "A serene mountain landscape", Size: "1024x1024", }) if err != nil { log.Fatal(err) } // Save first image if len(result.Images) > 0 { os.WriteFile("output.png", result.Images[0], 0644) } } ``` ## See Also - [GenerateImage](https://goaisdk.com/docs/reference/ai/generate-image.md) - High-level image generation - [Custom Provider](https://goaisdk.com/docs/reference/providers/custom-provider.md) - Implementing custom providers --- # LanguageModel Interface > API reference for the LanguageModel interface in the Go AI SDK, defining the methods and GenerateOptions every language model implementation must satisfy. Canonical URL: https://goaisdk.com/docs/reference/providers/language-model Documentation index: https://goaisdk.com/llms.txt Interface that all language model implementations must satisfy. ## Interface Definition ```go type LanguageModel interface { // Metadata methods SpecificationVersion() string Provider() string ModelID() string // Capability methods SupportsTools() bool SupportsStructuredOutput() bool SupportsImageInput() bool // Generation methods DoGenerate(ctx context.Context, opts *GenerateOptions) (*types.GenerateResult, error) DoStream(ctx context.Context, opts *GenerateOptions) (TextStream, error) } ``` ## Methods ### Metadata Methods | Method | Returns | Description | |--------|---------|-------------| | SpecificationVersion() | string | Returns "v3" for V3 models | | Provider() | string | Provider name (e.g., "openai", "anthropic") | | ModelID() | string | Model ID (e.g., "gpt-6-astra", "claude-opus-5-5") | ### Capability Methods | Method | Returns | Description | |--------|---------|-------------| | SupportsTools() | bool | Whether model supports tool calling | | SupportsStructuredOutput() | bool | Whether model supports structured output (JSON mode) | | SupportsImageInput() | bool | Whether model accepts image inputs | ### Generation Methods | Method | Parameters | Returns | Description | |--------|------------|---------|-------------| | DoGenerate() | ctx, *GenerateOptions | *types.GenerateResult, error | Non-streaming generation | | DoStream() | ctx, *GenerateOptions | TextStream, error | Streaming generation | ## GenerateOptions ```go type GenerateOptions struct { Prompt types.Prompt Temperature *float64 MaxTokens *int TopP *float64 TopK *int FrequencyPenalty *float64 PresencePenalty *float64 StopSequences []string Tools []types.Tool ToolChoice types.ToolChoice ResponseFormat *ResponseFormat Seed *int Headers map[string]string MaxSteps *int ProviderOptions map[string]interface{} } ``` ## Examples ### Using a Language Model ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { // Get a language model p := openai.New(openai.Config{ APIKey: "your-api-key", }) model, err := p.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } // Check capabilities fmt.Printf("Model: %s\n", model.ModelID()) fmt.Printf("Provider: %s\n", model.Provider()) fmt.Printf("Supports tools: %v\n", model.SupportsTools()) fmt.Printf("Supports structured output: %v\n", model.SupportsStructuredOutput()) fmt.Printf("Supports image input: %v\n", model.SupportsImageInput()) // Generate text result, err := model.DoGenerate(context.Background(), &provider.GenerateOptions{ Prompt: types.Prompt{ Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Hello!"}, }, }, }, }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Check Model Capabilities ```go func checkModelCapabilities(model provider.LanguageModel) { fmt.Printf("Model: %s (%s)\n", model.ModelID(), model.Provider()) if model.SupportsTools() { fmt.Println("✓ Tool calling supported") } if model.SupportsStructuredOutput() { fmt.Println("✓ Structured output supported") } if model.SupportsImageInput() { fmt.Println("✓ Image input supported") } } ``` ### Streaming Generation ```go result, err := model.DoStream(ctx, &provider.GenerateOptions{ Prompt: types.Prompt{ Messages: messages, }, }) if err != nil { log.Fatal(err) } defer result.Close() for { chunk, err := result.Next() if err == io.EOF { break } if err != nil { log.Fatal(err) } if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } } ``` ## See Also - [EmbeddingModel](https://goaisdk.com/docs/reference/providers/embedding-model.md) - Embedding model interface - [ImageModel](https://goaisdk.com/docs/reference/providers/image-model.md) - Image generation interface - [Custom Provider](https://goaisdk.com/docs/reference/providers/custom-provider.md) - Implementing custom providers - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - High-level text generation --- # OpenAI configuration > Reference for the OpenAI provider Config struct and model option types, generated from the Go source. Canonical URL: https://goaisdk.com/docs/reference/providers/openai-config Documentation index: https://goaisdk.com/llms.txt Package `github.com/digitallysavvy/go-ai/pkg/providers/openai`. Create a provider with `openai.New(openai.Config{...})`. When `APIKey` is empty, the provider reads `OPENAI_API_KEY`. When `BaseURL` is empty, it reads `OPENAI_BASE_URL`. The tables on this page are generated from the Go source. For usage, see the [OpenAI provider page](https://goaisdk.com/docs/providers/openai.md). ## Config {/* gen:fields providers/openai.Config */} | Field | Type | Description | | --- | --- | --- | | `APIKey` | `string` | APIKey is the OpenAI API key | | `Name` | `string` | Name overrides the provider name returned by Provider.Name(). Defaults to "openai". | | `BaseURL` | `string` | BaseURL is the base URL for the OpenAI API (default: https://api.openai.com/v1) | | `Organization` | `string` | Organization is the optional organization ID | | `Project` | `string` | Project is the optional project ID | | `HTTPClient` | `*stdhttp.Client` | HTTPClient overrides the HTTP client used for all requests. | | `Headers` | `map[string]string` | Headers are custom HTTP headers to include in requests. | | `ChatProviderName` | `string` | ChatProviderName overrides the provider identifier returned by ChatModel. Defaults to "openai.chat". | | `CompletionProviderName` | `string` | CompletionProviderName overrides the provider identifier returned by CompletionModel. Defaults to "openai.completion". | | `CompletionProviderOptionsName` | `string` | CompletionProviderOptionsName selects the providerOptions/providerMetadata namespace used by completion models. Defaults to "azure" when the completion provider name contains "azure", otherwise "openai". | | `CompletionQuery` | `map[string]string` | CompletionQuery contains query parameters added to Completions API requests. | | `ResponsesProviderName` | `string` | ResponsesProviderName overrides the provider identifier returned by ResponsesModel. Defaults to "openai.responses". | | `ResponsesProviderOptionsName` | `string` | ResponsesProviderOptionsName selects the providerOptions/providerMetadata namespace used by Responses models. Defaults to "azure" when the Responses provider name contains "azure", otherwise "openai". | | `ResponsesQuery` | `map[string]string` | ResponsesQuery contains query parameters added to Responses API requests. | | `FileIDPrefixes` | `[]string` | FileIDPrefixes identifies deprecated string file-data values that should be sent to the Responses API as file_id references instead of base64 data. Nil defaults to []string\{"file-"\} to match the TypeScript OpenAI provider; set an empty non-nil slice to disable this compatibility path. | | `SupportsWebSearchSourcesInclude` | `*bool` | SupportsWebSearchSourcesInclude controls whether the Responses API request automatically includes "web_search_call.action.sources" when a web_search tool is present. Defaults to true (nil). Set to a pointer to false for backends that reject that include value (e.g. Amazon Bedrock Mantle). Callers can also override this per-call via the Responses provider option "includeWebSearchSources". | | `ExplicitMessageItemType` | `bool` | ExplicitMessageItemType adds an explicit `"type":"message"` field to system/developer/user Responses input items. Azure AI Foundry projects require this. | | `AllowVideo` | `bool` | AllowVideo emits a "video_url" content part for a video/\* FileContent (prompt.ToOpenAIMessagesOptions.AllowVideo). OpenAI's own Chat Completions API has no video support, so this defaults to false and must stay false for the openai package's own provider construction. Set true only by wrapper providers whose TS counterpart is an OpenAICompatibleChatLanguageModel subclass reusing Go's openai.Provider as its OpenAI-compatible base (baseten, cerebras, deepinfra). | | `UserAgentName` | `string` | UserAgentName selects the `ai-sdk-/VERSION` User-Agent tag this provider construction adds (version.ProviderUserAgent). Every wrapper provider whose TS counterpart is its own distinct npm package (and therefore its own `ai-sdk-` tag) but is implemented in Go by reusing openai.New as an OpenAI-Chat-Completions-compatible transport (cerebras, deepinfra, baseten, vercel, amazon-bedrock's Mantle gateway, google-vertex's MaaS models) must set this to that TS package's name; leaving it empty here would wrongly tag those providers' requests "ai-sdk-openai". Defaults to "openai" — unless Headers already carries a "user-agent" entry (case-insensitive), meaning the caller (e.g. the azure package, whose own TS package already applies its own `ai-sdk-azure` tag before reaching here) has already tagged the request and no further tag should be appended. | | `TransformRequestBody` | `func(body map[string]interface{}) map[string]interface{}` | TransformRequestBody can rewrite the Chat Completions request body before it is sent, mirroring TS OpenAICompatibleChatLanguageModel's `transformRequestBody` config hook (e.g. cerebras-provider.ts's transformCerebrasRequestBody, which renames max_tokens -> max_completion_tokens and reasoning_content -> reasoning). It runs inside buildRequestBodyWithWarnings, before the body is captured for both the outgoing HTTP request and the optional provider.StreamRequestBody / types.StepRequest.Body exposure, so RequestBody() reflects the same post-transform shape TS's `request: { body }` does. | {/* /gen:fields */} ## Model factories | Method | Returns | | --- | --- | | `LanguageModel(id)` | The default language model (Chat Completions API). | | `ChatModel(id)` | A Chat Completions language model. | | `ResponsesModel(id)` | A language model that uses the Responses API (`/v1/responses`). | | `CompletionModel(id)` | A legacy completions model. | | `EmbeddingModel(id)` | An embedding model. | | `ImageModel(id)` | An image generation model. | | `SpeechModel(id)` | A speech synthesis model. | | `TranscriptionModel(id)` | A transcription model. | | `RerankingModel(id)` | A reranking model. | | `Files()` | The Files API client. | | `Skills()` | The Skills API client. | ## OpenAIImageModelOptions Options for image generation. {/* gen:fields providers/openai.OpenAIImageModelOptions */} | Field | Type | Description | | --- | --- | --- | | `Quality` | `string` | | | `Size` | `string` | | | `Style` | `string` | | | `Background` | `string` | | | `Moderation` | `string` | | | `OutputFormat` | `string` | | | `OutputCompression` | `*int` | | | `InputFidelity` | `string` | | | `User` | `string` | | {/* /gen:fields */} ## OpenAIRealtimeModelOptions Options for realtime sessions. {/* gen:fields providers/openai.OpenAIRealtimeModelOptions */} | Field | Type | Description | | --- | --- | --- | | `API` | `*string` | API overrides model ID routing: "live" or "realtime". Nil (unset) routes known Live model IDs (e.g. "gpt-live-1") to Live and everything else to the GA Realtime API, matching TS resolveRealtimeApi, where an omitted (undefined) api is distinct from an explicit invalid value. A non-nil value that is neither "live" nor "realtime" (including "") is rejected, matching TS. | {/* /gen:fields */} --- # SpeechModel Interface > API reference for the SpeechModel interface in the Go AI SDK, defining the methods every text-to-speech model implementation must satisfy. Canonical URL: https://goaisdk.com/docs/reference/providers/speech-model Documentation index: https://goaisdk.com/llms.txt Interface that all speech synthesis (text-to-speech) model implementations must satisfy. ## Interface Definition ```go type SpeechModel interface { // Metadata SpecificationVersion() string Provider() string ModelID() string // Speech synthesis DoGenerate(ctx context.Context, opts *SpeechGenerateOptions) (*types.SpeechResult, error) } ``` ## Methods | Method | Parameters | Returns | Description | |--------|------------|---------|-------------| | SpecificationVersion() | - | string | Specification version | | Provider() | - | string | Provider name | | ModelID() | - | string | Model identifier | | DoGenerate() | ctx, *SpeechGenerateOptions | *types.SpeechResult, error | Generate speech audio | ## SpeechGenerateOptions ```go type SpeechGenerateOptions struct { Text string Voice string Speed *float64 } ``` ## Examples ### Using a Speech Model ```go package main import ( "context" "log" "os" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { p := openai.New(openai.Config{ APIKey: "your-api-key", }) model, err := p.SpeechModel("tts-1") if err != nil { log.Fatal(err) } result, err := model.DoGenerate(context.Background(), &provider.SpeechGenerateOptions{ Text: "Hello, world!", Voice: "alloy", }) if err != nil { log.Fatal(err) } os.WriteFile("output.mp3", result.Audio, 0644) } ``` ## See Also - [GenerateSpeech](https://goaisdk.com/docs/reference/ai/generate-speech.md) - High-level speech generation - [TranscriptionModel](https://goaisdk.com/docs/reference/providers/transcription-model.md) - Speech-to-text interface - [Custom Provider](https://goaisdk.com/docs/reference/providers/custom-provider.md) - Implementing custom providers --- # TranscriptionModel Interface > API reference for the TranscriptionModel interface in the Go AI SDK, defining the methods every speech-to-text model implementation must satisfy. Canonical URL: https://goaisdk.com/docs/reference/providers/transcription-model Documentation index: https://goaisdk.com/llms.txt Interface that all speech-to-text transcription model implementations must satisfy. ## Interface Definition ```go type TranscriptionModel interface { // Metadata SpecificationVersion() string Provider() string ModelID() string // Transcription DoTranscribe(ctx context.Context, opts *TranscriptionOptions) (*types.TranscriptionResult, error) } ``` ## Methods | Method | Parameters | Returns | Description | |--------|------------|---------|-------------| | SpecificationVersion() | - | string | Specification version | | Provider() | - | string | Provider name | | ModelID() | - | string | Model identifier | | DoTranscribe() | ctx, *TranscriptionOptions | *types.TranscriptionResult, error | Transcribe audio to text | ## TranscriptionOptions ```go type TranscriptionOptions struct { Audio []byte MimeType string Language string Timestamps bool } ``` ## Examples ### Using a Transcription Model ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { p := openai.New(openai.Config{ APIKey: "your-api-key", }) model, err := p.TranscriptionModel("whisper-1") if err != nil { log.Fatal(err) } audioData, err := os.ReadFile("audio.mp3") if err != nil { log.Fatal(err) } result, err := model.DoTranscribe(context.Background(), &provider.TranscriptionOptions{ Audio: audioData, MimeType: "audio/mpeg", Language: "en", }) if err != nil { log.Fatal(err) } fmt.Println("Transcription:", result.Text) } ``` ## See Also - [Transcribe](https://goaisdk.com/docs/reference/ai/transcribe.md) - High-level transcription function - [SpeechModel](https://goaisdk.com/docs/reference/providers/speech-model.md) - Text-to-speech interface - [Custom Provider](https://goaisdk.com/docs/reference/providers/custom-provider.md) - Implementing custom providers --- # Provider Registry > API reference for the Go AI SDK's provider registry, which lets you register and access multiple AI providers globally through string IDs. Canonical URL: https://goaisdk.com/docs/reference/registry/provider-registry Documentation index: https://goaisdk.com/llms.txt The provider registry allows you to register and access AI providers globally. ## Registry Functions The package-level functions operate on a shared global registry: ```go // Register a provider in the global registry func RegisterProvider(name string, p provider.Provider) // Get a registered provider from the global registry func GetProvider(name string) (provider.Provider, error) // Resolve a "provider:model" string directly to a model, using the global registry func ResolveLanguageModel(model string) (provider.LanguageModel, error) func ResolveEmbeddingModel(model string) (provider.EmbeddingModel, error) func ResolveImageModel(model string) (provider.ImageModel, error) func ResolveSpeechModel(model string) (provider.SpeechModel, error) func ResolveTranscriptionModel(model string) (provider.TranscriptionModel, error) func ResolveRerankingModel(model string) (provider.RerankingModel, error) // GetGlobalRegistry returns the shared *Registry instance the above functions use func GetGlobalRegistry() *Registry ``` `NoSuchProviderError` is returned when a provider ID can't be resolved or doesn't support the requested model type: ```go type NoSuchProviderError struct { ProviderID string ModelID string ModelType string AvailableProviders []string Reason string } ``` ## Registry Type For an isolated registry (instead of the shared global one), use `NewRegistry`: ```go func NewRegistry(opts ...RegistryOption) *Registry func WithSeparator(separator string) RegistryOption // default ":" func WithLanguageModelMiddleware(mw ...*middleware.LanguageModelMiddleware) RegistryOption func WithImageModelMiddleware(mw ...*middleware.ImageModelMiddleware) RegistryOption ``` `*Registry` exposes the same operations as methods: `RegisterProvider`, `GetProvider`, `ListProviders`, `ResolveLanguageModel`, `ResolveEmbeddingModel`, `ResolveImageModel`, `ResolveSpeechModel`, `ResolveTranscriptionModel`, `ResolveRerankingModel`, plus alias and tool helpers (`RegisterAlias`, `ListAliases`, `RegisterTool`, `LookupTool`, `ListTools`). ## Examples ### Registering Providers ```go package main import ( "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/registry" ) func main() { // Register OpenAI registry.RegisterProvider("openai", openai.New(openai.Config{ APIKey: "openai-key", })) // Register Anthropic registry.RegisterProvider("anthropic", anthropic.New(anthropic.Config{ APIKey: "anthropic-key", })) } ``` ### Getting Providers ```go // Get a provider by name p, err := registry.GetProvider("openai") if err != nil { log.Fatal(err) } // Get a model model, err := p.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } ``` ### Resolving a "provider:model" String Directly ```go // Equivalent to GetProvider("openai") then LanguageModel("gpt-6-astra") model, err := registry.ResolveLanguageModel("openai:gpt-6-astra") if err != nil { log.Fatal(err) } ``` ### Listing Registered Providers ```go providers := registry.GetGlobalRegistry().ListProviders() fmt.Println("Available providers:") for _, name := range providers { fmt.Printf(" - %s\n", name) } ``` ### Checking Provider Existence There's no dedicated `Has`-style check; call `GetProvider` and inspect the error: ```go if _, err := registry.GetProvider("openai"); err == nil { fmt.Println("OpenAI is registered") } ``` ### Using with Configuration ```go type Config struct { ProviderName string ModelID string } func getModel(cfg Config) (provider.LanguageModel, error) { p, err := registry.GetProvider(cfg.ProviderName) if err != nil { return nil, fmt.Errorf("provider not found: %w", err) } model, err := p.LanguageModel(cfg.ModelID) if err != nil { return nil, fmt.Errorf("model not found: %w", err) } return model, nil } // Usage config := Config{ ProviderName: "openai", ModelID: "gpt-6-astra", } model, err := getModel(config) if err != nil { log.Fatal(err) } ``` ### Dynamic Provider Selection ```go func generateWithProvider(providerName, modelID, prompt string) (string, error) { // Get provider from registry p, err := registry.GetProvider(providerName) if err != nil { return "", fmt.Errorf("provider %s not found: %w", providerName, err) } // Get model model, err := p.LanguageModel(modelID) if err != nil { return "", fmt.Errorf("model %s not found: %w", modelID, err) } // Generate result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return "", err } return result.Text, nil } // Usage text, err := generateWithProvider("openai", "gpt-6-astra", "Hello!") ``` ### Multi-Provider Setup ```go func setupProviders(config map[string]string) error { providers := map[string]provider.Provider{ "openai": openai.New(openai.Config{ APIKey: config["openai_key"], }), "anthropic": anthropic.New(anthropic.Config{ APIKey: config["anthropic_key"], }), "google": google.New(google.Config{ APIKey: config["google_key"], }), } for name, p := range providers { registry.RegisterProvider(name, p) } return nil } // Usage config := map[string]string{ "openai_key": os.Getenv("OPENAI_API_KEY"), "anthropic_key": os.Getenv("ANTHROPIC_API_KEY"), "google_key": os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"), } if err := setupProviders(config); err != nil { log.Fatal(err) } ``` ### Provider Fallback ```go func generateWithFallback(modelConfigs []struct{ Provider, Model string }, prompt string) (string, error) { var lastErr error for _, cfg := range modelConfigs { p, err := registry.GetProvider(cfg.Provider) if err != nil { lastErr = err continue } model, err := p.LanguageModel(cfg.Model) if err != nil { lastErr = err continue } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { lastErr = err continue } return result.Text, nil } return "", fmt.Errorf("all providers failed: %w", lastErr) } // Usage with fallback chain configs := []struct{ Provider, Model string }{ {"openai", "gpt-6-astra"}, {"anthropic", "claude-opus-5-5"}, {"google", "gemini-2.5-pro"}, } text, err := generateWithFallback(configs, "Hello!") ``` ## Error Handling ```go _, err := registry.GetProvider("unknown") if err != nil { var notFound *registry.NoSuchProviderError if errors.As(err, ¬Found) { log.Printf("Provider %q is not registered (available: %v)", notFound.ProviderID, notFound.AvailableProviders) } else { log.Fatal("Registry error:", err) } } ``` ## See Also - [Custom Provider](https://goaisdk.com/docs/reference/providers/custom-provider.md) - Implementing custom providers - [LanguageModel Interface](https://goaisdk.com/docs/reference/providers/language-model.md) - Language model interface - [Provider Guide](https://goaisdk.com/docs/providers/overview.md) - Provider overview --- # JSON Schema > API reference for JSON Schema types in the Go AI SDK, covering the Schema and Validator interfaces, creating schemas, constraints, and circular refs. Canonical URL: https://goaisdk.com/docs/reference/schema/json-schema Documentation index: https://goaisdk.com/llms.txt Types and utilities for working with JSON Schema in structured output generation. ## Schema Interface ```go type Schema interface { Validator() Validator } ``` Interface for schema definitions that can be used with structured output. ## Validator Interface ```go type Validator interface { Validate(data interface{}) error JSONSchema() map[string]interface{} } ``` Interface for validating data against a schema. ## Creating Schemas ### JSONSchemaValidator ```go func NewJSONSchema(schema map[string]interface{}) *JSONSchemaValidator ``` Creates a validator from a JSON Schema object. ### SimpleJSONSchema ```go func NewSimpleJSONSchema(schema map[string]interface{}) *SimpleJSONSchema ``` Creates a simple JSON schema implementation. ### StructValidator ```go func NewStructSchema(targetType reflect.Type) *StructValidator ``` Creates a validator from a Go struct type. ## Examples ### Basic JSON Schema ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/schema" ) func main() { personSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{ "type": "string", "description": "Person's full name", }, "age": map[string]interface{}{ "type": "integer", "minimum": 0, "description": "Person's age in years", }, "email": map[string]interface{}{ "type": "string", "format": "email", "description": "Email address", }, }, "required": []string{"name", "age"}, "additionalProperties": false, }) p := openai.New(openai.Config{APIKey: "your-api-key"}) model, err := p.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } // Use with GenerateObject result, err := ai.GenerateObject(context.Background(), ai.GenerateObjectOptions{ Model: model, Prompt: "Generate a person", Schema: personSchema, }) if err != nil { log.Fatal(err) } fmt.Printf("%+v\n", result.Object) } ``` ### Array Schema ```go tasksSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "array", "items": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "id": map[string]interface{}{"type": "integer"}, "description": map[string]interface{}{"type": "string"}, "completed": map[string]interface{}{"type": "boolean"}, }, "required": []string{"id", "description", "completed"}, }, }) ``` ### Nested Schema ```go companySchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "address": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "street": map[string]interface{}{"type": "string"}, "city": map[string]interface{}{"type": "string"}, "country": map[string]interface{}{"type": "string"}, }, "required": []string{"city", "country"}, }, "employees": map[string]interface{}{ "type": "array", "items": map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, "position": map[string]interface{}{"type": "string"}, }, }, }, }, "required": []string{"name"}, }) ``` ### Enum Schema ```go statusSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "string", "enum": []string{"pending", "in_progress", "completed", "cancelled"}, }) ``` ### With Patterns and Constraints ```go userSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "username": map[string]interface{}{ "type": "string", "pattern": "^[a-zA-Z0-9_]{3,20}$", "minLength": 3, "maxLength": 20, }, "password": map[string]interface{}{ "type": "string", "minLength": 8, }, "age": map[string]interface{}{ "type": "integer", "minimum": 13, "maximum": 120, }, }, "required": []string{"username", "password"}, }) ``` ### Using with GenerateObject ```go recipeSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{ "type": "string", "description": "Name of the recipe", }, "ingredients": map[string]interface{}{ "type": "array", "description": "List of ingredients", "items": map[string]interface{}{"type": "string"}, }, "steps": map[string]interface{}{ "type": "array", "description": "Cooking steps", "items": map[string]interface{}{"type": "string"}, }, "cookTime": map[string]interface{}{ "type": "integer", "description": "Cooking time in minutes", }, }, "required": []string{"name", "ingredients", "steps"}, }) result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Prompt: "Generate a recipe for chocolate chip cookies", Schema: recipeSchema, }) if err != nil { log.Fatal(err) } ``` ### Validating Data ```go schema := schema.NewSimpleJSONSchema(personSchema) data := map[string]interface{}{ "name": "John Doe", "age": 30, } if err := schema.Validator().Validate(data); err != nil { log.Printf("Validation failed: %v", err) } ``` ### Getting JSON Schema ```go schema := schema.NewSimpleJSONSchema(personSchema) jsonSchema := schema.Validator().JSONSchema() // Returns the original schema map fmt.Printf("Schema: %+v\n", jsonSchema) ``` ## JSON Schema Types Common JSON Schema property types: - `"string"` - String values - `"integer"` - Integer numbers - `"number"` - Any numbers (including floats) - `"boolean"` - Boolean values - `"object"` - Nested objects - `"array"` - Arrays of items - `"null"` - Null values ## Common Constraints ### String Constraints ```go "minLength": 1, "maxLength": 100, "pattern": "^[A-Z][a-z]+$", "format": "email", // or "uri", "date-time", etc. ``` ### Number Constraints ```go "minimum": 0, "maximum": 100, "multipleOf": 5, ``` ### Array Constraints ```go "minItems": 1, "maxItems": 10, "uniqueItems": true, ``` ### Object Constraints ```go "required": []string{"field1", "field2"}, "additionalProperties": false, "minProperties": 1, "maxProperties": 10, ``` ## Applying Schema Defaults `schema.ApplyDefaults` returns a copy of a value with JSON Schema object defaults applied, walking object properties and array items recursively. The SDK runs this before validation wherever it validates externally-supplied input against a JSON Schema — MCP structured tool output, `ValidateUIMessages`, tool-approval revalidation, agent call options, harness tool `ContextSchema`, and `workflow.ValidateSerializableToolInput` — so a missing field that has a `default` in the schema no longer fails validation. ```go func ApplyDefaults(value interface{}, schema Schema) interface{} ``` ```go withDefaults := schema.ApplyDefaults(input, mySchema) if err := mySchema.Validator().Validate(withDefaults); err != nil { return err } ``` ## Circular $ref Detection A `$ref` cycle that never consumes input — for example two schemas that reference each other only through `allOf`/`anyOf`/`oneOf`/`not` with no concrete `type` in between — now returns a validation error instead of overflowing the stack. This applies to both `Validate` and `ApplyDefaults`. A `not`/`anyOf`/`oneOf` branch that simply fails to match its subschema is unaffected; only a genuine reference cycle is rejected. ## See Also - [GenerateObject](https://goaisdk.com/docs/reference/ai/generate-object.md) - Structured object generation - [StreamObject](https://goaisdk.com/docs/reference/ai/stream-object.md) - Streaming structured output - [Structured Output Guide](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) --- # Error Types > API reference for Go AI SDK error types, covering common errors returned by SDK functions and recommended error-handling patterns for applications. Canonical URL: https://goaisdk.com/docs/reference/types/errors Documentation index: https://goaisdk.com/llms.txt Error types returned by Go AI SDK functions. ## Typed Errors Many SDK functions return specific error types instead of (or in addition to) plain `errors.New`/`fmt.Errorf` values. Use `errors.As` to detect them. Most live in `pkg/ai`; a few tool/provider-level ones live in `pkg/provider/types` and `pkg/provider`. | Type | Package | Key Fields | When it's returned | |------|---------|------------|---------------------| | `NoSuchToolError` | ai | ToolName, AvailableTools | Model called a tool that isn't available | | `InvalidToolInputError` | ai | ToolName, ToolInput, Cause | Tool input doesn't satisfy the tool's schema | | `ToolCallRepairError` | ai | Cause, OriginalError | `RepairToolCall` failed to fix a bad tool call | | `ToolChoiceViolationError` | ai | ToolChoice, FinishReason, Provider, ModelID | Response didn't honor a required/specific `ToolChoice` | | `ToolCallNotFoundForApprovalError` | ai | ToolCallID, ApprovalID | Approval references a tool call missing from history | | `InvalidToolApprovalError` | ai | ApprovalID | Approval response has no matching request | | `InvalidToolApprovalSignatureError` | ai | ApprovalID, ToolCallID, Reason | Signed approval could not be verified | | `NoObjectGeneratedError` | ai | Message, Cause, Text, Response, Usage, FinishReason | `GenerateObject`/`StreamObject` produced no valid object | | `NoOutputGeneratedError` | ai | Message, Cause | Stream ended before any output or finish chunk | | `NoImageGeneratedError` | ai | Responses, Calls | `GenerateImage` produced zero images | | `NoSpeechGeneratedError` | ai | Responses | `GenerateSpeech` produced no audio | | `NoVideoGeneratedError` | ai | Responses | `GenerateVideo` produced no videos | | `NoTranscriptGeneratedError` | ai | Responses | `Transcribe` produced no text | | `NoTranslationGeneratedError` | ai | Response | Speech translation produced no audio or text | | `MessageConversionError` | ai | OriginalMessage, Message, Cause | A UI message couldn't convert to a model message | | `TimeoutError` | ai | Reason, Err | A configured `TimeoutConfig` deadline was hit | | `UIMessageStreamError` | ai | ChunkType, ChunkID, Message | A UI message stream received an out-of-sequence chunk | | `types.MissingToolResultError` | provider/types | ToolCallID, ToolName, Message | A provider-executed tool didn't return a result | | `types.MissingToolResultsError` | provider/types | ToolCallIDs, Message | Multiple provider-executed tools are missing results | | `types.ToolExecutionError` | provider/types | ToolCallID, ToolName, Err, ProviderExecuted | A tool's `Execute` function returned an error | | `provider.BatchError` | provider | Message, Type, Code, StatusCode | Serializable error info for a failed batch item | | `provider.NoSuchUploadAPIError` | provider | API | Provider has no upload API of the requested kind | The SDK also exposes `Is*Error(err) bool` helper functions for most of the types above as an alternative to `errors.As`, for example `ai.IsNoSuchToolError`, `ai.IsInvalidToolInputError`, `ai.IsToolChoiceViolationError`, `ai.IsNoImageGeneratedError`, `ai.IsNoOutputGeneratedError`, and `ai.IsMessageConversionError`. `NoObjectGeneratedError` and `TimeoutError` have no `Is*` helper; use `errors.As` for those. ### Example: errors.As ```go package main import ( "context" "errors" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { p := openai.New(openai.Config{APIKey: "your-api-key"}) model, _ := p.LanguageModel("gpt-6-astra") weatherTool := types.Tool{Name: "get_weather"} result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Use the weather tool", Tools: []types.Tool{weatherTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { var noSuchTool *ai.NoSuchToolError var invalidInput *ai.InvalidToolInputError switch { case errors.As(err, &noSuchTool): log.Printf("model tried calling unknown tool %q", noSuchTool.ToolName) case errors.As(err, &invalidInput): log.Printf("tool %q got invalid input: %v", invalidInput.ToolName, invalidInput.Cause) default: log.Fatal(err) } return } fmt.Println(result.Text) } ``` ## Untyped Errors Many lower-level validation and provider failures are still plain Go errors with descriptive messages rather than a dedicated type, for example: ```text "model is required" "prompt is required" provider-specific HTTP errors wrapped from the provider's API response ``` Check these with `strings.Contains(err.Error(), "...")` as shown below, or prefer the typed errors above and `ai.Is*Error` helpers where one exists. ## Error Handling Patterns ### Basic Error Handling ```go package main import ( "context" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { p := openai.New(openai.Config{APIKey: "your-api-key"}) model, _ := p.LanguageModel("gpt-6-astra") opts := ai.GenerateTextOptions{Model: model, Prompt: "Hello"} result, err := ai.GenerateText(context.Background(), opts) if err != nil { log.Fatal(err) } log.Println(result.Text) } ``` ### Timeout Errors ```go result, err := ai.GenerateText(ctx, opts) if err != nil { if errors.Is(err, context.DeadlineExceeded) { log.Println("Request timed out") return } log.Fatal(err) } ``` ### Cancellation Errors ```go ctx, cancel := context.WithCancel(context.Background()) go func() { time.Sleep(5 * time.Second) cancel() }() result, err := ai.GenerateText(ctx, opts) if err != nil { if errors.Is(err, context.Canceled) { log.Println("Request was canceled") return } log.Fatal(err) } ``` ### Validation Errors ```go result, err := ai.GenerateText(ctx, opts) if err != nil { switch { case strings.Contains(err.Error(), "model is required"): log.Println("Must provide a model") case strings.Contains(err.Error(), "prompt is required"): log.Println("Must provide a prompt or messages") default: log.Println("Unknown validation error:", err) } return } ``` ### Provider-Specific Errors ```go result, err := ai.GenerateText(ctx, opts) if err != nil { switch { case strings.Contains(err.Error(), "API key"): log.Println("Invalid or missing API key") case strings.Contains(err.Error(), "rate limit"): log.Println("Rate limited - retry later") time.Sleep(time.Minute) // Retry logic case strings.Contains(err.Error(), "content policy"): log.Println("Content policy violation") case strings.Contains(err.Error(), "invalid model"): log.Println("Model not found or not supported") default: log.Println("Provider error:", err) } return } ``` ### Retry Logic ```go func generateWithRetry(ctx context.Context, opts ai.GenerateTextOptions, maxRetries int) (*ai.GenerateTextResult, error) { var result *ai.GenerateTextResult var err error for attempt := 0; attempt < maxRetries; attempt++ { result, err = ai.GenerateText(ctx, opts) if err == nil { return result, nil } // Check if error is retryable if strings.Contains(err.Error(), "rate limit") { waitTime := time.Duration(attempt+1) * time.Second log.Printf("Rate limited, retrying in %v (attempt %d/%d)", waitTime, attempt+1, maxRetries) time.Sleep(waitTime) continue } // Non-retryable error return nil, err } return nil, fmt.Errorf("max retries exceeded: %w", err) } ``` ### Tool Execution Errors ```go weatherTool := types.Tool{ Name: "get_weather", Description: "Get weather", Parameters: weatherSchema, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location, ok := input["location"].(string) if !ok { return nil, fmt.Errorf("location must be a string") } weather, err := fetchWeather(location) if err != nil { return nil, fmt.Errorf("failed to fetch weather: %w", err) } return weather, nil }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather?", Tools: []types.Tool{weatherTool}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { if strings.Contains(err.Error(), "tool execution failed") { log.Println("A tool failed during execution:", err) } } ``` ### Stream Errors ```go result, err := ai.StreamText(ctx, opts) if err != nil { log.Fatal("Failed to start stream:", err) } defer result.Close() for { chunk, err := result.Stream().Next() if err == io.EOF { break } if err != nil { log.Printf("Stream error: %v", err) break } // Process chunk } // Check for accumulated errors if result.Err() != nil { log.Printf("Stream had errors: %v", result.Err()) } ``` ### Schema Validation Errors ```go result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Prompt: "Generate user data", Schema: schema, }) if err != nil { if strings.Contains(err.Error(), "validation failed") { log.Println("Generated object doesn't match schema:", err) } return } ``` ## See Also - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - Text generation - [Usage Types](https://goaisdk.com/docs/reference/types/usage.md) - Token usage tracking - [Tool Types](https://goaisdk.com/docs/reference/types/tools.md) - Tool-related types --- # Message Types > API reference for Go AI SDK message types, covering MessageRole, Message, and ContentPart implementations like TextContent, ImageContent, and FileContent. Canonical URL: https://goaisdk.com/docs/reference/types/messages Documentation index: https://goaisdk.com/llms.txt Types for representing messages in conversations between users, assistants, and tools. ## MessageRole ```go type MessageRole string const ( RoleSystem MessageRole = "system" RoleUser MessageRole = "user" RoleAssistant MessageRole = "assistant" RoleTool MessageRole = "tool" ) ``` Message roles define who is speaking in a conversation: - `RoleSystem`: System instructions that guide model behavior - `RoleUser`: User input or queries - `RoleAssistant`: Model-generated responses - `RoleTool`: Tool execution results ## Message ```go type Message struct { Role MessageRole `json:"role"` Content []ContentPart `json:"content"` Name string `json:"name,omitempty"` } ``` Represents a single message in a conversation. ### Fields | Field | Type | Description | |-------|------|-------------| | Role | MessageRole | Role of the message sender | | Content | []ContentPart | Content parts (text, images, tool results) | | Name | string | Optional sender name | ## ContentPart Interface ```go type ContentPart interface { ContentType() string } ``` Interface for different types of message content. ## TextContent ```go type TextContent struct { Text string `json:"text"` } ``` Text content in a message. ## ReasoningContent ```go type ReasoningContent struct { Text string `json:"text"` } ``` Reasoning or thinking content (for models that expose reasoning, like OpenAI o1). ## ImageContent ```go type ImageContent struct { Image []byte `json:"image"` MimeType string `json:"mimeType"` URL string `json:"url,omitempty"` } ``` Image content in a message. ## FileContent ```go type FileContent struct { FileData FileData `json:"fileData,omitempty"` Data []byte `json:"data,omitempty"` MimeType string `json:"mimeType,omitempty"` MediaType string `json:"mediaType,omitempty"` Filename string `json:"filename,omitempty"` URL string `json:"url,omitempty"` Reference string `json:"reference,omitempty"` Text string `json:"text,omitempty"` ProviderOptions map[string]interface{} `json:"providerOptions,omitempty"` ProviderMetadata json.RawMessage `json:"providerMetadata,omitempty"` } ``` File attachment in a message. ## GeneratedFileContent ```go type GeneratedFileContent struct { MediaType string `json:"mediaType"` FileData FileData `json:"fileData,omitempty"` Data []byte `json:"data,omitempty"` URL string `json:"url,omitempty"` ProviderOptions map[string]interface{} `json:"providerOptions,omitempty"` ProviderMetadata json.RawMessage `json:"providerMetadata,omitempty"` } ``` Generated file output from providers that return files or file URLs. Custom JSON encoding follows the AI SDK file-data shape: the emitted `data` field is a tagged object with `type: "data"` and base64 `data`, or `type: "url"` and `url`. ## ToolCallContent ```go type ToolCallContent struct { ToolCallID string `json:"toolCallId"` ToolName string `json:"toolName"` Title string `json:"title,omitempty"` Input string `json:"input,omitempty"` Arguments map[string]interface{} `json:"arguments,omitempty"` ProviderExecuted bool `json:"providerExecuted,omitempty"` ProviderOptions map[string]interface{} `json:"providerOptions,omitempty"` ProviderMetadata json.RawMessage `json:"providerMetadata,omitempty"` ToolMetadata map[string]interface{} `json:"toolMetadata,omitempty"` Dynamic bool `json:"dynamic,omitempty"` Invalid bool `json:"invalid,omitempty"` Error interface{} `json:"error,omitempty"` ThoughtSignature string `json:"thoughtSignature,omitempty"` } ``` Assistant tool-call content preserves ordered model output, including provider-executed, dynamic, invalid, and tool-metadata fields plus Google `thoughtSignature` metadata for round trips. `Input` stores the provider's raw streamed JSON when available, while `Arguments` stores the decoded Go map used by provider converters and tool execution. JSON serialization follows the TypeScript SDK public shape and emits `input`; it does not emit the Go-only `arguments` helper field. ## ToolResultContent ```go type ToolResultContent struct { ToolCallID string `json:"toolCallId"` ToolName string `json:"toolName"` Title string `json:"title,omitempty"` Input map[string]interface{} `json:"input,omitempty"` Result interface{} `json:"result,omitempty"` // deprecated; use Output Error string `json:"error,omitempty"` Output *ToolResultOutput `json:"output,omitempty"` ProviderExecuted bool `json:"providerExecuted,omitempty"` ProviderOptions map[string]interface{} `json:"providerOptions,omitempty"` ProviderMetadata json.RawMessage `json:"providerMetadata,omitempty"` ToolMetadata map[string]interface{} `json:"toolMetadata,omitempty"` Dynamic bool `json:"dynamic,omitempty"` Preliminary bool `json:"preliminary,omitempty"` } ``` Tool execution result content mirrors the TypeScript SDK `tool-result` part. Output content carries `ProviderMetadata`; when the part is replayed to a provider as a response message, that metadata is converted to provider-facing `ProviderOptions`. ## ToolErrorContent ```go type ToolErrorContent struct { ToolCallID string `json:"toolCallId"` ToolName string `json:"toolName"` Title string `json:"title,omitempty"` Input map[string]interface{} `json:"input,omitempty"` Error interface{} `json:"error"` ProviderExecuted bool `json:"providerExecuted,omitempty"` ProviderOptions map[string]interface{} `json:"providerOptions,omitempty"` ProviderMetadata json.RawMessage `json:"providerMetadata,omitempty"` ToolMetadata map[string]interface{} `json:"toolMetadata,omitempty"` Dynamic bool `json:"dynamic,omitempty"` } ``` Tool execution error content mirrors the TypeScript SDK `tool-error` part and follows the same `ProviderMetadata` to `ProviderOptions` replay behavior. ## ToolApprovalRequestContent ```go type ToolApprovalRequestContent struct { ApprovalID string `json:"approvalId"` ToolCallID string `json:"toolCallId"` // provider replay shape ToolCall ToolCall `json:"toolCall,omitempty"` // public result content shape IsAutomatic bool `json:"isAutomatic,omitempty"` } ``` Tool approval request content mirrors the TypeScript SDK split between public result content and replayed provider messages. In `GenerateTextResult.Content` and `StreamTextResult.Content()`, JSON contains `approvalId`, nested `toolCall`, and optional `isAutomatic`. In response messages replayed to a provider, JSON contains `approvalId`, `toolCallId`, and optional `isAutomatic`. ## ToolApprovalResponseContent ```go type ToolApprovalResponseContent struct { ApprovalID string `json:"approvalId"` ToolCallID string `json:"toolCallId,omitempty"` // internal lookup field ToolCall ToolCall `json:"toolCall,omitempty"` // public result content shape Approved bool `json:"approved"` Reason string `json:"reason,omitempty"` ProviderExecuted bool `json:"providerExecuted,omitempty"` } ``` Tool approval response content contains `approvalId`, nested `toolCall`, `approved`, optional `reason`, and optional `providerExecuted` in public result content. Provider replay strips the nested tool call and sends `approvalId`, `approved`, `reason`, and `providerExecuted`, matching the TypeScript SDK `toResponseMessages` behavior. ## Prompt ```go type Prompt struct { Messages []Message System string Text string } ``` Represents a prompt that can be either a simple string or a list of messages. ## Examples ### User Message ```go package main import ( "fmt" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func main() { message := types.Message{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Hello, how are you?"}, }, } fmt.Println(message.Role) } ``` ### Assistant Message ```go message := types.Message{ Role: types.RoleAssistant, Content: []types.ContentPart{ types.TextContent{Text: "I'm doing well, thank you!"}, }, } ``` ### Multi-turn Conversation ```go messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What is the capital of France?"}, }, }, { Role: types.RoleAssistant, Content: []types.ContentPart{ types.TextContent{Text: "The capital of France is Paris."}, }, }, { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What is its population?"}, }, }, } ``` ### Message with Image ```go imageData, _ := os.ReadFile("image.jpg") message := types.Message{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What's in this image?"}, types.ImageContent{ Image: imageData, MimeType: "image/jpeg", }, }, } ``` ### Tool Result Message ```go message := types.Message{ Role: types.RoleTool, Content: []types.ContentPart{ types.ToolResultContent{ ToolCallID: "call_abc123", ToolName: "get_weather", Result: map[string]interface{}{ "temperature": 72, "condition": "sunny", }, }, }, } ``` ### Assistant Tool Call Message ```go message := types.Message{ Role: types.RoleAssistant, Content: []types.ContentPart{ types.ToolCallContent{ ToolCallID: "call_abc123", ToolName: "get_weather", Input: `{"location":"Tokyo"}`, Arguments: map[string]interface{}{"location": "Tokyo"}, }, }, } ``` ### System Message ```go message := types.Message{ Role: types.RoleSystem, Content: []types.ContentPart{ types.TextContent{ Text: "You are a helpful assistant that answers questions concisely.", }, }, } ``` ## See Also - [Tool Types](https://goaisdk.com/docs/reference/types/tools.md) - Tool-related types - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - Using messages for generation --- # Tool Types > API reference for Go AI SDK tool-related types, covering Tool, ToolCall, ToolResult, ToolChoice, and the ToolExecutor function signature. Canonical URL: https://goaisdk.com/docs/reference/types/tools Documentation index: https://goaisdk.com/llms.txt Types related to tool definitions, tool calls, and tool execution. ## Tool See [Tool API Reference](https://goaisdk.com/docs/reference/ai/tool.md) for complete documentation. ## ToolCall ```go type ToolCall struct { ID string `json:"id"` ToolName string `json:"toolName"` Arguments map[string]interface{} `json:"arguments"` } ``` Represents a tool call made by the model. ### Fields | Field | Type | Description | |-------|------|-------------| | ID | string | Unique identifier for this tool call | | ToolName | string | Name of the tool to call | | Arguments | map[string]interface{} | Arguments to pass to the tool | ## ToolResult ```go type ToolResult struct { ToolCallID string `json:"toolCallId"` ToolName string `json:"toolName"` Result interface{} `json:"result"` Error error `json:"error,omitempty"` ProviderExecuted bool `json:"providerExecuted,omitempty"` } ``` Result of executing a tool. ### Fields | Field | Type | Description | |-------|------|-------------| | ToolCallID | string | ID of the tool call this result corresponds to | | ToolName | string | Name of the tool that was executed | | Result | interface{} | Result of the tool execution | | Error | error | Error if execution failed | | ProviderExecuted | bool | Whether provider executed this tool | ## ToolChoice ```go type ToolChoice struct { Type ToolChoiceType `json:"type"` ToolName string `json:"toolName,omitempty"` } ``` Specifies how the model should choose tools. ### Fields | Field | Type | Description | |-------|------|-------------| | Type | ToolChoiceType | Type of tool choice | | ToolName | string | Specific tool name (for ToolChoiceTool type) | ## ToolChoiceType ```go type ToolChoiceType string const ( ToolChoiceAuto ToolChoiceType = "auto" ToolChoiceNone ToolChoiceType = "none" ToolChoiceRequired ToolChoiceType = "required" ToolChoiceTool ToolChoiceType = "tool" ) ``` Types of tool choice strategies: - `ToolChoiceAuto`: Model decides whether to use tools - `ToolChoiceNone`: Model cannot use tools - `ToolChoiceRequired`: Model must use at least one tool - `ToolChoiceTool`: Model must use a specific tool ## ToolExecutor ```go type ToolExecutor func( ctx context.Context, input map[string]interface{}, options ToolExecutionOptions, ) (interface{}, error) ``` Function signature for tool execution. ## ToolExecutionOptions ```go type ToolExecutionOptions struct { ToolCallID string UserContext interface{} Usage *Usage Metadata map[string]interface{} } ``` Options passed to tool execution. ### Fields | Field | Type | Description | |-------|------|-------------| | ToolCallID | string | Unique ID of this tool call | | UserContext | interface{} | User-defined context | | Usage | *Usage | Current token usage | | Metadata | map[string]interface{} | Additional metadata | ## Examples ### Create Tool Call ```go package main import ( "fmt" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func main() { toolCall := types.ToolCall{ ID: "call_abc123", ToolName: "get_weather", Arguments: map[string]interface{}{ "location": "Tokyo", "units": "celsius", }, } fmt.Println(toolCall.ToolName) } ``` ### Auto Tool Choice ```go toolChoice := types.AutoToolChoice() // Equivalent to: // types.ToolChoice{Type: types.ToolChoiceAuto} ``` ### Required Tool Choice ```go toolChoice := types.RequiredToolChoice() // Model must call at least one tool ``` ### Specific Tool Choice ```go toolChoice := types.SpecificToolChoice("get_weather") // Model must call the "get_weather" tool ``` ### Handle Tool Result ```go toolResult := types.ToolResult{ ToolCallID: "call_abc123", ToolName: "get_weather", Result: map[string]interface{}{ "temperature": 22, "condition": "sunny", }, } if toolResult.Error != nil { log.Printf("Tool failed: %v", toolResult.Error) } ``` ## See Also - [Tool](https://goaisdk.com/docs/reference/ai/tool.md) - Tool definition - [Message Types](https://goaisdk.com/docs/reference/types/messages.md) - Message types - [Tool Calling Guide](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) --- # Usage Types > API reference for Go AI SDK usage types that track token counts and costs across API calls, including input and output token detail breakdowns. Canonical URL: https://goaisdk.com/docs/reference/types/usage Documentation index: https://goaisdk.com/llms.txt Types for tracking token usage and costs across API calls with detailed breakdowns for multimodal inputs, caching, and reasoning. ## Usage ```go type Usage struct { InputTokens *int64 // Total input tokens (prompt) OutputTokens *int64 // Total output tokens (completion) TotalTokens *int64 // Total tokens (input + output) InputDetails *InputTokenDetails // Detailed input token breakdown OutputDetails *OutputTokenDetails // Detailed output token breakdown Raw map[string]interface{} // Provider-specific raw usage data } ``` Token usage information for language model calls with v6.0 detailed tracking. ### Fields | Field | Type | Description | |-------|------|-------------| | InputTokens | *int64 | Total tokens in the input (prompt) | | OutputTokens | *int64 | Total tokens in the output (completion) | | TotalTokens | *int64 | Total tokens (input + output) | | InputDetails | *InputTokenDetails | Detailed breakdown of input tokens (cache, text/image) | | OutputDetails | *OutputTokenDetails | Detailed breakdown of output tokens (text/reasoning) | | Raw | map[string]interface{} | Provider-specific usage data | ### Methods ```go func (u Usage) Add(other Usage) Usage ``` Adds two usage objects together, merging all detailed breakdowns. ```go func (u Usage) GetInputTokens() int64 func (u Usage) GetOutputTokens() int64 func (u Usage) GetTotalTokens() int64 ``` Helper methods that return 0 if the pointer is nil. ## InputTokenDetails ```go type InputTokenDetails struct { NoCacheTokens *int64 // Tokens not served from cache CacheReadTokens *int64 // Tokens read from cache CacheWriteTokens *int64 // Tokens written to cache TextTokens *int64 // Text input tokens (for multimodal) ImageTokens *int64 // Image input tokens (for multimodal) } ``` Detailed breakdown of input token usage, particularly useful for: - Tracking prompt caching benefits (cost savings) - Differentiating text vs image costs in multimodal models ### Fields | Field | Type | Description | |-------|------|-------------| | NoCacheTokens | *int64 | Input tokens processed normally (not from cache) | | CacheReadTokens | *int64 | Input tokens read from cache (cheaper) | | CacheWriteTokens | *int64 | Input tokens written to cache for future use | | TextTokens | *int64 | Text input tokens (for multimodal cost tracking) | | ImageTokens | *int64 | Image input tokens (for multimodal cost tracking) | ## OutputTokenDetails ```go type OutputTokenDetails struct { TextTokens *int64 // Tokens for regular text generation ReasoningTokens *int64 // Tokens for internal reasoning (o1/o3) } ``` Detailed breakdown of output token usage for models with reasoning capabilities. ### Fields | Field | Type | Description | |-------|------|-------------| | TextTokens | *int64 | Tokens used for regular text generation | | ReasoningTokens | *int64 | Tokens used for internal reasoning (not shown in output) | ## Other Usage Types ### EmbeddingUsage ```go type EmbeddingUsage struct { InputTokens int TotalTokens int } ``` ### ImageUsage ```go type ImageUsage struct { ImageCount int } ``` ### TranscriptionUsage ```go type TranscriptionUsage struct { DurationSeconds float64 } ``` ### VideoUsage ```go type VideoUsage struct { VideoCount int TotalDurationSeconds float64 } ``` ## Examples ### Basic Usage Tracking ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { provider := openai.New(openai.Config{APIKey: "your-api-key"}) model, err := provider.LanguageModel("gpt-6-astra") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Hello, world!", }) if err != nil { log.Fatal(err) } usage := result.Usage fmt.Printf("Input tokens: %d\n", usage.GetInputTokens()) fmt.Printf("Output tokens: %d\n", usage.GetOutputTokens()) fmt.Printf("Total tokens: %d\n", usage.GetTotalTokens()) } ``` ### Track Text vs Image Tokens (Multimodal) ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Describe this image:"}, types.ImageContent{URL: imageURL}, }, }, }, }) if err != nil { log.Fatal(err) } // Check if provider supports text/image breakdown if result.Usage.InputDetails != nil { if result.Usage.InputDetails.TextTokens != nil { fmt.Printf("Text input tokens: %d\n", *result.Usage.InputDetails.TextTokens) } if result.Usage.InputDetails.ImageTokens != nil { fmt.Printf("Image input tokens: %d\n", *result.Usage.InputDetails.ImageTokens) } } ``` ### Estimate Cost with Multimodal Pricing ```go // Example pricing const ( textInputCost = 0.0000025 // $2.50 per 1M tokens imageInputCost = 0.0000075 // $7.50 per 1M tokens (approx) outputCost = 0.0000100 // $10.00 per 1M tokens ) usage := result.Usage var cost float64 // Calculate input cost (text + image) if usage.InputDetails != nil && usage.InputDetails.TextTokens != nil && usage.InputDetails.ImageTokens != nil { // Accurate multimodal cost cost += float64(*usage.InputDetails.TextTokens) * textInputCost cost += float64(*usage.InputDetails.ImageTokens) * imageInputCost } else { // Fallback to average rate if breakdown not available cost += float64(usage.GetInputTokens()) * textInputCost } // Add output cost cost += float64(usage.GetOutputTokens()) * outputCost fmt.Printf("Estimated cost: $%.6f\n", cost) ``` ### Track Cached Tokens ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: longPrompt, }) if err != nil { log.Fatal(err) } if result.Usage.InputDetails != nil && result.Usage.InputDetails.CacheReadTokens != nil { cacheRead := *result.Usage.InputDetails.CacheReadTokens if cacheRead > 0 { total := result.Usage.GetInputTokens() savings := float64(cacheRead) / float64(total) * 100 fmt.Printf("Cache hit! %d tokens from cache (%.1f%% savings)\n", cacheRead, savings) } } ``` ### Track Reasoning Tokens (o1/o3 Models) ```go model, _ := openaiProvider.LanguageModel("o3-mini") // OpenAI reasoning model result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Solve this complex problem...", }) if err != nil { log.Fatal(err) } if result.Usage.OutputDetails != nil && result.Usage.OutputDetails.ReasoningTokens != nil { reasoning := *result.Usage.OutputDetails.ReasoningTokens text := *result.Usage.OutputDetails.TextTokens fmt.Printf("Text tokens: %d\n", text) fmt.Printf("Reasoning tokens: %d (hidden)\n", reasoning) fmt.Printf("Reasoning overhead: %.1f%%\n", float64(reasoning)/float64(text)*100) } ``` ### Accumulate Usage Across Calls ```go var totalUsage types.Usage for i := 0; i < 5; i++ { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Question %d", i), }) if err != nil { log.Printf("Error on call %d: %v", i, err) continue } totalUsage = totalUsage.Add(result.Usage) } fmt.Printf("Total usage across all calls:\n") fmt.Printf(" Input tokens: %d\n", totalUsage.GetInputTokens()) fmt.Printf(" Output tokens: %d\n", totalUsage.GetOutputTokens()) fmt.Printf(" Total tokens: %d\n", totalUsage.GetTotalTokens()) ``` ### Provider-Specific Usage Data ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello", }) if err != nil { log.Fatal(err) } // Access raw provider-specific data if result.Usage.Raw != nil { if promptDetails, ok := result.Usage.Raw["prompt_tokens_details"]; ok { fmt.Printf("Provider details: %+v\n", promptDetails) } } ``` ### Usage Budget Enforcement ```go const maxTokens = 10000 var usedTokens int64 for _, task := range tasks { if usedTokens >= maxTokens { fmt.Println("Token budget exceeded") break } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: task, }) if err != nil { log.Printf("Error: %v", err) continue } usedTokens += result.Usage.GetTotalTokens() fmt.Printf("Used %d/%d tokens\n", usedTokens, maxTokens) } ``` ## Provider Support Different providers support different levels of usage detail: | Provider | Text/Image Tokens | Cache Tracking | Reasoning Tokens | |----------|-------------------|----------------|------------------| | OpenAI | ✅ Yes | ✅ Yes | ✅ Yes (o1/o3) | | XAI | ✅ Yes | ✅ Yes | ✅ Yes | | Google | ✅ Yes | ✅ Yes | ✅ Yes (thinking)| | Anthropic| ❌ No | ✅ Yes | ❌ No | | Azure | ✅ Yes | ✅ Yes | ✅ Yes | | Groq | ⚠️ Partial | ⚠️ Partial | ⚠️ Partial | | DeepSeek | ✅ Yes | ✅ Yes | ✅ Yes | | Together | ✅ Yes | ✅ Yes | ✅ Yes | | Fireworks| ✅ Yes | ✅ Yes | ✅ Yes | | Others | Varies | Varies | Varies | When detailed fields are not available from a provider, the pointers will be `nil`. ## See Also - [GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md) - Text generation with usage tracking - [Embed](https://goaisdk.com/docs/reference/ai/embed.md) - Embedding generation with usage tracking - [Error Types](https://goaisdk.com/docs/reference/types/errors.md) - Error handling - [Migration Guide](https://goaisdk.com/docs/migration-guides/token-usage-differentiation.md) - Migrating to new usage structure --- # Workflow > Reference for pkg/workflow in the Go AI SDK: WorkflowAgent, its options and results, the chat transports, serializable tools, step iteration and the durable harness runner. Canonical URL: https://goaisdk.com/docs/reference/workflow Documentation index: https://goaisdk.com/llms.txt Package `github.com/digitallysavvy/go-ai/pkg/workflow` is the Go port of the TypeScript `@ai-sdk/workflow` package. It provides an agent for durable execution, where one run spans several process executions, and the client and server pieces that stream a run to a chat UI. For a walkthrough of `WorkflowAgent`, see [WorkflowAgent](https://goaisdk.com/docs/agents/workflow-agent.md). ## WorkflowAgent `workflow.NewWorkflowAgent` validates a `WorkflowAgent` value and returns it. ```go func NewWorkflowAgent(agent WorkflowAgent) (*WorkflowAgent, error) ``` {/* gen:fields workflow.WorkflowAgent */} | Field | Type | Description | | --- | --- | --- | | `Model` | `provider.LanguageModel` | | | `System` | `string` | | | `Instructions` | `interface{}` | Instructions is the TypeScript-compatible name for system instructions. | | `AllowSystemInMessages` | `bool` | AllowSystemInMessages permits system-role messages in Messages. By default WorkflowAgent rejects system messages; use Instructions/System for default system prompts. | | `Tools` | `[]types.Tool` | | | `ToolSet` | `map[string]types.Tool` | | | `StopWhen` | `[]ai.StopCondition` | | | `Output` | `interface{}` | | | `Telemetry` | `*ai.TelemetrySettings` | | | `ID` | `string` | | | `Prompt` | `string` | | | `OnStart` | `StartCallback` | | | `OnStepStart` | `StepStartCallback` | | | `OnToolExecutionStart` | `ToolExecutionStartCallback` | | | `OnToolExecutionEnd` | `ToolExecutionEndCallback` | | | `OnStepEnd` | `StepEndCallback` | | | `OnStepFinish` | `StepFinishCallback` | Deprecated: use OnStepEnd. | | `OnEnd` | `EndCallback` | | | `OnFinish` | `FinishCallback` | Deprecated: use OnEnd. | | `OnError` | `ErrorCallback` | | | `OnAbort` | `AbortCallback` | | | `PrepareCall` | `PrepareCallHook` | | | `PrepareStep` | `PrepareStepHook` | | | `FilterActiveTools` | `FilterActiveToolsHook` | | | `CallOptionsSchema` | `schema.Schema` | | | `CallOptions` | `interface{}` | | | `ActiveTools` | `[]string` | | | `Temperature` | `*float64` | | | `MaxTokens` | `*int` | | | `TopP` | `*float64` | | | `TopK` | `*int` | | | `FrequencyPenalty` | `*float64` | | | `PresencePenalty` | `*float64` | | | `StopSequences` | `[]string` | | | `Seed` | `*int` | | | `Headers` | `map[string]string` | | | `Reasoning` | `*types.ReasoningLevel` | | | `SendReasoning` | `*bool` | | | `ProviderOptions` | `map[string]interface{}` | | | `RuntimeContext` | `interface{}` | | | `ToolsContext` | `map[string]interface{}` | | | `ToolChoice` | `types.ToolChoice` | | | `Include` | `*ai.IncludeOptions` | | | `ExperimentalSandbox` | `interface{}` | | | `ExperimentalRefineToolInput` | `map[string]ai.ToolInputRefiner` | | | `RepairToolCall` | `ai.ToolCallRepairFunction` | RepairToolCall attempts to repair tool calls that fail to parse. | | `ExperimentalRepairToolCall` | `ai.ToolCallRepairFunction` | ExperimentalRepairToolCall is a deprecated alias for RepairToolCall. Deprecated: use RepairToolCall. | | `ExperimentalToolApprovalSecret` | `[]byte` | ExperimentalToolApprovalSecret signs issued approval requests and verifies resumed approvals before approved tools execute. | | `MaxRetries` | `*int` | MaxRetries controls transient provider call retries for each model call, forwarded to agent.AgentConfig.MaxRetries. Defaults to 2 when nil, matching TS WorkflowAgent's `mergedGenerationSettings.maxRetries ?? 2`. | | `Timeout` | `*ai.TimeoutConfig` | Timeout provides granular timeout controls, forwarded to agent.AgentConfig.Timeout. | | `ExperimentalDownload` | `ai.DownloadFunction` | ExperimentalDownload customizes remote file URL downloads before model calls, forwarded to agent.AgentConfig.ExperimentalDownload. | {/* /gen:fields */} | Method | Description | | --- | --- | | `Generate(ctx, prompt string, opts *agent.AgentGenerateOptions) (*WorkflowResult, error)` | Runs the agent and returns the final result. | | `GenerateWithOptions(ctx, WorkflowGenerateOptions) (*WorkflowResult, error)` | Runs with workflow-native call options. | | `Stream(ctx, prompt string, opts *agent.AgentStreamOptions) (*WorkflowStreamResult, error)` | Runs in streaming mode. | | `StreamWithOptions(ctx, WorkflowStreamOptions) (*WorkflowStreamResult, error)` | Streams with workflow-native stream options. | ### WorkflowGenerateOptions {/* gen:fields workflow.WorkflowGenerateOptions */} | Field | Type | Description | | --- | --- | --- | | `Prompt` | `string` | | | `Messages` | `[]types.Message` | | | `System` | `string` | | | `Instructions` | `interface{}` | | | `AllowSystemInMessages` | `bool` | | | `Tools` | `[]types.Tool` | | | `ToolSet` | `map[string]types.Tool` | | | `StopWhen` | `[]ai.StopCondition` | | | `Telemetry` | `*ai.TelemetrySettings` | | | `RuntimeContext` | `interface{}` | | | `ToolsContext` | `map[string]interface{}` | | | `Include` | `*ai.IncludeOptions` | | | `ExperimentalSandbox` | `interface{}` | | | `ExperimentalRefineToolInput` | `map[string]ai.ToolInputRefiner` | | | `RepairToolCall` | `ai.ToolCallRepairFunction` | RepairToolCall attempts to repair tool calls that fail to parse. | | `ExperimentalRepairToolCall` | `ai.ToolCallRepairFunction` | ExperimentalRepairToolCall is a deprecated alias for RepairToolCall. Deprecated: use RepairToolCall. | | `ExperimentalToolApprovalSecret` | `[]byte` | ExperimentalToolApprovalSecret overrides the agent's approval secret. | | `MaxRetries` | `*int` | MaxRetries overrides the agent's MaxRetries for this call. | | `Timeout` | `*ai.TimeoutConfig` | Timeout overrides the agent's Timeout for this call. | | `ExperimentalDownload` | `ai.DownloadFunction` | ExperimentalDownload overrides the agent's ExperimentalDownload for this call. | | `OnStart` | `StartCallback` | | | `OnStepStart` | `StepStartCallback` | | | `OnToolExecutionStart` | `ToolExecutionStartCallback` | | | `OnToolExecutionEnd` | `ToolExecutionEndCallback` | | | `OnStepEnd` | `StepEndCallback` | | | `OnStepFinish` | `StepFinishCallback` | Deprecated: use OnStepEnd. | | `OnEnd` | `EndCallback` | | | `OnFinish` | `FinishCallback` | Deprecated: use OnEnd. | | `OnError` | `ErrorCallback` | | | `OnAbort` | `AbortCallback` | | {/* /gen:fields */} ### WorkflowStreamOptions {/* gen:fields workflow.WorkflowStreamOptions */} | Field | Type | Description | | --- | --- | --- | | `Prompt` | `string` | | | `Messages` | `[]types.Message` | | | `System` | `string` | | | `Instructions` | `interface{}` | | | `AllowSystemInMessages` | `bool` | | | `Tools` | `[]types.Tool` | | | `ToolSet` | `map[string]types.Tool` | | | `StopWhen` | `[]ai.StopCondition` | | | `Telemetry` | `*ai.TelemetrySettings` | | | `ActiveTools` | `[]string` | | | `RuntimeContext` | `interface{}` | | | `ToolsContext` | `map[string]interface{}` | | | `Include` | `*ai.IncludeOptions` | | | `ExperimentalSandbox` | `interface{}` | | | `ExperimentalRefineToolInput` | `map[string]ai.ToolInputRefiner` | | | `RepairToolCall` | `ai.ToolCallRepairFunction` | RepairToolCall attempts to repair tool calls that fail to parse. | | `ExperimentalRepairToolCall` | `ai.ToolCallRepairFunction` | ExperimentalRepairToolCall is a deprecated alias for RepairToolCall. Deprecated: use RepairToolCall. | | `ExperimentalToolApprovalSecret` | `[]byte` | ExperimentalToolApprovalSecret overrides the agent's approval secret. | | `MaxRetries` | `*int` | MaxRetries overrides the agent's MaxRetries for this call. | | `Timeout` | `*ai.TimeoutConfig` | Timeout overrides the agent's Timeout for this call. | | `ExperimentalDownload` | `ai.DownloadFunction` | ExperimentalDownload overrides the agent's ExperimentalDownload for this call. | | `ExperimentalTransform` | `[]ai.StreamTransformFunc` | ExperimentalTransform is an ordered list of transforms applied to raw model stream chunks before they reach OnChunk or the returned WorkflowStreamResult, matching TS WorkflowAgent's `experimental_transform` stream option (hash 165455d). Previously declared-but-unused in TS; this forwards it through to the underlying ai.StreamText call so it actually runs. | | `OnChunk` | `func(chunk provider.StreamChunk)` | | | `OnStart` | `StartCallback` | | | `OnStepStart` | `StepStartCallback` | | | `OnToolExecutionStart` | `ToolExecutionStartCallback` | | | `OnToolExecutionEnd` | `ToolExecutionEndCallback` | | | `OnStepEnd` | `StepEndCallback` | | | `OnStepFinish` | `StepFinishCallback` | Deprecated: use OnStepEnd. | | `OnEnd` | `EndCallback` | | | `OnFinish` | `FinishCallback` | Deprecated: use OnEnd. | | `OnError` | `ErrorCallback` | | | `OnAbort` | `AbortCallback` | | {/* /gen:fields */} ### Results `workflow.WorkflowResult` is the final non-streaming result. `IsLoopFinished()` reports whether the run ended naturally. {/* gen:fields workflow.WorkflowResult */} | Field | Type | Description | | --- | --- | --- | | (embedded) | `*agent.AgentResult` | Embedded. | {/* /gen:fields */} `workflow.WorkflowStreamResult` wraps the streaming result. {/* gen:fields workflow.WorkflowStreamResult */} | Field | Type | Description | | --- | --- | --- | | (embedded) | `*ai.StreamTextResult` | Embedded. | {/* /gen:fields */} ### Hooks and callbacks The hook and callback function types mirror the `ToolLoopAgent` ones. | Type | Description | | --- | --- | | `workflow.PrepareStepHook` | Mutates per-step options using the current history. | | `workflow.PrepareCallHook` | Mutates per-step call options before model invocation. | | `workflow.FilterActiveToolsHook` | Reduces the available tools for a step. | | `workflow.LanguageModelCallOptions` | The per-step call configuration, as in `ToolLoopAgent`. | | `workflow.StartCallback` | Fires once before the first step. | | `workflow.StepStartCallback` | Fires before each model step. | | `workflow.StepEndCallback`, `workflow.StepFinishCallback` | Fire after each completed step. | | `workflow.ToolExecutionStartCallback`, `workflow.ToolExecutionEndCallback` | Fire around local tool execution. | | `workflow.EndCallback`, `workflow.FinishCallback` | Fire once after the run completes. | | `workflow.ErrorCallback` | Fires when execution returns an error. | | `workflow.AbortCallback` | Fires when the context cancels the run. | {/* gen:fields workflow.LanguageModelCallOptions */} | Field | Type | Description | | --- | --- | --- | | `StepNumber` | `int` | | | `System` | `string` | | | `AllowSystemInMessages` | `bool` | | | `Messages` | `[]types.Message` | | | `Tools` | `[]types.Tool` | | | `ToolChoice` | `types.ToolChoice` | | | `CallOptions` | `interface{}` | | | `Temperature` | `*float64` | | | `MaxTokens` | `*int` | | | `TopP` | `*float64` | | | `TopK` | `*int` | | | `FrequencyPenalty` | `*float64` | | | `PresencePenalty` | `*float64` | | | `StopSequences` | `[]string` | | | `Seed` | `*int` | | | `Headers` | `map[string]string` | | | `Reasoning` | `*types.ReasoningLevel` | | | `SendReasoning` | `*bool` | | | `ProviderOptions` | `map[string]interface{}` | | | `RuntimeContext` | `interface{}` | | | `ToolsContext` | `map[string]interface{}` | | | `ExperimentalSandbox` | `interface{}` | | | `PreviousSteps` | `[]types.StepResult` | | | `AccumulatedUsage` | `types.Usage` | | | `CustomData` | `interface{}` | | | `StopWhen` | `[]ai.StopCondition` | StopWhen, ActiveTools and ExperimentalDownload mirror TS prepareCall's per-call stopWhen/activeTools/download settings parity (d56638a): a PrepareCall/PrepareStep hook can read the current effective value here and mutate it to override the call. | | `ActiveTools` | `[]string` | | | `ExperimentalDownload` | `ai.DownloadFunction` | | | `MaxRetries` | `*int` | MaxRetries and Timeout mirror TS's `maxRetries`/`abortSignal` prepareCall parity (419adc7). Timeout is the Go stand-in for TS's AbortSignal. | | `Timeout` | `*ai.TimeoutConfig` | | | `InitialInstructions` | `string` | InitialInstructions and InitialMessages are the original (unmutated) instructions/messages the call was invoked with, matching TS prepareCall's initial-inputs parity (b666f57). They are read-only: a hook should mutate System/Messages to change what is sent, not these. | | `InitialMessages` | `[]types.Message` | | {/* /gen:fields */} ## Iterating steps `workflow.StreamTextIterator` iterates the step results of a run. `Next(ctx)` returns the next `*types.StepResult` and `io.EOF` when the steps are exhausted. `Close()` marks the iterator closed. | Constructor | Description | | --- | --- | | `workflow.NewStreamTextIterator(steps []types.StepResult)` | Iterates a slice you already have. | | `workflow.NewStreamTextIteratorFromChannel(ch <-chan types.StepResult)` | Iterates live steps from a channel. | | `workflow.NewStreamTextIteratorFromResult(result *WorkflowResult)` | Iterates the steps of a result. | ## Serializable tools A durable run cannot persist a Go function. These helpers carry the tool definitions across an execution boundary, and you attach `Execute` functions again by name afterward. | Function | Description | | --- | --- | | `workflow.SerializeToolSet(tools []types.Tool) (map[string]SerializableToolDef, error)` | Converts tools to a JSON-safe map keyed by name. | | `workflow.ResolveSerializableTools(defs map[string]SerializableToolDef) []types.Tool` | Rebuilds non-executable tool descriptors. | | `workflow.ValidateSerializableToolInput(def SerializableToolDef, input interface{}) (interface{}, error)` | Validates input against the serialized schema. It applies schema defaults first, and returns the defaulted value. Pass that value to the tool, not the raw input. | | `workflow.MarshalSchema(schema interface{}) (map[string]interface{}, error)` | Converts a JSON-serializable schema to a map. | | `workflow.UnmarshalSchema(data map[string]interface{}) (interface{}, error)` | Rebuilds a schema as a generic JSON value. | {/* gen:fields workflow.SerializableToolDef */} | Field | Type | Description | | --- | --- | --- | | `Name` | `string` | | | `Description` | `string` | | | `Title` | `string` | | | `Parameters` | `map[string]interface{}` | | | `InputSchema` | `map[string]interface{}` | | | `Type` | `string` | | | `ID` | `string` | | | `Args` | `map[string]interface{}` | | | `Strict` | `*bool` | | | `ProviderExecuted` | `bool` | | | `IsProviderExecuted` | `bool` | | | `SupportsDeferredResults` | `bool` | | | `InputExamples` | `[]types.ToolInputExample` | | | `ProviderMetadata` | `map[string]interface{}` | | {/* /gen:fields */} ## Chat transports ### WorkflowChatTransport `workflow.WorkflowChatTransport` is a client-side `ai.ChatTransport`. It POSTs the message history to a workflow chat endpoint and reads the response as a stream of UI message chunks. When the response ends without a `finish` chunk, for example after a network drop or a function timeout, it reconnects with a GET to `//stream?startIndex=N`. It also repairs UI message stream framing and, on a resume with a negative start index, drops deltas and ends whose start fell outside the resumed window. ```go func NewWorkflowChatTransport(opts WorkflowChatTransportOptions) *WorkflowChatTransport ``` `SendMessages(ctx, ai.ChatTransportSendMessagesRequest)` and `ReconnectToStream(ctx, ai.ChatTransportReconnectToStreamRequest)` implement `ai.ChatTransport`. See [Stream transport helpers](https://goaisdk.com/docs/reference/ai/stream-transport-helpers.md#chattransport). {/* gen:fields workflow.WorkflowChatTransportOptions */} | Field | Type | Description | | --- | --- | --- | | `API` | `string` | API is the chat endpoint. Defaults to "/api/chat". | | `HTTPClient` | `*http.Client` | HTTPClient is used for both the initial POST and any reconnect GETs. Defaults to http.DefaultClient. | | `OnChatSendMessage` | `func(resp *http.Response, req ai.ChatTransportSendMessagesRequest) error` | OnChatSendMessage is invoked after the initial POST completes, useful for inspecting response headers or tracking chat history. | | `OnChatEnd` | `func(WorkflowChatTransportEndEvent) error` | OnChatEnd is invoked once, after a "finish" chunk is observed. | | `MaxConsecutiveErrors` | `int` | MaxConsecutiveErrors bounds reconnect attempts. Defaults to 3. | | `InitialStartIndex` | `int` | InitialStartIndex is the default startIndex used by ReconnectToStream when it is called directly (not as part of a SendMessages recovery). Negative values are tail-relative (e.g. -10 reads the last 10 chunks), useful for resuming a chat UI after a page refresh without replaying the whole conversation. Defaults to 0 (replay from the beginning). | | `PrepareSendMessagesRequest` | `func(ai.ChatTransportSendMessagesRequest) (PreparedChatRequest, error)` | PrepareSendMessagesRequest customizes the API endpoint, body, and headers used for the initial POST. | | `PrepareReconnectToStreamRequest` | `func(WorkflowChatTransportReconnectContext) (PreparedChatRequest, error)` | PrepareReconnectToStreamRequest customizes the API endpoint and headers used for reconnect GETs. | {/* /gen:fields */} {/* gen:fields workflow.SendMessagesOptions */} | Field | Type | Description | | --- | --- | --- | | `ChatID` | `string` | | | `MessageID` | `string` | | | `Trigger` | `string` | | | `Messages` | `interface{}` | | | `Body` | `map[string]interface{}` | | | `Headers` | `map[string]string` | | {/* /gen:fields */} {/* gen:fields workflow.ReconnectToStreamOptions */} | Field | Type | Description | | --- | --- | --- | | `RunID` | `string` | | | `ChatID` | `string` | | | `StartIndex` | `int` | | | `Headers` | `map[string]string` | | {/* /gen:fields */} `workflow.WorkflowChatTransportReconnectContext` is passed to `PrepareReconnectToStreamRequest`. `workflow.WorkflowChatTransportEndEvent` is passed to `OnChatEnd`. `workflow.ChatEndEvent` is the equivalent event for the multiplexer's `OnChatEnd`. It carries `ChatID`, `ChunkIndex`, `RunID` and the SSE `Events`. {/* gen:fields workflow.WorkflowChatTransportEndEvent */} | Field | Type | Description | | --- | --- | --- | | `ChatID` | `string` | | | `ChunkIndex` | `int` | | {/* /gen:fields */} ### WorkflowRunMultiplexer `workflow.WorkflowRunMultiplexer` is a Go-only server and client helper that predates `WorkflowChatTransport`. It serves run-scoped SSE and resume handlers. Its client helpers return raw SSE events, not UI message chunks, so it does not implement `ai.ChatTransport`. For new client code, use `WorkflowChatTransport`. ```go func NewWorkflowRunMultiplexer(opts ...WorkflowRunMultiplexerOptions) *WorkflowRunMultiplexer ``` | Method | Description | | --- | --- | | `ServeHTTP(w, r)` | Starts a new run stream and emits lifecycle events. | | `Resume(runID string) http.Handler` | Returns a handler that attaches an SSE reader to a run in progress. | | `SendMessages(ctx, SendMessagesOptions) ([]*streaming.SSEEvent, error)` | POSTs chat messages and returns parsed SSE events. | | `ReconnectToStream(ctx, ReconnectToStreamOptions) ([]*streaming.SSEEvent, error)` | GETs the API to resume a run-scoped SSE stream. | {/* gen:fields workflow.WorkflowRunMultiplexerOptions */} | Field | Type | Description | | --- | --- | --- | | `API` | `string` | | | `HTTPClient` | `*http.Client` | | | `MaxConsecutiveErrors` | `int` | | | `InitialStartIndex` | `int` | | | `OnChatSendMessage` | `func(*http.Response, SendMessagesOptions) error` | | | `OnChatEnd` | `func(ChatEndEvent) error` | | | `PrepareSendMessagesRequest` | `func(SendMessagesOptions) (PreparedChatRequest, error)` | | | `PrepareReconnectToStreamRequest` | `func(ReconnectToStreamOptions) (PreparedChatRequest, error)` | | {/* /gen:fields */} {/* gen:fields workflow.PreparedChatRequest */} | Field | Type | Description | | --- | --- | --- | | `API` | `string` | | | `Body` | `map[string]interface{}` | | | `Headers` | `map[string]string` | | {/* /gen:fields */} ## Durable harness runner These functions run a [harness agent](https://goaisdk.com/docs/reference/harness/agent.md) across process executions. Each execution returns a serializable `workflow.HarnessWorkflowState`. Your durable-workflow runtime stores it as the checkpoint, and you pass it to the next execution. | Function | Description | | --- | --- | | `workflow.CreateHarnessWorkflowState(HarnessWorkflowInput) HarnessWorkflowState` | Builds the initial state for one user turn. | | `workflow.RunHarnessAgent(ctx, RunHarnessAgentOptions) (HarnessWorkflowState, error)` | Runs one durable execution. It resumes or starts the session and streams the turn's chunks to `Writable`. When `TimeSliceSeconds` is positive, it races the turn against that budget and suspends the turn when the slice ends. | | `workflow.RunHarnessAgentTimeSlice(ctx, RunHarnessAgentTimeSliceOptions) (HarnessWorkflowState, error)` | Runs one time-boxed slice. When the slice ends before the turn, the status is `ready_for_next_step`. | | `workflow.RunHarnessAgentStep(ctx, RunHarnessAgentStepOptions) (HarnessWorkflowState, error)` | Runs until the next step boundary. Set `StopWhen` on the agent, for example `ai.IsStepCount(1)`. | | `workflow.RunHarnessAgentSlice(ctx, RunHarnessAgentSliceOptions)` | Deprecated alias of `RunHarnessAgentTimeSlice` that reports `timed_out` where the new function reports `ready_for_next_step`. | | `workflow.FinalizeHarnessWorkflow(state) (HarnessWorkflowFinalResult, error)` | Collapses a terminal state into its result. It returns an error if the run failed. | `workflow.DefaultTimeSliceSeconds` is 750. A Vercel Fluid Compute instance is recycled at about 800 seconds, so a 750-second slice leaves time for the next execution to reattach to the sandbox that is still running. ### Status {/* gen:consts workflow.HarnessWorkflowStatus */} | Constant | Value | Description | | --- | --- | --- | | `HarnessWorkflowStatusNotStarted` | `"not_started"` | HarnessWorkflowStatusNotStarted is the fresh state before any execution has run. | | `HarnessWorkflowStatusReadyForNextStep` | `"ready_for_next_step"` | HarnessWorkflowStatusReadyForNextStep means the turn remains unfinished and ContinueFrom carries the cursor for the next execution. | | `HarnessWorkflowStatusAwaitingToolApproval` | `"awaiting_tool_approval"` | HarnessWorkflowStatusAwaitingToolApproval means the turn emitted one or more tool approval/result requests and ContinueFrom carries the suspended turn. | | `HarnessWorkflowStatusFinished` | `"finished"` | HarnessWorkflowStatusFinished means the agent turn completed on its own; FinalResult is set. | | `HarnessWorkflowStatusFailed` | `"failed"` | HarnessWorkflowStatusFailed means the turn errored; Error is set. | | `HarnessWorkflowStatusTimedOut` | `"timed_out"` | HarnessWorkflowStatusTimedOut is returned by the deprecated RunHarnessAgentSlice in place of HarnessWorkflowStatusReadyForNextStep. Deprecated: use HarnessWorkflowStatusReadyForNextStep. | {/* /gen:consts */} ### Types `workflow.HarnessWorkflowAgent` is the subset of `*harness.Agent` the runner drives: `HasOutput`, `CreateSession`, `Stream` and `ContinueStream`. `*harness.Agent` satisfies it. {/* gen:fields workflow.HarnessWorkflowInput */} | Field | Type | Description | | --- | --- | --- | | `Prompt` | `harness.Prompt` | | | `Messages` | `[]types.Message` | | | `SessionID` | `string` | | | `ResumeFrom` | `*harness.ResumeSessionState` | | | `ContinueFrom` | `*harness.ContinueTurnState` | | {/* /gen:fields */} {/* gen:fields workflow.HarnessWorkflowState */} | Field | Type | Description | | --- | --- | --- | | `SessionID` | `string` | SessionID is the stable harness session id; doubles as the sandbox name across processes. | | `Prompt` | `harness.Prompt` | Prompt is the new user turn for this run. Sent once, on the execution that starts the turn. | | `Messages` | `[]types.Message` | Messages carries full model messages for continuing a suspended approval turn (e.g. tool-approval-response content). When non-nil, the next execution sends these instead of Prompt/ContinueFrom. | | `Status` | `HarnessWorkflowStatus` | | | `ResumeFrom` | `*harness.ResumeSessionState` | ResumeFrom carries resume coordinates for the next user turn. | | `ContinueFrom` | `*harness.ContinueTurnState` | ContinueFrom carries continuation coordinates for this run's current suspended turn. | | `StreamContext` | `*HarnessWorkflowStreamContext` | | | `FinalResult` | `*HarnessWorkflowFinalResult` | | | `Error` | `string` | | {/* /gen:fields */} {/* gen:fields workflow.HarnessWorkflowFinalResult */} | Field | Type | Description | | --- | --- | --- | | `SessionID` | `string` | | | `FinishReason` | `string` | | | `Usage` | `*HarnessWorkflowUsageSummary` | | | `Output` | `any` | Output is the agent's parsed and schema-validated output when the agent has an output specification (HarnessWorkflowAgent.HasOutput). | {/* /gen:fields */} {/* gen:fields workflow.RunHarnessAgentOptions */} | Field | Type | Description | | --- | --- | --- | | `Agent` | `HarnessWorkflowAgent` | | | `State` | `HarnessWorkflowState` | | | `SandboxSession` | `providerutils.SandboxSession` | SandboxSession, when set, is forwarded to HarnessWorkflowAgent. CreateSession as a caller-owned sandbox (harness.CreateSessionOptions. SandboxSession) instead of letting the agent's own SandboxProvider create/resume one. Mirrors TS `RunHarnessAgentOptions.sandboxSession`. | | `TimeSliceSeconds` | `float64` | TimeSliceSeconds is the wall-clock budget for this execution. Zero means no time slice: the run continues until the harness's own turn (or, with a StopWhen-configured agent, step) boundary. | | `DestroyOnFinish` | `bool` | DestroyOnFinish controls whether to destroy the sandbox when the run finishes or fails. Defaults to false: the session is parked and a fresh resume state is returned in ResumeFrom, so the next user turn reattaches to the same conversation (multi-turn chat). Set true for a one-shot run that should release the sandbox when the run ends. | | `Writable` | `HarnessWorkflowWriter` | Writable is where to write the turn's UI-message chunks. Required — TS's default resolveWorkflowWritable() (a Workflow DevKit getWritable()) has no Go equivalent. | {/* /gen:fields */} {/* gen:fields workflow.RunHarnessAgentTimeSliceOptions */} | Field | Type | Description | | --- | --- | --- | | `Agent` | `HarnessWorkflowAgent` | | | `State` | `HarnessWorkflowState` | | | `SandboxSession` | `providerutils.SandboxSession` | | | `TimeSliceSeconds` | `float64` | TimeSliceSeconds defaults to DefaultTimeSliceSeconds when zero. | | `DestroyOnFinish` | `bool` | | | `Writable` | `HarnessWorkflowWriter` | | {/* /gen:fields */} {/* gen:fields workflow.RunHarnessAgentStepOptions */} | Field | Type | Description | | --- | --- | --- | | `Agent` | `HarnessWorkflowAgent` | | | `State` | `HarnessWorkflowState` | | | `SandboxSession` | `providerutils.SandboxSession` | | | `DestroyOnFinish` | `bool` | | | `Writable` | `HarnessWorkflowWriter` | | {/* /gen:fields */} {/* gen:fields workflow.RunHarnessAgentSliceOptions */} | Field | Type | Description | | --- | --- | --- | | `Agent` | `HarnessWorkflowAgent` | | | `State` | `HarnessWorkflowState` | | | `SandboxSession` | `providerutils.SandboxSession` | | | `TimeSliceSeconds` | `float64` | | | `SliceTimeoutSeconds` | `float64` | SliceTimeoutSeconds is used when TimeSliceSeconds is zero. Deprecated: use TimeSliceSeconds. | | `DestroyOnFinish` | `bool` | | | `Writable` | `HarnessWorkflowWriter` | | {/* /gen:fields */} `workflow.HarnessWorkflowStreamContext` is the serializable subset of in-flight UI message state carried across an execution boundary. `workflow.HarnessWorkflowActiveToolInput` is a tool input that was partly streamed when the execution ended. `workflow.HarnessWorkflowUsageSummary` is a minimal token usage summary. {/* gen:fields workflow.HarnessWorkflowStreamContext */} | Field | Type | Description | | --- | --- | --- | | `ActiveTextParts` | `map[string]ai.UIMessageChunk` | | | `ActiveReasoningParts` | `map[string]ai.UIMessageChunk` | | | `ActiveToolInputs` | `map[string]HarnessWorkflowActiveToolInput` | | | `PendingToolInputs` | `map[string]ai.UIMessageChunk` | | {/* /gen:fields */} {/* gen:fields workflow.HarnessWorkflowUsageSummary */} | Field | Type | Description | | --- | --- | --- | | `InputTokens` | `*int64` | | | `OutputTokens` | `*int64` | | {/* /gen:fields */} ### Writer `workflow.HarnessWorkflowWriter` receives the UI message chunks of one execution: ```go type HarnessWorkflowWriter interface { Write(chunk ai.UIMessageChunk) error Close() error } ``` `Write` is called once per chunk, in order. `Close` is called only when the run reaches `finished`. The other statuses leave the writer open because a later execution keeps writing, or the failure propagates. `workflow.NewChanHarnessWorkflowWriter(ch)` returns a `*workflow.ChanHarnessWorkflowWriter` that forwards every chunk to `ch` and closes `ch` on `Close`. --- # Migrating from the TypeScript AI SDK > Map Vercel AI SDK (ai@7.0.x) calls to their Go SDK equivalents, and see where the Go port has to diverge. Canonical URL: https://goaisdk.com/docs/migration-guides/from-typescript-ai-sdk Documentation index: https://goaisdk.com/llms.txt This guide is for developers who know the [Vercel AI SDK](https://ai-sdk.dev) and are porting an app to Go, or building a new Go backend behind an existing TypeScript frontend. It assumes no Go SDK experience. The Go SDK (`github.com/digitallysavvy/go-ai`) targets 1:1 parity with TypeScript `ai@7.0.127`. Most concepts map directly: same function names in spirit, same option names where Go syntax allows, same wire formats for anything that crosses a network boundary (provider requests, UI message streams). The core package is `pkg/ai`, mirroring the `ai` package; individual providers live under `pkg/providers/*`, mirroring `@ai-sdk/*`. If you're coming from TypeScript v5 or v6 rather than v7, read [Coming from TS v5 or v6](#coming-from-ts-v5-or-v6) first — several TS names changed on the way to v7, and Go only implements the current (v7) names. ## Core API mapping Import paths below all live under `github.com/digitallysavvy/go-ai/pkg/...`. Options in Go are structs (`ai.GenerateTextOptions`, not an object literal); results are pointers (`*ai.GenerateTextResult`) plus an `error`, not a `Promise`. ### Text generation | TypeScript | Go | Note | |---|---|---| | `generateText({model, prompt})` | `ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "..."})` | `pkg/ai/generate.go`. Returns `(*ai.GenerateTextResult, error)`. | | `streamText({model, prompt})` | `ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "..."})` | `pkg/ai/stream.go`. Returns `(*ai.StreamTextResult, error)` before the first model request, matching TS `streamText`; only option-validation errors come back from the call itself, everything else surfaces through `Err()`/`ReadAll()`/`Stream()`. | | `instructions` (`system` deprecated) | `System string` / `Instructions *string` | Go's canonical field is still the plain string `System`; `Instructions` is a pointer alias for code ported verbatim from TS. Either works; `Instructions` wins if both are set. | | `stopWhen: isStepCount(n)` | `StopWhen: []ai.StopCondition{ai.IsStepCount(n)}` | `pkg/ai/stop_condition.go`. Also `ai.HasToolCall(names...)`, `ai.IsLoopFinished()`. | | `prepareStep` | `PrepareStep func(ctx context.Context, step ai.PrepareStepOptions) ai.PrepareStepOptions` | `pkg/ai/generate.go`. | | `result.text` | `result.Text` (GenerateText) / `result.Text()` (StreamText) | StreamText exposes accessor **methods** because text accumulates as chunks arrive. | | `result.usage.inputTokens` | `result.Usage.GetInputTokens()` (or `.InputTokens` — a `*int64`) | `pkg/provider/types/usage.go`. Prefer the `Get*Tokens()` helpers; the struct fields are pointers and can be `nil`. | | `result.toolCalls` / `.toolResults` | `result.ToolCalls` / `.ToolResults` (GenerateText) or `.ToolCalls()` / `.ToolResults()` (StreamText) | | | `result.steps` / `.finalStep` | `result.Steps` / `.FinalStep` (GenerateText) or `.Steps()` / `.FinalStep()` (StreamText) | | | `result.stream` (full event stream) | `result.FullStream()` — every chunk, including `start`/`start-step`/`finish-step`/`finish` lifecycle chunks | `pkg/ai/stream.go`. `ChunkTypeFinish` fires once per call with total usage; each step is bracketed by `ChunkTypeStartStep`/`ChunkTypeFinishStep`. See [StreamText reference](https://goaisdk.com/docs/reference/ai/stream-text.md#full-stream-lifecycle). | Consuming a stream is the one place the two SDKs look structurally different — see [Promises and async iterables vs. return values and channels](#promises-and-async-iterables-vs-return-values-and-channels). ### Structured output: generateObject / streamObject / Output Go ships two mechanisms. Prefer the `Output` one; `GenerateObject`/`StreamObject` are the legacy path (TS deprecated `generateObject`/`streamObject` the same way in 6.0, in favor of `generateText`/`streamText` with an `output` option). | TypeScript | Go | Note | |---|---|---| | `Output.object({schema, name, description})` | `ai.ObjectOutput[T](ai.ObjectOutputOptions{Schema: ai.SchemaFor[T](), Name: "...", Description: "..."})` | `pkg/ai/output.go`. `T` is a real Go generic type parameter — see [Generics](#generics-typed-tool-input-and-output). | | `Output.array({schema})` | `ai.ArrayOutput[ELEMENT](ai.ArrayOutputOptions[ELEMENT]{...})` | `output.go`. | | `Output.choice({values})` | `ai.ChoiceOutput[CHOICE](ai.ChoiceOutputOptions[CHOICE]{...})` | `output.go`. `CHOICE` must be a `~string` type. | | `Output.text()` | `ai.TextOutput()` | `output.go`. | | `Output.json()` | `ai.JSONOutput(ai.JSONOutputOptions{...})` | `output.go`. | | `generateText({output})` | `ai.GenerateTextOptions{Output: ai.ObjectOutput[T](...)}` | Result is `result.Output any` — type-assert to `T`. | | `streamText({output})`, `result.partialOutputStream` | `ai.StreamTextOptions{Output: ...}`, `result.PartialOutput()` | `PartialOutput()` returns the latest partial parse; `Output()` returns the final one after the stream drains. | | `generateObject({model, schema, prompt})` (deprecated) | `ai.GenerateObject(ctx, ai.GenerateObjectOptions{Model: model, Schema: mySchema, Prompt: "..."})` | `pkg/ai/object.go`. `Schema` is a `schema.Schema`, not a Go generic — `result.Object` is `interface{}`. | | — | `ai.GenerateObjectInto(ctx, opts, &target)` | Go-only convenience: unmarshals straight into a struct pointer via reflection, no type assertion needed. | `StreamObject` is one place Go and TS genuinely diverge: see [StreamObject is not lazy](#streamobject-is-not-lazy) below. ### Embeddings and reranking | TypeScript | Go | Note | |---|---|---| | `embed({model, value})` | `ai.Embed(ctx, ai.EmbedOptions{Model: model, Input: "..."})` | `pkg/ai/embed.go`. Result: `EmbedResult.Embedding []float64`. `MaxRetries` is `*int` (nil = TS default of 2). | | `embedMany({model, values})` | `ai.EmbedMany(ctx, ai.EmbedManyOptions{Model: model, Inputs: []string{...}})` | `embed.go`. Result: `EmbedManyResult.Embeddings [][]float64`. An empty `Inputs` now returns an empty result instead of an error, matching TS. | | `rerank({model, query, documents})` | `ai.Rerank(ctx, ai.RerankOptions{Model: model, Query: "...", Documents: []string{...}})` | `pkg/ai/rerank.go`. `Documents` is `interface{}` (accepts `[]string` or `[]map[string]interface{}`); result is `RerankResult.Ranking []RerankItem`. | | `cosineSimilarity(a, b)` | `ai.CosineSimilarity(a, b)` | `pkg/ai/embed.go`. A zero or empty vector returns `(0, nil)`, not an error; only a length mismatch errors. | ### Media: image, speech, video, transcription | TypeScript | Go | Note | |---|---|---| | `generateImage({model, prompt})` | `ai.GenerateImage(ctx, ai.GenerateImageOptions{Model: model, Prompt: "..."})` | `pkg/ai/generate_image.go`. Result: `.Images []types.GeneratedFile`, `.Image` (first one). | | `generateSpeech({model, text})` | `ai.GenerateSpeech(ctx, ai.GenerateSpeechOptions{Model: model, Text: "..."})` | `pkg/ai/generate_speech.go`. Result: `.Audio ai.GeneratedAudioFile`. | | `transcribe({model, audio})` | `ai.Transcribe(ctx, ai.TranscribeOptions{Model: model, Audio: audioBytes})` | `pkg/ai/transcribe.go`. Accepts `Audio []byte`, `AudioBase64 string`, or `AudioURL string`. | | `experimental_generateVideo({model, prompt})` | `ai.GenerateVideo(ctx, ai.GenerateVideoOptions{Model: model, Prompt: ai.VideoPrompt{Text: "..."}})` | `pkg/ai/generate_video.go`. Blocks until the video is ready. For the async job API, see the next row. | | `experimental_startVideo` / `experimental_getVideoStatus` | `ai.ExperimentalStartVideo(ctx, ai.StartVideoOptions{...})` / `ai.ExperimentalGetVideoStatus(ctx, model, ai.GetVideoStatusOptions{Operation: started.Operation})` | `pkg/ai/start_video.go`, `pkg/ai/get_video_status.go`. Implemented for fal, Google, Google Vertex, Replicate, and xAI. | | `experimental_streamTranscribe` | `ai.ExperimentalStreamTranscribe(ctx, ai.StreamTranscribeOptions{...})` | `pkg/ai/stream_transcribe.go`. Also `ai.ExperimentalStreamTranslate` for speech-to-speech translation. | ### Tools and agents | TypeScript | Go | Note | |---|---|---| | `tool({description, inputSchema, execute})` | `types.Tool{Name: "...", Description: "...", Parameters: schemaValue, Execute: fn}` | `pkg/provider/types/tool.go`. No `tool()` constructor — you build the struct directly, and `Name` is a real field (TS infers it from the object key). | | `dynamicTool({...})` | `types.Tool{..., Type: types.ToolTypeDynamic}` | `tool.go`. Set `Type` explicitly instead of calling a separate constructor. | | `execute: async ({location}, {toolCallId}) => ...` | `Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error)` | `tool.go`. Three parameters, always: context, raw input map, and call metadata (`opts.ToolCallID`, `opts.RuntimeContext`, `opts.ToolContext`). | | `tools: {weather: weatherTool}` | `Tools: []types.Tool{weatherTool}` | Tools are a slice, not a map — the tool's `Name` field does the job the object key does in TS. | | `tool-search`, `experimental_toolCallers` | `ai.ToolSearch(...)`, `types.Tool.DeferLoading`, `GenerateTextOptions.ExperimentalToolCallers` | `pkg/ai/tool_search.go`, `pkg/provider/types/tool.go`. Native tool search with tools loaded on demand; tools may be called by other tools (Anthropic code execution, OpenAI programmatic tool calling act as callers). Messages a caller adds persist across steps, and `PrepareStep`'s `InitialMessages` stays the raw prompt, as in TS. | | `new ToolLoopAgent({model, tools, instructions, stopWhen})` | `agent.NewToolLoopAgent(agent.AgentConfig{Model: model, Tools: tools, System: "...", StopWhen: ...})` | `pkg/agent/toolloop.go`, `pkg/agent/agent.go`. | | `agent.generate({prompt})` | `a.Generate(ctx, agent.AgentGenerateOptions{Prompt: "..."})` | `pkg/agent/agent.go` (interface method). Returns `*ai.GenerateTextResult`, matching TS exactly. Go also has `a.Execute(ctx, prompt)`, a convenience that returns a flatter `*agent.AgentResult` — not a TS export. | | `agent.stream({prompt})` | `a.Stream(ctx, agent.AgentStreamOptions{Prompt: "..."})` | Returns `*ai.StreamTextResult`. | | `prepareCall` (agent option) | `AgentConfig.PrepareCall func(ctx, agent.PrepareCallConfig) agent.PrepareCallConfig` | `agent.go`. Same name as TS; don't confuse with `PrepareStep` on the plain `GenerateText`/`StreamText` options. | ### Middleware | TypeScript | Go | Note | |---|---|---| | `wrapLanguageModel({model, middleware})` | `middleware.WrapLanguageModel(model, []*middleware.LanguageModelMiddleware{mw}, nil, nil)` | `pkg/middleware/language_model_middleware.go`. The trailing `nil, nil` are optional model-ID/provider-ID overrides. | | `wrapImageModel({model, middleware})` | `middleware.WrapImageModel(model, []*middleware.ImageModelMiddleware{mw}, nil, nil)` | `pkg/middleware/wrap_image_model.go`, added this cycle. `WrapProvider(..., WithImageModelMiddleware(...))` applies it provider-wide. | | `middleware.transformParams` | `LanguageModelMiddleware.TransformParams func(ctx, callType string, params *provider.GenerateOptions, model) (*provider.GenerateOptions, error)` | `language_model_middleware.go`. | | `middleware.wrapGenerate` / `wrapStream` | `LanguageModelMiddleware.WrapGenerate` / `.WrapStream` | Same two-hook split as TS (no combined `wrapResult`). | | `middleware.overrideSupportedUrls` | `LanguageModelMiddleware.OverrideSupportedURLs func(model provider.LanguageModel) map[string][]string` | `language_model_middleware.go`. A wrapped model without this set forwards the underlying model's `SupportedURLs()` unchanged. | | `extractReasoningMiddleware`, `simulateStreamingMiddleware`, `defaultSettingsMiddleware` | `middleware.ExtractReasoningMiddleware(...)`, `middleware.SimulateStreamingMiddleware(...)`, `middleware.DefaultSettingsMiddleware(...)` | `pkg/middleware/*.go`, one file per built-in middleware. `DefaultSettingsMiddleware` now deep-merges `ProviderOptions` and preserves all call fields instead of rebuilding from a fixed field list. | ### Providers and the registry | TypeScript | Go | Note | |---|---|---| | `import {openai} from '@ai-sdk/openai'; openai('gpt-4o')` | `p := openai.New(openai.Config{APIKey: "..."}); model, err := p.LanguageModel("gpt-4o")` | `pkg/providers/openai/provider.go`. Provider construction and model lookup are separate calls, and model lookup returns an error. | | `createProviderRegistry({openai, anthropic})` | `r := registry.NewRegistry(); r.RegisterProvider("openai", openai.New(...))` | `pkg/registry/registry.go`. `NewRegistry` also accepts `registry.WithSeparator`, `registry.WithLanguageModelMiddleware`, `registry.WithImageModelMiddleware`. | | `registry.languageModel('openai:gpt-4o')` | `r.ResolveLanguageModel("openai:gpt-4o")` | `registry.go`. Same `provider:model` string format. An unregistered or malformed ID returns a typed `*errors.NoSuchModelError`. | | `customProvider({languageModels: {...}, fallbackProvider})` | `registry.NewCustomProvider(registry.CustomProviderOptions{LanguageModels: map[string]provider.LanguageModel{...}, Fallback: p})` | `pkg/registry/custom_provider.go`. | ### MCP client | TypeScript | Go | Note | |---|---|---| | `createMCPClient({transport})` | `mcp.CreateStdioMCPClient(cmd, args)` / `mcp.CreateHTTPMCPClient(url, oauth)` / `mcp.CreateSSEMCPClient(url, oauth)` | `pkg/mcp/integration.go`. One constructor per transport instead of a transport object. `experimental_createMCPClient` in TS is now a deprecated alias for the same thing. | | `await mcpClient.tools()` | `mcp.GetMCPToolsForAgent(ctx, client)` | `integration.go`. Returns `[]types.Tool`, ready to drop into `Tools:`. | | `client.on('elicitation/create', handler)` / server elicitation | `client.OnElicitationRequest(func(ctx, req mcp.ElicitationRequest) (mcp.ElicitResult, error) { ... })` | `pkg/mcp/client.go`. Register before `Connect`; an unregistered handler causes the client to reject `elicitation/create` requests automatically. | | `client.listResourceTemplates()` | `client.ListResourceTemplates(ctx)` | `pkg/mcp/client.go`. | | Stdio server env inheritance | `mcp.StdioTransportConfig.Env []string` | The child process no longer inherits the full parent environment — only `Env` plus a small safe allowlist (`PATH`, `HOME`, `USER`, `LOGNAME`, `SHELL`, `TERM`) are passed through. | ### UI message streams | TypeScript | Go | Note | |---|---|---| | `validateUIMessages({messages})` | `ai.ValidateUIMessages(ctx, ai.ValidateUIMessagesOptions{Messages: messages})` | `pkg/ai/validate_ui_messages.go`. | | `await convertToModelMessages(uiMessages)` | `ai.ConvertToModelMessages(ctx, uiMessages)` | `pkg/ai/convert_to_model_messages.go`. Both are effectively async/fallible; Go returns `([]types.Message, error)` instead of a rejected promise. | | `result.toUIMessageStream()` (method, deprecated) / standalone `toUIMessageStream(result.stream)` | `result.ToUIMessageStream(ctx)` or `ai.ToUIMessageStream(ctx, stream)` | `pkg/ai/stream.go`, `pkg/ai/ui_message_stream.go`. | | `createUIMessageStreamResponse(...)` | `ai.CreateUIMessageStreamResponse(ctx, result)` | `ui_message_stream.go`. Returns `(*http.Response, error)`. To write to an `http.ResponseWriter` directly, use `ai.PipeUIMessageStreamToResponse` or `ai.PipeUIMessageChunksToResponse`, which set the headers and flush each chunk (see [Serving a TS frontend](#serving-a-ts-frontend-from-a-go-backend)). | | `pipeUIMessageStreamToResponse(response, ...)` | `ai.PipeUIMessageStreamToResponse(ctx, result, w)` | `ui_message_stream.go`. Writes SSE frames straight to any `io.Writer`, flushing after every event. | | `createUIMessageStreamResponse({ stream })` / `pipeUIMessageStreamToResponse({ response, stream })` | `ai.CreateUIMessageChunksResponse(chunks, init)` / `ai.PipeUIMessageChunksToResponse(chunks, w, init)` | For any UI message chunk stream, such as one from `CreateUIMessageStreamWithOptions`. | | `UI_MESSAGE_STREAM_HEADERS` | `ai.UIMessageStreamHeaders()` | The `Pipe*` helpers set these on an `http.ResponseWriter` for you. | | `createAgentUIStreamResponse({ agent, uiMessages })` / `pipeAgentUIStreamToResponse({ response, agent, uiMessages })` | `agent.CreateAgentUIStreamResponseFromUIMessages` / `agent.PipeAgentUIStreamFromUIMessagesToResponse` | Validate and convert `useChat` messages, then stream the agent's reply. See [Agent UI Stream Helpers](https://goaisdk.com/docs/reference/ai/agent-ui-stream-helpers.md). | `readUIMessageStream(stream)` | `ai.ReadUIMessageStream(r io.Reader)` | `ui_message_stream.go`. | | chunk-sequence error passed to `onError` | `ai.UIMessageStreamError` / `ai.IsUIMessageStreamError(err)` | Typed error for malformed UI message stream sequences. | | `lastAssistantMessageIsCompleteWithToolCalls` | `ai.LastAssistantMessageIsCompleteWithToolCalls(messages)` | Also `ai.LastAssistantMessageIsCompleteWithApprovalResponses`. | | `generateId()` / `createIdGenerator({...})` | `ai.GenerateID()` / `ai.CreateIDGenerator(ai.CreateIDGeneratorOptions{...})` | `pkg/ai/generate_id.go`. See [GenerateID reference](https://goaisdk.com/docs/reference/ai/generate-id.md). | ### Stream transforms and message pruning | TypeScript | Go | Note | |---|---|---| | `smoothStream({delayInMs, chunking})` | `ai.SmoothStream(ai.SmoothStreamOptions{DelayInMs: ..., Chunking: "word"})` | `pkg/ai/smooth_stream.go`. Returns `(ai.StreamTransformFunc, error)`. An invalid `chunking` value now returns a typed `*errors.InvalidArgumentError`. | | `streamText({experimental_transform: [smoothStream(...)]})` | `ai.StreamTextOptions{ExperimentalTransform: []ai.StreamTransformFunc{transform}}` | `pkg/ai/stream.go`. Still `experimental_` in **both** SDKs — TS hasn't graduated this option as of 7.0.127 either, so this isn't a Go gap. | | `pruneMessages({messages, toolCalls: [...]})` | `ai.PruneModelMessages(messages, ai.PruneModelMessagesOptions{ToolCalls: []ai.PruneToolCallsRule{{Type: "before-last-message"}}})` | `pkg/ai/prune_messages.go`. Doc comment explicitly mirrors TS `generate-text/prune-messages.ts`, quirks included. | ### Telemetry | TypeScript | Go | Note | |---|---|---| | `registerTelemetry(new OpenTelemetry())` (`ai.* span names, from `@ai-sdk/otel`) | `telemetry.RegisterTelemetryIntegration(telemetry.NewLegacyOpenTelemetry(telemetry.LegacyOpenTelemetryOptions{}))` | `pkg/telemetry/registry.go`. `LegacyOpenTelemetry` is the current name; `OTelTelemetryIntegration` remains only as a deprecated type alias. | | `registerTelemetry(new OpenTelemetry({semconv: 'genai'}))` (`gen_ai.*` span names) | `telemetry.RegisterTelemetryIntegration(telemetry.NewOpenTelemetry(telemetry.OpenTelemetryOptions{}))` | Added this cycle: a second integration following the OpenTelemetry GenAI semantic conventions. Both integrations can be registered together. | | `generateText({telemetry: {functionId}})` | `ai.GenerateTextOptions{Telemetry: &telemetry.Options{FunctionID: "..."}}` | `pkg/telemetry/settings.go`. `telemetry.Settings` is a type alias for `telemetry.Options`; use `Options` in new code. Setting a tracer per call (`telemetry.Options.Tracer`) is removed — construct the integration with its tracer instead. | ### Errors | TypeScript | Go | Note | |---|---|---| | `try { ... } catch (e) { if (e instanceof APICallError) ... }` | `result, err := ai.GenerateText(...); var apiErr *providererrors.ProviderError; if errors.As(err, &apiErr) { ... }` | `pkg/provider/errors/errors.go`. `ProviderError` is the Go name for TS `APICallError`. | | `RateLimitError`, `InvalidArgumentError`, `NoSuchModelError`, ... | `providererrors.RateLimitError`, `providererrors.InvalidArgumentError`, `providererrors.NoSuchModelError`, ... plus one `Is*Error(err) bool` helper per type | `errors.go` — every helper is a thin `errors.As` wrapper, e.g. `IsRateLimitError`. | | model ignores a forced/specific `toolChoice` | `providererrors.ToolChoiceViolationError` | New this cycle: returned from `GenerateText`/`StreamText` when the model ignores a required or specific tool choice. | ### Harness | TypeScript | Go | Note | |---|---|---| | `@ai-sdk/harness` + `HarnessV1Session` | `pkg/harness`, `harness.Session` | `pkg/harness/spec.go`. The low-level adapter protocol: adapts an external coding-agent runtime (Claude Code, Codex, Cursor, ...) into the shared `harness-v1` session/lifecycle/bootstrap protocol; the in-sandbox bridge scripts are the unchanged TS ones. | | `session.readHistory({ since })` / `HarnessV1Session.doReadHistory` | `AgentSession.ReadHistory` / `harness.HistoryReader` | Added this cycle (contract only — no adapter implements it yet): reads the conversation history the runtime itself persisted, normalized to `[]harness.HistoryMessage` plus an opaque `Cursor` for incremental reads. Returns `CapabilityUnsupportedError` when the adapter doesn't implement `HistoryReader`, or `HistoryUnavailableError` when it does but can't reach the runtime's store. | | `Agent` / `AgentSession` (harness runtime) | `harness.Agent` / `harness.AgentSession` | Added this cycle: a higher-level session runtime on top of `harness.Session` — host/builtin tool approvals, client-side tool pauses, continuations, `StopWhen`, active/inactive tools, structured output, `ExperimentalSteer`. | | adapter packages (`@ai-sdk/harness-claude-code`, `-codex`, `-opencode`, `-deepagents`, `-acp`, `-cursor`, `-fx`, `-github-copilot`, `-grok-build`) | `pkg/harness/claudecode`, `codex`, `opencode`, `deepagents`, `acp`, and ACP-based Cursor/fx/GitHub Copilot/Grok Build adapters | File-based subscription credentials with OAuth refresh; disk-log replay and rerun reconnection. | | `createClaudeCode({ agentProgressSummaries, forwardSubagentText })` | `claudecode.Settings{AgentProgressSummaries: true, ForwardSubagentText: true}` | Added this cycle: sub-agent tool activity, text, and background-task notifications (previously discarded for messages carrying `parent_tool_use_id`) are now forwarded as `harness.RawPart` stream parts, alongside Claude's own tool-progress and per-response `message_stop` usage boundaries. | | `claude-opus-5-5` (`TurnSettings.Model`) | `harness.TurnSettings{Model: "claude-opus-5-5"}` | Added this cycle: the embedded Claude Code bridge's `@anthropic-ai/claude-code`/`@anthropic-ai/claude-agent-sdk` dependencies were bumped to 2.1.281/0.3.281, which know this model. | | `resolveSandboxCredentialEnvironment()` | `harnessutil.ResolveSandboxCredentialEnvironment` | Changed this cycle (replaces `CreateSandboxCredentialEnvironment`, removed): resuming a session now resolves each credential variable individually — a name already saved in the resumed lifecycle state's `PreviousSandboxCredentialEnvironment` is reused verbatim, while a new or removed credential variable is still generated or dropped correctly, instead of the old all-or-nothing "reuse the whole saved map or regenerate everything" behavior. | | `sandboxConfig.workDir: '.'` | `harness.SandboxConfig{WorkDir: "."}` | Added this cycle: `"."` is the literal alias for the sandbox's own default working directory — sessions and lifecycle callbacks (`OnBootstrap`, `OnSession`) then run directly there instead of a generated `-` or fixed subdirectory. Any other input that merely normalizes to the sandbox root (`"./"`, `"repo/.."`, ...) is still rejected. | | `HarnessAgentSettings.runtimeContext` | `harness.AgentSettings.RuntimeContext` | Fixed this cycle: `Agent.Generate`/`Stream`/`ContinueGenerate`/`ContinueStream` now forward the agent's configured `RuntimeContext` (previously always dropped in favor of an empty value) to lifecycle callbacks and telemetry, unless a per-call `agent.AgentGenerateOptions.RuntimeContext` or a `PrepareCallResult{HasRuntimeContext: true}` override takes precedence. `CreateSessionOptions.RuntimeContext` rebinds it when resuming an unfinished turn from `ContinueFrom`/`ResumeFrom`, mirroring the existing `ToolsContext` rebind. | | `google/*` models via `GOOGLE_GENERATIVE_AI_API_KEY` | `opencode.AuthGoogle` | Fixed this cycle: OpenCode now resolves `google/*` models and a supplied/ambient `GOOGLE_GENERATIVE_AI_API_KEY` to the Google provider, forwards and brokers it (`x-goog-api-key` against `generativelanguage.googleapis.com`), and includes it in the sandbox credential allowlist. Automatic startup inference selects Google only for a non-empty key with no competing non-empty direct-provider credential; an explicit `anthropic`/`openai` auth mode stays authoritative. | | Tool lifecycle timing within a step | (no Go-side API change) | Fixed this cycle: `Callbacks.OnToolExecutionStart` (and the OpenTelemetry `execute_tool` span) now fires as a tool's execution actually begins — synchronously, before a host tool's own goroutine is even spawned, so a step's telemetry span state is never read concurrently with the main goroutine advancing past it — instead of being held until the tool's outcome is already known, and each tool's result (`OnToolExecutionEnd`, the matching span's end, and its streamed `tool-result` chunk) now publishes as soon as that tool finishes instead of waiting for every tool in the same step — a fast tool's result and end callback no longer wait behind a slower sibling. | | workflow harness detach failures | `pkg/workflow.RunHarnessAgent`/`RunHarnessAgentTimeSlice` | Fixed this cycle: a finished turn's `AgentSession.Detach` failure now propagates instead of being swallowed in favor of the previous (now stale) `ResumeFrom`, which would otherwise resume from the wrong point. `AgentSession.Detach` itself already only marked the session detached after a successful `Detach`, so a failed detach already left the local handle usable for `Stop`-based cleanup — a regression test now locks that in. | | Codex MCP tool-call names | (embedded bridge; no Go-side API change) | Fixed this cycle: a Codex MCP tool call now gets a `mcp____` name when the item carries a server identity (so tool calls from different MCP servers are distinguishable), falling back to the bare tool name only when no server is present. | | OpenCode turn cancellation | (no Go-side API change) | Fixed this cycle: cancelling an OpenCode turn (ctx cancellation on `DoPromptTurn`) now actually stops OpenCode generating (`session.abort`) and a replacement `DoPromptTurn`/`DoContinueTurn` call on the same session waits for the aborted turn to fully drain (stray events dropped, a draining error swallowed, then its trailing `finish`) before starting, instead of racing it. | | ACP host-tool relay request | (embedded bridge; no Go-side API change) | Fixed this cycle: the embedded ACP bridge's host-tool relay now requires a one-use authorization from an independently observed ACP tool call (matching tool name and input) before executing a host tool, closing a gap where the relay only checked its sandbox bearer token. | | bare LangChain tool lifecycles (`pkg/langchain` UI message stream) | `pkg/langchain` | Fixed this cycle: a `tool-input-start`/`tool-output-available` pair for a bare (non-ToolMessage) tool call is now tracked as unfinished per LangGraph namespace until its output arrives, so a provider reusing the same tool-call ID for a later, unrelated call in a different namespace or step is no longer mistaken for a delayed output of the earlier one. | | sandbox providers (`@vercel/sandbox`, ...) | `pkg/harness/sandbox/vercel` | Added this cycle: runs harness sandboxes on Vercel Sandbox. Credentials from `VERCEL_OIDC_TOKEN` or explicit `Token`/`TeamID`/`ProjectID` — see [Known differences](https://goaisdk.com/docs/migration-guides/known-differences.md#vercel-sandbox-no-oidc-token-refresh-loop). | | model-written-code tool execution (no direct TS equivalent shipped) | `pkg/codemode` | Added this cycle, experimental: runs model-written JavaScript in a QuickJS-on-WebAssembly sandbox with TS execution-policy limits. See [Known differences](https://goaisdk.com/docs/migration-guides/known-differences.md#code-mode-never-settling-promises-and-in-flight-limits). | ### Batch, evaluate, and files | TypeScript | Go | Note | |---|---|---| | `experimental_startBatch` / `getBatchStatus` / `getBatchResults` / `cancelBatch` / `listBatches` | `ai.ExperimentalStartBatch` / `ExperimentalGetBatchStatus` / `ExperimentalGetBatchResults` / `ExperimentalCancelBatch` / `ExperimentalListBatches` | `pkg/ai/batch.go`. Implemented by Anthropic, OpenAI, and Google (including image requests), plus AI Gateway batch passthrough. | | `experimental_evaluate` | `ai.ExperimentalEvaluate` | `pkg/ai/evaluate.go`. Native `EvaluationModel()` on Anthropic, OpenAI, and Google; `ai.NewEvaluationLanguageModel` evaluates with any language model. | | `getFileMetadata` / `downloadFile` / `deleteFile` | `ai.GetFileMetadata` / `ai.DownloadFile` / `ai.DeleteFile` | Files API v4, with streamed multipart uploads; implemented for OpenAI (new this cycle), Anthropic, Google, DeepSeek. | ## Where Go differs ### Promises and async iterables vs. return values and channels TS awaits a promise and iterates `result.stream` with `for await`. Go returns `(*Result, error)` immediately and hands you a `provider.TextStream` you drain in a loop, or a `<-chan provider.StreamChunk` you `range` over. `provider.TextStream` only has `Next()`, `Err()`, and `Close()` — it does **not** implement `io.Reader` (`pkg/provider/language_model.go`). ```typescript const result = streamText({ model, prompt: 'Count to five.' }); for await (const chunk of result.stream) { if (chunk.type === 'text-delta') process.stdout.write(chunk.text); } ``` ```go stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Count to five.", }) if err != nil { log.Fatal(err) } for chunk := range stream.Chunks() { if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } } if err := stream.Err(); err != nil { // check AFTER the range, not during log.Fatal(err) } ``` ### AbortSignal vs. context.Context TS cancels a call by passing an `AbortSignal`. Go threads a `context.Context` as the first argument everywhere and cancels or times it out the normal Go way; there's no separate signal object to construct or pass down. ```typescript const controller = new AbortController(); setTimeout(() => controller.abort(), 5000); await generateText({ model, prompt, abortSignal: controller.signal }); ``` ```go ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() _, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt}) ``` ### zod / Standard Schema vs. pkg/schema TS accepts any [Standard Schema](https://standardschema.dev) value (zod, valibot, arktype) for `schema`/`inputSchema` and infers the TypeScript type from it. Go has no schema-inference library; `pkg/schema.Schema` (`pkg/schema/validator.go`) is a thin interface — `schema.NewSimpleJSONSchema(map[string]interface{}{...})` for a hand-written JSON Schema, or `ai.SchemaFor[T]()` (`pkg/ai/output.go`) to reflect one from a Go struct's `json` tags. `schema.ApplyDefaults` applies JSON Schema `default` values before validation, and a `$ref` cycle that never consumes input now returns a validation error instead of overflowing the stack — see [JSON Schema reference](https://goaisdk.com/docs/reference/schema/json-schema.md). ```typescript const weatherTool = tool({ description: 'Get the weather', inputSchema: z.object({ location: z.string() }), execute: async ({ location }) => fetchWeather(location), }); ``` ```go weatherTool := types.Tool{ Name: "get_weather", Description: "Get the weather", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{"location": map[string]interface{}{"type": "string"}}, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return fetchWeather(input["location"].(string)) }, } ``` ### Generics: typed tool input and Output Go does have generics, and the modern `Output` API uses them: `ai.ObjectOutput[T any]`, `ai.ArrayOutput[ELEMENT any]`, `ai.ChoiceOutput[CHOICE ~string]` (`pkg/ai/output.go`). The gap is narrower than "no generics" — it's that generics give you a type-safe **spec**, not a type-safe **result**: `GenerateTextResult.Output` is still `any`, so you type-assert once after the call. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate a lasagna recipe.", Output: ai.ObjectOutput[Recipe](ai.ObjectOutputOptions{ Schema: ai.SchemaFor[Recipe](), Name: "recipe", }), }) recipe, ok := result.Output.(Recipe) // one assertion, not compile-time inference ``` Tool `Execute` functions have no generic input type at all — `input` is always `map[string]interface{}`, unmarshaled by hand or via `pkg/schema`'s struct validator. ### StreamObject is not lazy TS's `streamObject` returns lazy streams immediately. Go's `StreamObject` blocks until the stream is fully consumed, reporting progress through `OnChunk` instead. Prefer `StreamText` with `Output` (`PartialOutput()` / `ai.ElementStreamWithOutput`), which streams incrementally in Go too and is also what TS recommends going forward, since `generateObject`/`streamObject` are deprecated in both SDKs. See [Known differences: StreamObject is not lazy](https://goaisdk.com/docs/migration-guides/known-differences.md#streamobject-is-not-lazy) for the full example. ### Option structs and pointers for optional fields TS distinguishes "not passed" from "passed as zero" with `?`/`undefined`. Go structs give every field a zero value, so any option where `0`/`""`/`false` is a meaningful value the caller might explicitly want uses a pointer: `MaxRetries *int`, `Temperature *float64`, `TopP *float64`, `Instructions *string`, `types.Tool.Strict *bool`, Anthropic's `DisableParallelToolUse *bool`. Fields where zero is never meaningful (`Model`, `Prompt`, `Tools`) stay plain values. ```go maxRetries := 0 // explicitly disable retries — nil would mean "use the default" ai.GenerateTextOptions{Model: model, Prompt: "...", MaxRetries: &maxRetries} ``` ### camelCase JSON on the wire vs. PascalCase Go fields Anything that crosses a wire — provider requests, UI message stream chunks — keeps camelCase JSON tags for compatibility, even though the Go field name is PascalCase: ```go // pkg/provider/types/message.go type TextContent struct { Text string `json:"text"` ProviderOptions map[string]interface{} `json:"providerOptions,omitempty"` } // pkg/provider/types/usage.go type Usage struct { InputTokens *int64 `json:"inputTokens,omitempty"` // ... } ``` One field genuinely renamed rather than just re-cased: TS `toolCall.toolCallId` is Go `ToolCall.ID` (`json:"id"`) on the internal result type — but the UI-message-stream wire type (`ToolUIPart`) does use `json:"toolCallId"` (`pkg/ai/ui_message.go`), because that one has to match the protocol a TS frontend expects. Option structs (`GenerateTextOptions` and friends) carry no JSON tags at all — they're never serialized, so there's nothing to keep camelCase for. The one notable exception: `anthropic.ModelOptions` (provider-specific options that do get marshaled into `ProviderOptions` maps and round-tripped) uses camelCase JSON tags (`budgetTokens`, `contextManagement`, `automaticCaching`, ...) as of v0.5.0 — v0.4.0 used snake_case. ### Callbacks vs. Go func fields Same shape, different syntax: TS object-literal callbacks become named `func` fields on the options struct. ```typescript generateText({ model, prompt, onStepFinish: (step) => console.log(step.text) }); ``` ```go ai.GenerateTextOptions{ Model: model, Prompt: prompt, OnStepEnd: func(ctx context.Context, step types.StepResult, userContext interface{}) { fmt.Println(step.Text) }, } ``` `GenerateTextOptions` has the cleanest 7.0-parity names: `OnStart`, `OnStepStart`, `OnStepEnd`, `OnEnd`, `OnToolExecutionStart`, `OnToolExecutionEnd` (`pkg/ai/generate.go`), with `OnFinish`/`OnStepFinish`/`OnToolCallStart` kept as deprecated aliases, same as TS. `StreamTextOptions` and `AgentConfig` haven't fully caught up: the structured step/finish events there are named `OnStepEndEvent`/`OnFinishEvent` rather than `OnStepEnd`/`OnEnd` (`pkg/ai/stream.go`, `pkg/agent/agent.go`), and `StreamTextOptions.OnEnd` itself is a separate, simpler callback (`func(result *StreamTextResult)`, no `ctx`) rather than the structured-event one. `Embed`/`EmbedMany`/`Rerank` are further behind still — their lifecycle callbacks are `ExperimentalOnStart`/`ExperimentalOnEnd` (`pkg/ai/embed.go`, `pkg/ai/rerank.go`) even though TS 7.0 renamed the equivalents to plain `onStart`/`onEnd`. Check the option struct you're actually using rather than assuming the name carries over. ### Error handling: thrown errors vs. returned errors TS throws typed error classes you catch with `instanceof`. Go returns `error` from every call and gives you two ways to check the type: an `Is*Error(err) bool` helper per error type, or `errors.As` directly against the concrete type (`pkg/provider/errors/errors.go`). ```go result, err := ai.GenerateText(ctx, opts) if err != nil { var apiErr *providererrors.ProviderError if errors.As(err, &apiErr) { log.Printf("provider %s: HTTP %d: %s", apiErr.Provider, apiErr.StatusCode, apiErr.Message) } else if providererrors.IsRateLimitError(err) { // back off and retry } return err } ``` ### experimental_ prefixes → Go naming Where TS 7.0 graduated an `experimental_` option to a stable name, Go uses the stable name too: `telemetry` (not `experimental_telemetry`), `output` (not `experimental_output`), `onStart`/`onEnd` on `generateText`/`streamText`. Where TS 7.0 has **not** graduated something, Go keeps the same prefix rather than getting ahead of upstream: `ExperimentalTransform` (`smoothStream` input), `ExperimentalRepairText`, the `Embed`/`Rerank` `ExperimentalOnStart`/`ExperimentalOnEnd` callbacks, and the newer additions this cycle — `ai.ExperimentalStartBatch`/`ExperimentalEvaluate`/`ExperimentalStartVideo`/`ExperimentalStreamTranscribe`/`ExperimentalToolCallers` — all mirror TS's own `experimental_` state rather than being stale Go naming. ### Dependency injection of fetch vs. http.Client TS provider factories take a `fetch` function to override the HTTP layer (proxies, custom TLS, test doubles). Go providers take an `*http.Client` on their `Config` struct instead. ```typescript createOpenAI({ apiKey, fetch: myCustomFetch }); ``` ```go openai.New(openai.Config{ APIKey: "sk-...", HTTPClient: myCustomClient, // *http.Client }) ``` ## Coming from TS v5 or v6 Go implements the current (v7) names only. If your TS code still uses a v5 or v6 name, here's where it landed and the matching Go name to reach for. See the TS SDK's own [v5](https://ai-sdk.dev/docs/migration-guides/migration-guide-5-0), [v6](https://ai-sdk.dev/docs/migration-guides/migration-guide-6-0), and [v7](https://ai-sdk.dev/docs/migration-guides/migration-guide-7-0) guides for the full TS-side detail. | TS v5/v6 name | Current TS (v7) name | Go name | |---|---|---| | `maxTokens` | `maxOutputTokens` | `GenerateTextOptions.MaxTokens *int` (Go kept the shorter name) | | `maxSteps` | `stopWhen: isStepCount(n)` | `StopWhen: []ai.StopCondition{ai.IsStepCount(n)}` | | `stepCountIs` (was the canonical name in v6) | `isStepCount` (canonical in v7; `stepCountIs` now a deprecated alias) | `ai.IsStepCount` (`ai.IsStepCount` exists too, as a deprecated alias) | | `Experimental_Agent` | `ToolLoopAgent` | `agent.ToolLoopAgent` / `agent.NewToolLoopAgent` | | `system` | `instructions` | `System string` (canonical in Go) / `Instructions *string` (alias) | | `experimental_telemetry` | `telemetry` | `Telemetry *telemetry.Options` | | `experimental_prepareStep` | `prepareStep` (the `experimental_` alias was removed, not just deprecated) | `PrepareStep` | | `experimental_output` + `result.experimental_output` | `output` + `result.output` (old names removed entirely in 7.0) | `Output` + `result.Output` / `result.Output()` | | `experimental_createProviderRegistry` | `createProviderRegistry` | `registry.NewRegistry` | | `experimental_customProvider` (removed entirely in 7.0) | `customProvider` | `registry.NewCustomProvider` | | `experimental_createMCPClient` | `createMCPClient` (moved to `@ai-sdk/mcp`) | `mcp.CreateStdioMCPClient` / `CreateHTTPMCPClient` / `CreateSSEMCPClient` | | `onFinish` / `onStepFinish` | `onEnd` / `onStepEnd` | `OnEnd` / `OnStepEnd` (Go keeps `OnFinish`/`OnStepFinish` as deprecated aliases too) | | `convertToModelMessages()` (sync in v5) | `async convertToModelMessages()` | `ai.ConvertToModelMessages(ctx, messages)` returns `([]types.Message, error)` — always "awaited" | | `CoreMessage` | `ModelMessage` | `types.Message` | | tool `parameters` | `inputSchema` | `types.Tool.Parameters` (Go kept the v5 field name) | | tool `args` / `result` (call site) | `input` / `output` | `types.ToolCall.Arguments` / `types.ToolResult.Result` — Go did not rename the result field to `Output`; the JSON tag is still `"result"` | | `ToolCallOptions` | `ToolExecutionOptions` | `types.ToolExecutionOptions` | | system messages allowed inline in `prompt`/`messages` | rejected by default; opt in with `allowSystemInMessages` | `AllowSystemMessages` / `AllowSystemInMessages bool` | | `fullStream` | `stream` (`fullStream` kept as a deprecated alias) | `StreamTextResult.FullStream()` and `.Stream()` both exist and return the same thing | ## Not available in Go These are TS-only by design — the Go SDK is a backend/CLI library, not a UI framework, and doesn't ship a browser runtime: - **React/Vue/Angular/Svelte hooks** (`useChat`, `useCompletion` from `@ai-sdk/react`, `@ai-sdk/vue`, `@ai-sdk/angular`, `@ai-sdk/svelte`) — client-side UI state management. Go serves the same UI message stream protocol these hooks consume; see [Serving a TS frontend](#serving-a-ts-frontend-from-a-go-backend). - **React Server Components streaming** (`@ai-sdk/rsc`) — an RSC-specific streaming mechanism with no Go analogue. - **`@ai-sdk/devtools`** — a browser-based UI for inspecting LLM requests/responses. Go's own [debugging guide](https://goaisdk.com/docs/troubleshooting/debugging.md) covers the same goal — wrapping the HTTP client to log requests/responses, plus `ai.SetLogWarnings` for provider warnings — by different means. - **`@ai-sdk/codemod`** — jscodeshift-based automated upgrade transforms for TS source files; doesn't apply to Go source. - **`@ai-sdk/harness-cline`, `@ai-sdk/harness-pi`** — these load their vendor Node SDKs in-process with no TS bridge script to port; a Go port would need original work with no TS reference to mirror. Revisit on demand. - **Browser-only realtime transport** (`BrowserRealtimeTransport`, `BrowserRealtimeAudio`, `useRealtime`, OpenAI Live over WebRTC) — WebRTC/browser-audio glue. Go's [Realtime Sessions](https://goaisdk.com/docs/ai-sdk-core/realtime.md) support the same server-side WebSocket session API (`ai.ConnectRealtime`); only the in-browser/WebRTC transport is out of scope. - **StreamObject as a lazy stream** — see [StreamObject is not lazy](#streamobject-is-not-lazy) above; this is a behavior difference, not a missing feature. - **Durable webhook video suspension inside Vercel Workflow DevKit** — depends on the Node-only Workflow DevKit runtime. Go's video generation polls until the job is done, or you drive it yourself with `ai.ExperimentalStartVideo`/`ExperimentalGetVideoStatus` (which do support webhooks, just not durable workflow suspension). See [Known differences from the TypeScript AI SDK](https://goaisdk.com/docs/migration-guides/known-differences.md) for the complete, current list, including narrower runtime-level differences (Vercel Sandbox OIDC refresh, code-mode approval batching). ## Recently landed (previously tracked as gaps) As of v0.5.0, all of the following shipped and are no longer gaps: the **Batch API** (`ai.ExperimentalStartBatch` and friends), **`evaluate`** (`ai.ExperimentalEvaluate`), **async video jobs** (`ai.ExperimentalStartVideo`/`ExperimentalGetVideoStatus`), **streaming transcription** (`ai.ExperimentalStreamTranscribe`, plus `ai.ExperimentalStreamTranslate` for speech translation), and **tool search / tool callers** (`ai.ToolSearch`, `types.Tool.DeferLoading`, `ExperimentalToolCallers`). See the mapping tables above for each. ## Serving a TS frontend from a Go backend A `useChat` frontend posts `{ messages: UIMessage[] }` and expects a UI message stream back. The Go AI SDK writes that stream, so the TypeScript frontend needs no changes. This handler runs an agent on the incoming messages and streams the reply, the way TS `createAgentUIStreamResponse` does: ```go package main import ( "encoding/json" "log" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { model, err := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}). LanguageModel(openai.ModelGPT6Astra) if err != nil { log.Fatal(err) } assistant := agent.NewToolLoopAgent(agent.AgentConfig{Model: model}) http.HandleFunc("POST /api/chat", func(w http.ResponseWriter, r *http.Request) { var req struct { Messages json.RawMessage `json:"messages"` } if err := json.NewDecoder(r.Body).Decode(&req); err != nil { http.Error(w, err.Error(), http.StatusBadRequest) return } // Validates the UI messages and converts them to model messages. chunks, _, err := agent.CreateAgentUIStreamFromUIMessages(r.Context(), assistant, agent.CreateAgentUIStreamFromUIMessagesOptions{UIMessages: []byte(req.Messages)}) if err != nil { http.Error(w, err.Error(), http.StatusBadRequest) return } // Sets the stream headers, writes 200 and flushes every chunk. if err := ai.PipeUIMessageChunksToResponse(chunks, w, nil); err != nil { log.Printf("write stream: %v", err) } }) log.Fatal(http.ListenAndServe(":8080", nil)) } ``` On an `http.ResponseWriter`, the `Pipe*` helpers set the status and the UI message stream headers (`Content-Type: text/event-stream`, `X-Vercel-AI-UI-Message-Stream: v1` and so on) and flush after every chunk, like `pipeUIMessageStreamToResponse` on a Node `ServerResponse`. Do not copy bytes into the writer with `io.Copy`: a `ResponseWriter` does not flush on each write, so the browser would see the output in large blocks. If you already have a `*ai.StreamTextResult` and no UI messages, use `ai.PipeUIMessageStreamToResponse(ctx, result, w)`. `agent.PipeAgentUIStreamFromUIMessagesToResponse` does the validate-and-pipe steps in one call. For keep-alive, CORS, proxies, custom data parts and tool approval, see [Serve a useChat frontend from Go](https://goaisdk.com/docs/build-a-chat-app/serve-usechat-from-go.md) and [Tool approval end to end](https://goaisdk.com/docs/build-a-chat-app/tool-approval.md). ## See Also - [Migrating from v0.4.x to v0.5.0](https://goaisdk.com/docs/migration-guides/from-v0.4-to-v0.5.md) - [Known differences from the TypeScript AI SDK](https://goaisdk.com/docs/migration-guides/known-differences.md) --- # Migrating from v0.1.x to v0.2.0 > Breaking changes in Go AI SDK v0.2.0, the TypeScript AI SDK v6.0 parity release. Canonical URL: https://goaisdk.com/docs/migration-guides/from-v0.1-to-v0.2 Documentation index: https://goaisdk.com/llms.txt Go AI SDK v0.2.0 aligned the public API with TypeScript AI SDK v6.0: pointer-based usage tracking, a `ToolExecutionOptions` parameter on tool execution, request-scoped context in callbacks, and a typed `Output` system. This guide is for anyone still on v0.1.0. If you're upgrading from a later version instead, see [migrating from v0.3.x to v0.4.0](https://goaisdk.com/docs/migration-guides/from-v0.3-to-v0.4.md) or [migrating from v0.4.x to v0.5.0](https://goaisdk.com/docs/migration-guides/from-v0.4-to-v0.5.md), which cover the `RuntimeContext`/`ToolsContext` split and other renames introduced after this release. ## Recommended migration process 1. Update: `go get github.com/digitallysavvy/go-ai@v0.2.0` 2. Update `Usage` field access to handle `nil`, or switch to the `Get*Tokens()` helpers. 3. Add the `opts types.ToolExecutionOptions` parameter to every tool `Execute` function. 4. Add `ctx context.Context` and `userContext interface{}` parameters to `OnStepFinish`/`OnFinish` callbacks on `GenerateText` and `GenerateObject`. 5. Build and test: `go build ./...` and `go test ./...`. ## Breaking changes ### Usage fields are pointers **Impact:** breaking `Usage.InputTokens`, `OutputTokens`, and `TotalTokens` changed from plain integers to `*int64`, so `nil` can mean "not reported" instead of `0`. `Usage` also gained `InputDetails`, `OutputDetails`, and `Raw`. **Before (v0.1.0 — no longer compiles):** ```go skip-compile usage := types.Usage{ InputTokens: 10, OutputTokens: 20, TotalTokens: 30, } if result.Usage.TotalTokens > 0 { fmt.Printf("Used %d tokens\n", result.Usage.TotalTokens) } ``` **After (v0.2.0):** ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello", }) if err != nil { log.Fatal(err) } if result.Usage.GetTotalTokens() > 0 { fmt.Printf("Used %d tokens\n", result.Usage.GetTotalTokens()) } ``` `GetInputTokens()`, `GetOutputTokens()`, and `GetTotalTokens()` return `0` for `nil`, so most call sites only need this one change. ### ToolExecutor takes a ToolExecutionOptions parameter **Impact:** breaking `types.ToolExecutor` gained a third parameter carrying the tool call ID, accumulated usage, and any value passed through `ExperimentalContext`. **Before (v0.1.0 — no longer compiles):** ```go weatherTool := types.Tool{ Name: "weather", Description: "Get weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{"type": "string"}, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}) (interface{}, error) { location := input["location"].(string) return map[string]interface{}{"temperature": 72, "condition": "sunny"}, nil }, } ``` **After (v0.2.0):** ```go weatherTool := types.Tool{ Name: "weather", Description: "Get weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{"type": "string"}, }, "required": []string{"location"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { location := input["location"].(string) log.Printf("executing tool call %s for %s", opts.ToolCallID, location) return map[string]interface{}{"temperature": 72, "condition": "sunny"}, nil }, } ``` ### OnStepFinish and OnFinish add ctx and userContext **Impact:** breaking `GenerateTextOptions.OnStepFinish`, `OnFinish`, and the matching callbacks on `GenerateObjectOptions` now take a leading `ctx context.Context` and a trailing `userContext interface{}`. `userContext` is whatever you pass through the new `ExperimentalContext` option. **`StreamTextOptions.OnFinish` did not change** — it still takes only `*StreamTextResult`. Don't add parameters to it. **Before (v0.1.0 — no longer compiles):** ```go skip-compile result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Solve this problem", Tools: tools, OnStepFinish: func(step types.StepResult) { fmt.Printf("Step %d complete\n", step.StepNumber) }, }) ``` **After (v0.2.0):** ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Solve this problem", Tools: tools, ExperimentalContext: map[string]interface{}{ "userId": "user123", }, OnStepFinish: func(ctx context.Context, step types.StepResult, userContext interface{}) { fmt.Printf("Step %d complete\n", step.StepNumber) }, OnFinish: func(ctx context.Context, result *ai.GenerateTextResult, userContext interface{}) { if data, ok := userContext.(map[string]interface{}); ok { log.Printf("user %v finished", data["userId"]) } }, }) ``` `OnStepFinish` and `OnFinish` are now deprecated aliases for `OnStepEnd` and `OnEnd`, added in a later release — both spellings take the same arguments. ## New in v0.2.0 These are additive; existing code keeps compiling without them. **Detailed usage.** `Usage.InputDetails` and `OutputDetails` expose cache and reasoning token counts when a provider reports them; `Usage.Raw` carries the provider's raw usage payload. ```go if result.Usage.InputDetails != nil && result.Usage.InputDetails.CacheReadTokens != nil { fmt.Printf("cache read tokens: %d\n", *result.Usage.InputDetails.CacheReadTokens) } ``` **New tool fields.** `Tool.Title`, `Strict`, `NeedsApproval`, and `InputExamples` (`[]types.ToolInputExample`, not `[]map[string]interface{}`): ```go weatherTool := types.Tool{ Name: "weather", Title: "Weather Lookup", Strict: types.BoolPtr(true), InputExamples: []types.ToolInputExample{ {Input: map[string]interface{}{"location": "San Francisco, CA"}}, }, // Parameters, Execute, ... } ``` **Typed `Output` system.** `ai.ObjectOutput[T]`, `ai.ArrayOutput[T]`, and `ai.ChoiceOutput[T]` build an `Output` value for `GenerateTextOptions.Output` — an alternative to `ai.GenerateObject` that returns a typed result directly from `GenerateText`. Note that the options types are themselves generic: ```go tasksOutput := ai.ArrayOutput[Task](ai.ArrayOutputOptions[Task]{ ElementSchema: taskSchema, Name: "tasks", }) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Generate 3 tasks", Output: tasksOutput, }) if err != nil { log.Fatal(err) } tasks := result.Output.([]Task) ``` ## Common compile errors | Error | Fix | |---|---| | `cannot use 10 (untyped int constant) as *int64 value in struct literal` | Take the address of a variable: `n := int64(10); usage.InputTokens = &n`. | | `cannot use func(ctx context.Context, input map[string]interface{}) (interface{}, error) {…} as types.ToolExecutor value` | Add the third parameter: `func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error)`. | | `cannot use func(step types.StepResult) {…} as func(ctx context.Context, step types.StepResult, userContext interface{}) value` | Add `ctx context.Context` and `userContext interface{}` to the callback signature. | | `unknown field Output in struct literal of type ai.GenerateObjectOptions` | `Output` is a field on `GenerateTextOptions` (the new typed-output system), not `GenerateObjectOptions`. `GenerateObjectOptions` still takes `Schema`. | | `cannot use generic type ai.ArrayOutputOptions[ELEMENT any] without instantiation` | Supply the type parameter: `ai.ArrayOutputOptions[Task]{...}` (same for `ChoiceOutputOptions[T]`). | --- # Migrating from v0.3.x to v0.4.0 > Runtime behavior changes in Go AI SDK v0.4.0 — deferred streaming tool execution, xAI's Responses API default, and stricter MCP redirect handling. Canonical URL: https://goaisdk.com/docs/migration-guides/from-v0.3-to-v0.4 Documentation index: https://goaisdk.com/llms.txt Go AI SDK v0.4.0 changes when streaming tools execute, switches xAI's default API, removes two retired xAI model IDs, and rejects MCP server redirects by default. None of these rename or remove a Go type or function signature, so v0.3.x code keeps compiling — what changes is runtime behavior. This guide is for anyone still on v0.3.x. ## Recommended migration process 1. Update: `go get github.com/digitallysavvy/go-ai@v0.4.0` 2. If a `StreamText` consumer expects tool-result chunks interleaved with text, move that handling to run after the stream ends. 3. Replace `grok-2` and `grok-2-vision-1212` model IDs with a `grok-3` or later model. 4. If you rely on xAI's Chat Completions API, call `ChatCompletionsLanguageModel()` explicitly. 5. If an MCP server needs redirects, opt in with `Redirect: mcp.MCPRedirectFollow`. 6. Build and test: `go build ./...` and `go test ./...`. ## Behavior changes ### Streaming tool execution moves to after the stream **Impact:** runtime only — no compile change. `StreamText` forwards every provider chunk (text, reasoning, custom content) to `OnChunk` before running any tool's `Execute` callback; `ChunkTypeToolResult` chunks now arrive only after the stream ends. In v0.3.x they could interleave with text chunks. This applies to tools with a Go `Execute` callback — provider-executed results (`ProviderExecuted: true`, e.g. Google code execution) still arrive inline as the provider produces them. ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "What's the weather?", Tools: []types.Tool{weatherTool}, // StreamText defaults to a single step; allow more so the model gets // a turn to respond in text after the tool result. StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, OnChunk: func(chunk provider.StreamChunk) { fmt.Print(chunk.Text) }, }) if err != nil { log.Fatal(err) } defer result.Close() for range result.Chunks() { } ``` This code is unchanged between v0.3.x and v0.4.0. What changes is the order `OnChunk` sees: in v0.3.x a `ChunkTypeToolResult` chunk could arrive between text chunks; in v0.4.0 it always arrives after the stream's text and reasoning chunks have finished. ### XAI's default language model uses the Responses API **Impact:** runtime only — no compile change. `xai.Provider.LanguageModel()` now returns a model backed by xAI's Responses API instead of Chat Completions. The Responses API adds image input and finer-grained reasoning effort levels. Call `ChatCompletionsLanguageModel()` for the v0.3.x behavior. > **v0.5.0:** `ChatCompletionsLanguageModel()` has since been removed (the TS SDK dropped xAI's Chat Completions API). On v0.5.0, `LanguageModel()` is the only xAI language model. See [Migrating from v0.4 to v0.5](https://goaisdk.com/docs/migration-guides/from-v0.4-to-v0.5.md). ```go provider := xai.New(xai.Config{APIKey: os.Getenv("XAI_API_KEY")}) model, _ := provider.LanguageModel("grok-3") // Responses API (new default) // v0.4 only, removed in v0.5.0: Chat Completions API (v0.3.x behavior) // chatModel, _ := provider.ChatCompletionsLanguageModel("grok-3") ``` ### XAI removed grok-2 model IDs **Impact:** runtime only. `LanguageModel()` never validated model ID strings at compile time or call time, so this doesn't produce a Go error either — the request reaches xAI, and xAI rejects the retired model. ```go model, _ := provider.LanguageModel("grok-2") // xAI: model not found model, _ := provider.LanguageModel("grok-2-vision-1212") // xAI: model not found // Use instead: model, _ := provider.LanguageModel("grok-3") model, _ := provider.LanguageModel("grok-3-mini") ``` ### MCP HTTP transport rejects redirects by default **Impact:** runtime only — no compile change; may break an MCP integration that relied on a redirect. `MCPRedirectMode`'s zero value is `MCPRedirectError`, so `mcp.NewHTTPTransport` now rejects a redirect instead of following it. This closes an SSRF path where a malicious MCP server redirects requests to an internal service. ```go transport := mcp.NewHTTPTransport(mcp.HTTPTransportConfig{ URL: "https://mcp.example.com", }) // v0.4.0: a redirect from this server now returns an error. // To allow redirects (only for a server you trust): transport := mcp.NewHTTPTransport(mcp.HTTPTransportConfig{ URL: "https://mcp.example.com", Config: mcp.TransportConfig{ Redirect: mcp.MCPRedirectFollow, }, }) ``` ### Telemetry: register an integration instead of passing a Tracer per call **Impact:** low. `telemetry.Options.Tracer` still works; `RegisterTelemetryIntegration` is the supported path for new code. ```go telemetry.RegisterTelemetryIntegration(telemetry.OTelTelemetryIntegration{}) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello", Telemetry: &telemetry.Options{ IsEnabled: telemetry.Bool(true), RecordInputs: true, FunctionID: "my-chat", }, }) ``` If you relied on the default global tracer rather than a custom `Tracer`, register `telemetry.OTelTelemetryIntegration{}` — it uses `otel.Tracer("ai-sdk")` when no custom tracer is set. `AddTelemetryIntegration()` fans out to multiple backends, `ClearTelemetryIntegrations()` resets the registry, and `telemetry.NoopTelemetryIntegration{}` disables telemetry globally. ## New in v0.4.0 These are additive; existing code keeps compiling without them. - **Top-level `Reasoning`** on `GenerateTextOptions`/`StreamTextOptions` — see [Generating text](https://goaisdk.com/docs/ai-sdk-core/generating-text.md) - **Deferred provider tool results** via `Tool.SupportsDeferredResults` — see [Tools and tool calling](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - **`CustomContent` and `ReasoningFileContent`** content types - **Tool-level timeouts** via `TimeoutConfig.ToolMs` / `TimeoutConfig.Tools` - **Embed/rerank `OnStart`/`OnFinish` callbacks** — see [Embeddings](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) - **[KlingAI v3.0 motion control](https://goaisdk.com/docs/providers/klingai.md)** — new video generation capability - **[Prodia](https://goaisdk.com/docs/providers/prodia.md) language and video models** — Prodia's existing image generation provider gains a language model and a video model --- # Migrating from v0.4.x to v0.5.0 > Migration guide for upgrading Go AI SDK apps from v0.4.x to v0.5.0, covering runtime context changes, tool approval, provider updates, and breaking changes. Canonical URL: https://goaisdk.com/docs/migration-guides/from-v0.4-to-v0.5 Documentation index: https://goaisdk.com/llms.txt v0.4.0 is the last tagged release, so this guide covers everything that shipped after it: the May/June 2026 cycle (splitting runtime context from tool context, call-level tool approval, the provider v4 tagged-union file data shape, MCP Apps and Responses tool options) and the September 2026 parity cycle (removing the xAI Chat Completions API, rebuilding Bedrock-Anthropic and Bedrock on shared code paths, changing `StreamText`'s full-stream chunk lifecycle, and a long list of provider-specific parity fixes). Both cycles land together in v0.5.0. TS SDK parity now tracks `ai@7.0.127`. This guide is organized by area rather than by cycle, since most applications only care about what changed in the areas they use. See [Known differences from the TypeScript AI SDK](https://goaisdk.com/docs/migration-guides/known-differences.md) for behavior that stays intentionally different, and the repository's `release_notes/` and `CHANGELOG.md` for the complete, cycle-by-cycle list this guide is built from. ## Recommended migration process 1. Make sure you build with **Go 1.26 or later**. v0.5.0's `go.mod` declares `go 1.26.0`, up from 1.25 in v0.4.x, because Go 1.25 is end-of-life and the current `golang.org/x/*` modules require Go 1.26. 2. Upgrade on a clean branch and run `go vet ./...` and `go test ./...` before editing. 3. Search your codebase for `xai.ChatCompletionsLanguageModel`, `xai.NewLanguageModel`, and Bedrock `CacheConfig` usage — these no longer compile. 4. Search for `Strict: true` (tools), `DisableParallelToolUse:` (Anthropic), and `MaxRetries:` (Embed/EmbedMany/Rerank) — these fields changed from value to pointer types. 5. Search for `ExperimentalContext`, `NeedsApproval`, `ToolNeedsApproval`, and `types.ImageContent` — these still compile as deprecated aliases, but new code should move to `RuntimeContext`/`ToolsContext`, call-level `ToolApproval`, and `types.FileContent`/`FileData`. 6. If your code or wrappers still use TypeScript-style `experimental_*` option names (`experimental_output`, `experimental_activeTools`, `experimental_prepareStep`), map them to `Output`, `FilterActiveTools`, and `PrepareStep`. 7. If you read/write `anthropic.ModelOptions` as JSON (config files, serialized state), update field names to camelCase. 8. If you detect step boundaries in `StreamText`'s full stream by watching for `provider.ChunkTypeFinish`, switch to `provider.ChunkTypeFinishStep`. 9. If you set `telemetry.Options.Tracer` per call, move the tracer to a registered `telemetry.TelemetryIntegration` instead. 10. Move multi-step limits from `MaxSteps` to `StopWhen` where practical (`MaxSteps` still works as a compatibility shorthand). 11. Re-run `go vet ./...` and `go test ./...`. For examples that use `//go:build ignore`, compile the specific example file you changed, e.g. `go build ./examples/upload-file/provider-reference/main.go`. ## Breaking changes ### Runtime context, tool execution, and control flow #### Runtime and tool context are split **Impact:** breaking (in spirit) / low (in practice — the old field still compiles as a deprecated alias). `ExperimentalContext` and generic `Context`-style values are replaced by `RuntimeContext`. Per-tool state belongs in `ToolsContext`, keyed by tool name. Tool execution receives `ToolExecutionOptions.RuntimeContext` and `ToolExecutionOptions.ToolContext`. `ExperimentalContext` remains as a deprecated alias field on `GenerateTextOptions` for now. **Before:** ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Check the weather.", ExperimentalContext: map[string]interface{}{"user_id": "u_123", "city": "Paris"}, Tools: []types.Tool{weatherTool}, }) ``` **After:** ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Check the weather.", RuntimeContext: map[string]interface{}{"user_id": "u_123"}, ToolsContext: map[string]interface{}{ "weather": map[string]interface{}{"city": "Paris"}, }, Tools: []types.Tool{weatherTool}, }) ``` #### Tool execution options replace tool call options **Impact:** medium Tool executors receive `types.ToolExecutionOptions`. Use `ToolCallID`, `RuntimeContext`, and `ToolContext` from that struct. Do not depend on removed `ToolCallOptions` naming in examples or wrappers. **Before:** ```go Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return lookup(input, opts.UserContext), nil } ``` **After:** ```go Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return lookup(input, opts.RuntimeContext, opts.ToolContext), nil } ``` #### Tool approval is call-level **Impact:** breaking (in spirit) / low (in practice — `NeedsApproval` still compiles as a deprecated alias). `ToolNeedsApproval` was renamed to `ToolApproval`. Per-tool `NeedsApproval` is deprecated and is only used when no call-level approval is configured. `ToolApproval` can be a single function or a per-tool map with `approved`, `denied`, or `user-approval`. **Before:** ```go deleteTool := types.Tool{ Name: "delete_account", Description: "Delete an account.", Parameters: schema.NewSimpleJSONSchema(map[string]interface{}{"type": "object"}), NeedsApproval: true, Execute: deleteAccount, } ``` **After:** ```go reason := "account deletion requires review" result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Delete account acct_123.", Tools: []types.Tool{deleteTool}, ToolApproval: map[string]ai.ToolApprovalValue{ "delete_account": ai.ToolApprovalResult{ Status: ai.ToolApprovalStatusUserApproval, Reason: &reason, }, }, }) ``` #### Tool calls: invalid calls are marked, not silently dropped or executed **Impact:** behavior change for tool-call error handling, including `WorkflowAgent`'s default path (September cycle). Calls to unknown tools, or calls with input that fails schema validation, are now marked invalid (TS `parseToolCall`) instead of being silently skipped or executed with bad input. Tool calls made without any tools configured are also invalid. An invalid call is not executed — it carries a `NoSuchToolError` (unknown tool) or `InvalidToolInputError` (invalid input) result, unless a configured `RepairToolCall` fixes it first. #### Experimental generation options were removed **Impact:** breaking Use stable names for output, active tools, and step preparation. If your code or wrappers still expose TypeScript-style experimental option names, map `experimental_output` to `Output`, `experimental_activeTools` to `FilterActiveTools` or provider `ToolChoice`, and `experimental_prepareStep` to `PrepareStep`. **After:** ```go opts := ai.GenerateTextOptions{ Model: model, Prompt: "Return JSON.", Output: ai.JSONOutput(ai.JSONOutputOptions{}), PrepareStep: func(ctx context.Context, step ai.PrepareStepOptions) ai.PrepareStepOptions { return step }, } ``` #### FilterActiveTools is the stable helper **Impact:** low `ExperimentalFilterActiveTools` remains only as a compatibility alias. Use `FilterActiveTools` in new code. **Before:** ```go active := ai.ExperimentalFilterActiveTools(tools, []string{"search"}) ``` **After:** ```go active := ai.FilterActiveTools(tools, []string{"search"}) ``` #### MaxSteps is deprecated **Impact:** medium `MaxSteps *int` is still accepted as a compatibility shorthand, but `StopWhen` is the stable multi-step control. If both are set, `StopWhen` wins. **Before:** ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Use tools if needed.", Tools: tools, MaxSteps: intPtr(5), }) ``` **After:** ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Use tools if needed.", Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) ``` #### Tools: Strict is now *bool **Impact:** compile error if you set `Strict` as a plain bool (September cycle). `types.Tool.Strict` and Open Responses `FunctionTool.Strict` changed from `bool` to `*bool` (tri-state, TS parity): `nil` leaves strict mode unset, and an explicit `false` is now forwarded to providers (previously indistinguishable from unset). **Before:** ```go skip-compile tool := types.Tool{Name: "search", Strict: true, /* ... */} ``` **After:** ```go tool := types.Tool{Name: "search", Strict: types.BoolPtr(true), /* ... */} ``` #### StreamText is asynchronous **Impact:** breaking for code that only checks `StreamText`'s returned error (September cycle). `StreamText` now returns before the first model request, matching TS `streamText`. Only option-validation errors (nil model, invalid `MaxRetries`, invalid `InstructionMessages`) are returned from the call itself. Everything else — including prompt errors, approval-signature failures on resume, and failures starting the first provider stream — surfaces through `Err()` / `ReadAll()` / `Stream()` / `Chunks()`. **Before:** ```go result, err := ai.StreamText(ctx, opts) if err != nil { // This used to catch prompt errors and first-request failures too. log.Fatal(err) } ``` **After:** ```go result, err := ai.StreamText(ctx, opts) if err != nil { // Only option-validation errors land here now. log.Fatal(err) } defer result.Close() text, err := result.ReadAll() if err != nil { // Prompt errors, approval failures, and first-request failures land here. log.Fatal(err) } ``` #### StreamText full stream: chunk lifecycle changed **Impact:** breaking for code that used `ChunkTypeFinish` to detect step boundaries (September cycle). In v0.4.0, `provider.ChunkTypeFinish` was forwarded once per step. It is now emitted exactly once per call, carrying total usage across all steps. Each step is now bracketed by `provider.ChunkTypeStartStep` / `provider.ChunkTypeFinishStep` (carrying that step's own request, response, usage, finish reason, and provider metadata), and `provider.ChunkTypeStart` opens the stream once, before the first step. `OnChunk` receives all of these chunks. **Before:** ```go for chunk := range result.Chunks() { if chunk.Type == provider.ChunkTypeFinish { // Used to fire once per step. recordStep(chunk) } } ``` **After:** ```go for chunk := range result.Chunks() { switch chunk.Type { case provider.ChunkTypeFinishStep: recordStep(chunk) // fires once per step, as before case provider.ChunkTypeFinish: recordCallTotals(chunk) // fires once per call, with total usage } } ``` See [StreamText reference](https://goaisdk.com/docs/reference/ai/stream-text.md#full-stream-lifecycle) for the complete chunk-ordering description. ### Telemetry #### Telemetry: tracers belong to integrations **Impact:** medium, compile error if you set a tracer per call. `telemetry.Options.Tracer` and `WithTracer` are removed. Configure the tracer on the integration and register it once. `OTelTelemetryIntegration` is renamed `LegacyOpenTelemetry` (the old name remains as a deprecated alias), and `telemetry.NewOpenTelemetry` adds a GenAI semantic-convention integration (September cycle). `TelemetryIntegration.OnLanguageModelCallStart` now returns a `context.Context` — the model call runs in the returned context, so its span becomes the parent of provider work. Custom integrations must update this method to return a context. **Before:** ```go skip-compile Telemetry: &telemetry.Options{IsEnabled: telemetry.Bool(true), Tracer: tracer} ``` **After:** ```go telemetry.RegisterTelemetryIntegration( telemetry.NewLegacyOpenTelemetry(telemetry.LegacyOpenTelemetryOptions{Tracer: tracer}), ) // per call: Telemetry: &telemetry.Options{IsEnabled: telemetry.Bool(true)} ``` #### Telemetry names are stable **Impact:** low `ExperimentalTelemetry` is deprecated. Use `Telemetry`. `telemetry.Settings` remains as a compatibility alias for `telemetry.Options`, but new code should use `telemetry.Options`. Telemetry is active by default when at least one global or local integration is registered; set `IsEnabled: telemetry.Bool(false)` to disable a call. **Before:** ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Summarize the changelog.", ExperimentalTelemetry: &telemetry.Settings{ IsEnabled: telemetry.Bool(true), }, }) ``` **After:** ```go telemetry.RegisterTelemetryIntegration(telemetry.NewLegacyOpenTelemetry(telemetry.LegacyOpenTelemetryOptions{})) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Summarize the changelog.", Telemetry: &telemetry.Options{ FunctionID: "release-summary", }, }) ``` Also carried over from the May/June cycle: telemetry lifecycle method names are now `OnToolExecutionStart` / `OnToolExecutionEnd` on `TelemetryIntegration` (previously different naming). #### Callbacks / events: deprecated tool-execution fields excluded from JSON **Impact:** low — fields still exist but are no longer serialized (September cycle). Tool execution event fields `StepNumber`, `ModelProvider`, `ModelID`, `Args`, `Result`, `Error`, `DurationMs` are deprecated and excluded from JSON marshaling. If you serialize these events, read the replacement fields documented on the event types instead. ### Messages, files, and prompts #### File data is a tagged union **Impact:** breaking (in spirit) / low (in practice — `types.ImageContent` still works as a compatibility input). Provider v4 file parts use `types.FileData` with a `Type` discriminator. `types.ImageContent` still works as a compatibility input, but new multimodal code should use `types.FileContent` with `FileDataTypeData`, `FileDataTypeURL`, `FileDataTypeReference`, or `FileDataTypeText`. **Before:** ```go messages := []types.Message{{ Role: types.RoleUser, Content: []types.ContentPart{ types.ImageContent{ URL: "https://example.com/chart.png", MimeType: "image/png", }, }, }} ``` **After:** ```go messages := []types.Message{{ Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Explain this chart."}, types.FileContent{ FileData: types.FileData{ Type: types.FileDataTypeURL, URL: "https://example.com/chart.png", MediaType: "image/png", }, }, }, }} ``` #### System messages in Messages are rejected by default **Impact:** breaking Put system instructions in `GenerateTextOptions.System` or `StreamTextOptions.System`. If you intentionally need provider-native system-role messages in `Messages`, opt in with `AllowSystemMessages` or the TypeScript-compatible alias `AllowSystemInMessages`. **Before:** ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ {Role: types.RoleSystem, Content: []types.ContentPart{types.TextContent{Text: "Be brief."}}}, {Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "Hello"}}}, }, }) ``` **After:** ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "Be brief.", Messages: []types.Message{ {Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "Hello"}}}, }, }) ``` **Opt-in:** ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, AllowSystemMessages: true, Messages: providerNativeMessages, }) ``` #### Upload APIs use provider capabilities **Impact:** medium Use `ai.UploadFile` and `ai.UploadSkill` with a provider that exposes `Files()` or `Skills()`. Upload results return `ProviderReference`, which can be passed back as `types.FileDataTypeReference`. **Before:** ```go // Provider-specific upload helper or manual HTTP request. ``` **After:** ```go uploaded, err := ai.UploadFile(ctx, ai.UploadFileOptions{ API: openaiProvider, Data: []byte("hello"), Filename: "hello.txt", MediaType: "text/plain", }) if err != nil { return err } file := types.FileContent{ FileData: types.FileData{ Type: types.FileDataTypeReference, Reference: uploaded.ProviderReference, }, } ``` ### Workflow #### Models can be serialized for workflows **Impact:** medium Models that implement `provider.SerializableModel` can cross workflow boundaries. Use `provider.SerializeModel` and `provider.DeserializeModel`; provider deserializers must be registered before deserialization. Static JSON-compatible provider headers are preserved during serialization, matching the TypeScript SDK workflow behavior. API keys, HTTP clients, functions, and other non-serializable values are omitted; pass replacement auth through the provider factory or environment when restoring a workflow. **Before:** ```go workflowState.Model = model ``` **After:** ```go serialized, err := provider.SerializeModel(model) if err != nil { return err } restored, err := provider.DeserializeModel(serialized) ``` As of the September cycle, this same `Serialize*`/`Deserialize*`/ `Register*Deserializer` triplet exists for every other model kind a provider can serialize (image, video, speech, transcription, embedding, evaluation), for every provider model TS marks `WORKFLOW_SERIALIZE`. See [WorkflowAgent: Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization) for the full list. None of this is wired in automatically — your application calls these functions explicitly. #### Workflow: transport renames and per-call context precedence **Impact:** breaking for code using the old transport name; behavior change on approval resume (September cycle). `workflow.WorkflowChatTransport` is now the client-side `ai.ChatTransport` (the Go port of TS `WorkflowChatTransport`, built with `workflow.NewWorkflowChatTransport`). The previous server-side run multiplexer (`ServeHTTP` / `Resume`, SSE events) is renamed `workflow.WorkflowRunMultiplexer` (`workflow.NewWorkflowRunMultiplexer`). On approval resume, per-call runtime/tools context and sandbox settings now override the agent's defaults (previously the agent defaults always won). **Before:** ```go skip-compile transport := &workflow.WorkflowChatTransport{} http.Handle("/api/workflow", transport) ``` **After:** ```go multiplexer := workflow.NewWorkflowRunMultiplexer() http.Handle("/api/workflow", multiplexer) // Client-side transport is now a separate type: clientTransport := workflow.NewWorkflowChatTransport(workflow.WorkflowChatTransportOptions{ API: "https://example.com/api/chat", }) ``` ### xAI #### xAI Chat Completions API removed **Impact:** compile error if you use it. `xai.Provider.ChatCompletionsLanguageModel()`, `xai.NewLanguageModel`, and the Chat Completions `SearchParameters` option are removed, matching the TypeScript SDK. `LanguageModel()` (Responses API) is the only xAI language model. For live search, use the provider-executed `xai.WebSearch` / `xai.XSearch` tools instead of `SearchParameters`. **Before:** ```go model, _ := provider.ChatCompletionsLanguageModel("grok-3") result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What happened in the news today?", ProviderOptions: map[string]interface{}{ "xai": map[string]interface{}{ "searchParameters": map[string]interface{}{"mode": "auto"}, }, }, }) ``` **After:** ```go model, _ := provider.LanguageModel("grok-3") result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What happened in the news today?", Tools: []types.Tool{xai.WebSearch(xai.WebSearchConfig{})}, }) ``` #### Gateway: xai/* models renamed to spacexai/* **Impact:** compile error / runtime 404 if you reference the old IDs. All `xai/*` Gateway model IDs are renamed to `spacexai/*` (`GatewayLanguageModelXai*` constants become `...Spacexai...`); 44 language, 4 image, and 4 video IDs were removed upstream. `hipaaCompliant` is removed from `gateway.Config` and `GatewayProviderOptions` — it is no longer sent as `providerOptions.gateway.hipaaCompliant`. ### Anthropic **Impact:** breaking for direct Anthropic API inspection, custom base URLs, and JSON-marshaled `ModelOptions`. - Request `system` is now an array of text blocks; user content is always an array. If you asserted on the raw request JSON in tests, update the shape. - `BaseURL` now includes `/v1`. A bare `https://api.anthropic.com` is normalized automatically, but a custom base URL (proxy, test server) must include `/v1` itself. - `DisableParallelToolUse` changed from `bool` to `*bool`. - Model-level `Thinking` (on `anthropic.ModelOptions`) now takes precedence over call-level `Reasoning` when both are set. - `anthropic.ModelOptions` JSON tags are camelCase (`budgetTokens`, `contextManagement`, `automaticCaching`, ...), matching TS provider options. v0.4.0 used snake_case. - `anthropic.ContainerSkill` (JSON `skillId`, was `skill_id`) with `Type: "custom"` now requires `ProviderReference`; without it the call fails with `NoSuchProviderReferenceError`. - The exported `anthropic.ToAnthropicFormatWithCache` is removed (it was unused by the request path). - A spliced stream (a second `message_start` while a message is open) is now reported as a `ChunkTypeError` chunk, and the rest of the stream is dropped — `Next()` no longer returns it as a Go `error`. **Before:** ```go skip-compile // DisableParallelToolUse was a plain bool on anthropic.ModelOptions: model := anthropic.NewLanguageModel(prov, anthropic.ClaudeSonnet4_6, &anthropic.ModelOptions{ DisableParallelToolUse: true, }) ``` **After:** ```go disable := true model := anthropic.NewLanguageModel(prov, anthropic.ClaudeSonnet4_6, &anthropic.ModelOptions{ DisableParallelToolUse: &disable, }) ``` ```go // v0.4.0 ModelOptions JSON (snake_case) — no longer accepted: // {"budget_tokens": 4096, "context_management": {...}, "automatic_caching": true} // v0.5.0 ModelOptions JSON (camelCase): // {"budgetTokens": 4096, "contextManagement": {...}, "automaticCaching": true} ``` ### Bedrock #### Bedrock-Anthropic: rebuilt on anthropic.LanguageModel **Impact:** breaking — several exported APIs removed. Bedrock-Anthropic is now built on `anthropic.LanguageModel`, as the TypeScript SDK does, instead of a Go-only Converse-shaped implementation. Removed: - `CacheConfig` API: `WithSystemCache`, `WithToolCache`, `WithMessageCacheIndices`, `CacheTTL1Hour`, `CacheTTL5Minutes`, `NewCacheConfig`. It inserted Converse-style `cachePoint` blocks into Messages API requests, which is invalid for this code path. Use `anthropic.ModelOptions{AutomaticCaching: ...}` or `CacheControl` instead. - `PrepareTools`, `UpgradeToolVersion`, `MapToolName`, `GetBetaHeaders`, `IsComputerUseTool`, and the old stream reader types. - The unused exported `bedrock/anthropic.BaseURLFormat` constant. Explicit `AccessKeyID` / `SecretAccessKey` no longer pick up `AWS_SESSION_TOKEN` from the environment — pass `SessionToken` explicitly if you need it alongside explicit credentials. `anthropic-aws`'s default base URL now includes `/v1`. **Before:** ```go cache := bedrockanthropic.NewCacheConfig().WithSystemCache(true) ``` **After:** ```go model := anthropic.NewLanguageModel(prov, anthropic.ClaudeSonnet4_6, &anthropic.ModelOptions{ AutomaticCaching: true, }) ``` #### Bedrock (Converse): invoke API replaced with Converse **Impact:** breaking for code inspecting raw request/response bodies. `bedrock.LanguageModel` now calls the Converse API (`/model/{id}/converse`, `/converse-stream`) instead of `/invoke`, matching the TypeScript SDK. `RawRequest` / `RawResponse` on results now carry Converse-shaped JSON — code that parsed the old per-model invoke bodies (Claude Messages format, Titan format, etc.) must be updated to the single Converse shape. Forced tool choice now matches provider-defined Anthropic tools by their wire name. #### Bedrock: other changes - **Embeddings:** Cohere embedding models' `MaxEmbeddingsPerCall` is now 96 (was 1); `EmbedMany` batches Cohere inputs accordingly. - **Mantle:** the underlying OpenAI provider is now built per call, so `openai/v1` vs `v1` routing can follow the model ID. No API change. - **Endpoints:** when `BaseURL` is unset, `AWS_ENDPOINT_URL_BEDROCK_RUNTIME` (or `_AGENT_RUNTIME`) and then `AWS_ENDPOINT_URL` are now honored, and China/ISO regions get their partition DNS suffix. If your environment sets these variables for other AWS tooling, Bedrock calls now route through them too — set `BaseURL` explicitly to opt out. - `ProviderError` from Bedrock calls now carries response headers and body. ### OpenAI and Azure #### OpenAI and Azure chat: reasoning-model parameter handling **Impact:** behavior change for reasoning models; no signature change. For reasoning models, `Temperature`, `TopP`, frequency/presence penalties, `LogitBias`, and `Logprobs` are now dropped with warnings (instead of being sent and rejected by the API), and the output limit is sent as `max_completion_tokens` instead of `max_tokens`. `ReasoningNone` is now sent as `reasoning_effort: "none"` (was `"disabled"`); `minimal` and `xhigh` are passed through verbatim. `whisper-1` transcription always requests `response_format: "verbose_json"`. Replayed tool-call arguments (`RawArguments`) that aren't a JSON object are now sent as `{}` for OpenAI and Azure chat specifically (TS `serializeToolCallArguments`); other OpenAI-compatible providers (Groq, DeepSeek, Together, Fireworks, Mistral, Ollama, ...) are unaffected and still send the raw arguments. #### Azure: base URL classification **Impact:** behavior change for non-standard base URLs. Azure base URLs are now classified like TS: hosts that don't match a recognized Azure OpenAI / Foundry pattern are treated as custom gateways and used as-is (no `/v1` or `api-version` appended). If you use a custom-domain gateway in front of Azure, verify the resulting request URL still works. ### Google / Vertex - Vertex provider options are now read in TS order: `googleVertex` → `vertex` → `google` (was a different order in v0.4.0). - Gemini 2.5 thinking budget is now `min(model max, round(65536 × pct))`, with no 1024-token floor. - Imagen (non-`gemini-*`) image models are removed from Google and Vertex, matching TS. Such model IDs now fail with: `"Google image models other than Gemini are no longer supported. Use a model ID that starts with gemini-."` - Gemini image models ignore `N` (one image per call, auto-batched) instead of erroring when `N > 1`. - The Interactions API now sends `input` as a flat array of steps (`user_input` / `model_output` / tool steps) instead of role/content turns — a wire-format fix, not an API change. - `ToGoogleMessages` is deprecated in favor of the new converter; it still works but new code should not call it directly. ### Embed / EmbedMany / Rerank: MaxRetries is now a pointer **Impact:** compile error if you set `MaxRetries` as a plain int. `MaxRetries` changed from `int` to `*int` (nil now means the TS default of 2, matching `GenerateTextOptions.MaxRetries`). `EmbedMany` with no input values now returns an empty result instead of an error. **Before:** ```go skip-compile result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: model, Input: "hello", MaxRetries: 5, }) ``` **After:** ```go retries := 5 result, err := ai.Embed(ctx, ai.EmbedOptions{ Model: model, Input: "hello", MaxRetries: &retries, }) ``` ### Schema: additionalProperties: false is enforced **Impact:** behavior change — previously-accepted objects with extra properties now fail validation. The `pkg/schema` validator now enforces `additionalProperties: false` when a schema sets it. If you relied on extra properties silently passing validation, either remove `additionalProperties: false` from the schema or update the produced/validated data to match the schema exactly. See the [JSON Schema reference](https://goaisdk.com/docs/reference/schema/json-schema.md) for the defaults-before-validate and `$ref` cycle guard behavior added alongside this. ### UI message streams: isAborted / isCancelled semantics **Impact:** behavior change for consumers of `CreateUIMessageStream`'s end callback. The end callback no longer reports `isAborted: true` just because the request context was cancelled. Consumer cancellation before an outcome is declared now sets `isCancelled: true`, and the event carries a new `Outcome` field (`completed` / `failed` / `aborted` / `unknown`), matching TS. `ValidateUIMessages` now rejects tool parts missing `input` (and `output` in `output-available`) for states where TS requires them; tool parts in `input-streaming` / `output-error` omit the `input` key instead of sending `null`. `finish-step` no longer closes active text/reasoning parts. ### SmoothStream: invalid chunking value returns a typed error **Impact:** low. An invalid `chunking` value now returns `*errors.InvalidArgumentError` instead of a generic error. ### Core: structured output, tool choice violation, custom content - `GenerateText` / `StreamText` return `ToolChoiceViolationError` when the model ignores a required or specific tool choice. - `GenerateText` structured output parse semantics follow TS: a truncated response now yields `NoObjectGeneratedError` (previously may have returned a partial value). - `StreamText` with `Output` always parses the final step. - `CustomContent.Kind` values use the `provider.type` form (for example `xai.citation`, `openai.compaction`), not `provider-type`. Update any code that pattern-matches on the old hyphenated form. - Core no longer supplies a default denial reason for denied tool approvals; providers that need text now apply their own default. ### LegacyOpenTelemetry span shape (affects dashboards and queries) **Impact:** breaking for existing dashboards/alerts built on span names or `gen_ai.*` root attributes. - Span names no longer carry the `.` suffix; filter on the `operation.name` / `resource.name` attributes instead. - Tool call spans are named `ai.toolCall` (was `ai.toolCall.`); the tool name is in the `ai.toolCall.name` attribute. - Root and evaluate spans no longer set `gen_ai.system` / `gen_ai.request.model`; use `ai.model.provider` / `ai.model.id`. - `ai.prompt` is JSON `{system, messages}` (object operations: `{system, prompt, messages}`). - Embed and rerank spans use `ai.value` / `ai.values` / `ai.documents` instead of `ai.prompt`; `ai.values.count` is removed. - Nested step spans are named `ai..doGenerate` / `ai..doStream` (were `" step N"`). - Nested embed/rerank spans are named `ai.embed.doEmbed` / `ai.embedMany.doEmbed` / `ai.rerank.doRerank` (were `"embeddings "` / `"reranking "`) and carry no `gen_ai.*` attributes. - Root span end attributes are now per operation: rerank sets none, embed sets `ai.embedding` / `ai.embeddings` and `ai.usage.tokens`, object operations set `ai.response.object`; `gen_ai.usage.*` is no longer emitted on root spans. - The nested `"chat "` span is removed — the `ai..doGenerate` / `ai..doStream` step span is now the model-call span. The GenAI `OpenTelemetry` integration's root span is now named ` ` (for example `invoke_agent gpt-5`), step spans are 1-indexed, and `ai.values.count` is removed there too. ### Provider-specific breaking changes - **Perplexity** (catch-up to TS): the provider now uses the Agent API (`/v1/agent`) instead of `/chat/completions`. Every Sonar-era `providerOptions.perplexity` key is removed (`search_recency_filter`, `search_domain_filter`, `web_search_options`, `return_images`, `return_related_questions`, `disable_search`, `reasoning_effort`, `media_response`, `stream_mode`, ...); new keys are `instructions`, `tools`, `models`, `max_steps`, `max_tool_calls`, `previous_response_id`, `store`, `language_preference`, `reasoning.effort` (adds `xhigh`), and `skills`. `PerplexityMetadata.Images` is always nil; `Cost` gains `Currency`, `CacheCreationCost`, `CacheReadCost`, `ToolCallsCost`; new `ToolCalls` map; `Usage.CitationTokens` is always nil; `NumSearchQueries` comes from `tool_calls_details`. Reasoning tokens are now a subset of `outputTokens.total` (previously additive). PDF input is rejected (images only); tools, structured output, and image input are now supported. - **MCP stdio**: the child process no longer inherits the whole parent environment — it gets `StdioTransportConfig.Env` plus a small allowlist (`PATH`, `HOME`, `USER`, `LOGNAME`, `SHELL`, `TERM`; Windows equivalents on Windows). Pass API keys and other secrets explicitly through `Env`. - **Harness**: `sandboxConfig.OnBootstrap` now actually runs for a caller-owned `SandboxSession` (it previously never ran). - **HuggingFace**: the provider is now a port of TS's Responses-API provider. Base URL is `router.huggingface.co/v1`, calling `POST /responses` instead of `/models/{id}`. `EmbeddingModel` and `ImageModel` always return an error (use the HF Inference API directly) — the exported `huggingface.EmbeddingModel` / `NewEmbeddingModel` and `huggingface.ImageModel` / `NewImageModel` types are removed. `Provider()` is now `"huggingface.responses"`. An empty model ID is no longer rejected. - **CosineSimilarity**: a zero (or empty) vector now returns `(0, nil)` instead of an error. A length mismatch returns a typed `*errors.InvalidArgumentError`. See [CosineSimilarity reference](https://goaisdk.com/docs/reference/ai/cosine-similarity.md). - **fal**: the default base URL is now `https://fal.run` (was `https://fal.run/fal-ai`, which doubled the path for fully-qualified model IDs). Pass fully-qualified IDs, e.g. `fal-ai/fast-sdxl`, not `fast-sdxl`. New defaults are `fal-ai/fast-sdxl` (image) and `fal-ai/luma-ray` (video). - **Middleware / Registry**: `AddToolInputExamplesOptions.Remove` changed from `bool` to `*bool` — `nil` now means remove (the TS default); before, the zero value `false` silently kept the examples. A model ID without the registry separator now returns a typed `*errors.NoSuchModelError`. See [Provider & Model Management](https://goaisdk.com/docs/ai-sdk-core/provider-management.md) and [Middleware Aliases](https://goaisdk.com/docs/reference/ai/middleware-aliases.md). - **OpenAI-compatible streams**: a stream that ends without a `finish_reason` now emits an `InvalidResponseDataError` error chunk and a finish chunk with reason `error`, instead of ending silently. Assistant messages with tool calls: Groq now sends the text content verbatim (empty string, not `null`), and Cohere omits `content` entirely. - **Moonshot**: default base URL is now `https://api.moonshot.ai/v1` (was `api.moonshot.cn`); set `Config.BaseURL` for the China endpoint. Provider metadata moved from `providerMetadata.moonshot` to `providerMetadata.moonshotai`; `providerOptions.moonshotai` is now honored alongside `moonshot`. - **Baseten**: default base URL is now `https://inference.baseten.co/v1` (was `bridge.baseten.co`), and the chat provider name is `baseten.chat`. With a custom `/sync/v1` `ModelURL`, an empty model ID now defaults to `placeholder` instead of `chat`. - **Cerebras**: removed model constants for retired models (`ModelLlama31_8B`, `ModelQwen3_235BA22BInstruct2507`, `ModelQwen3_235BA22BThinking2507`, `ModelZaiGLM4_6`, `ModelZaiGLM4_7`). - **DeepSeek**: `SupportsImageInput()` now returns `true` (V4 vision) — image parts are sent instead of being rejected. - **Mistral embeddings**: `MaxEmbeddingsPerCall` is 32 (was 2048) and `SupportsParallelCalls` is false; `EmbedMany` sends more, sequential requests. - **Together**: `providerOptions.togetherai` is now honored for chat options alongside `together`. ### Every provider request now sends a User-Agent header **Impact:** low v0.4.x deliberately sent no custom `User-Agent` header. v0.5.0 reverses that decision to match the TypeScript SDK: every provider request now carries `ai-sdk-/ go/` (for example `ai-sdk-openai/0.5.0 go/go1.25.1`), appended to any `User-Agent` value you already set via `Headers` rather than replacing it. `pkg/ai` call-level functions (`GenerateText`, `GenerateObject`, `GenerateImage`, `Embed`, `EmbedMany`, `Transcribe`, `GenerateSpeech`, `GenerateVideo`, `Batch`, `Evaluate`) additionally tag their own headers with `ai/` before the request reaches the model, matching the TS SDK's `ai` package. `StreamText`, `StreamObject`, and `Rerank` do not add this `ai/` tag, matching TS. Each product identifier carries exactly one `/`, per [RFC 9110's `product` grammar](https://www.rfc-editor.org/rfc/rfc9110.html#section-10.1.5) (`token ["/" version]`); some servers, including Azure, reject a `User-Agent` with more than one `/` in a single identifier. Earlier builds sent `ai-sdk// runtime/go/` (two slashes in each identifier); update any code that parses or matches on the old format. If you filter or allowlist outbound headers at a proxy or egress gateway, update that configuration to permit `User-Agent`, since requests that previously had no custom value now do. ## Provider updates to review The following areas gained substantial new functionality without breaking existing calls — review them if they're relevant to your application: - **OpenAI**: GPT-5.5/GPT-6 family model IDs, `gpt-image-2`, function-call namespace preservation, image detail forwarding, null assistant content only for tool-call messages, custom headers, Responses metadata. **OpenAI Responses**: computer tool, async and programmatic tool calling, GPT-6 reasoning config, prompt cache options, `RawFinishReason` on stream chunks, `serviceTier` on finish, citations/annotations as source parts, MCP approvals answered in earlier turns. - **Anthropic**: Claude Opus 4.7, `InferenceGeo`, `TaskBudget`, schema sanitization, obsolete beta header removal, request-level `Compaction` model option, `ForwardContainerIDFromLastStep`, citations as source parts, `container_upload` custom content, corrected finish-reason mapping (`pause_turn`→`stop`, `refusal`→`content-filter`, `model_context_window_exceeded`→`length`). - **Google and Vertex**: model refresh, per-modality token usage, thought signatures, no-arg tool calls, auth reuse, MaaS auth overrides, Vertex MaaS Grok models, Interactions API (including managed agents) on Vertex, Chirp 3 HD text-to-speech, Gemini 3.5 Transcribe, Google Vertex video (Veo), realtime lifecycle events. - **Bedrock**: explicit credential precedence, reasoning block filtering, Opus 4.7 behavior, native structured output, and the Go-specific note that TypeScript lazy `fetch` resolution has no direct runtime equivalent. - **XAI**: non-image file URLs in Responses, cost metadata, ZDR encrypted reasoning, file uploads, `referenceVoiceIds` limit. - **Gateway**: reranking, routing sort, quota entity ID, HIPAA filtering, unknown model type resilience. - **MCP**: secure JSON parsing, server instructions, TypeScript-compatible MCP provider metadata, MCP Apps helpers, negotiated protocol headers, custom transports, `MCPClient.OnElicitationRequest`, `MCPClient.ListResourceTemplates`, a standing inbound SSE listener for legacy-protocol servers. See [MCP](https://goaisdk.com/docs/ai-sdk-core/mcp-tools.md). - **Voyage, Cohere, Mistral, and DeepSeek**: new provider or model-feature parity updates (Voyage AI embedding/rerank is a new provider). - **Async video**: `ai.ExperimentalStartVideo` / `ExperimentalGetVideoStatus` for fal, Google, Google Vertex, Replicate, and xAI. See [Video Generation](https://goaisdk.com/docs/ai-sdk-core/video-generation.md). - **Speech translation (experimental)**: Google and OpenAI `SpeechTranslationModel`. See [Speech](https://goaisdk.com/docs/ai-sdk-core/speech.md#speech-translation-experimental). - **Batch API / Evaluation (experimental)**: `ai.ExperimentalStartBatch` family and `ai.ExperimentalEvaluate`, implemented by Anthropic, OpenAI, and Google. - **Code-mode (experimental)**: `pkg/codemode` runs model-written JavaScript in a QuickJS-on-WebAssembly sandbox. See [pkg/codemode reference](https://goaisdk.com/docs/reference/ai/code-mode.md) and [Known differences](https://goaisdk.com/docs/migration-guides/known-differences.md#code-mode-never-settling-promises-and-in-flight-limits). - **Harness**: `pkg/harness` (Go port of `@ai-sdk/harness`) with adapters for Claude Code, Codex, OpenCode, Deep Agents, ACP, Cursor, GitHub Copilot, Grok Build, and a Vercel Sandbox provider. See [pkg/harness/sandbox/vercel reference](https://goaisdk.com/docs/reference/ai/harness-sandbox-vercel.md). - **New providers**: GMI Cloud, Z.AI, MiniMax, TypeSafe AI, Fish Audio, Cartesia, Rev.ai, Hume, Luma. - **LlamaIndex**: `pkg/llamaindex.ToUIMessageStream` adapts a LlamaIndex chat-engine delta stream to UI message stream chunks. See [pkg/llamaindex reference](https://goaisdk.com/docs/reference/ai/llamaindex.md). ## Appendix: May 2026 parity cycle details This appendix keeps the detailed notes from the May 2026 migration guide, which covers the first of the two cycles that land in v0.5.0. Use it for before and after code on agents, tool approval, sandboxes, registries, and upload APIs. ### Agent interface (May 2026 detail) The `agent.Agent` interface now follows the TypeScript SDK agent contract. Custom agent implementations must expose identity, tool, generate, and stream methods in addition to the legacy Go convenience execution methods. Before: ```go type MyAgent struct{} func (a *MyAgent) Execute(ctx context.Context, prompt string) (*agent.AgentResult, error) { /* ... */ } func (a *MyAgent) ExecuteWithMessages(ctx context.Context, messages []types.Message) (*agent.AgentResult, error) { /* ... */ } ``` After: ```go type MyAgent struct{} func (a *MyAgent) Version() string { return "agent-v1" } func (a *MyAgent) ID() string { return "support-agent" } func (a *MyAgent) Tools() []types.Tool { return nil } func (a *MyAgent) Generate(ctx context.Context, opts agent.AgentGenerateOptions) (*ai.GenerateTextResult, error) { /* ... */ } func (a *MyAgent) Stream(ctx context.Context, opts agent.AgentStreamOptions) (*ai.StreamTextResult, error) { /* ... */ } func (a *MyAgent) Execute(ctx context.Context, prompt string) (*agent.AgentResult, error) { /* ... */ } func (a *MyAgent) ExecuteWithMessages(ctx context.Context, messages []types.Message) (*agent.AgentResult, error) { /* ... */ } ``` ### RuntimeContext and ToolsContext (May 2026 detail) `RuntimeContext` is now the primary call-level context. `ExperimentalContext` remains as a deprecated alias and is used only when `RuntimeContext` is not set. Per-tool context belongs in `ToolsContext` and is validated against each tool's `ContextSchema` before approval callbacks or tool execution. Before: ```go agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, ExperimentalContext: map[string]any{"requestID": "req_123"}, }) ``` After: ```go agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, RuntimeContext: map[string]any{"requestID": "req_123"}, ToolsContext: map[string]any{ "lookup": map[string]any{"tenant": "acme"}, }, }) ``` Dynamic tool descriptions receive the same per-tool context used by the TypeScript SDK: ```go types.Tool{ Name: "lookup", DescriptionFunc: func(ctx context.Context, opts types.ToolDescriptionOptions) string { return fmt.Sprintf("Lookup data for %v", opts.Context) }, } ``` ### Tool approval (May 2026 detail) Tool approval callbacks can now return a zero-value `types.ToolApprovalResult{}` to mean `not-applicable`. Denials preserve the approval status and reason in `types.ToolResult` and use a default message that includes the tool name. Per-tool approval callbacks run only after the matching `ContextSchema` validates successfully; validated JSON-schema defaults are applied before `ToolContext` is passed to approval or execution callbacks. Tool-defined approval callbacks can use the options-based `types.ToolNeedsApprovalFunc` to receive the same data as the TypeScript SDK: tool call ID, pre-call messages, and validated tool context. The older `types.NeedsApprovalFunc` remains available as a deprecated input-only compatibility path. Before: ```go ToolApprovalRequired: true, ToolApprover: func(call types.ToolCall) bool { return false }, ``` After: ```go ToolApproval: types.ToolApprovalFunc(func( call types.ToolCall, tools []types.Tool, messages []types.Message, runtimeCtx any, toolsCtx map[string]any, ) types.ToolApprovalResult { return types.ToolApprovalResult{Status: types.ToolApprovalStatusDenied} }), ``` ### MaxSteps and StopWhen (May 2026 detail) `MaxSteps` is deprecated for agents. Existing callers still work: `MaxSteps: n` is routed to `StopWhen: []ai.StopCondition{ai.IsStepCount(n)}` and emits a deprecation warning. New code should use `StopWhen` directly. Before: ```go agent.NewToolLoopAgent(agent.AgentConfig{Model: model, MaxSteps: 5}) ``` After: ```go agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) ``` ### Prompt field (May 2026 detail) `Prompt` is available as a TypeScript-compatible alias for the agent system prompt. If `System` is empty, `Prompt` is used as the system prompt sent to the model. ```go agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Prompt: "You are a careful support agent.", }) ``` ### Sandbox API (May 2026 detail) Sandbox implementations now use the TypeScript SDK v7 naming and file helper surface. Implement `ai.Experimental_SandboxSession` or `providerutils.SandboxSession`. Before (old API, will not compile): ```go skip-compile // Removed API: ai.SandboxExecuteOptions no longer exists. result, err := sandbox.Execute(ctx, "go test ./...", ai.SandboxExecuteOptions{ WorkingDirectory: "/workspace", }) ``` After: ```go result, err := sandbox.Run(ctx, ai.SandboxProcessOptions{ Command: "go test ./...", WorkingDirectory: "/workspace", Env: map[string]string{ "GOFLAGS": "-mod=mod", }, }) ``` The session interface also includes `ReadFile`, `ReadBinaryFile`, `ReadTextFile`, `WriteFile`, `WriteBinaryFile`, and `WriteTextFile`. The old Go-only `Sandbox`, `SandboxExecuteOptions`, and `SandboxExecuteResult` names are removed; use `Experimental_SandboxSession`, `SandboxProcessOptions`, and `SandboxRunResult`. `SandboxProcessOptions.Env` is forwarded to both `Run` and `Spawn`, matching the TypeScript `Experimental_SandboxSession` process options shape. ### OpenAI Responses files (May 2026 detail) OpenAI Responses now matches the TypeScript SDK default for file parts: images and PDFs are accepted by default, while unsupported file media types require an explicit opt-in. ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: promptWithCSVFile, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "passThroughUnsupportedFiles": true, }, }, }) ``` ### FilterActiveTools (May 2026 detail) Use `FilterActiveTools` to restrict tool availability per step. ```go agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: tools, FilterActiveTools: func(ctx context.Context, step int, tools []types.Tool) []types.Tool { if step == 1 { return tools[:1] } return tools }, }) ``` ### Telemetry event names (May 2026 detail) Telemetry integrations should use the stable end-event names from the May 2026 TypeScript SDK cycle: | Previous Go name | New Go name | | --- | --- | | `OnToolCallStart` / `FireOnToolCallStart` | `OnToolExecutionStart` / `FireOnToolCallStart` | | `OnToolCallFinish` / `FireOnToolCallFinish` | `OnToolExecutionEnd` / `FireOnToolCallFinish` | | `OnFinish` / `FireOnFinish` | `OnEnd` / `FireOnEnd` | | `OnEmbedFinish` / `FireOnEmbedFinish` | `OnEmbedEnd` / `FireOnEmbedEnd` | | `OnRerankFinish` / `FireOnRerankFinish` | `OnRerankEnd` / `FireOnRerankEnd` | `TelemetryIntegration` now includes the TypeScript-aligned `OnToolExecutionStart` and `OnToolExecutionEnd` methods. `OnToolCallStart` and `OnToolCallFinish` remain as deprecated aliases. Telemetry no longer emits per-chunk `OnChunk` events; stream consumers should continue using `StreamTextOptions.OnChunk` for application-level chunk handling. `GenerateTextOptions`, `StreamTextOptions`, `agent.AgentGenerateOptions`, and `agent.AgentConfig` also prefer `OnToolExecutionStart` and `OnToolExecutionEnd` for tool execution callbacks. The older `OnToolCallStart` and `OnToolCallFinish` option fields remain deprecated aliases. ### Result content and step performance (May 2026 detail) `GenerateTextResult.Content` now matches the TypeScript SDK `result.content` getter: it is the ordered aggregate of every step's content parts, including text, reasoning, files, sources, tool calls, tool results, tool errors, and tool approval request/response parts. `types.StepResult` now includes `Performance` statistics matching the TypeScript SDK: `StepTimeMs`, `ResponseTimeMs`, `ToolExecutionMs`, `EffectiveOutputTokensPerSecond`, `EffectiveTotalTokensPerSecond`, and streaming-only `OutputTokensPerSecond`, `InputTokensPerSecond`, `TimeToFirstOutputMs`, and `TimeBetweenOutputChunksMs`. Use `FinalStep.Performance.EffectiveOutputTokensPerSecond` for final-step output throughput and `Steps[i].Performance` for multi-step workflows. `StepTimeMs` is measured as wall-clock step duration, including client-side tool execution. `ToolExecutionMs` is keyed by `toolCallID`. ### Sandbox description (May 2026 detail) `ai.Experimental_SandboxSession` includes `Description() string`. When a sandbox is active, non-empty descriptions are appended to system instructions before provider calls, matching the TypeScript SDK sandbox instruction behavior. ```go sandbox := ai.NewShellSandbox( ai.WithShellSandboxDescription("Ubuntu 22.04, root: /workspace"), ) ``` ### CallOptionsSchema (May 2026 detail) `CallOptionsSchema` is enforced before any model call. Validation errors include the `options` schema path and are returned as provider validation errors. ```go agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, CallOptions: map[string]any{"mode": "safe"}, CallOptionsSchema: schema.NewSimpleJSONSchema(map[string]any{ "type": "object", "required": []any{"mode"}, }), }) ``` ### Registry files and skills (May 2026 detail) Registries now resolve upload APIs directly. Providers that do not expose the requested API return `NoSuchProviderError`-equivalent errors. ```go files, err := registry.GetGlobalRegistry().Files("anthropic") skills, err := registry.GetGlobalRegistry().Skills("anthropic") ``` ### Provider v4 file data and reasoning files (May 2026 detail) Middleware wrappers preserve the wrapped model specification version and pass provider file data, custom content, and reasoning-file content through without downgrading model metadata. ### System message validation (May 2026 detail) System instructions should be passed via `System` or agent `Prompt`. Passing system-role messages still requires `AllowSystemMessages` / `AllowSystemInMessages` on lower-level generation calls. ### Upload APIs (May 2026 detail) Use `ai.UploadFile` and `ai.UploadSkill` with providers or registry-resolved `Files()` / `Skills()` APIs. Unsupported providers return explicit capability errors rather than nil API panics. ## Appendix: May 31 stream, telemetry, and provider details This appendix maps the May 31 TypeScript AI SDK changes to Go AI SDK API names. ### Stream helpers (May 31 detail) Use standalone stream helpers with `result.Stream()` for new code: ```go result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a concise release note.", }) if err != nil { log.Fatal(err) } defer result.Close() textStream, errStream := ai.ToTextStream(ctx, result.Stream()) for text := range textStream { fmt.Print(text) } if err := <-errStream; err != nil { log.Fatal(err) } ``` For HTTP responses, prefer helpers that accept a raw `provider.TextStream`: ```go response, err := ai.CreateTextStreamResponseFromStream(ctx, result.Stream(), &ai.TextStreamResponseInit{ Headers: map[string]string{"Cache-Control": "no-cache"}, }) ``` `StreamTextResult.ToTextStreamResponse`, `StreamTextResult.ToUIMessageStream`, `StreamTextResult.ToUIMessageStreamResponse`, `StreamTextResult.PipeTextStreamToResponse`, and `StreamTextResult.PipeUIMessageStreamToResponse` are retained as deprecated compatibility wrappers. ### Stream and FullStream (May 31 detail) `Stream()` is the canonical Go accessor for the underlying provider stream. `FullStream()` remains available as a deprecated alias for TypeScript `fullStream` migration. ```go stream := result.Stream() // Deprecated compatibility: legacy := result.FullStream() _ = legacy ``` ### UI message conversion (May 31 detail) Use `ToUIMessageChunk` for one chunk and `ToUIMessageStream` for a full stream: ```go sendStart := true sendFinish := true uiChunks, uiErrs := ai.ToUIMessageStream(ctx, result.Stream(), ai.UIMessageStreamResultOptions{ SendStart: &sendStart, SendFinish: &sendFinish, }) for chunk := range uiChunks { fmt.Printf("%v\n", chunk["type"]) } if err := <-uiErrs; err != nil { log.Fatal(err) } ``` The standalone stream helper is the Go equivalent of TypeScript's helper surface. Result-bound methods delegate to these helpers only for source compatibility. ### Telemetry (May 31 detail) `Telemetry` is the canonical field. `ExperimentalTelemetry` still works as a deprecated alias. Telemetry integrations now receive abort events through `OnAbort`, and OpenTelemetry spans are closed when generation is canceled. Model calls run inside the telemetry context, so integrations that attach spans or values in `OnStart` can make them visible downstream. ```go ctx, cancel := context.WithCancel(context.Background()) defer cancel() enabled := true result, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Stream until canceled.", Telemetry: &ai.TelemetrySettings{ IsEnabled: &enabled, }, }) if err != nil { log.Fatal(err) } defer result.Close() cancel() _ = result.ConsumeStream() ``` Output timing fields use the TypeScript-compatible names: - `TimeToFirstOutputMs` - `TimeBetweenOutputChunksMs` File outputs and tool calls count as first output events, matching the TypeScript SDK. Earlier token-specific naming is not canonical in Go. Nested runtime/tool context objects are split into separate telemetry attributes when opted in through `IncludeRuntimeContext` or `IncludeToolsContext`. ## See Also - [Known differences from the TypeScript AI SDK](https://goaisdk.com/docs/migration-guides/known-differences.md) - [Migrating from the TypeScript AI SDK](https://goaisdk.com/docs/migration-guides/from-typescript-ai-sdk.md) - [StreamText reference](https://goaisdk.com/docs/reference/ai/stream-text.md) --- # Known differences from the TypeScript AI SDK > Behavior that is intentionally different in Go, and features that only exist in the TypeScript SDK. Canonical URL: https://goaisdk.com/docs/migration-guides/known-differences Documentation index: https://goaisdk.com/llms.txt The Go AI SDK tracks the TypeScript AI SDK (`ai@7.0.127` as of v0.5.0) closely, but a few areas are intentionally different, either because the Go runtime has no equivalent primitive or because a decision was made to diverge. This page is the reference for both. ## StreamObject is not lazy `StreamObject` is deprecated in both SDKs. In TypeScript, `streamObject` returns lazy streams immediately, before the model call completes. In Go, `StreamObject` returns only after the stream is fully consumed, and reports progress through `OnChunk` as the partial object is refined. For incremental partial objects or array elements, use `StreamText` with `Output` (`PartialOutput` / `ElementStream`) instead — this is also what the TypeScript SDK recommends going forward. ```go // Not recommended: StreamObject blocks until the stream finishes. result, err := ai.StreamObject(ctx, ai.StreamObjectOptions{ Model: model, Prompt: "Generate a recipe", Schema: recipeSchema, OnChunk: func(partial interface{}) { // Called incrementally as the partial object is refined. }, }) // Recommended: StreamText + Output returns immediately and lets you read // the incrementally parsed value as chunks arrive, matching TS's lazy // streamObject. textResult, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Generate a recipe", Output: ai.ObjectOutput[Recipe](ai.ObjectOutputOptions{Schema: recipeSchema}), }) for chunk := range textResult.Chunks() { if chunk.Type == provider.ChunkTypeText { partial := textResult.PartialOutput() // latest parsed partial value _ = partial } } ``` ## Video generation webhook suspension in workflows TypeScript's `experimental_generateVideo` can durably suspend a Vercel Workflow DevKit run until a provider webhook fires, resuming the workflow from durable storage when the video finishes. This depends on the Node-only Workflow DevKit runtime and has no Go equivalent. In Go, video generation polls until the job is done (`GenerateVideo`), or you drive the job yourself with `ai.ExperimentalStartVideo` / `ai.ExperimentalGetVideoStatus`, which support both polling and webhook completion (fal and Replicate) outside of a workflow. ## Vercel Sandbox: no OIDC token refresh loop `pkg/harness/sandbox/vercel` (the Go port of `@vercel/sandbox`'s harness provider) does not run a background `@vercel/oidc` token refresh loop. `VERCEL_OIDC_TOKEN` is re-read from the environment on each call instead; explicit `Token` / `TeamID` / `ProjectID` credentials also work and bypass OIDC entirely. ## Code-mode: never-settling promises and in-flight limits `pkg/codemode` runs model-written JavaScript in a QuickJS-on-WebAssembly sandbox (wazero). Concurrent tool calls that need approval (for example inside one `Promise.all`) are batched into a single interrupt, as in TS. Two details differ: - A promise that never settles and has no pending tool call fails fast with a `ProtocolError`. TypeScript waits for the execution timeout. - `MaxInFlightBridgeRequests` counts the calls in an approval batch. ## Features that are TypeScript-only (not ported) - **OpenAI Live over WebRTC** — the browser transport for OpenAI's realtime API. The Go SDK only implements the server-side WebSocket transport (`gpt-live-1` via `ai.ConnectRealtime`), which covers server-side and backend use cases; WebRTC is a browser-only transport with no Go runtime equivalent. - **`@ai-sdk/harness-cline` and `@ai-sdk/harness-pi`** — these harness adapters run vendor Node SDKs in-process (they embed a Node.js runtime and the vendor's own TypeScript SDK). There is no Go equivalent; use the TypeScript SDK for Cline or Pi harness integrations. - **DeepSeek batch API / extra Files methods** — TS's `@ai-sdk/deepseek` does not implement these either, so this is parity, not a gap; the existing DeepSeek file support already matches TS. ## Streaming pattern: Next() instead of io.Reader This is a longstanding, deliberate Go-idiom difference (not new in v0.5.0): `provider.TextStream` exposes `Next()` / `Err()` / `Close()`, not an `io.Reader`. See the [Streaming guide](https://goaisdk.com/docs/guides/STREAMING.md) for the rationale. ## See Also - [Migrating from v0.4.x to v0.5.0](https://goaisdk.com/docs/migration-guides/from-v0.4-to-v0.5.md) - [Migrating from the TypeScript AI SDK](https://goaisdk.com/docs/migration-guides/from-typescript-ai-sdk.md) --- # Token usage differentiation > Explains the TextTokens and ImageTokens fields added to types.InputTokenDetails in v0.2.0, letting you cost multimodal requests accurately in Go. Canonical URL: https://goaisdk.com/docs/migration-guides/token-usage-differentiation Documentation index: https://goaisdk.com/llms.txt `types.InputTokenDetails` reports separate text and image input token counts on multimodal requests, so you can cost a vision call accurately instead of billing every input token at the text rate. The fields shipped in v0.2.0 and are optional — code written against v0.1.x keeps compiling and running unchanged. This applies to anyone billing or budgeting multimodal calls per token. ## TextTokens and ImageTokens on InputTokenDetails **Impact:** additive — existing code is unaffected. ```go type InputTokenDetails struct { NoCacheTokens *int64 CacheReadTokens *int64 CacheWriteTokens *int64 TextTokens *int64 // added in v0.2.0 ImageTokens *int64 // added in v0.2.0 } ``` Support varies by provider: | Provider | Populates TextTokens / ImageTokens | |---|---| | OpenAI, XAI, Azure, DeepSeek, Together, Fireworks, Perplexity, Groq, Mistral | Yes — parsed from the OpenAI-compatible `prompt_tokens_details.text_tokens` / `image_tokens` wire fields | | Google | Yes — computed from Gemini's per-modality token counts | | Anthropic | No — both fields stay `nil`; the API doesn't report a text/image split | **Before (text-rate cost for every input token):** ```go result, err := ai.GenerateText(ctx, opts) if err != nil { return err } inputCost := float64(result.Usage.GetInputTokens()) * textInputRate outputCost := float64(result.Usage.GetOutputTokens()) * outputTokenRate totalCost := inputCost + outputCost ``` **After (accurate multimodal cost, with a fallback for providers that don't report the split):** ```go result, err := ai.GenerateText(ctx, opts) if err != nil { return err } var inputCost float64 details := result.Usage.InputDetails if details != nil && details.TextTokens != nil && details.ImageTokens != nil { inputCost = float64(*details.TextTokens)*textInputRate + float64(*details.ImageTokens)*imageInputRate } else { inputCost = float64(result.Usage.GetInputTokens()) * textInputRate } outputCost := float64(result.Usage.GetOutputTokens()) * outputTokenRate totalCost := inputCost + outputCost ``` `GetInputTokens()`, `GetOutputTokens()`, and `GetTotalTokens()` are unaffected — they keep returning the total (text + image) regardless of whether the detailed breakdown is available, and return `0` instead of panicking when the underlying pointer is `nil`. ## Example ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) const ( textInputRate = 0.0000025 // $2.50 per 1M tokens (GPT-4o) imageInputRate = 0.0000075 // ~$7.50 per 1M tokens (varies by image size) outputTokenRate = 0.0000100 ) func main() { ctx := context.Background() prov := openai.New(openai.Config{APIKey: "your-api-key"}) model, err := prov.LanguageModel("gpt-4o") if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "What's in this image?"}, types.ImageContent{URL: "https://example.com/image.jpg"}, }, }, }, }) if err != nil { log.Fatal(err) } details := result.Usage.InputDetails if details != nil && details.TextTokens != nil && details.ImageTokens != nil { fmt.Printf("text tokens: %d, image tokens: %d\n", *details.TextTokens, *details.ImageTokens) } } ``` ## Reference - Full field list: [Usage types reference](https://goaisdk.com/docs/reference/types/usage.md) --- # Common Errors > Covers the most frequent Go AI SDK errors and fixes, including API key issues, nil pointer errors, channel/streaming errors, and import/type errors. Canonical URL: https://goaisdk.com/docs/troubleshooting/common-errors Documentation index: https://goaisdk.com/llms.txt This guide covers the most common errors you'll encounter when using the Go AI SDK and how to fix them. ## API Key Issues ### Error: "Invalid API Key" **Symptoms:** ``` Error: provider error (401): Invalid API key provided Error: authentication failed ``` **Cause:** The API key is missing, incorrect, or not properly loaded from environment variables. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { // Check if API key is set apiKey := os.Getenv("OPENAI_API_KEY") if apiKey == "" { log.Fatal("OPENAI_API_KEY environment variable not set") } // Verify key is not empty or whitespace if len(apiKey) < 20 { log.Fatal("OPENAI_API_KEY appears to be invalid (too short)") } provider := openai.New(openai.Config{ APIKey: apiKey, }) model, err := provider.LanguageModel("gpt-6-astra") if err != nil { log.Fatalf("Failed to create model: %v", err) } ctx := context.Background() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { log.Fatalf("Failed to generate text: %v", err) } fmt.Println(result.Text) } ``` **Best Practices:** - Use environment variables for API keys (never hardcode) - Use `.env` files with `godotenv` for local development - Validate API keys on application startup - Use separate keys for development and production ### Error: "Model Not Found" **Symptoms:** ``` Error: model "gpt-5" not found Error: invalid model specified ``` **Cause:** Trying to use a model that doesn't exist or isn't available in your account. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) // Use valid model names validModels := []string{ "gpt-6-astra", "gpt-6-sol", "gpt-6-astra", "gpt-5.4-mini", "gpt-5.4-nano", } log.Printf("known-good model names: %v", validModels) // Try to create model with error handling model, err := provider.LanguageModel("gpt-6-astra") if err != nil { log.Printf("Failed to create gpt-6-astra model: %v", err) log.Println("Falling back to gpt-5.4-mini...") model, err = provider.LanguageModel("gpt-5.4-mini") if err != nil { log.Fatalf("Failed to create fallback model: %v", err) } } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { log.Fatalf("Failed to generate text: %v", err) } fmt.Println(result.Text) } ``` ## Nil Pointer Errors ### Error: "nil pointer dereference" **Symptoms:** ``` panic: runtime error: invalid memory address or nil pointer dereference ``` **Cause:** Not checking for errors before using returned values, or accessing fields on nil structs. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) // ALWAYS check errors before using the returned value model, err := provider.LanguageModel("gpt-6-astra") if err != nil { log.Fatalf("Failed to create model: %v", err) } // Now it's safe to use model // Check result is not nil before accessing fields result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { log.Fatalf("Failed to generate text: %v", err) } // Safe to access result.Text now if result != nil { fmt.Println(result.Text) } } ``` ## Channel and Streaming Errors ### Error: "Channel Already Closed" **Symptoms:** ``` panic: send on closed channel panic: close of closed channel ``` **Cause:** Trying to read from a stream channel multiple times or not properly handling channel closure. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a short story", }) if err != nil { log.Fatal(err) } // Read from channel until it closes // The range loop automatically handles channel closure for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } // Channel is now closed - DO NOT try to read again // Check for errors after stream completes if err := stream.Err(); err != nil { log.Printf("Stream error: %v", err) } // DO NOT try to read from Chunks again - it's closed } ``` ### Error: "Goroutine Leak / Stream Not Consumed" **Symptoms:** - Application hangs indefinitely - Memory usage grows over time - Goroutines never terminate **Cause:** Not consuming all values from a stream channel, causing the goroutine to block. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func processStream(ctx context.Context, model provider.LanguageModel, prompt string) error { stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return err } // ALWAYS consume the entire stream, even if you don't need all the data for chunk := range stream.Chunks() { fmt.Print(chunk.Text) // If you need to exit early, the context will cancel the stream select { case <-ctx.Done(): // Stream will be automatically cleaned up return ctx.Err() default: // Continue processing } } return stream.Err() } func main() { // Use context with timeout to prevent infinite hangs ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") if err := processStream(ctx, model, "Write a story"); err != nil { log.Printf("Error: %v", err) } } ``` ## Import Errors ### Error: "Package Not Found" **Symptoms:** ``` cannot find package "github.com/digitallysavvy/go-ai/pkg/ai" ``` **Cause:** Module not downloaded or `go.mod` not properly initialized. **Solution:** ```bash # Initialize module (if not already done) go mod init your-project-name # Install the Go AI SDK go get github.com/digitallysavvy/go-ai # Tidy up dependencies go mod tidy # Verify installation go list -m github.com/digitallysavvy/go-ai ``` ### Error: "Ambiguous Import" **Symptoms:** ``` ambiguous import: found package in multiple locations ``` **Cause:** Conflicting package names or vendored dependencies. **Solution:** ```go package main import ( "context" // Use full import paths with aliases if needed goai "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" // If you have naming conflicts, use aliases aitools "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func main() { ctx := context.Background() _ = aitools.RoleUser // use the aliased import, same as any other import model, _ := openai.New(openai.Config{}).LanguageModel("gpt-6-astra") // Use aliases in your code result, err := goai.GenerateText(ctx, goai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) _ = result _ = err } ``` ## Type Errors ### Error: "Type Mismatch" **Symptoms:** ``` cannot use X (type string) as type interface in argument cannot convert X to type Y ``` **Cause:** Passing wrong types to SDK functions. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // CORRECT: Pass messages as proper types.Message slice, with Content // as a []types.ContentPart, not a bare string messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "Hello!"}}, }, } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Messages: messages, // Not strings }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## JSON Schema Errors ### Error: "Schema Validation Failed" **Symptoms:** ``` Error: schema validation failed: invalid type for field X Error: generated object does not match schema ``` **Cause:** JSON schema doesn't match your Go struct, or schema is invalid. **Solution:** ```go skip-compile package main import ( "context" "encoding/json" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" goschema "github.com/digitallysavvy/go-ai/pkg/schema" "github.com/invopop/jsonschema" ) // Define struct with proper JSON tags type Recipe struct { Name string `json:"name" jsonschema:"required,description=Name of the recipe"` Ingredients []string `json:"ingredients" jsonschema:"required,minItems=1,description=List of ingredients"` Steps []string `json:"steps" jsonschema:"required,minItems=1,description=Cooking steps"` PrepTime int `json:"prepTime" jsonschema:"minimum=0,description=Preparation time in minutes"` } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Generate a JSON schema from the struct with invopop/jsonschema, then // wrap the resulting map in goschema.Schema — ai.GenerateObjectOptions.Schema // takes a schema.Schema, not a *jsonschema.Schema directly. reflector := jsonschema.Reflector{ AllowAdditionalProperties: false, DoNotReference: true, } reflected := reflector.Reflect(&Recipe{}) schemaBytes, err := json.Marshal(reflected) if err != nil { log.Fatalf("Invalid schema: %v", err) } log.Printf("Using schema:\n%s", schemaBytes) var schemaMap map[string]interface{} if err := json.Unmarshal(schemaBytes, &schemaMap); err != nil { log.Fatalf("Failed to decode schema: %v", err) } var recipe Recipe if err := ai.GenerateObjectInto(ctx, ai.GenerateObjectOptions{ Model: model, Schema: goschema.NewSimpleJSONSchema(schemaMap), Prompt: "Generate a lasagna recipe", }, &recipe); err != nil { log.Fatalf("Failed to generate object: %v", err) } fmt.Printf("Recipe: %s\n", recipe.Name) } ``` ## Context Errors ### Error: "Context Deadline Exceeded" **Symptoms:** ``` Error: context deadline exceeded Request timed out ``` **Cause:** Operation took longer than the context timeout allows. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Increase timeout for long operations ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute) defer cancel() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a detailed 5000-word essay on quantum computing", }) if err != nil { if ctx.Err() == context.DeadlineExceeded { log.Println("Operation timed out - increase the timeout or use streaming") // Alternative: Use streaming for long responses stream, err := ai.StreamText(context.Background(), ai.StreamTextOptions{ Model: model, Prompt: "Write a detailed 5000-word essay on quantum computing", }) if err != nil { log.Fatal(err) } for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } if err := stream.Err(); err != nil { log.Fatal(err) } return } log.Fatalf("Error: %v", err) } fmt.Println(result.Text) } ``` ## Best Practices ### 1. Always Check Errors ```go // Bad result, _ := ai.GenerateText(ctx, options) // Good result, err := ai.GenerateText(ctx, options) if err != nil { return fmt.Errorf("generation failed: %w", err) } ``` ### 2. Use Context Timeouts ```go // Always use context with timeout for production ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() ``` ### 3. Consume Streams Completely ```go // Always read all values from stream channels for chunk := range stream.Chunks() { // Process chunk } // Check error after stream closes if err := stream.Err(); err != nil { log.Printf("Stream error: %v", err) } ``` ### 4. Validate Configuration ```go // Check configuration at startup if apiKey := os.Getenv("OPENAI_API_KEY"); apiKey == "" { log.Fatal("Missing OPENAI_API_KEY") } ``` ### 5. Use Proper Types ```go // Use SDK types, not raw strings messages := []types.Message{ {Role: types.RoleUser, Content: []types.ContentPart{types.TextContent{Text: "Hello"}}}, } ``` ## See Also - [Provider Errors](https://goaisdk.com/docs/troubleshooting/provider-errors.md) - [Context Cancellation](https://goaisdk.com/docs/troubleshooting/context-cancellation.md) - [Streaming Issues](https://goaisdk.com/docs/troubleshooting/streaming-issues.md) - [Error Handling Guide](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) --- # Provider Errors > Troubleshooting guide for provider-specific Go AI SDK errors across OpenAI, Anthropic, Google Gemini, Azure OpenAI, and AWS Bedrock integrations. Canonical URL: https://goaisdk.com/docs/troubleshooting/provider-errors Documentation index: https://goaisdk.com/llms.txt This guide covers common errors specific to different AI providers and how to resolve them. ## OpenAI Provider Errors ### Error: "Insufficient Quota" **Symptoms:** ``` Error: You exceeded your current quota, please check your plan and billing details Status Code: 429 ``` **Cause:** You've exceeded your OpenAI usage quota or haven't set up billing. **Solution:** ```go package main import ( "context" "errors" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func generateWithQuotaFallback(ctx context.Context, prompt string) (string, error) { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) // Try primary model model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) && providerErr.StatusCode == 429 { log.Println("Quota exceeded, trying alternative provider or model...") // Option 1: Use a cheaper model model, _ = provider.LanguageModel("gpt-5.4-mini") result, err = ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return "", fmt.Errorf("fallback failed: %w", err) } return result.Text, nil } return "", err } return result.Text, nil } func main() { ctx := context.Background() text, err := generateWithQuotaFallback(ctx, "Explain quantum computing") if err != nil { log.Fatalf("Failed: %v", err) } fmt.Println(text) } ``` **Prevention:** 1. Monitor your OpenAI usage dashboard 2. Set up billing alerts 3. Implement usage tracking in your application 4. Use cheaper models for non-critical tasks ### Error: "Model Permission Denied" **Symptoms:** ``` Error: The model 'gpt-6-astra' does not exist or you do not have access to it Status Code: 404 ``` **Cause:** Trying to access a model not available in your OpenAI account tier. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" goprovider "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func getAvailableModel(provider *openai.Provider, preferredModels []string) (goprovider.LanguageModel, error) { ctx := context.Background() maxTokens := 5 // Try models in order of preference for _, modelName := range preferredModels { model, err := provider.LanguageModel(modelName) if err != nil { log.Printf("Model %s not available: %v", modelName, err) continue } // Test if model is actually accessible with a cheap request _, err = ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hi", MaxTokens: &maxTokens, }) if err == nil { log.Printf("Using model: %s", modelName) return model, nil } log.Printf("Model %s access denied: %v", modelName, err) } return nil, fmt.Errorf("no available models found") } func main() { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) // List models in order of preference preferredModels := []string{ "gpt-6-astra", "gpt-6-sol", "gpt-6-astra", "gpt-5.4-nano", } model, err := getAvailableModel(provider, preferredModels) if err != nil { log.Fatal(err) } ctx := context.Background() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Anthropic Provider Errors ### Error: "Invalid Authentication" **Symptoms:** ``` Error: invalid x-api-key Status Code: 401 ``` **Cause:** Invalid or missing Anthropic API key. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "strings" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func validateAnthropicKey(key string) error { if key == "" { return fmt.Errorf("API key is empty") } // Anthropic keys should start with "sk-ant-" if !strings.HasPrefix(key, "sk-ant-") { return fmt.Errorf("invalid Anthropic API key format (should start with 'sk-ant-')") } return nil } func main() { apiKey := os.Getenv("ANTHROPIC_API_KEY") // Validate key format before using if err := validateAnthropicKey(apiKey); err != nil { log.Fatalf("API key validation failed: %v", err) } provider := anthropic.New(anthropic.Config{ APIKey: apiKey, }) model, err := provider.LanguageModel("claude-sonnet-5-5") if err != nil { log.Fatalf("Failed to create model: %v", err) } ctx := context.Background() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { log.Fatalf("Failed to generate text: %v", err) } fmt.Println(result.Text) } ``` ### Error: "Overloaded Error" **Symptoms:** ``` Error: Overloaded Status Code: 529 ``` **Cause:** Anthropic's servers are temporarily overloaded. **Solution:** ```go package main import ( "context" "errors" "fmt" "log" "math" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func generateWithRetry(ctx context.Context, model provider.LanguageModel, prompt string, maxRetries int) (string, error) { var lastErr error for attempt := 0; attempt <= maxRetries; attempt++ { if attempt > 0 { // Exponential backoff with jitter backoff := time.Duration(math.Pow(2, float64(attempt-1))) * time.Second jitter := time.Duration(float64(backoff) * 0.1) wait := backoff + jitter log.Printf("Attempt %d/%d failed, waiting %v before retry...", attempt, maxRetries, wait) select { case <-time.After(wait): case <-ctx.Done(): return "", ctx.Err() } } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err == nil { return result.Text, nil } lastErr = err // Check if error is retryable var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { // Retry on 529 (overloaded) and 5xx errors if providerErr.StatusCode == 529 || (providerErr.StatusCode >= 500 && providerErr.StatusCode < 600) { log.Printf("Retryable error (status %d): %s", providerErr.StatusCode, providerErr.Message) continue } // Don't retry on 4xx errors (except 429) if providerErr.StatusCode >= 400 && providerErr.StatusCode < 500 && providerErr.StatusCode != 429 { return "", fmt.Errorf("non-retryable error: %w", err) } } } return "", fmt.Errorf("failed after %d attempts: %w", maxRetries+1, lastErr) } func main() { ctx := context.Background() provider := anthropic.New(anthropic.Config{ APIKey: os.Getenv("ANTHROPIC_API_KEY"), }) model, _ := provider.LanguageModel("claude-sonnet-5-5") text, err := generateWithRetry(ctx, model, "Explain quantum computing", 5) if err != nil { log.Fatalf("Failed: %v", err) } fmt.Println(text) } ``` ## Google (Gemini) Provider Errors ### Error: "API Key Not Valid" **Symptoms:** ``` Error: API key not valid. Please pass a valid API key Status Code: 400 ``` **Cause:** Invalid Google AI API key or using the wrong type of credentials. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/google" ) func main() { // Google AI Studio API key (not GCP service account) apiKey := os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY") if apiKey == "" { log.Fatal("GOOGLE_GENERATIVE_AI_API_KEY not set") } // For Google AI Studio (free tier) provider := google.New(google.Config{ APIKey: apiKey, }) model, err := provider.LanguageModel("gemini-2.5-pro") if err != nil { log.Fatalf("Failed to create model: %v", err) } ctx := context.Background() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { log.Fatalf("Failed to generate text: %v", err) } fmt.Println(result.Text) } ``` ### Error: "Resource Exhausted" **Symptoms:** ``` Error: Resource has been exhausted (e.g. check quota) Status Code: 429 ``` **Cause:** Exceeded Google AI API quota (RPM or TPM limits). **Solution:** ```go package main import ( "context" "errors" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" "github.com/digitallysavvy/go-ai/pkg/providers/google" "golang.org/x/time/rate" ) // Rate limiter for Google AI (15 RPM for free tier) var googleLimiter = rate.NewLimiter(rate.Limit(15.0/60.0), 1) func generateWithRateLimit(ctx context.Context, model provider.LanguageModel, prompt string) (string, error) { // Wait for rate limiter if err := googleLimiter.Wait(ctx); err != nil { return "", fmt.Errorf("rate limiter error: %w", err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { var rateLimitErr *providererrors.RateLimitError if errors.As(err, &rateLimitErr) { // Wait and retry if rate limited if rateLimitErr.RetryAfterSeconds != nil { waitDuration := time.Duration(*rateLimitErr.RetryAfterSeconds) * time.Second log.Printf("Rate limited, waiting %v...", waitDuration) select { case <-time.After(waitDuration): return generateWithRateLimit(ctx, model, prompt) case <-ctx.Done(): return "", ctx.Err() } } } return "", err } return result.Text, nil } func main() { ctx := context.Background() provider := google.New(google.Config{ APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY"), }) model, _ := provider.LanguageModel("gemini-2.5-flash") text, err := generateWithRateLimit(ctx, model, "Explain quantum computing") if err != nil { log.Fatalf("Failed: %v", err) } fmt.Println(text) } ``` ## Azure OpenAI Provider Errors ### Error: "Deployment Not Found" **Symptoms:** ``` Error: The API deployment for this resource does not exist Status Code: 404 ``` **Cause:** Using incorrect deployment name or endpoint URL. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/azure" ) func main() { // Azure OpenAI requires resource, key, and deployment configuration. resourceName := os.Getenv("AZURE_RESOURCE_NAME") apiKey := os.Getenv("AZURE_API_KEY") deployment := os.Getenv("AZURE_OPENAI_DEPLOYMENT") // e.g., "gpt-5-mini" if resourceName == "" || apiKey == "" || deployment == "" { log.Fatal("Missing Azure OpenAI configuration (resource name, key, or deployment)") } provider, err := azure.New(azure.Config{ ResourceName: resourceName, APIKey: apiKey, UseDeploymentBasedURLs: true, }) if err != nil { log.Fatalf("Failed to create provider: %v", err) } // Use the deployment name, not the model name model, err := provider.LanguageModel(deployment) if err != nil { log.Fatalf("Failed to create model: %v", err) } ctx := context.Background() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { log.Fatalf("Failed to generate text: %v", err) } fmt.Println(result.Text) } ``` ## AWS Bedrock Provider Errors ### Error: "Access Denied" **Symptoms:** ``` Error: User is not authorized to perform: bedrock:InvokeModel Status Code: 403 ``` **Cause:** IAM credentials don't have permission to access Bedrock. **Solution:** ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/bedrock" ) func main() { ctx := context.Background() // bedrock.Config signs requests itself (SigV4), so there is no AWS SDK // aws.Config to load. With AWSAccessKeyID/AWSSecretAccessKey left unset, // credentials and region fall back to the environment (AWS_REGION, // AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY), the shared credentials file, // or an instance role — the same chain the AWS CLI uses. provider := bedrock.New(bedrock.Config{ Region: "us-east-1", }) // Use full model ARN or ID model, err := provider.LanguageModel("anthropic.claude-sonnet-5-5") if err != nil { log.Fatalf("Failed to create model: %v", err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { log.Fatalf("Failed to generate text: %v", err) } fmt.Println(result.Text) } ``` **Required IAM Policy:** ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream" ], "Resource": "*" } ] } ``` ## Network and Connection Errors ### Error: "Connection Timeout" **Symptoms:** ``` Error: context deadline exceeded Error: dial tcp: i/o timeout ``` **Cause:** Network connectivity issues or firewall blocking API requests. **Solution:** ```go package main import ( "context" "fmt" "log" "net/http" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { // Create HTTP client with custom timeout httpClient := &http.Client{ Timeout: 2 * time.Minute, Transport: &http.Transport{ MaxIdleConns: 100, MaxIdleConnsPerHost: 10, IdleConnTimeout: 90 * time.Second, }, } provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), HTTPClient: httpClient, }) model, _ := provider.LanguageModel("gpt-6-astra") // Use context with longer timeout ctx, cancel := context.WithTimeout(context.Background(), 3*time.Minute) defer cancel() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a long essay", }) if err != nil { log.Fatalf("Failed: %v", err) } fmt.Println(result.Text) } ``` ## Best Practices ### 1. Implement Provider Fallback ```go providers := []struct { name string provider interface{} model string }{ {"openai", openaiProvider, "gpt-5.4-mini"}, {"anthropic", anthropicProvider, "claude-haiku-4-5"}, {"google", googleProvider, "gemini-2.5-flash"}, } for _, p := range providers { result, err := tryGenerate(p.provider, p.model, prompt) if err == nil { return result } log.Printf("Provider %s failed: %v", p.name, err) } ``` ### 2. Validate Configuration at Startup ```go func validateConfig() error { required := map[string]string{ "OPENAI_API_KEY": os.Getenv("OPENAI_API_KEY"), } for key, value := range required { if value == "" { return fmt.Errorf("missing required env var: %s", key) } } return nil } ``` ### 3. Log Provider Errors with Context ```go if err != nil { var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { log.Printf("Provider %s error (status %d): %s [code: %s]", providerErr.Provider, providerErr.StatusCode, providerErr.Message, providerErr.ErrorCode, ) } } ``` ### 4. Monitor Provider Health ```go type ProviderHealth struct { Name string Healthy bool LastCheck time.Time ErrorCount int } func checkProviderHealth(model provider.LanguageModel) ProviderHealth { ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second) defer cancel() maxTokens := 5 _, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "test", MaxTokens: &maxTokens, }) return ProviderHealth{ Healthy: err == nil, LastCheck: time.Now(), } } ``` ## See Also - [Rate Limits](https://goaisdk.com/docs/troubleshooting/rate-limits.md) - [Common Errors](https://goaisdk.com/docs/troubleshooting/common-errors.md) - [Error Handling Guide](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [Provider Documentation](https://goaisdk.com/docs/providers.md) --- # Rate Limit Issues > Covers handling rate limiting and quota exhaustion in the Go AI SDK, including client-side rate limiting, concurrent request management, and monitoring. Canonical URL: https://goaisdk.com/docs/troubleshooting/rate-limits Documentation index: https://goaisdk.com/llms.txt This guide covers how to handle rate limiting errors from AI providers and implement effective rate limiting strategies. ## Understanding Rate Limits AI providers enforce rate limits to prevent abuse and ensure fair resource allocation: - **RPM (Requests Per Minute)**: Maximum API calls per minute - **TPM (Tokens Per Minute)**: Maximum tokens processed per minute - **TPD (Tokens Per Day)**: Maximum tokens processed per day - **Concurrent Requests**: Maximum simultaneous requests ## Common Rate Limit Errors ### Error: "Rate Limit Exceeded" **Symptoms:** ``` Error: Rate limit exceeded Status Code: 429 X-RateLimit-Limit: 3500 X-RateLimit-Remaining: 0 Retry-After: 60 ``` **Cause:** Exceeded the provider's rate limit for requests or tokens. **Solution:** ```go package main import ( "context" "errors" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func generateWithRateLimitHandling(ctx context.Context, model provider.LanguageModel, prompt string) (string, error) { maxRetries := 3 var lastErr error for attempt := 0; attempt <= maxRetries; attempt++ { if attempt > 0 { log.Printf("Retry attempt %d/%d", attempt, maxRetries) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err == nil { return result.Text, nil } lastErr = err // Check if it's a rate limit error var rateLimitErr *providererrors.RateLimitError if errors.As(err, &rateLimitErr) { log.Printf("Rate limit hit for %s", rateLimitErr.Provider) // Use Retry-After header if provided if rateLimitErr.RetryAfterSeconds != nil { waitDuration := time.Duration(*rateLimitErr.RetryAfterSeconds) * time.Second log.Printf("Waiting %v before retry (from Retry-After header)...", waitDuration) select { case <-time.After(waitDuration): continue case <-ctx.Done(): return "", ctx.Err() } } // Default exponential backoff if no Retry-After backoff := time.Duration(1<= time.Minute { tb.tokens = tb.maxTokensPerMinute tb.lastRefill = now } else { // Proportional refill tokensToAdd := int(float64(tb.maxTokensPerMinute) * elapsed.Seconds() / 60.0) tb.tokens = min(tb.tokens+tokensToAdd, tb.maxTokensPerMinute) if tokensToAdd > 0 { tb.lastRefill = now } } } func (tb *TokenBucket) Wait(ctx context.Context, estimatedTokens int) error { tb.mu.Lock() defer tb.mu.Unlock() for { tb.refill() if tb.tokens >= estimatedTokens { tb.tokens -= estimatedTokens return nil } // Calculate wait time tokensNeeded := estimatedTokens - tb.tokens waitSeconds := float64(tokensNeeded) / float64(tb.maxTokensPerMinute) * 60.0 waitDuration := time.Duration(waitSeconds * float64(time.Second)) log.Printf("Not enough tokens, waiting %v...", waitDuration) // Release lock while waiting tb.mu.Unlock() select { case <-time.After(waitDuration): tb.mu.Lock() case <-ctx.Done(): tb.mu.Lock() return ctx.Err() } } } func min(a, b int) int { if a < b { return a } return b } // TokenAwareGenerator manages token-based rate limiting type TokenAwareGenerator struct { model provider.LanguageModel bucket *TokenBucket } func NewTokenAwareGenerator(model provider.LanguageModel, tokensPerMinute int) *TokenAwareGenerator { return &TokenAwareGenerator{ model: model, bucket: NewTokenBucket(tokensPerMinute), } } func (g *TokenAwareGenerator) Generate(ctx context.Context, prompt string, maxTokens int) (string, error) { // Estimate total tokens (prompt + response) estimatedPromptTokens := len(prompt) / 4 // Rough estimate: 1 token ~= 4 chars estimatedTotalTokens := estimatedPromptTokens + maxTokens // Wait for token availability if err := g.bucket.Wait(ctx, estimatedTotalTokens); err != nil { return "", err } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: g.model, Prompt: prompt, MaxTokens: &maxTokens, }) if err != nil { return "", err } // Update bucket with actual token usage actualTokens := int(result.Usage.GetTotalTokens()) difference := actualTokens - estimatedTotalTokens if difference > 0 { // We used more than estimated, remove additional tokens g.bucket.mu.Lock() g.bucket.tokens -= difference g.bucket.mu.Unlock() } return result.Text, nil } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Limit to 10,000 tokens per minute generator := NewTokenAwareGenerator(model, 10000) text, err := generator.Generate(ctx, "Explain quantum computing", 500) if err != nil { log.Fatal(err) } fmt.Println(text) } ``` ## Concurrent Request Management ### Limiting Concurrent Requests Use a semaphore to limit concurrent API calls: ```go package main import ( "context" "fmt" "log" "os" "sync" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // Semaphore limits concurrent operations type Semaphore struct { ch chan struct{} } func NewSemaphore(maxConcurrent int) *Semaphore { return &Semaphore{ ch: make(chan struct{}, maxConcurrent), } } func (s *Semaphore) Acquire(ctx context.Context) error { select { case s.ch <- struct{}{}: return nil case <-ctx.Done(): return ctx.Err() } } func (s *Semaphore) Release() { <-s.ch } // ConcurrentGenerator limits concurrent API requests type ConcurrentGenerator struct { model provider.LanguageModel semaphore *Semaphore } func NewConcurrentGenerator(model provider.LanguageModel, maxConcurrent int) *ConcurrentGenerator { return &ConcurrentGenerator{ model: model, semaphore: NewSemaphore(maxConcurrent), } } func (g *ConcurrentGenerator) Generate(ctx context.Context, prompt string) (string, error) { // Acquire semaphore if err := g.semaphore.Acquire(ctx); err != nil { return "", err } defer g.semaphore.Release() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: g.model, Prompt: prompt, }) if err != nil { return "", err } return result.Text, nil } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Limit to 5 concurrent requests generator := NewConcurrentGenerator(model, 5) prompts := []string{ "Explain quantum computing", "What is machine learning?", "Describe photosynthesis", "How do computers work?", "What is DNA?", "Explain relativity", "What is electricity?", "Describe the water cycle", } var wg sync.WaitGroup results := make([]string, len(prompts)) for i, prompt := range prompts { wg.Add(1) go func(index int, p string) { defer wg.Done() text, err := generator.Generate(ctx, p) if err != nil { log.Printf("Request %d failed: %v", index, err) return } results[index] = text fmt.Printf("Completed request %d\n", index) }(i, prompt) } wg.Wait() for i, result := range results { if result != "" { fmt.Printf("\n=== Result %d ===\n%s\n", i, result) } } } ``` ## Provider-Specific Rate Limits ### OpenAI Rate Limits ```go // Typical OpenAI limits: // - Free tier: 3 RPM, 40,000 TPM // - Tier 1: 500 RPM, 30,000 TPM // - Tier 2: 5,000 RPM, 450,000 TPM const ( OpenAITier1RPM = 500 OpenAITier1TPM = 30000 ) limiter := rate.NewLimiter(rate.Limit(float64(OpenAITier1RPM)/60.0), 5) ``` ### Anthropic Rate Limits ```go // Claude typical limits: // - Free tier: 5 RPM, 20,000 TPM // - Tier 1: 50 RPM, 40,000 TPM // - Tier 2: 1,000 RPM, 80,000 TPM const ( AnthropicTier1RPM = 50 AnthropicTier1TPM = 40000 ) limiter := rate.NewLimiter(rate.Limit(float64(AnthropicTier1RPM)/60.0), 3) ``` ### Google AI Rate Limits ```go // Gemini typical limits: // - Free tier: 15 RPM, 32,000 TPM // - Paid tier: 1,000 RPM, 4,000,000 TPM const ( GoogleFreeRPM = 15 GoogleFreTPM = 32000 ) limiter := rate.NewLimiter(rate.Limit(float64(GoogleFreeRPM)/60.0), 1) ``` ## Monitoring Rate Limit Usage ### Track and Log Rate Limits ```go package main import ( "context" "fmt" "log" "os" "sync" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // RateLimitTracker monitors rate limit usage type RateLimitTracker struct { requestCount int tokenCount int periodStart time.Time mu sync.Mutex } func NewRateLimitTracker() *RateLimitTracker { return &RateLimitTracker{ periodStart: time.Now(), } } func (t *RateLimitTracker) RecordRequest(tokens int) { t.mu.Lock() defer t.mu.Unlock() now := time.Now() if now.Sub(t.periodStart) >= time.Minute { // New period log.Printf("Rate limit stats - Last minute: %d requests, %d tokens", t.requestCount, t.tokenCount) t.requestCount = 0 t.tokenCount = 0 t.periodStart = now } t.requestCount++ t.tokenCount += tokens } func (t *RateLimitTracker) GetStats() (requests, tokens int, elapsed time.Duration) { t.mu.Lock() defer t.mu.Unlock() return t.requestCount, t.tokenCount, time.Since(t.periodStart) } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") tracker := NewRateLimitTracker() // Start monitoring goroutine go func() { ticker := time.NewTicker(10 * time.Second) defer ticker.Stop() for range ticker.C { requests, tokens, elapsed := tracker.GetStats() rpm := float64(requests) / elapsed.Minutes() tpm := float64(tokens) / elapsed.Minutes() log.Printf("Current rate: %.1f RPM, %.1f TPM", rpm, tpm) } }() // Make requests for i := 0; i < 20; i++ { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Request %d", i), }) if err != nil { log.Printf("Request failed: %v", err) continue } tracker.RecordRequest(int(result.Usage.GetTotalTokens())) time.Sleep(2 * time.Second) // Space out requests } } ``` ## Best Practices ### 1. Implement Retry with Backoff Always respect `Retry-After` headers and use exponential backoff: ```go if rateLimitErr.RetryAfterSeconds != nil { time.Sleep(time.Duration(*rateLimitErr.RetryAfterSeconds) * time.Second) } else { time.Sleep(time.Duration(1< 0.8 * maxRPM { log.Println("WARN: approaching rate limit threshold") } ``` ### 4. Use Appropriate Batch Sizes Don't send too many requests at once: ```go batchSize := 5 for i := 0; i < len(prompts); i += batchSize { batch := prompts[i:min(i+batchSize, len(prompts))] processBatch(batch) time.Sleep(time.Minute) // Wait between batches } ``` ### 5. Implement Circuit Breakers Stop making requests after repeated rate limit errors: ```go if consecutiveRateLimitErrors > 3 { log.Println("ERROR: too many rate limit errors, circuit breaker activated") time.Sleep(5 * time.Minute) consecutiveRateLimitErrors = 0 } ``` ## See Also - [Provider Errors](https://goaisdk.com/docs/troubleshooting/provider-errors.md) - [Performance Optimization](https://goaisdk.com/docs/troubleshooting/performance.md) - [Rate Limiting Guide](https://goaisdk.com/docs/advanced/rate-limiting.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) --- # Context Cancellation > Debugging guide for Go context timeout and cancellation issues in the Go AI SDK, covering propagation, goroutines, and streaming context problems. Canonical URL: https://goaisdk.com/docs/troubleshooting/context-cancellation Documentation index: https://goaisdk.com/llms.txt This guide covers common issues with Go contexts in the AI SDK, including timeouts, cancellations, and proper context handling. ## Understanding Go Contexts Go's `context.Context` is used for: - **Cancellation**: Stop operations that are no longer needed - **Timeouts**: Enforce maximum duration for operations - **Deadlines**: Set absolute time limits - **Values**: Pass request-scoped data (use sparingly) ## Common Context Errors ### Error: "Context Deadline Exceeded" **Symptoms:** ``` Error: context deadline exceeded Operation timed out after 30s ``` **Cause:** Operation took longer than the context timeout allows. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // PROBLEM: Timeout too short for long operations ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a comprehensive 5000-word essay on quantum physics", }) if err != nil { if ctx.Err() == context.DeadlineExceeded { log.Println("Operation timed out - need longer timeout or streaming") // SOLUTION 1: Increase timeout ctx2, cancel2 := context.WithTimeout(context.Background(), 2*time.Minute) defer cancel2() result, err = ai.GenerateText(ctx2, ai.GenerateTextOptions{ Model: model, Prompt: "Write a comprehensive 5000-word essay on quantum physics", }) if err != nil { log.Printf("Still failed with longer timeout: %v", err) // SOLUTION 2: Use streaming instead ctx3 := context.Background() stream, err := ai.StreamText(ctx3, ai.StreamTextOptions{ Model: model, Prompt: "Write a comprehensive 5000-word essay on quantum physics", }) if err != nil { log.Fatal(err) } for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } if err := stream.Err(); err != nil { log.Fatal(err) } return } } else { log.Fatal(err) } } fmt.Println(result.Text) } ``` ### Error: "Context Canceled" **Symptoms:** ``` Error: context canceled Stream terminated early ``` **Cause:** Context was explicitly canceled before operation completed. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "os/signal" "syscall" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Create cancellable context for graceful shutdown ctx, cancel := context.WithCancel(context.Background()) defer cancel() // Handle OS signals for graceful shutdown sigCh := make(chan os.Signal, 1) signal.Notify(sigCh, os.Interrupt, syscall.SIGTERM) go func() { <-sigCh log.Println("Received interrupt signal, canceling operations...") cancel() }() stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long story about space exploration", }) if err != nil { log.Fatal(err) } for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } // Check if canceled or completed normally if err := stream.Err(); err != nil { if ctx.Err() == context.Canceled { log.Println("\nOperation was canceled by user") // Perform cleanup savePartialResults() } else { log.Printf("\nStream error: %v", err) } } else { log.Println("\nOperation completed successfully") } } func savePartialResults() { // Save any partial results before exiting log.Println("Saving partial results...") } ``` ## Context Propagation ### Error: Missing Context Propagation **Symptoms:** - Child operations don't cancel when parent is canceled - Timeouts not enforced in nested calls **Cause:** Not passing context through the call chain. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // BAD: Creating new context instead of propagating func processPromptBad(prompt string) (string, error) { // This ignores any parent cancellation or timeout! ctx := context.Background() provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return "", err } return result.Text, nil } // GOOD: Propagating context through call chain func processPromptGood(ctx context.Context, prompt string) (string, error) { // Use the provided context so parent cancellation/timeout applies provider := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return "", err } return result.Text, nil } func main() { // Parent context with timeout ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() // CORRECT: Pass context to children text, err := processPromptGood(ctx, "Explain quantum computing") if err != nil { if ctx.Err() == context.DeadlineExceeded { log.Println("Entire operation timed out") } else { log.Printf("Error: %v", err) } return } fmt.Println(text) } ``` ## Streaming Context Issues ### Error: Stream Doesn't Respect Context **Symptoms:** - Stream continues after context canceled - Goroutine leak when stream not fully consumed **Cause:** Not checking context in stream processing loop. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func streamWithTimeout(ctx context.Context, model provider.LanguageModel, prompt string) error { stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return err } // Process stream with context checking for { select { case <-ctx.Done(): // Context canceled or timed out log.Println("Context done, stopping stream processing") return ctx.Err() case chunk, ok := <-stream.Chunks(): if !ok { // Channel closed, stream finished return stream.Err() } fmt.Print(chunk.Text) } } } func main() { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Set timeout for entire stream ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second) defer cancel() if err := streamWithTimeout(ctx, model, "Write a very long story"); err != nil { if ctx.Err() == context.DeadlineExceeded { log.Println("\nStream timed out after 15 seconds") } else if ctx.Err() == context.Canceled { log.Println("\nStream was canceled") } else { log.Printf("\nStream error: %v", err) } } } ``` ## Context in Goroutines ### Error: Context Not Passed to Goroutines **Symptoms:** - Goroutines don't stop when parent context canceled - Memory leaks from orphaned goroutines **Cause:** Not passing context to goroutines or creating new contexts. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "sync" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func processPrompts(ctx context.Context, model provider.LanguageModel, prompts []string) []string { results := make([]string, len(prompts)) var wg sync.WaitGroup for i, prompt := range prompts { wg.Add(1) // CORRECT: Pass context to goroutine go func(index int, p string) { defer wg.Done() // Check if context already canceled before starting work select { case <-ctx.Done(): log.Printf("Goroutine %d: context canceled before starting", index) return default: } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: p, }) if err != nil { if ctx.Err() != nil { log.Printf("Goroutine %d: canceled during execution", index) } else { log.Printf("Goroutine %d: error: %v", index, err) } return } results[index] = result.Text log.Printf("Goroutine %d: completed", index) }(i, prompt) } // Wait for all goroutines with timeout awareness done := make(chan struct{}) go func() { wg.Wait() close(done) }() select { case <-done: log.Println("All goroutines completed") case <-ctx.Done(): log.Println("Context canceled, waiting for goroutines to cleanup...") <-done // Still wait for goroutines to finish cleanup } return results } func main() { provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Context with timeout ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() prompts := []string{ "Explain AI", "Explain ML", "Explain DL", "Explain NLP", "Explain CV", } results := processPrompts(ctx, model, prompts) for i, result := range results { if result != "" { fmt.Printf("\n=== Result %d ===\n%s\n", i, result) } } } ``` ## Context Value Issues ### Error: Misusing Context Values **Symptoms:** - Type assertions panic - Values not found in context **Cause:** Incorrect use of context.WithValue. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // Define custom type for context keys (prevents collisions) type contextKey string const ( userIDKey contextKey = "userID" requestIDKey contextKey = "requestID" ) // Helper functions to set/get values safely func WithUserID(ctx context.Context, userID string) context.Context { return context.WithValue(ctx, userIDKey, userID) } func GetUserID(ctx context.Context) (string, bool) { userID, ok := ctx.Value(userIDKey).(string) return userID, ok } func WithRequestID(ctx context.Context, requestID string) context.Context { return context.WithValue(ctx, requestIDKey, requestID) } func GetRequestID(ctx context.Context) (string, bool) { requestID, ok := ctx.Value(requestIDKey).(string) return requestID, ok } func processWithContext(ctx context.Context, prompt string) (string, error) { // Safely extract context values userID, hasUser := GetUserID(ctx) requestID, hasRequest := GetRequestID(ctx) if hasUser { log.Printf("Processing request for user: %s", userID) } if hasRequest { log.Printf("Request ID: %s", requestID) } provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return "", err } return result.Text, nil } func main() { ctx := context.Background() // Add request metadata to context ctx = WithUserID(ctx, "user-123") ctx = WithRequestID(ctx, "req-456") text, err := processWithContext(ctx, "Explain quantum computing") if err != nil { log.Fatal(err) } fmt.Println(text) } ``` ## Best Practices ### 1. Always Pass Context as First Parameter ```go // Good func processPrompt(ctx context.Context, prompt string) (string, error) // Bad func processPrompt(prompt string) (string, error) ``` ### 2. Never Store Context in Struct ```go // Bad - don't store context type Processor struct { ctx context.Context // NEVER do this } // Good - pass context to methods type Processor struct { // other fields } func (p *Processor) Process(ctx context.Context) error { // use ctx parameter } ``` ### 3. Use Appropriate Timeouts ```go // Short operations ctx, cancel := context.WithTimeout(ctx, 10*time.Second) defer cancel() // Long operations ctx, cancel := context.WithTimeout(ctx, 2*time.Minute) defer cancel() // Streaming - no timeout or very long timeout ctx := context.Background() // or ctx, cancel := context.WithTimeout(ctx, 10*time.Minute) defer cancel() ``` ### 4. Always Defer cancel() ```go ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() // Always defer, even if context times out naturally ``` ### 5. Check Context Error After Operations ```go result, err := ai.GenerateText(ctx, options) if err != nil { if ctx.Err() == context.DeadlineExceeded { // Handle timeout } else if ctx.Err() == context.Canceled { // Handle cancellation } else { // Handle other errors } } ``` ### 6. Don't Ignore Context in OnFinish ```go stream, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "...", OnFinish: func(ctx context.Context, result *ai.StreamTextResult, userContext interface{}) { // OnFinish receives the same context - check it! if ctx.Err() != nil { log.Println("Context was canceled, skipping post-processing") return } // Process result... }, }) ``` ### 7. Use Context for Graceful Shutdown ```go ctx, cancel := context.WithCancel(context.Background()) // Handle shutdown signals sigCh := make(chan os.Signal, 1) signal.Notify(sigCh, os.Interrupt, syscall.SIGTERM) go func() { <-sigCh log.Println("Shutting down...") cancel() }() // All operations use this context processRequests(ctx) ``` ## See Also - [Common Errors](https://goaisdk.com/docs/troubleshooting/common-errors.md) - [Streaming Issues](https://goaisdk.com/docs/troubleshooting/streaming-issues.md) - [Performance Optimization](https://goaisdk.com/docs/troubleshooting/performance.md) - [Error Handling Guide](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) --- # Streaming Issues > Solves common Go AI SDK streaming problems, covering channel-related issues, stream interruption, concurrent streaming, and buffer and memory issues. Canonical URL: https://goaisdk.com/docs/troubleshooting/streaming-issues Documentation index: https://goaisdk.com/llms.txt This guide covers common problems with streaming AI responses and their solutions. ## Channel-Related Issues ### Error: "Goroutine Leak - Stream Not Consumed" **Symptoms:** - Application hangs - Memory usage grows continuously - `runtime.NumGoroutine()` keeps increasing **Cause:** Not consuming all values from stream channels, causing producer goroutines to block. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // BAD: Not consuming entire stream func badStreamHandling(ctx context.Context, model provider.LanguageModel) error { stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long story", }) if err != nil { return err } // Only reading first chunk - LEAK! chunk := <-stream.Chunks() fmt.Println(chunk.Text) // Returning without consuming rest of channel // Producer goroutine will block forever return nil } // GOOD: Always consume entire stream func goodStreamHandling(ctx context.Context, model provider.LanguageModel) error { stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long story", }) if err != nil { return err } // Range loop consumes all values until channel closes for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } // Check for errors after stream completes return stream.Err() } // ALSO GOOD: Early exit with context cancellation func earlyExitStreamHandling(ctx context.Context, model provider.LanguageModel) error { stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long story", }) if err != nil { return err } count := 0 for chunk := range stream.Chunks() { fmt.Print(chunk.Text) count++ // If you need to exit early, cancel the context // This will cause the stream to close cleanly if count >= 10 { // Stream will be canceled and channel will close return nil } } return stream.Err() } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") if err := goodStreamHandling(ctx, model); err != nil { log.Fatal(err) } } ``` ### Error: "Send on Closed Channel" **Symptoms:** ``` panic: send on closed channel ``` **Cause:** Trying to read from a stream channel multiple times or after it's closed. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a short story", }) if err != nil { log.Fatal(err) } // First read - OK for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } // Channel is now closed // BAD: Don't try to read again // for chunk := range stream.Chunks() { // Channel is closed! // fmt.Print(chunk.Text) // } // GOOD: Only check error after stream closes if err := stream.Err(); err != nil { log.Printf("Stream error: %v", err) } // Don't reuse the stream object } ``` ## Stream Interruption Issues ### Error: "Stream Stops Prematurely" **Symptoms:** - Stream closes before all content received - Incomplete responses - No error reported **Cause:** Network issues, provider errors, or context cancellation. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "strings" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func robustStreaming(ctx context.Context, model provider.LanguageModel, prompt string) error { stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return fmt.Errorf("failed to start stream: %w", err) } var fullText strings.Builder var chunkCount int var finishReason string // Use Chunks to get detailed chunk information for chunk := range stream.Chunks() { switch chunk.Type { case provider.ChunkTypeText: fmt.Print(chunk.Text) fullText.WriteString(chunk.Text) chunkCount++ case provider.ChunkTypeToolCall: log.Printf("Tool call: %s", chunk.ToolCall.ToolName) case provider.ChunkTypeError: log.Printf("Error in stream: %v", chunk.Err) case provider.ChunkTypeFinish: finishReason = string(chunk.FinishReason) log.Printf("\nStream finished: %s", finishReason) } } // Check for errors if err := stream.Err(); err != nil { log.Printf("Stream error (received %d chunks): %v", chunkCount, err) return err } // Validate stream completed properly if chunkCount == 0 { return fmt.Errorf("stream produced no chunks") } if finishReason == "" { log.Println("Warning: Stream ended without finish reason") } log.Printf("Received %d chunks, %d bytes", chunkCount, fullText.Len()) return nil } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") if err := robustStreaming(ctx, model, "Write a story about AI"); err != nil { log.Fatalf("Streaming failed: %v", err) } } ``` ### Error: "Missing Final Chunks" **Symptoms:** - Last few chunks not received - Response appears incomplete **Cause:** Not waiting for channel to close naturally. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func completeStreamRead(ctx context.Context, model provider.LanguageModel) error { stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Count from 1 to 100", }) if err != nil { return err } timeout := time.After(5 * time.Minute) for { select { case chunk, ok := <-stream.Chunks(): if !ok { // Channel properly closed, all chunks received log.Println("\nStream completed") return stream.Err() } fmt.Print(chunk.Text) case <-timeout: log.Println("\nTimeout waiting for stream") return fmt.Errorf("stream timeout") case <-ctx.Done(): log.Println("\nContext canceled") return ctx.Err() } } } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") if err := completeStreamRead(ctx, model); err != nil { log.Fatalf("Error: %v", err) } } ``` ## Concurrent Streaming Issues ### Error: "Race Condition in Stream Processing" **Symptoms:** - Panic: concurrent map read/write - Data corruption - Unpredictable behavior **Cause:** Multiple goroutines accessing stream data without synchronization. **Solution:** ```go package main import ( "context" "fmt" "log" "os" "sync" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // SafeStreamProcessor handles concurrent stream processing safely type SafeStreamProcessor struct { mu sync.Mutex chunks []string done bool } func (p *SafeStreamProcessor) AddChunk(chunk string) { p.mu.Lock() defer p.mu.Unlock() p.chunks = append(p.chunks, chunk) } func (p *SafeStreamProcessor) GetChunks() []string { p.mu.Lock() defer p.mu.Unlock() result := make([]string, len(p.chunks)) copy(result, p.chunks) return result } func (p *SafeStreamProcessor) MarkDone() { p.mu.Lock() defer p.mu.Unlock() p.done = true } func (p *SafeStreamProcessor) IsDone() bool { p.mu.Lock() defer p.mu.Unlock() return p.done } func processStreamSafely(ctx context.Context, model provider.LanguageModel, prompt string) error { processor := &SafeStreamProcessor{} stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return err } var wg sync.WaitGroup // Goroutine 1: Read and store chunks wg.Add(1) go func() { defer wg.Done() for chunk := range stream.Chunks() { processor.AddChunk(chunk.Text) fmt.Print(chunk.Text) } processor.MarkDone() }() // Goroutine 2: Monitor progress wg.Add(1) go func() { defer wg.Done() for !processor.IsDone() { chunks := processor.GetChunks() log.Printf("Progress: %d chunks received", len(chunks)) select { case <-ctx.Done(): return case <-time.After(2 * time.Second): } } }() wg.Wait() return stream.Err() } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") if err := processStreamSafely(ctx, model, "Write a story"); err != nil { log.Fatal(err) } } ``` ## Buffer and Memory Issues ### Error: "Out of Memory with Streaming" **Symptoms:** - High memory usage - OOM errors - Application crashes **Cause:** Accumulating all chunks in memory instead of processing them incrementally. **Solution:** ```go package main import ( "bufio" "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // BAD: Accumulating all chunks in memory func memoryHungryStreaming(ctx context.Context, model provider.LanguageModel) error { stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a very long essay", }) if err != nil { return err } var allText string // Accumulates all chunks - can use lots of memory! for chunk := range stream.Chunks() { allText += chunk.Text // Bad: growing string } fmt.Println(allText) return stream.Err() } // GOOD: Process chunks incrementally func memoryEfficientStreaming(ctx context.Context, model provider.LanguageModel, outputFile string) error { stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a very long essay", }) if err != nil { return err } // Write directly to file instead of accumulating in memory file, err := os.Create(outputFile) if err != nil { return err } defer file.Close() writer := bufio.NewWriter(file) defer writer.Flush() for chunk := range stream.Chunks() { // Write chunk immediately, don't accumulate if _, err := writer.WriteString(chunk.Text); err != nil { return err } // Also print to console fmt.Print(chunk.Text) } return stream.Err() } // ALSO GOOD: Process chunks with fixed buffer func bufferedStreaming(ctx context.Context, model provider.LanguageModel) error { stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long story", }) if err != nil { return err } const maxBufferSize = 1000 // Process in chunks of 1000 chars buffer := make([]byte, 0, maxBufferSize) for chunk := range stream.Chunks() { buffer = append(buffer, []byte(chunk.Text)...) // Process buffer when full if len(buffer) >= maxBufferSize { // Process the buffer (e.g., write to file, send to client, etc.) processBuffer(buffer) buffer = buffer[:0] // Reset buffer } } // Process remaining buffer if len(buffer) > 0 { processBuffer(buffer) } return stream.Err() } func processBuffer(data []byte) { // Process data (write to file, database, etc.) fmt.Print(string(data)) } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") if err := memoryEfficientStreaming(ctx, model, "output.txt"); err != nil { log.Fatal(err) } } ``` ## Best Practices ### 1. Always Consume Full Stream ```go // Use range to ensure all chunks consumed for chunk := range stream.Chunks() { process(chunk) } // Channel is now properly closed ``` ### 2. Check Errors After Stream ```go for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } if err := stream.Err(); err != nil { log.Printf("Stream error: %v", err) } ``` ### 3. Use Chunks for Detailed Info ```go for chunk := range stream.Chunks() { switch chunk.Type { case provider.ChunkTypeText: // Handle text case provider.ChunkTypeError: // Handle errors mid-stream case provider.ChunkTypeFinish: // Handle completion } } ``` ### 4. Process Incrementally ```go // Don't accumulate all chunks var allText string // BAD // Process each chunk as it arrives for chunk := range stream.Chunks() { processImmediately(chunk) // GOOD } ``` ### 5. Use Timeouts for Streams ```go ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute) defer cancel() stream, _ := ai.StreamText(ctx, options) ``` ### 6. Handle Context Cancellation ```go for { select { case chunk, ok := <-stream.Chunks(): if !ok { return stream.Err() } process(chunk) case <-ctx.Done(): return ctx.Err() } } ``` ## See Also - [Context Cancellation](https://goaisdk.com/docs/troubleshooting/context-cancellation.md) - [Performance Optimization](https://goaisdk.com/docs/troubleshooting/performance.md) - [Common Errors](https://goaisdk.com/docs/troubleshooting/common-errors.md) - [Streaming Guide](https://goaisdk.com/docs/foundations/streaming.md) --- # Tool Calling Errors > Debugging guide for Go AI SDK tool execution and schema issues, covering tool definition errors, tool choice problems, and async tool execution. Canonical URL: https://goaisdk.com/docs/troubleshooting/tool-calling-errors Documentation index: https://goaisdk.com/llms.txt This guide covers common errors when using tools with language models and how to resolve them. ## Tool Definition Errors ### Error: "Invalid Tool Schema" **Symptoms:** ``` Error: tool schema validation failed Error: invalid parameter type in tool definition ``` **Cause:** Tool parameters schema doesn't match expected JSON schema format. **Solution:** ```go skip-compile package main import ( "context" "encoding/json" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/invopop/jsonschema" ) // Define input struct with proper JSON tags type WeatherInput struct { Location string `json:"location" jsonschema:"required,description=City name"` Unit string `json:"unit" jsonschema:"enum=celsius,enum=fahrenheit,description=Temperature unit"` } // BAD: Manually creating schema (error-prone) func badToolDefinition() types.Tool { return types.Tool{ Name: "get_weather", Description: "Get weather for a location", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "location": map[string]interface{}{ "type": "string", // Missing "description" and other fields }, }, // Missing "required" array }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return "Sunny", nil }, } } // GOOD: Using jsonschema to generate schema func goodToolDefinition() types.Tool { reflector := jsonschema.Reflector{ AllowAdditionalProperties: false, DoNotReference: true, } schema := reflector.Reflect(&WeatherInput{}) return types.Tool{ Name: "get_weather", Description: "Get the current weather for a location", Parameters: schema, // Execute always takes (ctx, input map[string]interface{}, opts // types.ToolExecutionOptions) — the SDK decodes the model's JSON // arguments into the map before calling Execute. Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { inputBytes, err := json.Marshal(input) if err != nil { return nil, fmt.Errorf("invalid input: %w", err) } var params WeatherInput if err := json.Unmarshal(inputBytes, ¶ms); err != nil { return nil, fmt.Errorf("invalid input: %w", err) } // Simulate weather lookup return map[string]interface{}{ "location": params.Location, "temperature": 72, "unit": params.Unit, "condition": "sunny", }, nil }, } } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") weatherTool := goodToolDefinition() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What's the weather in San Francisco?", Tools: []types.Tool{weatherTool}, // GenerateText defaults to a single step; allow more so the model // can see the tool result and produce a final text answer. StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Tool Execution Errors ### Error: "Tool Execution Failed" **Symptoms:** ``` Error: tool execution failed: panic in tool function Error: tool returned invalid type ``` **Cause:** Runtime error in tool execution function or invalid return value. **Solution:** ```go skip-compile package main import ( "context" "encoding/json" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/invopop/jsonschema" ) type CalculatorInput struct { Operation string `json:"operation" jsonschema:"required,enum=add,enum=subtract,enum=multiply,enum=divide"` A float64 `json:"a" jsonschema:"required"` B float64 `json:"b" jsonschema:"required"` } // BAD: No error handling in tool func unsafeCalculatorTool() types.Tool { reflector := jsonschema.Reflector{ DoNotReference: true, } schema := reflector.Reflect(&CalculatorInput{}) return types.Tool{ Name: "calculator", Description: "Perform basic arithmetic", Parameters: schema, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { inputBytes, _ := json.Marshal(input) var params CalculatorInput json.Unmarshal(inputBytes, ¶ms) // Ignoring error! // No validation - can panic on divide by zero! var result float64 switch params.Operation { case "divide": result = params.A / params.B // Panic if B is 0! } return result, nil }, } } // GOOD: Proper error handling and validation func safeCalculatorTool() types.Tool { reflector := jsonschema.Reflector{ AllowAdditionalProperties: false, DoNotReference: true, } schema := reflector.Reflect(&CalculatorInput{}) return types.Tool{ Name: "calculator", Description: "Perform basic arithmetic operations", Parameters: schema, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Validate input JSON inputBytes, err := json.Marshal(input) if err != nil { return nil, fmt.Errorf("invalid input: %w", err) } var params CalculatorInput if err := json.Unmarshal(inputBytes, ¶ms); err != nil { return nil, fmt.Errorf("invalid input JSON: %w", err) } // Validate operation validOps := map[string]bool{ "add": true, "subtract": true, "multiply": true, "divide": true, } if !validOps[params.Operation] { return nil, fmt.Errorf("invalid operation: %s", params.Operation) } // Perform calculation with error handling var result float64 switch params.Operation { case "add": result = params.A + params.B case "subtract": result = params.A - params.B case "multiply": result = params.A * params.B case "divide": if params.B == 0 { return nil, fmt.Errorf("division by zero") } result = params.A / params.B } // Return structured result return map[string]interface{}{ "operation": params.Operation, "result": result, "inputs": map[string]float64{"a": params.A, "b": params.B}, }, err }, } } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") calcTool := safeCalculatorTool() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What is 144 divided by 12?", Tools: []types.Tool{calcTool}, // GenerateText defaults to a single step; allow more so the model // can see the tool result and produce a final text answer. StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Tool Choice Issues ### Error: "Tool Not Called When Expected" **Symptoms:** - Model doesn't call the tool - Model hallucinates answers instead of using tool **Cause:** Poor tool description or model doesn't understand when to use the tool. **Solution:** ```go skip-compile package main import ( "context" "encoding/json" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/invopop/jsonschema" ) type TimeInput struct { Timezone string `json:"timezone" jsonschema:"required,description=IANA timezone name (e.g. America/New_York)"` } // BAD: Vague description func vagueTimeTool() types.Tool { reflector := jsonschema.Reflector{DoNotReference: true} schema := reflector.Reflect(&TimeInput{}) return types.Tool{ Name: "time", Description: "Gets time", // Too vague! Parameters: schema, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { inputBytes, _ := json.Marshal(input) var params TimeInput json.Unmarshal(inputBytes, ¶ms) return time.Now().Format(time.RFC3339), nil }, } } // GOOD: Clear, specific description func clearTimeTool() types.Tool { reflector := jsonschema.Reflector{ AllowAdditionalProperties: false, DoNotReference: true, } schema := reflector.Reflect(&TimeInput{}) return types.Tool{ Name: "get_current_time", Description: `Get the current time for a specific timezone. Use this tool when the user asks for the current time, date, or what time it is in a specific location. Always use this tool instead of guessing the current time.`, Parameters: schema, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { inputBytes, err := json.Marshal(input) if err != nil { return nil, err } var params TimeInput if err := json.Unmarshal(inputBytes, ¶ms); err != nil { return nil, err } loc, err := time.LoadLocation(params.Timezone) if err != nil { return nil, fmt.Errorf("invalid timezone: %w", err) } now := time.Now().In(loc) return map[string]interface{}{ "timezone": params.Timezone, "time": now.Format(time.RFC3339), "formatted": now.Format("Monday, January 2, 2006 at 3:04 PM MST"), }, nil }, } } // Force tool usage with ToolChoice func forceToolUsage() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") timeTool := clearTimeTool() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "What time is it in Tokyo?", Tools: []types.Tool{timeTool}, // Force the model to use the tool ToolChoice: types.ToolChoice{Type: types.ToolChoiceRequired}, // GenerateText defaults to a single step; allow more so the model // can see the tool result and produce a final text answer. StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } func main() { forceToolUsage() } ``` ## Async Tool Execution Errors ### Error: "Tool Execution Timeout" **Symptoms:** - Tool execution takes too long - Application hangs waiting for tool **Cause:** Tool performs slow operations (API calls, database queries) without timeout. **Solution:** ```go skip-compile package main import ( "context" "encoding/json" "fmt" "log" "net/http" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/invopop/jsonschema" ) type SearchInput struct { Query string `json:"query" jsonschema:"required"` } // BAD: No timeout on external calls func slowSearchTool() types.Tool { reflector := jsonschema.Reflector{DoNotReference: true} schema := reflector.Reflect(&SearchInput{}) return types.Tool{ Name: "web_search", Description: "Search the web", Parameters: schema, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { inputBytes, _ := json.Marshal(input) var params SearchInput json.Unmarshal(inputBytes, ¶ms) // BAD: No timeout, can hang forever resp, err := http.Get("https://api.example.com/search?q=" + params.Query) if err != nil { return nil, err } defer resp.Body.Close() return "results", nil }, } } // GOOD: Timeout on external calls func timeoutSearchTool() types.Tool { reflector := jsonschema.Reflector{ AllowAdditionalProperties: false, DoNotReference: true, } schema := reflector.Reflect(&SearchInput{}) return types.Tool{ Name: "web_search", Description: "Search the web for current information", Parameters: schema, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { inputBytes, err := json.Marshal(input) if err != nil { return nil, err } var params SearchInput if err := json.Unmarshal(inputBytes, ¶ms); err != nil { return nil, err } // Create context with timeout for external call ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second) defer cancel() // Create request with context req, err := http.NewRequestWithContext(ctx, "GET", "https://api.example.com/search?q="+params.Query, nil) if err != nil { return nil, err } client := &http.Client{ Timeout: 10 * time.Second, } resp, err := client.Do(req) if err != nil { if ctx.Err() == context.DeadlineExceeded { return nil, fmt.Errorf("search timeout after 10s") } return nil, fmt.Errorf("search failed: %w", err) } defer resp.Body.Close() // Process response... return map[string]interface{}{ "query": params.Query, "results": []string{"result1", "result2"}, }, nil }, } } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") searchTool := timeoutSearchTool() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Search for the latest news about AI", Tools: []types.Tool{searchTool}, // GenerateText defaults to a single step; allow more so the model // can see the tool result and produce a final text answer. StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Multiple Tool Calls ### Error: "Tool Execution Order Issues" **Symptoms:** - Tools called in wrong order - Dependencies between tools not handled **Cause:** Not managing tool call dependencies properly. **Solution:** ```go skip-compile package main import ( "context" "encoding/json" "fmt" "log" "os" "sync" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/invopop/jsonschema" ) // Tool execution coordinator type ToolCoordinator struct { mu sync.Mutex results map[string]interface{} } func NewToolCoordinator() *ToolCoordinator { return &ToolCoordinator{ results: make(map[string]interface{}), } } func (tc *ToolCoordinator) StoreResult(toolName string, result interface{}) { tc.mu.Lock() defer tc.mu.Unlock() tc.results[toolName] = result } func (tc *ToolCoordinator) GetResult(toolName string) (interface{}, bool) { tc.mu.Lock() defer tc.mu.Unlock() result, ok := tc.results[toolName] return result, ok } type GetUserInput struct { UserID string `json:"user_id" jsonschema:"required"` } type GetOrdersInput struct { UserID string `json:"user_id" jsonschema:"required"` } func createToolsWithDependencies(coordinator *ToolCoordinator) []types.Tool { reflector := jsonschema.Reflector{ AllowAdditionalProperties: false, DoNotReference: true, } // Tool 1: Get user info userSchema := reflector.Reflect(&GetUserInput{}) getUserTool := types.Tool{ Name: "get_user", Description: "Get user information by user ID", Parameters: userSchema, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { inputBytes, err := json.Marshal(input) if err != nil { return nil, err } var params GetUserInput if err := json.Unmarshal(inputBytes, ¶ms); err != nil { return nil, err } result := map[string]interface{}{ "user_id": params.UserID, "name": "John Doe", "email": "john@example.com", } // Store result for dependent tools coordinator.StoreResult("get_user", result) return result, nil }, } // Tool 2: Get orders (depends on user info) ordersSchema := reflector.Reflect(&GetOrdersInput{}) getOrdersTool := types.Tool{ Name: "get_orders", Description: "Get orders for a user", Parameters: ordersSchema, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { inputBytes, err := json.Marshal(input) if err != nil { return nil, err } var params GetOrdersInput if err := json.Unmarshal(inputBytes, ¶ms); err != nil { return nil, err } // Check if user info is available from previous tool call userInfo, ok := coordinator.GetResult("get_user") if ok { log.Printf("Using cached user info: %v", userInfo) } return map[string]interface{}{ "user_id": params.UserID, "orders": []map[string]interface{}{ {"order_id": "123", "total": 99.99}, {"order_id": "456", "total": 149.99}, }, }, nil }, } return []types.Tool{getUserTool, getOrdersTool} } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") coordinator := NewToolCoordinator() tools := createToolsWithDependencies(coordinator) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Get all orders for user ID user-123", Tools: tools, // GenerateText defaults to a single step; this scenario needs // multiple tool-call rounds (get_user, then get_orders) plus a // final text response, so allow more steps. StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Best Practices ### 1. Use JSON Schema Generation ```go reflector := jsonschema.Reflector{ AllowAdditionalProperties: false, DoNotReference: true, } schema := reflector.Reflect(&YourStruct{}) ``` ### 2. Validate Tool Inputs ```go Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { inputBytes, err := json.Marshal(input) if err != nil { return nil, fmt.Errorf("invalid input: %w", err) } var params InputType if err := json.Unmarshal(inputBytes, ¶ms); err != nil { return nil, fmt.Errorf("invalid input: %w", err) } // Validate params... return result, nil } ``` ### 3. Write Clear Tool Descriptions ```go Description: "Get current weather for a location. Use this tool when the user asks about weather, temperature, or climate conditions. Required parameter: location (city name)." ``` ### 4. Handle Tool Errors Gracefully ```go Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { result, err := performOperation() if err != nil { return nil, fmt.Errorf("operation failed: %w", err) } return result, nil } ``` ### 5. Add Timeouts to External Calls ```go ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second) defer cancel() ``` ## See Also - [Tool Calling Guide](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - [Schema Validation](https://goaisdk.com/docs/troubleshooting/schema-validation.md) - [Common Errors](https://goaisdk.com/docs/troubleshooting/common-errors.md) - [Tools Reference](https://goaisdk.com/docs/reference/types/tools.md) --- # Schema Validation Errors > Fixes JSON schema validation errors when generating structured data with the Go AI SDK, covering type mismatches, nested structures, and arrays. Canonical URL: https://goaisdk.com/docs/troubleshooting/schema-validation Documentation index: https://goaisdk.com/llms.txt This guide covers JSON schema validation errors when generating structured data with the Go AI SDK. ## Schema Definition Errors ### Error: "Invalid Schema Format" **Symptoms:** ``` Error: schema validation failed: invalid type definition Error: missing required field in schema ``` **Cause:** Incorrectly formatted JSON schema or missing required fields. **Solution:** ```go skip-compile package main import ( "context" "encoding/json" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" goschema "github.com/digitallysavvy/go-ai/pkg/schema" "github.com/invopop/jsonschema" ) // Define struct with comprehensive JSON schema tags type Recipe struct { Name string `json:"name" jsonschema:"required,minLength=1,maxLength=100,description=Recipe name"` Description string `json:"description" jsonschema:"required,minLength=10,description=Recipe description"` Ingredients []string `json:"ingredients" jsonschema:"required,minItems=1,description=List of ingredients"` Steps []string `json:"steps" jsonschema:"required,minItems=1,description=Cooking instructions"` PrepTime int `json:"prepTime" jsonschema:"required,minimum=1,description=Preparation time in minutes"` CookTime int `json:"cookTime" jsonschema:"required,minimum=1,description=Cooking time in minutes"` Servings int `json:"servings" jsonschema:"required,minimum=1,maximum=20,description=Number of servings"` Difficulty string `json:"difficulty" jsonschema:"required,enum=easy,enum=medium,enum=hard,description=Recipe difficulty level"` Cuisine string `json:"cuisine" jsonschema:"description=Type of cuisine"` IsVegetarian bool `json:"isVegetarian" jsonschema:"description=Whether the recipe is vegetarian"` } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Generate schema with proper configuration reflector := jsonschema.Reflector{ AllowAdditionalProperties: false, // Strict validation DoNotReference: true, // Inline definitions RequiredFromJSONSchemaTags: true, // Use jsonschema tags } reflected := reflector.Reflect(&Recipe{}) // Validate schema before using schemaBytes, err := json.MarshalIndent(reflected, "", " ") if err != nil { log.Fatalf("Invalid schema: %v", err) } log.Printf("Generated schema:\n%s", schemaBytes) // ai.GenerateObjectOptions.Schema takes a schema.Schema, not a // *jsonschema.Schema directly — decode the reflected schema into a map // and wrap it with goschema.NewSimpleJSONSchema. var schemaMap map[string]interface{} if err := json.Unmarshal(schemaBytes, &schemaMap); err != nil { log.Fatalf("Failed to decode schema: %v", err) } var recipe Recipe if err := ai.GenerateObjectInto(ctx, ai.GenerateObjectOptions{ Model: model, Schema: goschema.NewSimpleJSONSchema(schemaMap), Prompt: "Create a lasagna recipe for 4 people", }, &recipe); err != nil { log.Fatalf("Failed to generate object: %v", err) } fmt.Printf("Recipe: %s\n", recipe.Name) fmt.Printf("Prep time: %d minutes\n", recipe.PrepTime) fmt.Printf("Servings: %d\n", recipe.Servings) } ``` ## Type Mismatch Errors ### Error: "Generated Object Type Mismatch" **Symptoms:** ``` Error: json: cannot unmarshal string into Go value of type int Error: field type mismatch ``` **Cause:** Model generates JSON that doesn't match your struct types. **Solution:** ```go skip-compile package main import ( "context" "encoding/json" "fmt" "log" "os" "strconv" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" goschema "github.com/digitallysavvy/go-ai/pkg/schema" "github.com/invopop/jsonschema" ) // Custom type with JSON unmarshaling type FlexibleInt int func (f *FlexibleInt) UnmarshalJSON(data []byte) error { // Try to unmarshal as int first var i int if err := json.Unmarshal(data, &i); err == nil { *f = FlexibleInt(i) return nil } // If that fails, try as string var s string if err := json.Unmarshal(data, &s); err != nil { return err } // Convert string to int i, err := strconv.Atoi(s) if err != nil { return fmt.Errorf("cannot convert %q to int: %w", s, err) } *f = FlexibleInt(i) return nil } // Product with flexible types type Product struct { Name string `json:"name" jsonschema:"required"` Price FlexibleInt `json:"price" jsonschema:"required,type=integer,description=Price in cents"` Quantity FlexibleInt `json:"quantity" jsonschema:"required,type=integer,description=Quantity available"` InStock bool `json:"inStock" jsonschema:"required"` Description string `json:"description" jsonschema:"required"` } // Validation after unmarshaling func (p *Product) Validate() error { if p.Name == "" { return fmt.Errorf("name is required") } if p.Price < 0 { return fmt.Errorf("price must be non-negative") } if p.Quantity < 0 { return fmt.Errorf("quantity must be non-negative") } return nil } func generateProduct(ctx context.Context, model provider.LanguageModel, prompt string) (*Product, error) { reflector := jsonschema.Reflector{ AllowAdditionalProperties: false, DoNotReference: true, } reflected := reflector.Reflect(&Product{}) schemaBytes, err := json.Marshal(reflected) if err != nil { return nil, fmt.Errorf("invalid schema: %w", err) } var schemaMap map[string]interface{} if err := json.Unmarshal(schemaBytes, &schemaMap); err != nil { return nil, fmt.Errorf("failed to decode schema: %w", err) } var product Product if err := ai.GenerateObjectInto(ctx, ai.GenerateObjectOptions{ Model: model, Schema: goschema.NewSimpleJSONSchema(schemaMap), Prompt: prompt, }, &product); err != nil { return nil, fmt.Errorf("generation failed: %w", err) } // Validate after unmarshaling if err := product.Validate(); err != nil { return nil, fmt.Errorf("validation failed: %w", err) } return &product, nil } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") product, err := generateProduct(ctx, model, "Create a laptop product listing") if err != nil { log.Fatal(err) } fmt.Printf("Product: %s - $%.2f\n", product.Name, float64(product.Price)/100) } ``` ## Nested Structure Errors ### Error: "Complex Nested Schema Validation Failed" **Symptoms:** ``` Error: nested object validation failed Error: array item schema mismatch ``` **Cause:** Complex nested structures not properly defined in schema. **Solution:** ```go skip-compile package main import ( "context" "encoding/json" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" goschema "github.com/digitallysavvy/go-ai/pkg/schema" "github.com/invopop/jsonschema" ) // Address is a nested structure type Address struct { Street string `json:"street" jsonschema:"required,description=Street address"` City string `json:"city" jsonschema:"required,description=City name"` State string `json:"state" jsonschema:"required,pattern=^[A-Z]{2}$,description=Two-letter state code"` ZipCode string `json:"zipCode" jsonschema:"required,pattern=^[0-9]{5}$,description=Five-digit ZIP code"` } // OrderItem represents a single item in an order type OrderItem struct { ProductID string `json:"productId" jsonschema:"required,description=Product identifier"` ProductName string `json:"productName" jsonschema:"required,description=Product name"` Quantity int `json:"quantity" jsonschema:"required,minimum=1,description=Quantity ordered"` UnitPrice float64 `json:"unitPrice" jsonschema:"required,minimum=0,description=Price per unit"` Subtotal float64 `json:"subtotal" jsonschema:"required,minimum=0,description=Line item total"` } // Order is the top-level structure with nested objects type Order struct { OrderID string `json:"orderId" jsonschema:"required,description=Unique order identifier"` CustomerName string `json:"customerName" jsonschema:"required,description=Customer full name"` CustomerEmail string `json:"customerEmail" jsonschema:"required,format=email,description=Customer email address"` ShippingAddr Address `json:"shippingAddress" jsonschema:"required,description=Shipping address"` BillingAddr Address `json:"billingAddress" jsonschema:"required,description=Billing address"` Items []OrderItem `json:"items" jsonschema:"required,minItems=1,description=Order items"` Subtotal float64 `json:"subtotal" jsonschema:"required,minimum=0,description=Subtotal before tax"` Tax float64 `json:"tax" jsonschema:"required,minimum=0,description=Tax amount"` ShippingCost float64 `json:"shippingCost" jsonschema:"required,minimum=0,description=Shipping cost"` Total float64 `json:"total" jsonschema:"required,minimum=0,description=Total amount"` Status string `json:"status" jsonschema:"required,enum=pending,enum=processing,enum=shipped,enum=delivered,description=Order status"` } // Validate ensures the order data is consistent func (o *Order) Validate() error { // Validate items if len(o.Items) == 0 { return fmt.Errorf("order must have at least one item") } // Calculate and verify subtotal calculatedSubtotal := 0.0 for i, item := range o.Items { if item.Quantity <= 0 { return fmt.Errorf("item %d: quantity must be positive", i) } if item.UnitPrice < 0 { return fmt.Errorf("item %d: price cannot be negative", i) } expectedSubtotal := float64(item.Quantity) * item.UnitPrice if item.Subtotal != expectedSubtotal { log.Printf("Warning: item %d subtotal mismatch (expected %.2f, got %.2f)", i, expectedSubtotal, item.Subtotal) } calculatedSubtotal += item.Subtotal } // Verify order totals if o.Subtotal != calculatedSubtotal { log.Printf("Warning: order subtotal mismatch (expected %.2f, got %.2f)", calculatedSubtotal, o.Subtotal) } expectedTotal := o.Subtotal + o.Tax + o.ShippingCost if o.Total != expectedTotal { log.Printf("Warning: order total mismatch (expected %.2f, got %.2f)", expectedTotal, o.Total) } return nil } func generateOrder(ctx context.Context, model provider.LanguageModel) (*Order, error) { reflector := jsonschema.Reflector{ AllowAdditionalProperties: false, DoNotReference: true, RequiredFromJSONSchemaTags: true, } reflected := reflector.Reflect(&Order{}) // Log schema for debugging schemaJSON, _ := json.MarshalIndent(reflected, "", " ") log.Printf("Schema:\n%s", schemaJSON) var schemaMap map[string]interface{} if err := json.Unmarshal(schemaJSON, &schemaMap); err != nil { return nil, fmt.Errorf("failed to decode schema: %w", err) } prompt := `Generate a sample order for Jane Smith (jane@example.com) who ordered: - 2x Laptop at $999.99 each - 1x Mouse at $29.99 Ship to: 123 Main St, Boston, MA 02101 Same billing address Tax rate: 6.25% Shipping: $15.00` var order Order if err := ai.GenerateObjectInto(ctx, ai.GenerateObjectOptions{ Model: model, Schema: goschema.NewSimpleJSONSchema(schemaMap), Prompt: prompt, }, &order); err != nil { return nil, fmt.Errorf("generation failed: %w", err) } // Validate the order if err := order.Validate(); err != nil { return nil, fmt.Errorf("validation failed: %w", err) } return &order, nil } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") order, err := generateOrder(ctx, model) if err != nil { log.Fatal(err) } fmt.Printf("Order %s for %s\n", order.OrderID, order.CustomerName) fmt.Printf("Items: %d\n", len(order.Items)) fmt.Printf("Total: $%.2f\n", order.Total) } ``` ## Array Validation Errors ### Error: "Array Items Don't Match Schema" **Symptoms:** ``` Error: array item validation failed at index 2 Error: array contains items of wrong type ``` **Cause:** Array items don't conform to the specified item schema. **Solution:** ```go skip-compile package main import ( "context" "encoding/json" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" goschema "github.com/digitallysavvy/go-ai/pkg/schema" "github.com/invopop/jsonschema" ) // Task with validation type Task struct { ID string `json:"id" jsonschema:"required,pattern=^TASK-[0-9]+$,description=Task ID in format TASK-XXX"` Title string `json:"title" jsonschema:"required,minLength=5,maxLength=100,description=Task title"` Description string `json:"description" jsonschema:"required,minLength=10,description=Detailed task description"` Priority string `json:"priority" jsonschema:"required,enum=low,enum=medium,enum=high,enum=urgent,description=Task priority"` Status string `json:"status" jsonschema:"required,enum=todo,enum=in_progress,enum=done,description=Task status"` Assignee string `json:"assignee" jsonschema:"required,minLength=1,description=Person assigned to task"` DueDate string `json:"dueDate" jsonschema:"required,format=date,description=Due date in YYYY-MM-DD format"` } // TaskList with array validation type TaskList struct { ProjectName string `json:"projectName" jsonschema:"required,description=Project name"` Tasks []Task `json:"tasks" jsonschema:"required,minItems=3,maxItems=10,description=List of tasks"` } // Validate all tasks in the list func (tl *TaskList) Validate() error { if len(tl.Tasks) == 0 { return fmt.Errorf("task list cannot be empty") } taskIDs := make(map[string]bool) for i, task := range tl.Tasks { // Check for duplicate IDs if taskIDs[task.ID] { return fmt.Errorf("duplicate task ID at index %d: %s", i, task.ID) } taskIDs[task.ID] = true // Validate individual task if err := validateTask(&task, i); err != nil { return err } } return nil } func validateTask(task *Task, index int) error { if task.Title == "" { return fmt.Errorf("task %d: title is required", index) } if task.Description == "" { return fmt.Errorf("task %d: description is required", index) } validPriorities := map[string]bool{"low": true, "medium": true, "high": true, "urgent": true} if !validPriorities[task.Priority] { return fmt.Errorf("task %d: invalid priority %q", index, task.Priority) } validStatuses := map[string]bool{"todo": true, "in_progress": true, "done": true} if !validStatuses[task.Status] { return fmt.Errorf("task %d: invalid status %q", index, task.Status) } return nil } func generateTaskList(ctx context.Context, model provider.LanguageModel) (*TaskList, error) { reflector := jsonschema.Reflector{ AllowAdditionalProperties: false, DoNotReference: true, RequiredFromJSONSchemaTags: true, } reflected := reflector.Reflect(&TaskList{}) schemaBytes, err := json.Marshal(reflected) if err != nil { return nil, fmt.Errorf("invalid schema: %w", err) } var schemaMap map[string]interface{} if err := json.Unmarshal(schemaBytes, &schemaMap); err != nil { return nil, fmt.Errorf("failed to decode schema: %w", err) } var taskList TaskList if err := ai.GenerateObjectInto(ctx, ai.GenerateObjectOptions{ Model: model, Schema: goschema.NewSimpleJSONSchema(schemaMap), Prompt: "Generate a task list for a website redesign project with 5 tasks", }, &taskList); err != nil { return nil, fmt.Errorf("generation failed: %w", err) } // Validate the task list if err := taskList.Validate(); err != nil { return nil, fmt.Errorf("validation failed: %w", err) } return &taskList, nil } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") taskList, err := generateTaskList(ctx, model) if err != nil { log.Fatal(err) } fmt.Printf("Project: %s\n", taskList.ProjectName) fmt.Printf("Tasks: %d\n", len(taskList.Tasks)) for i, task := range taskList.Tasks { fmt.Printf("%d. [%s] %s (Priority: %s, Assignee: %s)\n", i+1, task.Status, task.Title, task.Priority, task.Assignee) } } ``` ## Best Practices ### 1. Use Comprehensive JSON Schema Tags ```go type Field struct { Name string `json:"name" jsonschema:"required,minLength=1,maxLength=100,description=Field name"` } ``` ### 2. Always Validate After Unmarshaling ```go var result MyType json.Unmarshal(data, &result) if err := result.Validate(); err != nil { return fmt.Errorf("validation failed: %w", err) } ``` ### 3. Log Raw JSON on Errors When the model's output doesn't parse against the schema, `ai.GenerateObject` and `ai.GenerateObjectInto` return an `*ai.NoObjectGeneratedError`, whose `Text` field carries the model's raw (unparseable) JSON response: ```go result, err := ai.GenerateObject(ctx, opts) if err != nil { var noObjErr *ai.NoObjectGeneratedError if errors.As(err, &noObjErr) { log.Printf("Raw JSON: %s", noObjErr.Text) } return err } ``` ### 4. Use Flexible Types When Needed ```go type FlexibleType struct { Value interface{} `json:"value"` } func (f *FlexibleType) UnmarshalJSON(data []byte) error { // Custom unmarshaling logic } ``` ### 5. Test Schema Generation ```go schema := reflector.Reflect(&YourType{}) schemaJSON, _ := json.MarshalIndent(schema, "", " ") log.Printf("Generated schema:\n%s", schemaJSON) ``` ## See Also - [Generating Structured Data](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) - [Tool Calling Errors](https://goaisdk.com/docs/troubleshooting/tool-calling-errors.md) - [Common Errors](https://goaisdk.com/docs/troubleshooting/common-errors.md) - [Schema Reference](https://goaisdk.com/docs/reference/schema/json-schema.md) --- # Performance Optimization > Covers optimizing slow Go AI SDK operations, including slow response times, memory issues, caching for performance, and concurrent processing patterns. Canonical URL: https://goaisdk.com/docs/troubleshooting/performance Documentation index: https://goaisdk.com/llms.txt This guide covers performance issues and optimization techniques for Go AI applications. ## Slow Response Times ### Issue: High Latency **Symptoms:** - Responses take several seconds - Users experience delays **Causes & Solutions:** ```go package main import ( "context" "fmt" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // 1. Use Faster Models func useFasterModels() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) // SLOW: large model (high quality, slower) slowModel, _ := provider.LanguageModel("gpt-6-astra") start := time.Now() slowResult, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: slowModel, Prompt: "Explain quantum computing in one sentence", }) fmt.Printf("Slow model took: %v\n", time.Since(start)) fmt.Println(slowResult.Text) // FAST: small model (good quality, much faster) fastModel, _ := provider.LanguageModel("gpt-5.4-mini") start = time.Now() result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: fastModel, Prompt: "Explain quantum computing in one sentence", }) fmt.Printf("Fast model took: %v\n", time.Since(start)) fmt.Println(result.Text) } // 2. Use Streaming for Perceived Performance func useStreaming() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") // Streaming gives immediate feedback start := time.Now() stream, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a detailed explanation of quantum computing", }) firstChunkTime := time.Duration(0) chunkCount := 0 for chunk := range stream.Chunks() { if chunkCount == 0 { firstChunkTime = time.Since(start) fmt.Printf("Time to first chunk: %v\n", firstChunkTime) } fmt.Print(chunk.Text) chunkCount++ } fmt.Printf("\nTotal chunks: %d, Total time: %v\n", chunkCount, time.Since(start)) } // 3. Limit Token Output func limitTokens() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") start := time.Now() maxTokens := 100 // Limit response length result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing", MaxTokens: &maxTokens, }) fmt.Printf("Limited response took: %v\n", time.Since(start)) fmt.Println(result.Text) } func main() { fmt.Println("=== Fast Models ===") useFasterModels() fmt.Println("\n=== Streaming ===") useStreaming() fmt.Println("\n=== Limited Tokens ===") limitTokens() } ``` ## Memory Issues ### Issue: High Memory Usage **Symptoms:** - Application memory grows over time - Out of memory errors - Slow performance due to GC pressure **Solutions:** ```go package main import ( "context" "fmt" "log" "os" "runtime" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // BAD: Accumulating all results in memory func memoryHeavy(ctx context.Context, model provider.LanguageModel, prompts []string) []string { results := make([]string, 0, len(prompts)) for _, prompt := range prompts { result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) // Accumulates all results in memory results = append(results, result.Text) } return results } // GOOD: Process and discard results immediately func memoryEfficient(ctx context.Context, model provider.LanguageModel, prompts []string) { for i, prompt := range prompts { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { log.Printf("Error on prompt %d: %v", i, err) continue } // Process immediately (e.g., write to file, send to client) processResult(result.Text) // Result can be garbage collected after processing // Optional: Force GC periodically if i%10 == 0 { runtime.GC() } } } func processResult(text string) { // Process the result (save to DB, file, etc.) fmt.Printf("Processed: %s...\n", text[:min(50, len(text))]) } func min(a, b int) int { if a < b { return a } return b } // Monitor memory usage func monitorMemory() { var m runtime.MemStats runtime.ReadMemStats(&m) fmt.Printf("Alloc = %v MiB", m.Alloc/1024/1024) fmt.Printf("\tTotalAlloc = %v MiB", m.TotalAlloc/1024/1024) fmt.Printf("\tSys = %v MiB", m.Sys/1024/1024) fmt.Printf("\tNumGC = %v\n", m.NumGC) } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-5.4-mini") prompts := []string{ "Explain AI", "Explain ML", "Explain DL", } fmt.Println("Before processing:") monitorMemory() memoryEfficient(ctx, model, prompts) fmt.Println("\nAfter processing:") monitorMemory() } ``` ## Caching for Performance ### Implement Response Caching ```go package main import ( "context" "crypto/sha256" "encoding/hex" "fmt" "log" "os" "sync" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // CachedGenerator caches AI responses type CachedGenerator struct { model provider.LanguageModel cache map[string]*CacheEntry mu sync.RWMutex ttl time.Duration hitCount int missCount int } type CacheEntry struct { Response string Timestamp time.Time } func NewCachedGenerator(model provider.LanguageModel, ttl time.Duration) *CachedGenerator { return &CachedGenerator{ model: model, cache: make(map[string]*CacheEntry), ttl: ttl, } } func (cg *CachedGenerator) generateCacheKey(prompt string, opts ...string) string { data := prompt for _, opt := range opts { data += "|" + opt } hash := sha256.Sum256([]byte(data)) return hex.EncodeToString(hash[:]) } func (cg *CachedGenerator) Generate(ctx context.Context, prompt string) (string, error) { key := cg.generateCacheKey(prompt) // Check cache cg.mu.RLock() if entry, ok := cg.cache[key]; ok { if time.Since(entry.Timestamp) < cg.ttl { cg.hitCount++ cg.mu.RUnlock() log.Printf("Cache HIT for prompt: %s...", prompt[:min(30, len(prompt))]) return entry.Response, nil } } cg.mu.RUnlock() // Cache miss cg.mu.Lock() cg.missCount++ cg.mu.Unlock() log.Printf("Cache MISS for prompt: %s...", prompt[:min(30, len(prompt))]) // Generate response result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: cg.model, Prompt: prompt, }) if err != nil { return "", err } // Store in cache cg.mu.Lock() cg.cache[key] = &CacheEntry{ Response: result.Text, Timestamp: time.Now(), } cg.mu.Unlock() return result.Text, nil } func (cg *CachedGenerator) GetStats() (hits, misses int, hitRate float64) { cg.mu.RLock() defer cg.mu.RUnlock() total := cg.hitCount + cg.missCount hitRate = 0 if total > 0 { hitRate = float64(cg.hitCount) / float64(total) * 100 } return cg.hitCount, cg.missCount, hitRate } func (cg *CachedGenerator) ClearExpired() { cg.mu.Lock() defer cg.mu.Unlock() now := time.Now() for key, entry := range cg.cache { if now.Sub(entry.Timestamp) > cg.ttl { delete(cg.cache, key) } } } func min(a, b int) int { if a < b { return a } return b } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-5.4-mini") // Create cached generator with 5 minute TTL generator := NewCachedGenerator(model, 5*time.Minute) // Start cache cleanup goroutine go func() { ticker := time.NewTicker(1 * time.Minute) defer ticker.Stop() for range ticker.C { generator.ClearExpired() } }() prompts := []string{ "What is AI?", "Explain machine learning", "What is AI?", // Duplicate - should hit cache "What is AI?", // Duplicate - should hit cache "Explain deep learning", } for i, prompt := range prompts { start := time.Now() response, err := generator.Generate(ctx, prompt) elapsed := time.Since(start) if err != nil { log.Printf("Error: %v", err) continue } fmt.Printf("\n[%d] Prompt: %s\n", i+1, prompt) fmt.Printf("Response: %s...\n", response[:min(100, len(response))]) fmt.Printf("Time: %v\n", elapsed) } hits, misses, hitRate := generator.GetStats() fmt.Printf("\n=== Cache Statistics ===\n") fmt.Printf("Hits: %d\n", hits) fmt.Printf("Misses: %d\n", misses) fmt.Printf("Hit Rate: %.1f%%\n", hitRate) } ``` ## Concurrent Processing ### Optimize with Goroutines ```go package main import ( "context" "fmt" "log" "os" "sync" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // Process prompts concurrently func processConcurrently(ctx context.Context, model provider.LanguageModel, prompts []string, maxConcurrent int) []string { results := make([]string, len(prompts)) var wg sync.WaitGroup // Semaphore to limit concurrency sem := make(chan struct{}, maxConcurrent) for i, prompt := range prompts { wg.Add(1) go func(index int, p string) { defer wg.Done() // Acquire semaphore sem <- struct{}{} defer func() { <-sem }() result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: p, }) if err != nil { log.Printf("Error on prompt %d: %v", index, err) return } results[index] = result.Text }(i, prompt) } wg.Wait() return results } func benchmark(name string, fn func()) { start := time.Now() fn() fmt.Printf("%s took: %v\n", name, time.Since(start)) } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-5.4-mini") prompts := []string{ "Explain AI in one sentence", "Explain ML in one sentence", "Explain DL in one sentence", "Explain NLP in one sentence", "Explain CV in one sentence", } // Sequential processing benchmark("Sequential", func() { for _, prompt := range prompts { ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) } }) // Concurrent processing benchmark("Concurrent (3 workers)", func() { processConcurrently(ctx, model, prompts, 3) }) } ``` ## Best Practices ### 1. Use Appropriate Models ```go // Fast, cheap for simple tasks model, _ := provider.LanguageModel("gpt-5.4-mini") // Powerful for complex tasks model, _ := provider.LanguageModel("gpt-6-astra") ``` ### 2. Stream Long Responses ```go // Better UX with streaming stream, _ := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: prompt, }) ``` ### 3. Implement Caching ```go // Cache frequent queries generator := NewCachedGenerator(model, 5*time.Minute) ``` ### 4. Limit Concurrent Requests ```go // Don't overwhelm the API maxConcurrent := 5 sem := make(chan struct{}, maxConcurrent) ``` ### 5. Monitor Performance ```go start := time.Now() // ... operation ... log.Printf("Operation took: %v", time.Since(start)) ``` ## See Also - [Caching](https://goaisdk.com/docs/advanced/caching.md) - [Rate Limiting](https://goaisdk.com/docs/advanced/rate-limiting.md) - [Streaming](https://goaisdk.com/docs/foundations/streaming.md) --- # Debugging Guide > Comprehensive debugging techniques for Go AI applications, covering debug logging, goroutine debugging, pprof profiling, and stream debugging tools. Canonical URL: https://goaisdk.com/docs/troubleshooting/debugging Documentation index: https://goaisdk.com/llms.txt This guide covers debugging techniques and tools for troubleshooting Go AI applications. ## Enable Debug Logging ### SDK Warning Logging The SDK has no global debug/log-level switch. What it does provide is a warnings logger: every call that returns provider warnings (unsupported settings, compatibility fallbacks, etc.) is routed through `ai.LogWarnings`, which writes to stderr by default. Install your own sink with `ai.SetLogWarnings`, or silence it with `ai.DisableLogWarnings()` (or the `AI_SDK_LOG_WARNINGS=false` environment variable): ```go package main import ( "log" "github.com/digitallysavvy/go-ai/pkg/ai" ) func main() { // Route SDK warnings into your own logger instead of stderr. ai.SetLogWarnings(func(opts ai.LogWarningsOptions) { for _, w := range opts.Warnings { log.Printf("[%s/%s] %s", opts.Provider, opts.Model, ai.FormatWarning(w, opts.Provider, opts.Model)) } }) // Your code here... } ``` For full HTTP request/response visibility, wrap the HTTP client instead — see below. ### Request/Response Logging ```go package main import ( "bytes" "context" "fmt" "io" "log" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // Custom HTTP transport with logging type LoggingTransport struct { Transport http.RoundTripper } func (t *LoggingTransport) RoundTrip(req *http.Request) (*http.Response, error) { // Log request log.Printf("=== REQUEST ===") log.Printf("%s %s", req.Method, req.URL) log.Printf("Headers: %v", req.Header) if req.Body != nil { bodyBytes, _ := io.ReadAll(req.Body) req.Body = io.NopCloser(bytes.NewBuffer(bodyBytes)) log.Printf("Body: %s", string(bodyBytes)) } // Execute request resp, err := t.Transport.RoundTrip(req) if err != nil { log.Printf("Request error: %v", err) return nil, err } // Log response log.Printf("=== RESPONSE ===") log.Printf("Status: %d %s", resp.StatusCode, resp.Status) log.Printf("Headers: %v", resp.Header) if resp.Body != nil { bodyBytes, _ := io.ReadAll(resp.Body) resp.Body = io.NopCloser(bytes.NewBuffer(bodyBytes)) log.Printf("Body: %s", string(bodyBytes)) } return resp, nil } func main() { ctx := context.Background() // Create HTTP client with logging transport httpClient := &http.Client{ Transport: &LoggingTransport{ Transport: http.DefaultTransport, }, } provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), HTTPClient: httpClient, }) model, _ := provider.LanguageModel("gpt-5.4-mini") result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello!", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Debugging Goroutines ### Detect Goroutine Leaks ```go package main import ( "context" "fmt" "log" "os" "runtime" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func printGoroutineCount(label string) { count := runtime.NumGoroutine() log.Printf("[%s] Active goroutines: %d", label, count) } func detectLeaks() { printGoroutineCount("Start") ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-5.4-mini") // Start streaming stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a short story", }) if err != nil { log.Fatal(err) } printGoroutineCount("After stream start") // IMPORTANT: Consume the stream completely for chunk := range stream.Chunks() { fmt.Print(chunk.Text) } printGoroutineCount("After stream complete") // Wait for cleanup time.Sleep(100 * time.Millisecond) printGoroutineCount("After cleanup wait") // Force garbage collection runtime.GC() time.Sleep(100 * time.Millisecond) printGoroutineCount("After GC") } func main() { detectLeaks() } ``` ### Stack Traces for Debugging ```go package main import ( "fmt" "os" "os/signal" "runtime" "syscall" ) func dumpStackTraces() { buf := make([]byte, 1024*1024) stackSize := runtime.Stack(buf, true) fmt.Printf("=== Stack Traces ===\n%s\n", buf[:stackSize]) } func setupDebugHandlers() { // Dump stack traces on SIGUSR1 sigCh := make(chan os.Signal, 1) signal.Notify(sigCh, syscall.SIGUSR1) go func() { for range sigCh { dumpStackTraces() } }() } func main() { setupDebugHandlers() // Your application code... select {} } ``` ## Debugging with pprof ### CPU and Memory Profiling ```go package main import ( "context" "fmt" "log" "net/http" _ "net/http/pprof" "os" "runtime" "runtime/pprof" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func startProfiling() { // Start pprof HTTP server go func() { log.Println("Starting pprof server on :6060") log.Println(http.ListenAndServe("localhost:6060", nil)) }() // Or profile to files cpuFile, _ := os.Create("cpu.prof") pprof.StartCPUProfile(cpuFile) defer pprof.StopCPUProfile() memFile, _ := os.Create("mem.prof") defer func() { runtime.GC() pprof.WriteHeapProfile(memFile) memFile.Close() }() } func main() { startProfiling() ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-5.4-mini") // Run your code to profile for i := 0; i < 10; i++ { result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: fmt.Sprintf("Request %d", i), }) fmt.Println(result.Text) } // View profiles: // go tool pprof cpu.prof // go tool pprof mem.prof // Or visit http://localhost:6060/debug/pprof/ } ``` ## Debugging Streams ### Stream Debug Wrapper ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func debugStream(ctx context.Context, model provider.LanguageModel, prompt string) error { log.Printf("Starting stream for prompt: %s", prompt) stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: prompt, }) if err != nil { return fmt.Errorf("stream creation failed: %w", err) } startTime := time.Now() chunkCount := 0 totalBytes := 0 var firstChunkTime time.Duration log.Println("Waiting for chunks...") for chunk := range stream.Chunks() { if chunkCount == 0 { firstChunkTime = time.Since(startTime) log.Printf("Time to first chunk: %v", firstChunkTime) } chunkCount++ switch chunk.Type { case provider.ChunkTypeText: totalBytes += len(chunk.Text) log.Printf("Chunk %d: %d bytes (type: text)", chunkCount, len(chunk.Text)) fmt.Print(chunk.Text) case provider.ChunkTypeToolCall: log.Printf("Chunk %d: tool call %s (id: %s)", chunkCount, chunk.ToolCall.ToolName, chunk.ToolCall.ID) case provider.ChunkTypeError: log.Printf("Chunk %d: ERROR - %v", chunkCount, chunk.Err) case provider.ChunkTypeFinish: log.Printf("Chunk %d: FINISH - reason: %s", chunkCount, chunk.FinishReason) } } if err := stream.Err(); err != nil { log.Printf("Stream error: %v", err) return err } totalTime := time.Since(startTime) bytesPerSecond := float64(totalBytes) / totalTime.Seconds() log.Printf("\n=== Stream Statistics ===") log.Printf("Total chunks: %d", chunkCount) log.Printf("Total bytes: %d", totalBytes) log.Printf("Time to first chunk: %v", firstChunkTime) log.Printf("Total time: %v", totalTime) log.Printf("Throughput: %.2f bytes/sec", bytesPerSecond) return nil } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: os.Getenv("OPENAI_API_KEY"), }) model, _ := provider.LanguageModel("gpt-6-astra") if err := debugStream(ctx, model, "Write a short story about a robot"); err != nil { log.Fatalf("Debug stream failed: %v", err) } } ``` ## Error Tracking ### Structured Error Logging ```go package main import ( "context" "encoding/json" "errors" "log" "time" "github.com/digitallysavvy/go-ai/pkg/ai" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) type ErrorLog struct { Timestamp time.Time `json:"timestamp"` ErrorType string `json:"error_type"` Message string `json:"message"` Provider string `json:"provider,omitempty"` StatusCode int `json:"status_code,omitempty"` ErrorCode string `json:"error_code,omitempty"` Operation string `json:"operation"` Prompt string `json:"prompt,omitempty"` StackTrace string `json:"stack_trace,omitempty"` Additional map[string]interface{} `json:"additional,omitempty"` } func logError(err error, operation string, prompt string) { errorLog := ErrorLog{ Timestamp: time.Now(), Message: err.Error(), Operation: operation, Prompt: prompt, Additional: make(map[string]interface{}), } // Check for specific error types var providerErr *providererrors.ProviderError if errors.As(err, &providerErr) { errorLog.ErrorType = "provider_error" errorLog.Provider = providerErr.Provider errorLog.StatusCode = providerErr.StatusCode errorLog.ErrorCode = providerErr.ErrorCode } var rateLimitErr *providererrors.RateLimitError if errors.As(err, &rateLimitErr) { errorLog.ErrorType = "rate_limit_error" errorLog.Provider = rateLimitErr.Provider if rateLimitErr.RetryAfterSeconds != nil { errorLog.Additional["retry_after_seconds"] = *rateLimitErr.RetryAfterSeconds } } var validationErr *providererrors.ValidationError if errors.As(err, &validationErr) { errorLog.ErrorType = "validation_error" errorLog.Additional["field"] = validationErr.Field } // Log as JSON jsonLog, _ := json.MarshalIndent(errorLog, "", " ") log.Printf("ERROR LOG:\n%s", string(jsonLog)) } func main() { ctx := context.Background() provider := openai.New(openai.Config{ APIKey: "invalid-key", // Will cause auth error }) model, _ := provider.LanguageModel("gpt-6-astra") prompt := "Hello, world!" _, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { logError(err, "generate_text", prompt) } } ``` ## Testing and Debugging ### Mock Provider for Testing ```go package main import ( "context" "fmt" "testing" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/testutil" ) func int64Ptr(v int64) *int64 { return &v } func TestGeneration(t *testing.T) { mock := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { return &types.GenerateResult{ Text: "Test response", FinishReason: types.FinishReasonStop, Usage: types.Usage{ InputTokens: int64Ptr(10), OutputTokens: int64Ptr(20), TotalTokens: int64Ptr(30), }, }, nil }, } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: mock, Prompt: "Test prompt", }) if err != nil { t.Fatalf("Expected no error, got: %v", err) } if result.Text != "Test response" { t.Fatalf("Expected 'Test response', got: %s", result.Text) } } func TestGenerationError(t *testing.T) { mock := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { return nil, fmt.Errorf("mock error") }, } _, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: mock, Prompt: "Test prompt", }) if err == nil { t.Fatal("Expected error, got nil") } } ``` `testutil.MockLanguageModel` implements `provider.LanguageModel` for you — see [Testing](https://goaisdk.com/docs/ai-sdk-core/testing.md) for its other hooks (`DoStreamFunc`, call tracking, etc.). ## Best Practices ### 1. Use Structured Logging ```go log.Printf("[%s] %s: %v", level, operation, details) ``` ### 2. Log Request Context ```go log.Printf("Request: model=%s, prompt_length=%d, max_tokens=%d", modelName, len(prompt), maxTokens) ``` ### 3. Monitor Goroutines ```go log.Printf("Active goroutines: %d", runtime.NumGoroutine()) ``` ### 4. Profile Performance ```go import _ "net/http/pprof" go func() { log.Println(http.ListenAndServe("localhost:6060", nil)) }() ``` ### 5. Test with Mocks ```go mock := &testutil.MockLanguageModel{} result, _ := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: mock}) ``` ## See Also - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [Testing](https://goaisdk.com/docs/ai-sdk-core/testing.md) - [Common Errors](https://goaisdk.com/docs/troubleshooting/common-errors.md) - [Performance](https://goaisdk.com/docs/troubleshooting/performance.md) --- # Timeout Errors > Explains timeout errors in the Go AI SDK, covering how to detect, handle, and prevent them, plus recommended timeouts and async processing patterns. Canonical URL: https://goaisdk.com/docs/troubleshooting/timeout-errors Documentation index: https://goaisdk.com/llms.txt Timeout errors occur when an AI operation takes longer than the configured timeout duration. This guide covers how to detect, handle, and prevent timeout errors. ## Understanding Timeouts ### Client-Side vs Server-Side Timeouts **Client-Side Timeout**: Your application's context deadline is exceeded before receiving a response. ```go ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() options.Model = model result, err := ai.GenerateText(ctx, options) // If this takes longer than 30 seconds, you get a timeout error ``` **Server-Side Timeout**: The provider's API times out the request. ```go // Provider returns a timeout response (rare) // Usually appears as a 408 or 504 HTTP status code ``` ## Detecting Timeout Errors ### Gateway Provider The Gateway provider provides enhanced timeout error detection: ```go import gatewayerrors "github.com/digitallysavvy/go-ai/pkg/providers/gateway/errors" options.Model = videoModel result, err := ai.GenerateVideo(ctx, options) if err != nil { if gatewayerrors.IsGatewayTimeoutError(err) { // This is a timeout error with troubleshooting guidance fmt.Printf("Timeout: %v\n", err) } } ``` ### Standard Error Detection ```go import ( "context" "errors" ) options.Model = model result, err := ai.GenerateText(ctx, options) if err != nil { if errors.Is(err, context.DeadlineExceeded) { // Context timeout fmt.Println("Request timed out") } } ``` ## Common Timeout Scenarios ### Text Generation Timeouts ```go // ❌ Too short for complex generation ctx, cancel := context.WithTimeout(ctx, 5*time.Second) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a detailed essay about AI...", MaxTokens: ptr(4000), }) // Likely to timeout // ✅ Appropriate timeout ctx, cancel := context.WithTimeout(ctx, 2*time.Minute) result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Write a detailed essay about AI...", MaxTokens: ptr(4000), }) ``` ### Video Generation Timeouts Video generation typically takes several minutes: ```go // ❌ Way too short for video generation ctx, cancel := context.WithTimeout(ctx, 30*time.Second) options.Model = videoModel result, err := ai.GenerateVideo(ctx, options) // Will timeout // ✅ Appropriate timeout for video generation ctx, cancel := context.WithTimeout(ctx, 10*time.Minute) options.Model = videoModel result, err := ai.GenerateVideo(ctx, options) ``` ### Streaming Timeouts For streaming operations, timeouts apply to the entire stream: ```go // Timeout applies to full stream duration ctx, cancel := context.WithTimeout(ctx, 5*time.Minute) defer cancel() options.Model = model stream, err := ai.StreamText(ctx, options) for chunk := range stream.Chunks() { // Process chunks // If total time exceeds 5 minutes, stream stops } ``` ## Recommended Timeouts ### By Operation Type ```go // Text Generation // - Simple prompts: 30-60 seconds ctx, cancel := context.WithTimeout(ctx, 1*time.Minute) // - Complex generation: 2-5 minutes ctx, cancel := context.WithTimeout(ctx, 5*time.Minute) // Video Generation // - Short videos (2-4s): 5-10 minutes ctx, cancel := context.WithTimeout(ctx, 10*time.Minute) // - Long videos: 15-30 minutes ctx, cancel := context.WithTimeout(ctx, 30*time.Minute) // - Batch videos: 20-40 minutes ctx, cancel := context.WithTimeout(ctx, 40*time.Minute) // Image Generation // - Single image: 1-2 minutes ctx, cancel := context.WithTimeout(ctx, 2*time.Minute) // - Batch images: 3-5 minutes ctx, cancel := context.WithTimeout(ctx, 5*time.Minute) // Embeddings // - Single embedding: 10-30 seconds ctx, cancel := context.WithTimeout(ctx, 30*time.Second) // - Batch embeddings: 1-5 minutes ctx, cancel := context.WithTimeout(ctx, 5*time.Minute) ``` ### By Provider Different providers have different typical response times: ```go // Fast providers (OpenAI, Anthropic, Google) ctx, cancel := context.WithTimeout(ctx, 1*time.Minute) // Medium providers (Cohere, Mistral) ctx, cancel := context.WithTimeout(ctx, 2*time.Minute) // Slower providers (self-hosted, Ollama) ctx, cancel := context.WithTimeout(ctx, 5*time.Minute) // Gateway (automatic provider selection) // Use longer timeouts to allow for failover ctx, cancel := context.WithTimeout(ctx, 3*time.Minute) ``` ## Handling Timeout Errors ### Basic Error Handling ```go ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute) defer cancel() options.Model = model result, err := ai.GenerateText(ctx, options) if err != nil { if errors.Is(err, context.DeadlineExceeded) { log.Println("Request timed out - consider increasing timeout") return handleTimeout() } return fmt.Errorf("generation failed: %w", err) } ``` ### Retry with Longer Timeout ```go func generateWithRetry(model provider.LanguageModel, options ai.GenerateTextOptions) (*ai.GenerateTextResult, error) { timeouts := []time.Duration{ 1 * time.Minute, 3 * time.Minute, 5 * time.Minute, } for i, timeout := range timeouts { ctx, cancel := context.WithTimeout(context.Background(), timeout) options.Model = model result, err := ai.GenerateText(ctx, options) cancel() if err == nil { return result, nil } if !errors.Is(err, context.DeadlineExceeded) { // Non-timeout error, don't retry return nil, err } if i < len(timeouts)-1 { log.Printf("Attempt %d timed out after %v, retrying...", i+1, timeout) } } return nil, fmt.Errorf("all retry attempts timed out") } ``` ### Exponential Backoff ```go func generateWithBackoff(model provider.LanguageModel, options ai.GenerateTextOptions) (*ai.GenerateTextResult, error) { timeout := 1 * time.Minute maxTimeout := 10 * time.Minute maxAttempts := 5 for attempt := 0; attempt < maxAttempts; attempt++ { ctx, cancel := context.WithTimeout(context.Background(), timeout) options.Model = model result, err := ai.GenerateText(ctx, options) cancel() if err == nil { return result, nil } if !errors.Is(err, context.DeadlineExceeded) { return nil, err } // Double timeout for next attempt timeout *= 2 if timeout > maxTimeout { timeout = maxTimeout } log.Printf("Attempt %d timed out, next timeout: %v", attempt+1, timeout) } return nil, fmt.Errorf("all attempts timed out") } ``` ## Preventing Timeouts ### 1. Set Appropriate Timeouts ```go // Consider operation complexity func getTimeout(opts ai.GenerateTextOptions) time.Duration { baseTimeout := 30 * time.Second // Adjust for token count if opts.MaxTokens != nil && *opts.MaxTokens > 1000 { baseTimeout = 2 * time.Minute } // Adjust for tools if len(opts.Tools) > 0 { baseTimeout += 1 * time.Minute } return baseTimeout } ``` ### 2. Optimize Requests ```go // Reduce MaxTokens for faster responses result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing", MaxTokens: ptr(200), // Shorter = faster }) // Use streaming for long content stream, err := ai.StreamText(ctx, ai.StreamTextOptions{ Model: model, Prompt: "Write a long article...", // Get partial results immediately }) ``` ### 3. Use Streaming Streaming provides partial results immediately, reducing perceived latency: ```go ctx, cancel := context.WithTimeout(ctx, 5*time.Minute) defer cancel() options.Model = model stream, err := ai.StreamText(ctx, options) if err != nil { return err } for chunk := range stream.Chunks() { // Display partial results immediately fmt.Print(chunk.Text) } if err := stream.Err(); err != nil { if errors.Is(err, context.DeadlineExceeded) { log.Println("Stream timed out, but got partial results") } } ``` ### 4. Implement Progress Tracking ```go type ProgressTracker struct { StartTime time.Time Timeout time.Duration } func (p *ProgressTracker) Check() error { elapsed := time.Since(p.StartTime) if elapsed > p.Timeout { return fmt.Errorf("operation exceeded %v", p.Timeout) } return nil } tracker := &ProgressTracker{ StartTime: time.Now(), Timeout: 5 * time.Minute, } options.Model = model stream, err := ai.StreamText(ctx, options) for chunk := range stream.Chunks() { if err := tracker.Check(); err != nil { log.Printf("Progress timeout: %v", err) break } fmt.Print(chunk.Text) } ``` ## Async Processing For very long operations, consider async processing: ```go type VideoJob struct { ID string Status string Result *ai.GenerateVideoResult Error error } func startVideoGeneration(opts ai.GenerateVideoOptions) *VideoJob { job := &VideoJob{ ID: generateJobID(), Status: "pending", } go func() { ctx, cancel := context.WithTimeout(context.Background(), 30*time.Minute) defer cancel() result, err := ai.GenerateVideo(ctx, opts) job.Result = result job.Error = err job.Status = "completed" }() return job } // Usage job := startVideoGeneration(options) fmt.Printf("Video generation started: %s\n", job.ID) // Check status later time.Sleep(5 * time.Minute) if job.Status == "completed" { if job.Error != nil { log.Printf("Generation failed: %v", job.Error) } else { log.Printf("Video ready!") } } else { log.Printf("Still generating...") } ``` ## Timeout Error Messages ### Gateway Provider Timeout Error ``` Gateway request timed out: context deadline exceeded This is a client-side timeout. To resolve this: 1. Increase your context timeout: ctx, cancel := context.WithTimeout(ctx, 5*time.Minute) 2. For video generation, consider using longer timeouts (e.g., 10 minutes) 3. Check your network connectivity 4. Verify the Gateway API is accessible For more information, see: https://vercel.com/docs/ai-gateway/capabilities/video-generation#extending-timeouts ``` ### Standard Context Timeout ``` context deadline exceeded ``` ## Debugging Timeouts ### Add Logging ```go import "log" timeout := 2 * time.Minute log.Printf("Starting generation with %v timeout", timeout) startTime := time.Now() ctx, cancel := context.WithTimeout(context.Background(), timeout) defer cancel() options.Model = model result, err := ai.GenerateText(ctx, options) elapsed := time.Since(startTime) if err != nil { if errors.Is(err, context.DeadlineExceeded) { log.Printf("Timed out after %v (limit: %v)", elapsed, timeout) } else { log.Printf("Failed after %v: %v", elapsed, err) } } else { log.Printf("Completed in %v", elapsed) } ``` ### Monitor Performance ```go type OperationMetrics struct { TotalRequests int64 TimeoutCount int64 AverageDuration time.Duration } var metrics OperationMetrics func trackOperation(ctx context.Context, model provider.LanguageModel, opts ai.GenerateTextOptions) error { startTime := time.Now() metrics.TotalRequests++ opts.Model = model result, err := ai.GenerateText(ctx, opts) duration := time.Since(startTime) if errors.Is(err, context.DeadlineExceeded) { metrics.TimeoutCount++ log.Printf("Timeout: %d/%d requests (%.1f%%)", metrics.TimeoutCount, metrics.TotalRequests, float64(metrics.TimeoutCount)/float64(metrics.TotalRequests)*100) } // Update average metrics.AverageDuration = (metrics.AverageDuration*time.Duration(metrics.TotalRequests-1) + duration) / time.Duration(metrics.TotalRequests) return err } ``` ## Production Best Practices 1. **Set Conservative Timeouts**: Start with longer timeouts and optimize based on metrics 2. **Implement Retry Logic**: Retry with longer timeouts on failure 3. **Use Streaming**: Get partial results immediately 4. **Monitor Timeout Rates**: Track and alert on high timeout rates 5. **Adjust Based on Complexity**: Longer timeouts for complex operations 6. **Consider Async Processing**: For very long operations (video generation) 7. **Provide User Feedback**: Show progress indicators for long operations 8. **Test with Real Conditions**: Test timeouts under production-like network conditions ## Related Documentation - [Context Cancellation](https://goaisdk.com/docs/troubleshooting/context-cancellation.md) - [Performance Optimization](https://goaisdk.com/docs/troubleshooting/performance.md) - [Gateway Provider](https://goaisdk.com/docs/providers/gateway.md) - [Video Generation](https://goaisdk.com/docs/ai-sdk-core/video-generation.md) ## Examples - [Timeout Handling Example](https://github.com/digitallysavvy/go-ai/tree/main/examples/providers/gateway/timeout-handling) - [Video Generation Example](https://github.com/digitallysavvy/go-ai/tree/main/examples/providers/gateway/video-generation) ```go func ptr[T any](v T) *T { return &v } ``` --- # Troubleshooting > Landing page for Go AI SDK troubleshooting guides, covering common issues, a quick diagnostic guide, debug mode, and where to get further help. Canonical URL: https://goaisdk.com/docs/troubleshooting/ Documentation index: https://goaisdk.com/llms.txt This section provides solutions for common problems you may encounter when using the Go AI SDK. From provider errors to streaming issues, you'll find practical solutions with working code examples. ## Common Issues The most frequent issues developers face when building AI applications with Go: - **[Common Errors](https://goaisdk.com/docs/troubleshooting/common-errors.md)** - Solutions for the most frequent SDK errors - **[Provider Errors](https://goaisdk.com/docs/troubleshooting/provider-errors.md)** - Troubleshooting provider-specific issues - **[Rate Limits](https://goaisdk.com/docs/troubleshooting/rate-limits.md)** - Handling rate limiting and quota exhaustion - **[Context Cancellation](https://goaisdk.com/docs/troubleshooting/context-cancellation.md)** - Debugging context timeout and cancellation issues - **[Streaming Issues](https://goaisdk.com/docs/troubleshooting/streaming-issues.md)** - Solving problems with streaming responses - **[Tool Calling Errors](https://goaisdk.com/docs/troubleshooting/tool-calling-errors.md)** - Debugging tool execution and schema issues - **[Schema Validation](https://goaisdk.com/docs/troubleshooting/schema-validation.md)** - Fixing JSON schema validation errors - **[Performance](https://goaisdk.com/docs/troubleshooting/performance.md)** - Optimizing slow operations and memory usage - **[Debugging](https://goaisdk.com/docs/troubleshooting/debugging.md)** - Comprehensive debugging techniques ## Quick Diagnostic Guide ### Error Messages If you see specific error messages: - `context deadline exceeded` → See [Context Cancellation](https://goaisdk.com/docs/troubleshooting/context-cancellation.md) - `rate limit exceeded` → See [Rate Limits](https://goaisdk.com/docs/troubleshooting/rate-limits.md) - `invalid API key` → See [Provider Errors](https://goaisdk.com/docs/troubleshooting/provider-errors.md) - `schema validation failed` → See [Schema Validation](https://goaisdk.com/docs/troubleshooting/schema-validation.md) - `tool execution failed` → See [Tool Calling Errors](https://goaisdk.com/docs/troubleshooting/tool-calling-errors.md) ### Behavior Issues If you're experiencing specific behaviors: - Stream stops unexpectedly → See [Streaming Issues](https://goaisdk.com/docs/troubleshooting/streaming-issues.md) - Application hangs → See [Context Cancellation](https://goaisdk.com/docs/troubleshooting/context-cancellation.md) and [Debugging](https://goaisdk.com/docs/troubleshooting/debugging.md) - Slow responses → See [Performance](https://goaisdk.com/docs/troubleshooting/performance.md) - Memory leaks → See [Performance](https://goaisdk.com/docs/troubleshooting/performance.md) and [Common Errors](https://goaisdk.com/docs/troubleshooting/common-errors.md) ## Getting Help If you can't find a solution in this troubleshooting guide: 1. **Check the logs** - Enable debug logging to see detailed error information 2. **Review the documentation** - Check the [API Reference](https://goaisdk.com/docs/reference.md) for correct usage 3. **Search GitHub Issues** - Your issue may already be reported or solved 4. **Create a minimal reproduction** - Isolate the problem in a small example 5. **Report the issue** - Open a GitHub issue with your reproduction case ## Debug Mode The SDK has no global log-level switch. Its built-in logging hook is the warnings logger: route provider warnings to your own logger with `ai.SetLogWarnings`, or see [Debugging](https://goaisdk.com/docs/troubleshooting/debugging.md#sdk-warning-logging) for the full pattern, including wrapping the HTTP client to log raw requests and responses. ```go import ( "log" "github.com/digitallysavvy/go-ai/pkg/ai" ) func main() { // Route SDK warnings into your own logger instead of stderr. ai.SetLogWarnings(func(opts ai.LogWarningsOptions) { for _, w := range opts.Warnings { log.Printf("[%s/%s] %s", opts.Provider, opts.Model, ai.FormatWarning(w, opts.Provider, opts.Model)) } }) // Your code here... } ``` ## Next Steps - Start with [Common Errors](https://goaisdk.com/docs/troubleshooting/common-errors.md) for the most frequent issues - Review [Debugging](https://goaisdk.com/docs/troubleshooting/debugging.md) for comprehensive debugging techniques - Check [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) for proper error handling patterns --- # Serve a useChat frontend from Go > Build a Go HTTP endpoint that the stock useChat hook talks to, with validation, keep-alive, CORS, proxy settings and custom data parts. Canonical URL: https://goaisdk.com/docs/build-a-chat-app/serve-usechat-from-go Documentation index: https://goaisdk.com/llms.txt The `useChat` hook from `@ai-sdk/react` posts a list of UI messages to an endpoint and reads back a stream of Server-Sent Events (SSE). The Go AI SDK writes that stream, so the frontend needs no adapter. This guide builds the endpoint the way the [Shipyard demo](https://github.com/digitallysavvy/go-ai-demo) does. ## The protocol `DefaultChatTransport` sends a `POST` with a JSON body: ```json { "id": "chat-id", "messages": [ { "id": "m1", "role": "user", "parts": [{ "type": "text", "text": "Hello" }] } ], "trigger": "submit-message", "messageId": "m2" } ``` Anything you pass in the transport's `body` option is merged into the same object. The `messages` array holds UI messages, not model messages: each has `parts` such as text, tool calls with their state, files and your own data parts. The response is `text/event-stream`. Each event is `data: `, and the stream ends with `data: [DONE]`. The response also carries the header `X-Vercel-AI-UI-Message-Stream: v1`, which `useChat` checks. ## A minimal endpoint This program runs an agent on the messages `useChat` sends. It validates first, so a bad request gets a plain 400 before any streaming starts. ```go package main import ( "encoding/json" "log" "net/http" "os" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) type chatRequest struct { Messages json.RawMessage `json:"messages"` } func main() { model, err := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}). LanguageModel(anthropic.ClaudeSonnet5_5) if err != nil { log.Fatal(err) } assistant := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a concise assistant.", }) http.HandleFunc("POST /api/chat", func(w http.ResponseWriter, r *http.Request) { var req chatRequest if err := json.NewDecoder(http.MaxBytesReader(w, r.Body, 8<<20)).Decode(&req); err != nil { http.Error(w, "invalid request body", http.StatusBadRequest) return } // Validates the UI messages against the agent's tools and converts // them to model messages. Nothing is written to w yet. chunks, _, err := agent.CreateAgentUIStreamFromUIMessages(r.Context(), assistant, agent.CreateAgentUIStreamFromUIMessagesOptions{UIMessages: []byte(req.Messages)}) if err != nil { http.Error(w, err.Error(), http.StatusBadRequest) return } // Sets the stream headers, writes 200 and flushes every chunk. if err := ai.PipeUIMessageChunksToResponse(chunks, w, nil); err != nil { log.Printf("write stream: %v", err) } }) log.Fatal(http.ListenAndServe(":8080", nil)) } ``` Pass the request context. `r.Context()` is cancelled when the browser disconnects, which stops the agent and the provider call. ## Choose a helper | You have | Use | | --- | --- | | UI messages from `useChat`, and you want one call | `agent.PipeAgentUIStreamFromUIMessagesToResponse(ctx, agent, opts, w)` | | UI messages, and you want to answer 400 on bad input | `agent.CreateAgentUIStreamFromUIMessages`, then `ai.PipeUIMessageChunksToResponse` | | A stream you built with `ai.CreateUIMessageStreamWithOptions` | `ai.PipeUIMessageChunksToResponse` | | A `*ai.StreamTextResult` from `ai.StreamText` | `ai.PipeUIMessageStreamToResponse` | | A framework that wants an `*http.Response` | `agent.CreateAgentUIStreamResponseFromUIMessages` or `ai.CreateUIMessageChunksResponse` | The `Pipe*` helpers set the status and headers when the writer is an `http.ResponseWriter`, so do not call `WriteHeader` or set the SSE headers yourself. Headers you set earlier, such as CORS headers, stay in place. See [Stream transport helpers](https://goaisdk.com/docs/reference/ai/stream-transport-helpers.md) and [Agent UI stream helpers](https://goaisdk.com/docs/reference/ai/agent-ui-stream-helpers.md) for every signature. `PipeAgentUIStreamFromUIMessagesToResponse` returns an error for invalid messages before it writes anything, but it also returns write errors after streaming starts, and the two look the same to the caller. Use the two-step form above when you want to send a 400. ## Errors before and after streaming Two kinds of failure need different handling. **Before streaming starts**, the response is still open. Bad JSON, a `messages` value that fails validation, and messages that can't be converted all return an error from `CreateAgentUIStreamFromUIMessages`. Answer with a 4xx status. `useChat` sets its `error` state and `status` becomes `error`. **After streaming starts**, the status line has already gone out as `200`. A provider failure, a tool that panics, or an invalid approval signature surfaces as an `error` chunk in the stream: ```text data: {"type":"error","errorText":"An error occurred."} ``` `useChat` shows it through the same `error` state. The default text is generic so that provider messages, URLs and keys do not reach the browser. To send real text, set an `OnError` function that maps the error to the string you want users to see: ```go chunks, _, err := agent.CreateAgentUIStreamFromUIMessages(ctx, assistant, agent.CreateAgentUIStreamFromUIMessagesOptions{ UIMessages: []byte(req.Messages), UIMessageStream: ai.UIMessageStreamResultOptions{ OnError: func(err error) string { log.Printf("chat error: %v", err) return "Something went wrong. Try again." }, }, }) ``` If you build the stream yourself with `ai.CreateUIMessageStreamWithOptions`, the same `OnError` goes on `ai.UIMessageStreamOptions`. The demo returns `err.Error()` because it runs locally; do not do that in production. ## Keep-alive Coding agents and slow tools can leave a stream quiet for a minute. Idle connections get closed by proxies and load balancers. `UIMessageStreamResponseInit.KeepAliveMs` makes the pipe send an opening `: stream-open` comment right away, then a `: keep-alive` comment whenever the stream has been silent for that long. SSE clients ignore comments. ```go keepAlive := 15 * time.Second err := ai.PipeUIMessageChunksToResponse(chunks, w, &ai.UIMessageStreamResponseInit{ KeepAliveMs: &keepAlive, }) ``` The field is a `*time.Duration`, despite the `Ms` suffix, which it keeps for parity with the TypeScript option. Pick an interval shorter than your proxy's idle timeout. ## CORS When the frontend and the API live on different origins, such as `localhost:3000` and `localhost:8080`, the browser sends a preflight `OPTIONS` request because the body is JSON. Answer it, and name the one origin you trust: ```go func withCORS(origin string, next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { w.Header().Set("Access-Control-Allow-Origin", origin) w.Header().Set("Access-Control-Allow-Headers", "Content-Type") w.Header().Set("Access-Control-Allow-Methods", "GET, POST, OPTIONS") if r.Method == http.MethodOptions { w.WriteHeader(http.StatusNoContent) return } next.ServeHTTP(w, r) }) } ``` If you send an `Authorization` header from the client, add it to `Access-Control-Allow-Headers`. If you serve the frontend and the API from one origin, or proxy `/api` through your frontend server, you do not need CORS. ## Proxies and buffering SSE only works if every layer between Go and the browser passes bytes through as they arrive. The SDK flushes after every chunk and sends `Cache-Control: no-cache` and `X-Accel-Buffering: no`. The rest is up to your stack: - **Wrap the `ResponseWriter` carefully.** Logging, metrics and compression middleware often wrap the writer in a type that does not implement `http.Flusher`. The SDK then cannot flush, and the output arrives in large blocks. Implement `Flush` on your wrapper by delegating to the underlying writer, or skip the wrapper for this route. - **Do not compress the stream.** Gzip buffers until it has a block to emit. Exclude `text/event-stream` from compression. - **nginx.** The `X-Accel-Buffering: no` header already tells nginx not to buffer this response. If you still see batching, set `proxy_buffering off;` and `proxy_http_version 1.1;` on the location, and raise `proxy_read_timeout` above your keep-alive interval. - **CDNs and platform proxies.** Check that streaming responses are enabled and that the idle timeout is longer than your keep-alive interval. - **Server timeouts.** `http.Server.WriteTimeout` ends the stream when it expires. Leave it at zero for streaming routes, and set `ReadHeaderTimeout` as the demo does. ### Troubleshooting **`useChat` shows nothing.** Open the network tab and check the `/api/chat` response. A non-200 status or an HTML error page means the request failed before streaming; look at your server log. A CORS error in the console means the preflight was not answered. A 200 with the right content type but no events points to a proxy that holds the response. **Output arrives all at once at the end.** Something is buffering. Test the Go server directly with `curl -N -X POST localhost:8080/api/chat -H 'Content-Type: application/json' -d '{"messages":[...]}'`. If chunks appear one by one there, the problem is a proxy, compression or a wrapped `ResponseWriter` in front of it. If they arrive together even there, check for a middleware that wraps the writer. **It streams, but `useChat` reports a parse error.** The response is missing `X-Vercel-AI-UI-Message-Stream: v1`, or something else is writing to the body. Use the `Pipe*` helpers and write nothing yourself. **The stream stops after about a minute.** An idle timeout closed the connection. Turn on `KeepAliveMs`. ## Custom data parts Data parts let the server send typed side-channel data, such as progress or the activity log of a coding agent, as part of the assistant message. Build the stream with `ai.CreateUIMessageStreamWithOptions`, write data chunks through the writer, and merge the agent's stream into it: ```go package main import ( "encoding/json" "log" "net/http" "os" "time" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { model, err := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}). LanguageModel(anthropic.ClaudeSonnet5_5) if err != nil { log.Fatal(err) } assistant := agent.NewToolLoopAgent(agent.AgentConfig{Model: model}) http.HandleFunc("POST /api/chat", func(w http.ResponseWriter, r *http.Request) { var req struct { Messages json.RawMessage `json:"messages"` } if err := json.NewDecoder(http.MaxBytesReader(w, r.Body, 8<<20)).Decode(&req); err != nil { http.Error(w, "invalid request body", http.StatusBadRequest) return } ctx := r.Context() chunks, _ := ai.CreateUIMessageStreamWithOptions(ctx, ai.UIMessageStreamOptions{ OnError: func(err error) string { log.Printf("chat error: %v", err) return "Something went wrong." }, Execute: func(writer ai.UIMessageStreamWriter) { // A data part with an id is replaced in place when you write // the same id again. writer.Write(ai.UIMessageChunk{ "type": "data-status", "id": "status", "data": map[string]interface{}{"phase": "thinking"}, }) stream, errs, err := agent.CreateAgentUIStreamFromUIMessages(ctx, assistant, agent.CreateAgentUIStreamFromUIMessagesOptions{UIMessages: []byte(req.Messages)}) if err != nil { // The response is already open, so report it as a chunk. writer.Write(ai.UIMessageChunk{"type": "error", "errorText": err.Error()}) return } writer.Merge(stream) go func() { for err := range errs { if err != nil { log.Printf("agent stream: %v", err) } } }() }, }) keepAlive := 15 * time.Second if err := ai.PipeUIMessageChunksToResponse(chunks, w, &ai.UIMessageStreamResponseInit{KeepAliveMs: &keepAlive}); err != nil { log.Printf("write stream: %v", err) } }) log.Fatal(http.ListenAndServe(":8080", nil)) } ``` Rules for data chunks: - The `type` is `data-` plus a name you choose. The client sees it as a part of that type. - Writing the same `id` again replaces the part's `data` in place. The demo sends one `data-coder` part per coding run and rewrites it as work progresses, using the tool call ID as the part ID. - Add `"transient": true` for data that should reach the client's `onData` callback but not be stored in the message history. - Data must be JSON-serializable. Send empty slices, not nil, when the client expects an array: `null` breaks a client that calls `.map` on it. - `Execute` must return after it has set up the stream. Anything you pass to `writer.Merge` keeps streaming after `Execute` returns. Because the data parts live in the assistant message, they come back to the server on the next request. Validation accepts them. Pass `ConvertDataPart` in `CreateAgentUIStreamFromUIMessagesOptions` if you want any of them turned into text or file parts for the model; by default they are ignored. ## The client On the client, point `DefaultChatTransport` at your endpoint. `body` can be an object or a function; the function runs on every request, including the automatic resend after a tool approval. ```tsx 'use client'; import { useChat } from '@ai-sdk/react'; import { DefaultChatTransport, type UIMessage } from 'ai'; import { useMemo, useRef, useState } from 'react'; type ChatMessage = UIMessage; export default function Chat() { const [mode, setMode] = useState('default'); const modeRef = useRef(mode); modeRef.current = mode; const transport = useMemo( () => new DefaultChatTransport({ api: 'http://localhost:8080/api/chat', body: () => ({ mode: modeRef.current }), }), [], ); const { messages, sendMessage, status, error } = useChat({ transport }); const [input, setInput] = useState(''); return (
{messages.map((m) => (
{m.parts.map((part, i) => { if (part.type === 'text') return

{part.text}

; if (part.type === 'data-status') return {part.data.phase}; return null; })}
))} {error &&

{error.message}

}
{ e.preventDefault(); sendMessage({ text: input }); setInput(''); }} > setInput(e.target.value)} disabled={status !== 'ready'} />
); } ``` The `mode` field arrives in the JSON body next to `messages`. Read it by adding a field to your request struct, as the demo does with `agent`. The second type argument to `UIMessage` types your data parts, so `part.data` is checked against what the server sends. ## See it in the demo - [`server/chat.go`](https://github.com/digitallysavvy/go-ai-demo/blob/main/server/chat.go): the request struct, the 400 for a bad body, `CreateUIMessageStreamWithOptions` with `Execute`, `CreateAgentUIStreamFromUIMessages`, `Merge`, and the keep-alive pipe. - [`server/main.go`](https://github.com/digitallysavvy/go-ai-demo/blob/main/server/main.go): `withCORS`, the HTTP server timeouts and graceful shutdown. - [`web/app/page.tsx`](https://github.com/digitallysavvy/go-ai-demo/blob/main/web/app/page.tsx): `DefaultChatTransport` with a `body` function and `useChat`. - [`web/lib/types.ts`](https://github.com/digitallysavvy/go-ai-demo/blob/main/web/lib/types.ts): the typed `data-coder` part that mirrors the Go struct. Next: [Tool approval end to end](https://goaisdk.com/docs/build-a-chat-app/tool-approval.md). --- # Tool approval end to end > Let a model ask and a user decide. Mark tools as approval-gated, sign approval requests, and run the full round trip with useChat and a Go backend. Canonical URL: https://goaisdk.com/docs/build-a-chat-app/tool-approval Documentation index: https://goaisdk.com/llms.txt Some tools should not run just because a model asked. Deploying, deleting, paying and handing work to a coding agent all deserve a human decision. The Go AI SDK pauses the tool call, sends an approval request to the client, and runs the tool only after the client sends back a signed approval. This guide covers the server, the client, and what happens on each side. The [Serve a useChat frontend](https://goaisdk.com/docs/build-a-chat-app/serve-usechat-from-go.md) guide covers the endpoint itself. ## How it works 1. The model calls a tool that needs approval. The agent does not run it. It ends the step and streams a `tool-approval-request` chunk. 2. `useChat` shows the tool part in the `approval-requested` state. Your UI renders Approve and Deny buttons. 3. The user clicks. The client calls `addToolApprovalResponse`, and the tool part moves to `approval-responded`. 4. With `sendAutomaticallyWhen` set, the client sends a new request with the updated message list. 5. The server validates the approval against its signature. If the user approved, it runs the tool and streams the result, then the model continues. If the user denied, it skips the tool and tells the model. The server holds no state between steps 1 and 5. Everything it needs comes back in the message history, which is why the signature matters. ## Mark a tool as needing approval Set `ToolApproval` on the tool. It takes a bool or a function: ```go package main import ( "context" "fmt" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/schema" ) func deployTool() types.Tool { return types.Tool{ Name: "deploy", Description: "Deploy the service to an environment.", Parameters: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "env": map[string]interface{}{"type": "string", "enum": []string{"staging", "prod"}}, }, "required": []string{"env"}, }), // Always ask: // ToolApproval: true, // // Ask only for production. The conversion to the named function type // is required; see the note below. ToolApproval: types.ToolNeedsApprovalFunc(func(ctx context.Context, input map[string]interface{}, opts types.ToolNeedsApprovalOptions) bool { return input["env"] == "prod" }), Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{"deployed": input["env"]}, nil }, } } func main() { fmt.Println(deployTool().Name) } ``` `ToolApproval` accepts these values: | Value | Meaning | | --- | --- | | `true` | Always ask the user. | | `types.ToolNeedsApprovalFunc` | Ask when the function returns `true`. It receives the parsed input, the tool call ID, the messages and the tool's context. | | `types.ToolApprovalStatus` | A fixed outcome: `ToolApprovalStatusUserApproval`, `ToolApprovalStatusApproved` or `ToolApprovalStatusDenied`. | | `nil` or `false` | Run without asking. | `ToolApproval` is an `interface{}`, so the compiler can't check what you put in it. A bare function literal with one of the signatures above works the same as its named type. The SDK fails closed on anything else: a value of an unsupported type, or an unknown status string such as `"user_approval"`, becomes a tool error and the tool doesn't run. Writing the named type, as in the example, makes the intent clear and lets the compiler check the signature. `NeedsApproval` still works but is deprecated. Use `ToolApproval`. ## Set a policy for the whole call You can also decide approval outside the tools, on the agent or on a single call. A call-level policy wins over a tool's own setting. Use a map to cover chosen tools, or a function to decide for every call: ```go reason := "Production deploys are blocked outside business hours." assistant := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: tools, ToolApproval: map[string]types.ToolApprovalValue{ "read_logs": types.ToolApprovalStatusApproved, "deploy": types.ToolApprovalResult{Status: types.ToolApprovalStatusUserApproval}, "drop_db": types.ToolApprovalResult{Status: types.ToolApprovalStatusDenied, Reason: &reason}, }, }) ``` A map decides only for the tools it names. Other tools fall back to their own `ToolApproval`. A `types.GenericToolApprovalFunc` receives `types.ToolApprovalOptions` (the tool call, the tools, the messages and the contexts) and returns the final decision for every call: ```go ToolApproval: types.GenericToolApprovalFunc(func(opts types.ToolApprovalOptions) types.ToolApprovalResult { if opts.ToolCall.ToolName == "deploy" { return types.ToolApprovalResult{Status: types.ToolApprovalStatusUserApproval} } return types.ToolApprovalResult{Status: types.ToolApprovalStatusApproved} }), ``` A denied status needs no user. The tool does not run, and the model receives an execution-denied result with your reason. In a per-tool map you can also use `types.SingleToolApprovalFunc`, which receives that tool's parsed arguments. To change the policy between steps, set `config.ToolApproval` in the agent's `PrepareCall` function. `ai.GenerateTextOptions` and `ai.StreamTextOptions` have the same `ToolApproval` field. ## Sign approval requests When the user approves, the client sends the decision back inside the message history. A client controls that history. Without a check, anyone who can call your endpoint could send an `approval-responded` part for a call the server never requested, or change the tool's input after the user approved it. `ExperimentalToolApprovalSecret` closes that gap. With a secret set, the server signs each approval request with HMAC-SHA256 over the approval ID, the tool call ID, the tool name and a digest of the input. The signature travels to the client in the `tool-approval-request` chunk and comes back with the response. On resume, the server verifies it before running anything. A missing or wrong signature stops the call. ```go shipyard := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, Tools: tools, ExperimentalToolApprovalSecret: secret, // []byte }) ``` Handle the secret like any signing key: - Use at least 32 random bytes and load it from configuration, not source code. - Use the same secret on every instance. A request approved on one replica must verify on another. - A secret that changes, such as a new random value at every start, invalidates approvals that are still waiting. The demo generates a random secret per process unless `TOOL_APPROVAL_SECRET` is set, which is fine locally. - Without a secret, nothing is signed or verified. Set one for any endpoint that is reachable by users. The same field exists on `ai.GenerateTextOptions`, `ai.StreamTextOptions` and `PrepareCallConfig`. A signature failure returns an `*ai.InvalidToolApprovalSignatureError`; `ai.IsInvalidToolApprovalSignatureError` tests for it. Over HTTP it reaches the client as an `error` chunk, after the `200` status has gone out, so use `OnError` if you want to log or shape the message. ## The server This server has one approval-gated tool. Nothing in the handler is specific to approval: the tool's setting and the secret do the work. ```go package main import ( "context" "crypto/rand" "encoding/json" "log" "net/http" "os" "time" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/schema" ) func main() { model, err := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}). LanguageModel(anthropic.ClaudeSonnet5_5) if err != nil { log.Fatal(err) } secret := []byte(os.Getenv("TOOL_APPROVAL_SECRET")) if len(secret) == 0 { secret = make([]byte, 32) if _, err := rand.Read(secret); err != nil { log.Fatal(err) } } deploy := types.Tool{ Name: "deploy", Description: "Deploy the service to an environment. Requires user approval.", Parameters: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "env": map[string]interface{}{"type": "string", "enum": []string{"staging", "prod"}}, }, "required": []string{"env"}, }), ToolApproval: true, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{"deployed": input["env"]}, nil }, } assistant := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You deploy services. Call deploy when the user asks. Keep replies short.", Tools: []types.Tool{deploy}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, ExperimentalToolApprovalSecret: secret, }) http.HandleFunc("POST /api/chat", func(w http.ResponseWriter, r *http.Request) { var req struct { Messages json.RawMessage `json:"messages"` } if err := json.NewDecoder(http.MaxBytesReader(w, r.Body, 8<<20)).Decode(&req); err != nil { http.Error(w, "invalid request body", http.StatusBadRequest) return } chunks, _, err := agent.CreateAgentUIStreamFromUIMessages(r.Context(), assistant, agent.CreateAgentUIStreamFromUIMessagesOptions{UIMessages: []byte(req.Messages)}) if err != nil { http.Error(w, err.Error(), http.StatusBadRequest) return } keepAlive := 15 * time.Second if err := ai.PipeUIMessageChunksToResponse(chunks, w, &ai.UIMessageStreamResponseInit{KeepAliveMs: &keepAlive}); err != nil { log.Printf("write stream: %v", err) } }) log.Fatal(http.ListenAndServe(":8080", nil)) } ``` ## What goes over the wire The first response ends at the approval request. The model's finish reason stays `tool-calls`: ```text data: {"type":"start"} data: {"type":"start-step"} data: {"type":"tool-input-available","toolCallId":"call-1","toolName":"deploy","input":{"env":"prod"}} data: {"type":"tool-approval-request","toolCallId":"call-1","approvalId":"982897cfe028b8ec","signature":"-oKB..."} data: {"type":"finish-step"} data: {"type":"finish","finishReason":"tool-calls"} data: [DONE] ``` The follow-up request carries the same assistant message with the tool part updated. The part type is `tool-` plus the tool name: ```json { "id": "a1", "role": "assistant", "parts": [ { "type": "tool-deploy", "toolCallId": "call-1", "state": "approval-responded", "input": { "env": "prod" }, "approval": { "id": "982897cfe028b8ec", "approved": true, "signature": "-oKB..." } } ] } ``` If approved, the server runs the tool and streams `tool-output-available`, followed by the model's next step. The new chunks continue the same assistant message, because the `start` chunk reuses its ID. ## The client ```tsx 'use client'; import { useChat } from '@ai-sdk/react'; import { DefaultChatTransport, isToolUIPart, lastAssistantMessageIsCompleteWithApprovalResponses } from 'ai'; import { useState } from 'react'; export default function Chat() { const { messages, sendMessage, addToolApprovalResponse, status } = useChat({ transport: new DefaultChatTransport({ api: 'http://localhost:8080/api/chat' }), sendAutomaticallyWhen: lastAssistantMessageIsCompleteWithApprovalResponses, }); const [input, setInput] = useState(''); return (
{messages.map((message) => message.parts.map((part, i) => { if (part.type === 'text') return

{part.text}

; if (isToolUIPart(part) && part.state === 'approval-requested') { return (

The assistant wants to run a tool with {JSON.stringify(part.input)}.

); } if (isToolUIPart(part) && part.state === 'output-denied') { return

Denied.

; } return null; }), )}
{ e.preventDefault(); sendMessage({ text: input }); setInput(''); }} > setInput(e.target.value)} disabled={status !== 'ready'} />
); } ``` `addToolApprovalResponse` takes the approval `id`, not the tool call ID. Without `sendAutomaticallyWhen: lastAssistantMessageIsCompleteWithApprovalResponses`, the click only updates local state and nothing is sent until the user types again. ## Denial When the user denies, the client sends `approved: false` and an optional `reason`. The server does not run the tool. The stream carries a `tool-output-denied` chunk, the tool part ends in `output-denied`, and the model receives an execution-denied tool result containing the reason. The model then continues. It can explain that the action did not happen or try another approach. Write your system prompt with that in mind: tell the model to stop and report when a tool is denied. A call-level `ToolApprovalStatusDenied` works the same way without asking the user. If the tool's input fails validation again after approval, the model gets a tool error and the tool does not run. ## The legacy blocking approver (CLI only) Before UI approval existed, the agent had a blocking callback: ```go agent.AgentConfig{ ToolApprovalRequired: true, ToolApprover: func(call types.ToolCall) bool { fmt.Printf("Run %s? [y/N] ", call.ToolName) var answer string fmt.Scanln(&answer) return answer == "y" }, } ``` `ToolApprover` runs inline, on the agent's goroutine, and returns the answer before the loop moves on. That fits a terminal program where a person is waiting at the keyboard. It does not fit an HTTP handler: the request would block while the browser waits, with no way to show a button. It applies only to calls whose approval status is not-applicable, so a tool already marked with `ToolApproval` goes through the round trip above instead. For anything with a UI, use `ToolApproval`. See [Building agents](https://goaisdk.com/docs/agents/building-agents.md) for the CLI pattern. ## See it in the demo - [`server/tools.go`](https://github.com/digitallysavvy/go-ai-demo/blob/main/server/tools.go): `delegate_to_coding_agent` sets `ToolApproval: true`, so the model can ask and only the user can say yes. - [`server/chat.go`](https://github.com/digitallysavvy/go-ai-demo/blob/main/server/chat.go): `ExperimentalToolApprovalSecret` on the agent and `CreateAgentUIStreamFromUIMessages` handling the resumed request. - [`server/main.go`](https://github.com/digitallysavvy/go-ai-demo/blob/main/server/main.go): where the per-process secret and `TOOL_APPROVAL_SECRET` are loaded. - [`web/app/page.tsx`](https://github.com/digitallysavvy/go-ai-demo/blob/main/web/app/page.tsx): `addToolApprovalResponse` and `lastAssistantMessageIsCompleteWithApprovalResponses`. - [`web/components/ToolPart.tsx`](https://github.com/digitallysavvy/go-ai-demo/blob/main/web/components/ToolPart.tsx): the approval card for each tool part state. Next: [Coding agents with the harness](https://goaisdk.com/docs/build-a-chat-app/coding-agents-harness.md). --- # Coding agents with the harness > Run Claude Code or Codex from Go with harness.NewAgent. Covers sandboxes, sessions, permission modes, seeding files, streaming activity into a chat UI and reading changes back. Canonical URL: https://goaisdk.com/docs/build-a-chat-app/coding-agents-harness Documentation index: https://goaisdk.com/llms.txt The harness lets a Go program drive an existing coding agent, such as Claude Code or Codex, inside a sandbox. Your code gives it a task and a workspace. It edits files and runs commands. You stream what it does, then read the files it changed. This is different from the [hand-rolled tool loop](https://goaisdk.com/docs/advanced/coding-agents.md), where your model calls tools you wrote. With the harness, the coding agent brings its own tools, prompts and loop. ## The pieces | Piece | What it is | | --- | --- | | `harness.Agent` | Built with `harness.NewAgent`. Holds the settings: which coding agent, which model, which sandbox. | | Adapter | `claudecode.New` or `codex.New`. Translates between the harness protocol and the coding agent's own CLI. | | Sandbox provider | Where the coding agent runs. `local` runs on your machine. `vercel` runs on Vercel Sandbox. | | `harness.AgentSession` | One sandbox and one conversation. Created with `CreateSession`, cleaned up with `Destroy`. | The flow is: create the agent, create a session, stream a prompt into the session, read the files back, destroy the session. ## A complete program This program starts Claude Code in a local sandbox, seeds a file with a bug, asks the agent to fix it, prints its activity, and reads the fixed file back. ```go package main import ( "context" "errors" "fmt" "io" "log" "os" "path" "path/filepath" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/harness" "github.com/digitallysavvy/go-ai/pkg/harness/claudecode" "github.com/digitallysavvy/go-ai/pkg/harness/sandbox/local" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/providerutils" ) const brokenFile = `package main import "fmt" func add(a, b int) int { return a - b } // bug func main() { fmt.Println(add(2, 3)) } ` func main() { ctx := context.Background() cc, err := claudecode.New(claudecode.Settings{MaxTurns: 20}) if err != nil { log.Fatal(err) } coder, err := harness.NewAgent(harness.AgentSettings{ Harness: cc, Model: anthropic.ClaudeSonnet5_5, PermissionMode: harness.PermissionModeAllowAll, Instructions: "Fix the bug with a minimal change. End with one sentence describing it.", Sandbox: local.NewProvider(local.Options{ RootDir: filepath.Join(os.TempDir(), "recipe-sandboxes"), Ports: []int{4319}, }), SandboxConfig: harness.AgentSandboxConfig{ OnSession: func(ctx context.Context, sc harness.SandboxSessionContext) error { return sc.Session.WriteTextFile(ctx, providerutils.SandboxWriteTextFileOptions{ Path: path.Join(sc.SessionWorkDir, "main.go"), Content: brokenFile, }) }, }, }) if err != nil { log.Fatal(err) } session, err := coder.CreateSession(ctx, harness.CreateSessionOptions{SessionID: "guide-session"}) if err != nil { log.Fatal(err) } defer func() { _ = session.Destroy(context.WithoutCancel(ctx)) }() result, err := coder.Stream(ctx, agent.AgentStreamOptions{AgentGenerateOptions: agent.AgentGenerateOptions{ HarnessSession: session, Prompt: "add(2, 3) in main.go returns -1 but should return 5. Fix it.", }}) if err != nil { log.Fatal(err) } stream := result.Stream() for { chunk, err := stream.Next() if errors.Is(err, io.EOF) { break } if err != nil { log.Fatal(err) } switch chunk.Type { case provider.ChunkTypeText: fmt.Print(chunk.Text) case provider.ChunkTypeToolCall: if chunk.ToolCall != nil { fmt.Printf("\n[tool] %s\n", chunk.ToolCall.ToolName) } } } if err := stream.Err(); err != nil { log.Fatal(err) } content, err := session.GetSandboxSession().ReadTextFile(ctx, providerutils.SandboxReadTextFileOptions{ Path: path.Join(session.GetSessionWorkDir(), "main.go"), }) if err != nil { log.Fatal(err) } if content != nil { fmt.Println(*content) } } ``` The rest of this guide explains each part. ## Claude Code or Codex Both adapters implement the same interface, so switching is one line: ```go cc, err := claudecode.New(claudecode.Settings{MaxTurns: 30}) // or cx := codex.New(codex.Settings{}) ``` Set `harness.AgentSettings.Harness` to the adapter and `Model` to a model ID the coding agent understands. The demo reads the Codex model from an environment variable and leaves it empty by default, which lets Codex choose. Credentials are discovered from the host environment. Claude Code looks for `ANTHROPIC_API_KEY`, `ANTHROPIC_AUTH_TOKEN` or `CLAUDE_CODE_OAUTH_TOKEN`. Codex looks for `OPENAI_API_KEY` or `CODEX_API_KEY`. Set `Auth` in the adapter settings to choose a route explicitly, for example through the AI Gateway. Other adapters exist for OpenCode, Deep Agents, ACP-based agents and more. They share this shape. See [Migrating from the TypeScript AI SDK](https://goaisdk.com/docs/migration-guides/from-typescript-ai-sdk.md#harness) for the list. ## Sandboxes The sandbox provider decides where commands run. ### Local `local.NewProvider` runs everything on your machine, in one directory per session under `RootDir`: ```go local.NewProvider(local.Options{ RootDir: "/tmp/sandboxes", Ports: []int{4319}, }) ``` > **Warning:** the local sandbox is not isolated. Commands run as your user, with your files, network and environment. Use it for development and for work you trust. Use an isolated provider for anything else. `Ports` is required. The harness starts a small bridge process inside the sandbox and talks to it over a WebSocket on one port. A sandbox that exposes no port makes `CreateSession` fail with an error that asks for one. The local sandbox shares your machine's network, so two sessions on the same port at the same time collide. The demo gives Claude Code port 4319 and Codex port 4318, and runs one coding session at a time. The bridge runs on Node.js, so Node must be installed on the host. The first session installs the adapter's dependencies into the sandbox and takes longer than later ones. ### Vercel Sandbox For isolation, run the coding agent on [Vercel Sandbox](https://vercel.com/docs/vercel-sandbox). Create the sandbox session yourself, expose a port, and hand it to `CreateSession`. You then own its lifecycle: ```go import "github.com/digitallysavvy/go-ai/pkg/harness/sandbox/vercel" sandbox, err := vercel.CreateNetworkSandboxSession(ctx, vercel.CreateSessionOptions{ Runtime: "node24", Ports: []int{4319}, }) if err != nil { log.Fatal(err) } defer func() { _ = sandbox.Destroy(context.WithoutCancel(ctx)) }() session, err := coder.CreateSession(ctx, harness.CreateSessionOptions{ SessionID: "run-1", SandboxSession: sandbox, }) ``` Credentials come from `VERCEL_OIDC_TOKEN`, or from `Credentials` with a token, team ID and project ID. You do not need a `Sandbox` setting on the `harness.Agent` in this mode. See [Harness sandbox: Vercel](https://goaisdk.com/docs/reference/ai/harness-sandbox-vercel.md) for the full option list. ## Sessions A session is one sandbox plus one conversation with the coding agent. ```go session, err := coder.CreateSession(ctx, harness.CreateSessionOptions{SessionID: "run-1"}) defer session.Destroy(context.WithoutCancel(ctx)) ``` - `SessionID` names the session. Leave it empty to have one generated. Use a stable value when you may resume the session later. - `Destroy` stops the agent and deletes the sandbox. Defer it right after `CreateSession` succeeds. Wrap the context in `context.WithoutCancel`, so cleanup still runs when the request was cancelled. - `Detach` and `Stop` return a state value you can pass back in `CreateSessionOptions.ResumeFrom` to continue the same session later. - `GetSandboxSession()` returns the sandbox for file and command access. `GetSessionWorkDir()` returns the directory the agent works in. You can send several prompts to one session. Each call to `Stream` is a turn, and the agent remembers earlier turns. ## Permission modes `PermissionMode` sets the baseline for the agent's built-in tools: | Mode | Meaning | | --- | --- | | `harness.PermissionModeAllowAll` | Reads, edits and shell commands run without asking. The default. | | `harness.PermissionModeAllowEdits` | Reads and edits run. Shell commands need approval. | | `harness.PermissionModeAllowReads` | Reads run. Edits and shell commands need approval. | Codex supports only `allow-all`; starting a session in another mode fails with a capability error. Claude Code supports all three. With `allow-reads` or `allow-edits`, a call to a gated built-in tool surfaces as a tool approval request, which you answer with `Agent.ContinueStream`. For a coding agent that you have already put behind an approval in your own chat, as the demo does, `allow-all` inside a sandbox is the usual choice: the user approved the task, and the sandbox bounds the damage. For tools you define on the harness agent, use `harness.AgentSettings.ToolApproval`, a map from tool name to a static `ToolApprovalStatus`. ## Seed files with OnSession `SandboxConfig.OnSession` runs after the sandbox exists and the work directory has been created, and before the agent starts. Use it to copy the project in: ```go SandboxConfig: harness.AgentSandboxConfig{ OnSession: func(ctx context.Context, sc harness.SandboxSessionContext) error { for rel, content := range files { err := sc.Session.WriteTextFile(ctx, providerutils.SandboxWriteTextFileOptions{ Path: path.Join(sc.SessionWorkDir, rel), Content: content, }) if err != nil { return err } } // A baseline commit lets the agent use git status and git diff. res, err := sc.Session.Run(ctx, providerutils.SandboxProcessOptions{ Command: "git init -q && git add -A && git -c user.name=app -c user.email=app@localhost commit -qm baseline", WorkingDirectory: sc.SessionWorkDir, }) if err != nil { return err } if res.ExitCode != 0 { return fmt.Errorf("git init: %s", res.Stderr) } return nil }, }, ``` `sc.Session` is the sandbox, with `WriteTextFile`, `ReadTextFile` and `Run`. Always build paths from `sc.SessionWorkDir`; the sandbox's own root is not where the agent works. For work that every session repeats, such as installing dependencies, `SandboxConfig.OnBootstrap` with a `BootstrapHash` runs once and lets snapshot-capable providers reuse the result. ## Stream activity `coder.Stream` returns an `*ai.StreamTextResult`. Its `Stream()` yields `provider.StreamChunk` values for the agent's text, its tool calls and their results. (`FullStream()` is a deprecated alias for the same stream, and the demo still uses it.) | Chunk type | What to read | | --- | --- | | `provider.ChunkTypeText` | `chunk.Text`, a piece of the agent's message. `ChunkTypeTextEnd` closes it. | | `provider.ChunkTypeToolCall` | `chunk.ToolCall.ToolName`, `.ID` and `.Arguments`. The agent's own tool names, such as `Bash`, `Edit` or `Read`, and Codex's equivalents. | | `provider.ChunkTypeToolResult` | `chunk.ToolResult.ToolCallID` and `.Result`, which for shell commands carries `exitCode` and `output`. | The agents name their tools differently, so classify them by name. The demo's `describeToolCall` maps a shell tool to a `run` event, edit and write tools to `edit`, read tools to `read`, and everything else to a generic line. ## Turn activity into UI data parts To show the work in a chat, fold the chunks into one state value and write it as a data part each time it changes. Use the tool call ID as the part ID so each update replaces the last. This code follows the demo, with the classification reduced to the essentials: ```go type coderEvent struct { ID string `json:"id"` Kind string `json:"kind"` // say | run | edit | read | tool Text string `json:"text"` } type coderState struct { Status string `json:"status"` // running | done | error Events []coderEvent `json:"events"` } // runCoder streams the agent's activity into the chat as a data-coder part. func runCoder(ctx context.Context, result *ai.StreamTextResult, toolCallID string, writer ai.UIMessageStreamWriter) error { state := &coderState{Status: "running", Events: []coderEvent{}} emit := func() { snapshot := *state // Copy the slice, and never send nil: JSON null breaks a client that maps over it. snapshot.Events = append([]coderEvent{}, state.Events...) writer.Write(ai.UIMessageChunk{"type": "data-coder", "id": toolCallID, "data": snapshot}) } emit() stream := result.Stream() say := -1 // index of the event that receives text deltas for { chunk, err := stream.Next() if errors.Is(err, io.EOF) { break } if err != nil { return err } switch chunk.Type { case provider.ChunkTypeText: if say < 0 { state.Events = append(state.Events, coderEvent{ID: fmt.Sprintf("say-%d", len(state.Events)), Kind: "say"}) say = len(state.Events) - 1 } state.Events[say].Text += chunk.Text case provider.ChunkTypeTextEnd: say = -1 continue case provider.ChunkTypeToolCall: say = -1 if chunk.ToolCall == nil { continue } state.Events = append(state.Events, coderEvent{ID: chunk.ToolCall.ID, Kind: "tool", Text: chunk.ToolCall.ToolName}) default: continue } emit() } if err := stream.Err(); err != nil { return err } state.Status = "done" emit() return nil } ``` Write the part from inside the tool that runs the coding agent. The demo's `delegate_to_coding_agent` tool receives the request's stream writer through a closure, which is why `s.tools(kind, writer)` is built per request. Throttle `emit` for chatty agents: the demo skips updates that arrive less than 120 ms apart unless the chunk is not text. On the client, find the part by type and ID: ```tsx const coderById = new Map(); for (const part of message.parts) { if (part.type === 'data-coder' && part.id) coderById.set(part.id, part.data); } // Then render coderById.get(toolPart.toolCallId) next to the tool call. ``` ## Read changes back When the stream ends, the agent has finished its turn. Read the result through the sandbox API on the session: ```go sb := session.GetSandboxSession() workDir := session.GetSessionWorkDir() res, err := sb.Run(ctx, providerutils.SandboxProcessOptions{ Command: "git diff --stat && git status --short", WorkingDirectory: workDir, }) content, err := sb.ReadTextFile(ctx, providerutils.SandboxReadTextFileOptions{ Path: path.Join(workDir, "shortlink.go"), }) // content is a *string, nil when the file does not exist. ``` The demo lists every file with `find`, reads each one, and keeps the ones that differ from the snapshot it took before the run. Then it applies those files to its own workspace and builds a unified diff for the chat. `Run` is also the place to execute the project's tests after the agent finishes, before you accept its work. ## Practical notes - **One writer per workspace.** The demo guards coding runs with a mutex because of the fixed bridge port and because two agents editing one workspace would conflict. - **Pass a cancellable context.** Cancelling it stops the stream and the agent. Tie it to the HTTP request so a closed tab does not leave an agent running. - **Keep the stream open.** Coding runs go quiet for long stretches. Use `KeepAliveMs` on the chat response, as described in [Serve a useChat frontend](https://goaisdk.com/docs/build-a-chat-app/serve-usechat-from-go.md#keep-alive). - **Gate the start.** Put the call that starts a coding agent behind `ToolApproval`, as in [Tool approval end to end](https://goaisdk.com/docs/build-a-chat-app/tool-approval.md). ## See it in the demo - [`server/coder.go`](https://github.com/digitallysavvy/go-ai-demo/blob/main/server/coder.go): `newHarnessAgent` (Claude Code or Codex, local sandbox with `Ports`, `OnSession`), `Run` (session, stream, `Destroy`), `activityLog` (chunks to events), `collectChanges` (reads files back) and `unifiedDiff`. - [`server/tools.go`](https://github.com/digitallysavvy/go-ai-demo/blob/main/server/tools.go): the approval-gated tool that calls `coder.Run` with the request's writer. - [`web/components/CoderCard.tsx`](https://github.com/digitallysavvy/go-ai-demo/blob/main/web/components/CoderCard.tsx): rendering the `data-coder` part as a live log with a diff. - [`web/lib/types.ts`](https://github.com/digitallysavvy/go-ai-demo/blob/main/web/lib/types.ts): the TypeScript type that mirrors `coderState`. --- # Build a chat app > Guides for building a chat product with a useChat frontend, a Go backend, approval-gated tools and coding agents, based on the Shipyard reference app. Canonical URL: https://goaisdk.com/docs/build-a-chat-app/ Documentation index: https://goaisdk.com/llms.txt Three guides cover the path from a streaming endpoint to a chat product that can change code on a user's behalf. They follow the code in the Shipyard demo, a reference app you can run and read. 1. [Serve a useChat frontend from Go](https://goaisdk.com/docs/build-a-chat-app/serve-usechat-from-go.md): the request and response protocol, validation, keep-alive, CORS, proxies and custom data parts. 2. [Tool approval end to end](https://goaisdk.com/docs/build-a-chat-app/tool-approval.md): let a model ask and a user decide, with signed approval requests. 3. [Coding agents with the harness](https://goaisdk.com/docs/build-a-chat-app/coding-agents-harness.md): run Claude Code or Codex in a sandbox and show their work in the chat. ## The reference app [Shipyard](https://github.com/digitallysavvy/go-ai-demo) is a Next.js page that uses the stock `useChat` hook, talking to a Go server built with the Go AI SDK. A failing Go project sits in a workspace. The chat model investigates with its own tools, asks the user before handing the fix to Claude Code or Codex, and streams the coding agent's activity into the thread. | Piece | Where it lives | | --- | --- | | Chat endpoint, validation, keep-alive | [`server/chat.go`](https://github.com/digitallysavvy/go-ai-demo/blob/main/server/chat.go) | | Tools and the approval-gated tool | [`server/tools.go`](https://github.com/digitallysavvy/go-ai-demo/blob/main/server/tools.go) | | Claude Code and Codex through the harness | [`server/coder.go`](https://github.com/digitallysavvy/go-ai-demo/blob/main/server/coder.go) | | `useChat` setup and approval buttons | [`web/app/page.tsx`](https://github.com/digitallysavvy/go-ai-demo/blob/main/web/app/page.tsx) | ## Short versions If you only need the code, the [recipes](https://goaisdk.com/docs/recipes.md) have one page per task, each with a program you can run. --- # Serve a useChat frontend > Stream a Go agent to the stock useChat hook with validation, keep-alive and the right headers. Canonical URL: https://goaisdk.com/docs/recipes/usechat-backend Documentation index: https://goaisdk.com/llms.txt Serve an endpoint that the `useChat` hook from `@ai-sdk/react` can talk to. ```go package main import ( "encoding/json" "log" "net/http" "os" "time" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { model, err := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}). LanguageModel(anthropic.ClaudeSonnet5_5) if err != nil { log.Fatal(err) } assistant := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a concise assistant.", }) http.HandleFunc("POST /api/chat", func(w http.ResponseWriter, r *http.Request) { var req struct { Messages json.RawMessage `json:"messages"` } if err := json.NewDecoder(http.MaxBytesReader(w, r.Body, 8<<20)).Decode(&req); err != nil { http.Error(w, "invalid request body", http.StatusBadRequest) return } // Validate and convert the UI messages before anything is written, so // a bad request gets a plain 400. chunks, _, err := agent.CreateAgentUIStreamFromUIMessages(r.Context(), assistant, agent.CreateAgentUIStreamFromUIMessagesOptions{UIMessages: []byte(req.Messages)}) if err != nil { http.Error(w, err.Error(), http.StatusBadRequest) return } // Sets the SSE headers, writes 200 and flushes each chunk. keepAlive := 15 * time.Second if err := ai.PipeUIMessageChunksToResponse(chunks, w, &ai.UIMessageStreamResponseInit{KeepAliveMs: &keepAlive}); err != nil { log.Printf("write stream: %v", err) } }) log.Println("listening on :8080") log.Fatal(http.ListenAndServe(":8080", nil)) } ``` Run it: ```bash ANTHROPIC_API_KEY=... go run ./examples/recipes/usechat-backend ``` ## Notes - `CreateAgentUIStreamFromUIMessages` validates the UI messages and runs the agent. A bad request returns an error before anything is written, so the handler answers 400. - `PipeUIMessageChunksToResponse` sets the SSE headers, writes the status and flushes every chunk. Do not set them yourself. - `KeepAliveMs` (a `*time.Duration`) keeps idle proxies from closing the stream. - On the client, use `DefaultChatTransport({ api: "http://localhost:8080/api/chat" })`. Add a CORS wrapper if the frontend is on another origin. ## Go deeper - [Guide: Serve a useChat frontend from Go](https://goaisdk.com/docs/build-a-chat-app/serve-usechat-from-go.md) - [Reference: Stream transport helpers](https://goaisdk.com/docs/reference/ai/stream-transport-helpers.md) - [Reference: Agent UI stream helpers](https://goaisdk.com/docs/reference/ai/agent-ui-stream-helpers.md) The full program is at [`examples/recipes/usechat-backend/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/usechat-backend/main.go). --- # Ask before running a tool > Pause a tool call until a user approves it, with signed approval requests and the useChat round trip. Canonical URL: https://goaisdk.com/docs/recipes/tool-approval Documentation index: https://goaisdk.com/llms.txt Let the model request a tool, but run it only after the user says yes. ```go package main import ( "context" "crypto/rand" "encoding/json" "log" "net/http" "os" "time" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/schema" ) func main() { model, err := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}). LanguageModel(anthropic.ClaudeSonnet5_5) if err != nil { log.Fatal(err) } // Signs approval requests so a client cannot forge an approval. Use a // stable secret in production, the same on every instance. secret := []byte(os.Getenv("TOOL_APPROVAL_SECRET")) if len(secret) == 0 { secret = make([]byte, 32) if _, err := rand.Read(secret); err != nil { log.Fatal(err) } } deploy := types.Tool{ Name: "deploy", Description: "Deploy the service to an environment. Requires user approval.", Parameters: schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "env": map[string]interface{}{"type": "string", "enum": []string{"staging", "prod"}}, }, "required": []string{"env"}, }), // The model can ask; only the user can say yes. ToolApproval: true, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return map[string]interface{}{"deployed": input["env"]}, nil }, } assistant := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You deploy services. Call deploy when asked. Keep replies short.", Tools: []types.Tool{deploy}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, ExperimentalToolApprovalSecret: secret, }) http.HandleFunc("POST /api/chat", func(w http.ResponseWriter, r *http.Request) { var req struct { Messages json.RawMessage `json:"messages"` } if err := json.NewDecoder(http.MaxBytesReader(w, r.Body, 8<<20)).Decode(&req); err != nil { http.Error(w, "invalid request body", http.StatusBadRequest) return } chunks, _, err := agent.CreateAgentUIStreamFromUIMessages(r.Context(), assistant, agent.CreateAgentUIStreamFromUIMessagesOptions{UIMessages: []byte(req.Messages)}) if err != nil { http.Error(w, err.Error(), http.StatusBadRequest) return } keepAlive := 15 * time.Second if err := ai.PipeUIMessageChunksToResponse(chunks, w, &ai.UIMessageStreamResponseInit{KeepAliveMs: &keepAlive}); err != nil { log.Printf("write stream: %v", err) } }) log.Println("listening on :8080") log.Fatal(http.ListenAndServe(":8080", nil)) } ``` Run it: ```bash ANTHROPIC_API_KEY=... go run ./examples/recipes/tool-approval ``` ## Notes - `ToolApproval: true` on the tool makes the agent stop and send an approval request instead of running it. - `ExperimentalToolApprovalSecret` signs each request. The server verifies the signature when the approval comes back, so a client cannot forge one. - The client needs `addToolApprovalResponse` and `sendAutomaticallyWhen: lastAssistantMessageIsCompleteWithApprovalResponses`. - To ask only sometimes, set `ToolApproval` to a `types.ToolNeedsApprovalFunc`. Values the SDK doesn't recognize fail closed: the tool reports an error instead of running. ## Go deeper - [Guide: Tool approval end to end](https://goaisdk.com/docs/build-a-chat-app/tool-approval.md) - [Reference: Tool](https://goaisdk.com/docs/reference/ai/tool.md) - [Guide: Building agents](https://goaisdk.com/docs/agents/building-agents.md) The full program is at [`examples/recipes/tool-approval/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/tool-approval/main.go). --- # Use tools from an MCP server > Connect to a Model Context Protocol server over HTTP and give its tools to a model. Canonical URL: https://goaisdk.com/docs/recipes/mcp-tools Documentation index: https://goaisdk.com/llms.txt Call the tools of a Model Context Protocol (MCP) server from `GenerateText`. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/mcp" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { ctx := context.Background() transport := mcp.NewHTTPTransport(mcp.HTTPTransportConfig{ URL: os.Getenv("MCP_URL"), TimeoutMS: 30000, }) client := mcp.NewMCPClient(transport, mcp.MCPClientConfig{ ClientName: "recipe", ClientVersion: "1.0.0", }) if err := client.Connect(ctx); err != nil { log.Fatal(err) } defer func() { _ = client.Close() }() // Every tool the server lists becomes a go-ai tool. tools, err := mcp.NewMCPToolConverter(client).ConvertToGoAITools(ctx) if err != nil { log.Fatal(err) } model, err := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}). LanguageModel(anthropic.ClaudeSonnet5_5) if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Use your tools to answer: what can you do?", Tools: tools, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` Run it: ```bash MCP_URL=https://your-server.example/mcp ANTHROPIC_API_KEY=... go run ./examples/recipes/mcp-tools ``` ## Notes - `mcp.NewHTTPTransport` and `mcp.NewMCPClient` connect to the server. Use HTTP in production. Stdio is for local servers. - `mcp.NewMCPToolConverter(client).ConvertToGoAITools` turns every listed tool into a `types.Tool`. - Set `StopWhen` so the model can call a tool and then answer. - Close the client when you are done. For OAuth, pass an `OAuth` config to the transport. ## Go deeper - [Guide: Model Context Protocol (MCP)](https://goaisdk.com/docs/ai-sdk-core/mcp-tools.md) - [Reference: Generate text](https://goaisdk.com/docs/reference/ai/generate-text.md) The full program is at [`examples/recipes/mcp-tools/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/mcp-tools/main.go). --- # Get typed data from a model > Generate a validated Go struct with GenerateText and ObjectOutput. Canonical URL: https://goaisdk.com/docs/recipes/structured-output Documentation index: https://goaisdk.com/llms.txt Get a Go struct back from a model, not free text. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // The schema comes from the struct's JSON tags. type Recipe struct { Name string `json:"name"` Ingredients []string `json:"ingredients"` Steps []string `json:"steps"` } func main() { model, err := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}). LanguageModel(openai.ModelGPT6Astra) if err != nil { log.Fatal(err) } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Give me a lasagna recipe.", Output: ai.ObjectOutput[Recipe](ai.ObjectOutputOptions{ Schema: ai.SchemaFor[Recipe](), Name: "recipe", Description: "A recipe with a name, ingredients and steps", }), }) if err != nil { log.Fatal(err) } recipe := result.Output.(Recipe) fmt.Println(recipe.Name) for i, step := range recipe.Steps { fmt.Printf("%d. %s\n", i+1, step) } } ``` Run it: ```bash OPENAI_API_KEY=... go run ./examples/recipes/structured-output ``` ## Notes - `ai.SchemaFor[T]()` builds the JSON Schema from the struct's `json` tags. - `result.Output` holds the parsed value. Type-assert it to your struct. - Other output types: `ArrayOutput[T]` for a list, `ChoiceOutput[T]` for a classification and `JSONOutput()` for untyped JSON. - `ai.GenerateObject` still works but is deprecated. Prefer `GenerateText` with `Output`. ## Go deeper - [Guide: Generating structured data](https://goaisdk.com/docs/ai-sdk-core/generating-structured-data.md) - [Reference: Generate text](https://goaisdk.com/docs/reference/ai/generate-text.md) - [Reference: Generate object](https://goaisdk.com/docs/reference/ai/generate-object.md) The full program is at [`examples/recipes/structured-output/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/structured-output/main.go). --- # Switch providers > Resolve models from provider:model strings so one program runs on Anthropic, OpenAI or Google. Canonical URL: https://goaisdk.com/docs/recipes/switch-providers Documentation index: https://goaisdk.com/llms.txt Run the same code on different providers, chosen by configuration. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/providers/google" "github.com/digitallysavvy/go-ai/pkg/providers/openai" "github.com/digitallysavvy/go-ai/pkg/registry" ) func main() { // Register each provider once, at startup. registry.RegisterProvider("anthropic", anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")})) registry.RegisterProvider("openai", openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")})) registry.RegisterProvider("google", google.New(google.Config{APIKey: os.Getenv("GOOGLE_GENERATIVE_AI_API_KEY")})) // The rest of the program only sees a "provider:model" string. id := os.Getenv("PROVIDER_MODEL") if id == "" { id = "anthropic:" + anthropic.ClaudeSonnet5_5 } model, err := registry.ResolveLanguageModel(id) if err != nil { log.Fatal(err) } result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Say hello in one short sentence.", }) if err != nil { log.Fatal(err) } fmt.Printf("%s: %s\n", id, result.Text) } ``` Run it: ```bash PROVIDER_MODEL=openai:gpt-6-astra OPENAI_API_KEY=... go run ./examples/recipes/switch-providers ``` ## Notes - Register each provider once with `registry.RegisterProvider`. - `registry.ResolveLanguageModel("provider:model")` returns a `provider.LanguageModel`. Everything after that is provider-neutral. - Read the string from an environment variable or config file to switch without a code change. - Provider-specific features go in `ProviderOptions`, keyed by provider name. Portable settings such as `Reasoning` work across providers. ## Go deeper - [Guide: Provider and model management](https://goaisdk.com/docs/ai-sdk-core/provider-management.md) - [Guide: Providers and models](https://goaisdk.com/docs/foundations/providers-and-models.md) - [Reference: Generate text](https://goaisdk.com/docs/reference/ai/generate-text.md) The full program is at [`examples/recipes/switch-providers/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/switch-providers/main.go). --- # Use Claude with reasoning > Turn on reasoning for Claude with the portable Reasoning setting and read the reasoning blocks. Canonical URL: https://goaisdk.com/docs/recipes/claude-reasoning Documentation index: https://goaisdk.com/llms.txt Have Claude think before it answers, and read what it thought. ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { model, err := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}). LanguageModel(anthropic.ClaudeSonnet5_5) if err != nil { log.Fatal(err) } // One portable setting. The SDK maps it to Claude's thinking budget. effort := types.ReasoningMedium result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "A train leaves at 3:40 and the trip takes 2 hours 35 minutes. When does it arrive?", Reasoning: &effort, }) if err != nil { log.Fatal(err) } // Reasoning blocks are content parts of the final step. if len(result.Steps) > 0 { for _, part := range result.Steps[len(result.Steps)-1].Content { if rc, ok := part.(types.ReasoningContent); ok { fmt.Println("reasoning:", rc.Text) } } } fmt.Println("answer:", result.Text) } ``` Run it: ```bash ANTHROPIC_API_KEY=... go run ./examples/recipes/claude-reasoning ``` ## Notes - `Reasoning` is a `*types.ReasoningLevel`. Pass a pointer to `types.ReasoningMedium` or another level. Nil leaves the provider default. - The SDK maps the level to a thinking budget for the model you chose. - Reasoning comes back as `types.ReasoningContent` parts in the step content, not as a top-level field. - For provider-specific control, use `anthropic.ModelOptions` with `ThinkingConfig` through `LanguageModelWithOptions`. ## Go deeper - [Guide: Reasoning](https://goaisdk.com/docs/ai-sdk-core/reasoning.md) - [Provider: Anthropic](https://goaisdk.com/docs/providers/anthropic.md) - [Reference: Generate text](https://goaisdk.com/docs/reference/ai/generate-text.md) The full program is at [`examples/recipes/claude-reasoning/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/claude-reasoning/main.go). --- # Run Claude Code or Codex > Drive Claude Code or Codex from Go with the harness, in a sandbox, and read the changed files back. Canonical URL: https://goaisdk.com/docs/recipes/harness-coding-agent Documentation index: https://goaisdk.com/llms.txt Hand a coding task to Claude Code or Codex and read what it changed. ```go package main import ( "context" "errors" "fmt" "io" "log" "os" "path" "path/filepath" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/harness" "github.com/digitallysavvy/go-ai/pkg/harness/claudecode" "github.com/digitallysavvy/go-ai/pkg/harness/sandbox/local" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/providerutils" ) const brokenFile = `package main import "fmt" func add(a, b int) int { return a - b } // bug func main() { fmt.Println(add(2, 3)) } ` func main() { ctx := context.Background() cc, err := claudecode.New(claudecode.Settings{MaxTurns: 20}) if err != nil { log.Fatal(err) } coder, err := harness.NewAgent(harness.AgentSettings{ Harness: cc, Model: anthropic.ClaudeSonnet5_5, PermissionMode: harness.PermissionModeAllowAll, Instructions: "Fix the bug with a minimal change. End with one sentence describing it.", // The local provider needs Ports: the harness bridge listens on one. Sandbox: local.NewProvider(local.Options{ RootDir: filepath.Join(os.TempDir(), "recipe-sandboxes"), Ports: []int{4319}, }), // Runs once the sandbox exists, before the agent starts. Seed files here. SandboxConfig: harness.AgentSandboxConfig{ OnSession: func(ctx context.Context, sc harness.SandboxSessionContext) error { return sc.Session.WriteTextFile(ctx, providerutils.SandboxWriteTextFileOptions{ Path: path.Join(sc.SessionWorkDir, "main.go"), Content: brokenFile, }) }, }, }) if err != nil { log.Fatal(err) } session, err := coder.CreateSession(ctx, harness.CreateSessionOptions{SessionID: "recipe-session"}) if err != nil { log.Fatal(err) } defer func() { _ = session.Destroy(context.WithoutCancel(ctx)) }() result, err := coder.Stream(ctx, agent.AgentStreamOptions{AgentGenerateOptions: agent.AgentGenerateOptions{ HarnessSession: session, Prompt: "add(2, 3) in main.go returns -1 but should return 5. Fix it.", }}) if err != nil { log.Fatal(err) } // Print the agent's activity as it happens. stream := result.Stream() for { chunk, err := stream.Next() if errors.Is(err, io.EOF) { break } if err != nil { log.Fatal(err) } switch chunk.Type { case provider.ChunkTypeText: fmt.Print(chunk.Text) case provider.ChunkTypeToolCall: if chunk.ToolCall != nil { fmt.Printf("\n[tool] %s\n", chunk.ToolCall.ToolName) } } } if err := stream.Err(); err != nil { log.Fatal(err) } // Read the result back through the sandbox API. content, err := session.GetSandboxSession().ReadTextFile(ctx, providerutils.SandboxReadTextFileOptions{ Path: path.Join(session.GetSessionWorkDir(), "main.go"), }) if err != nil { log.Fatal(err) } if content != nil { fmt.Println("\n--- main.go ---") fmt.Println(*content) } } ``` Run it: ```bash ANTHROPIC_API_KEY=... go run ./examples/recipes/harness-coding-agent ``` ## Notes - Swap `claudecode.New` for `codex.New` to use Codex. Codex supports only `PermissionModeAllowAll`. - The local sandbox needs `Ports` for the harness bridge and Node.js on the host. It is not isolated: commands run as your user. - `SandboxConfig.OnSession` seeds files before the agent starts. `session.Destroy` cleans up. - Read results through `session.GetSandboxSession()`. For isolation, use the Vercel sandbox provider. ## Go deeper - [Guide: Coding agents with the harness](https://goaisdk.com/docs/build-a-chat-app/coding-agents-harness.md) - [Reference: Harness sandbox, Vercel](https://goaisdk.com/docs/reference/ai/harness-sandbox-vercel.md) The full program is at [`examples/recipes/harness-coding-agent/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/harness-coding-agent/main.go). --- # Retrieve context with embeddings > Embed documents, rank them by cosine similarity and answer a question from the best matches. Canonical URL: https://goaisdk.com/docs/recipes/rag-embeddings Documentation index: https://goaisdk.com/llms.txt Answer questions from your own documents. ```go package main import ( "context" "fmt" "log" "os" "sort" "strings" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) var documents = []string{ "Refunds are available within 30 days of purchase with a receipt.", "Support is open Monday to Friday, 9am to 5pm Eastern.", "Premium plans include priority support and a 99.9% uptime guarantee.", } func main() { ctx := context.Background() embedder, err := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}). EmbeddingModel(openai.ModelTextEmbedding3Small) if err != nil { log.Fatal(err) } // Embed the documents once. In a real app, store these in a vector database. docs, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{Model: embedder, Inputs: documents}) if err != nil { log.Fatal(err) } question := "How long do I have to return something?" q, err := ai.Embed(ctx, ai.EmbedOptions{Model: embedder, Input: question}) if err != nil { log.Fatal(err) } // Rank the documents by cosine similarity to the question. type scored struct { text string score float64 } var ranked []scored for i, e := range docs.Embeddings { score, err := ai.CosineSimilarity(q.Embedding, e) if err != nil { log.Fatal(err) } ranked = append(ranked, scored{documents[i], score}) } sort.Slice(ranked, func(i, j int) bool { return ranked[i].score > ranked[j].score }) var contextText strings.Builder for _, r := range ranked[:2] { contextText.WriteString("- " + r.text + "\n") } model, err := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}). LanguageModel(anthropic.ClaudeSonnet5_5) if err != nil { log.Fatal(err) } result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "Answer using only the context. If the context does not say, say so.", Prompt: fmt.Sprintf("Context:\n%s\nQuestion: %s", contextText.String(), question), }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` Run it: ```bash OPENAI_API_KEY=... ANTHROPIC_API_KEY=... go run ./examples/recipes/rag-embeddings ``` ## Notes - `ai.EmbedMany` embeds the documents in one batch. Store the vectors in a database in a real app. - `ai.Embed` embeds the question. `ai.CosineSimilarity` scores each document against it. - Put the top matches in the prompt and tell the model to answer only from them. - Embedding and chat models can come from different providers, as here. ## Go deeper - [Guide: Embeddings](https://goaisdk.com/docs/ai-sdk-core/embeddings.md) - [Reference: Embed many](https://goaisdk.com/docs/reference/ai/embed-many.md) - [Reference: Cosine similarity](https://goaisdk.com/docs/reference/ai/cosine-similarity.md) The full program is at [`examples/recipes/rag-embeddings/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/rag-embeddings/main.go). --- # Handle rate limits and retries > Set retries and timeouts, and inspect the error when the SDK gives up. Canonical URL: https://goaisdk.com/docs/recipes/rate-limits-retries Documentation index: https://goaisdk.com/llms.txt Make calls resilient to 429 and 5xx responses, and fail clearly when they are not. ```go package main import ( "context" "errors" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/ai" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { model, err := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}). LanguageModel(openai.ModelGPT6Astra) if err != nil { log.Fatal(err) } // The SDK retries 408, 409, 429 and 5xx responses with exponential // backoff and honors Retry-After. The default is 2 retries. maxRetries := 4 total := 60 * time.Second result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Name three Go proverbs.", MaxRetries: &maxRetries, Timeout: &ai.TimeoutConfig{Total: &total}, }) if err != nil { // RetryError means the SDK gave up. Its last error is the cause. var retryErr *providererrors.RetryError if errors.As(err, &retryErr) { fmt.Printf("gave up after %d attempts (%s)\n", len(retryErr.Errors), retryErr.Reason) } var perr *providererrors.ProviderError if errors.As(err, &perr) && perr.StatusCode == 429 { fmt.Println("rate limited; retry-after:", perr.ResponseHeaders["retry-after"]) } log.Fatal(err) } fmt.Println(result.Text) } ``` Run it: ```bash OPENAI_API_KEY=... go run ./examples/recipes/rate-limits-retries ``` ## Notes - `MaxRetries` is a `*int`. Nil means the default of 2, and 0 turns retries off. - The SDK retries 408, 409, 429 and 5xx with exponential backoff and honors `Retry-After`. - `ai.TimeoutConfig` bounds the whole call (`Total`), each step (`PerStep`) and, for streams, each chunk (`PerChunk`). - When retries run out you get a `*providererrors.RetryError`. Use `errors.As` to reach the underlying `*providererrors.ProviderError` and its status code and headers. ## Go deeper - [Troubleshooting: Rate limit issues](https://goaisdk.com/docs/troubleshooting/rate-limits.md) - [Troubleshooting: Timeout errors](https://goaisdk.com/docs/troubleshooting/timeout-errors.md) - [Guide: Error handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) The full program is at [`examples/recipes/rate-limits-retries/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/rate-limits-retries/main.go). --- # Test code that calls a model > Test model-calling code with testutil.MockLanguageModel, with no API key and no network. Canonical URL: https://goaisdk.com/docs/recipes/mock-model-tests Documentation index: https://goaisdk.com/llms.txt Write fast, deterministic tests for code that calls a language model. ```go package main import ( "context" "fmt" "log" "os" "strings" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) // Summarize is the code under test. It takes the model as a parameter, so a // test can pass a mock. func Summarize(ctx context.Context, model provider.LanguageModel, text string) (string, error) { result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, System: "Summarize the text in one sentence.", Prompt: text, }) if err != nil { return "", err } return strings.TrimSpace(result.Text), nil } func main() { model, err := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}). LanguageModel(openai.ModelGPT6Astra) if err != nil { log.Fatal(err) } summary, err := Summarize(context.Background(), model, "Go is a statically typed, compiled language designed at Google.") if err != nil { log.Fatal(err) } fmt.Println(summary) } ``` The tests, in `main_test.go` next to it: ```go // main_test.go package main import ( "context" "errors" "strings" "testing" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/testutil" ) func TestSummarize(t *testing.T) { model := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { return &types.GenerateResult{ Text: " Go is a compiled language. ", FinishReason: types.FinishReasonStop, }, nil }, } got, err := Summarize(context.Background(), model, "long text") if err != nil { t.Fatal(err) } if got != "Go is a compiled language." { t.Errorf("got %q", got) } // The mock records every call, so you can check what the model was sent. if len(model.GenerateCalls) != 1 { t.Fatalf("calls = %d, want 1", len(model.GenerateCalls)) } if !strings.Contains(promptText(model.GenerateCalls[0]), "long text") { t.Error("prompt did not contain the input text") } } func TestSummarizeError(t *testing.T) { boom := errors.New("provider unavailable") model := &testutil.MockLanguageModel{ DoGenerateFunc: func(ctx context.Context, opts *provider.GenerateOptions) (*types.GenerateResult, error) { return nil, boom }, } if _, err := Summarize(context.Background(), model, "x"); !errors.Is(err, boom) { t.Errorf("err = %v, want %v", err, boom) } } func promptText(opts *provider.GenerateOptions) string { var b strings.Builder for _, m := range opts.Prompt.Messages { for _, p := range m.Content { if tc, ok := p.(types.TextContent); ok { b.WriteString(tc.Text) } } } return b.String() } ``` Run it: ```bash go test ./examples/recipes/mock-model-tests ``` ## Notes - Take `provider.LanguageModel` as a parameter. Production code passes a real model and tests pass a mock. - `testutil.MockLanguageModel` lets you set `DoGenerateFunc` and `DoStreamFunc`, and records every call in `GenerateCalls` and `StreamCalls`. - Return an error from the mock to test your failure paths. - Keep a few integration tests against real providers, and skip them when the API key is not set. ## Go deeper - [Guide: Testing](https://goaisdk.com/docs/ai-sdk-core/testing.md) - [Reference: Generate text](https://goaisdk.com/docs/reference/ai/generate-text.md) The full program is at [`examples/recipes/mock-model-tests/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/mock-model-tests/main.go). --- # Run a tool loop with an agent > Give an agent a tool and let it call the tool and answer, with a step limit. Canonical URL: https://goaisdk.com/docs/recipes/agent-with-tools Documentation index: https://goaisdk.com/llms.txt Let a model call your Go functions until it has an answer. ```go package main import ( "context" "fmt" "log" "os" "time" "github.com/digitallysavvy/go-ai/pkg/agent" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" "github.com/digitallysavvy/go-ai/pkg/schema" ) func main() { model, err := anthropic.New(anthropic.Config{APIKey: os.Getenv("ANTHROPIC_API_KEY")}). LanguageModel(anthropic.ClaudeSonnet5_5) if err != nil { log.Fatal(err) } clock := types.Tool{ Name: "current_time", Description: "Returns the current time in RFC 3339 format.", Parameters: schema.NewSimpleJSONSchema(map[string]interface{}{"type": "object", "properties": map[string]interface{}{}}), Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return time.Now().Format(time.RFC3339), nil }, } assistant := agent.NewToolLoopAgent(agent.AgentConfig{ Model: model, System: "You are a helpful assistant. Use tools when they help.", Tools: []types.Tool{clock}, StopWhen: []ai.StopCondition{ai.IsStepCount(5)}, }) result, err := assistant.Generate(context.Background(), agent.AgentGenerateOptions{ Prompt: "What time is it right now? Answer in one sentence.", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` Run it: ```bash ANTHROPIC_API_KEY=... go run ./examples/recipes/agent-with-tools ``` ## Notes - A `types.Tool` has a name, a description, a JSON Schema for its input and an `Execute` function. - `ToolLoopAgent` calls the model, runs any tools it asks for, and repeats until the model answers or `StopWhen` fires. - `ai.IsStepCount(5)` caps the loop. Always set a stop condition. - Use `Stream` instead of `Generate` to stream the answer. ## Go deeper - [Guide: Building agents](https://goaisdk.com/docs/agents/building-agents.md) - [Reference: ToolLoopAgent](https://goaisdk.com/docs/reference/ai/tool-loop-agent.md) - [Reference: Tool](https://goaisdk.com/docs/reference/ai/tool.md) The full program is at [`examples/recipes/agent-with-tools/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/agent-with-tools/main.go). --- # Stream text as it arrives > Read a streaming response chunk by chunk with StreamText and the Next() loop. Canonical URL: https://goaisdk.com/docs/recipes/stream-text Documentation index: https://goaisdk.com/llms.txt Print a model's answer while it is still being written. ```go package main import ( "context" "errors" "fmt" "io" "log" "os" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/openai" ) func main() { model, err := openai.New(openai.Config{APIKey: os.Getenv("OPENAI_API_KEY")}). LanguageModel(openai.ModelGPT6Astra) if err != nil { log.Fatal(err) } result, err := ai.StreamText(context.Background(), ai.StreamTextOptions{ Model: model, Prompt: "Write a haiku about concurrency.", }) if err != nil { log.Fatal(err) } stream := result.Stream() for { chunk, err := stream.Next() if errors.Is(err, io.EOF) { break } if err != nil { log.Fatal(err) } if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } } if err := stream.Err(); err != nil { log.Fatal(err) } fmt.Println() } ``` Run it: ```bash OPENAI_API_KEY=... go run ./examples/recipes/stream-text ``` ## Notes - `ai.StreamText` returns a result right away. `result.Stream()` yields chunks. - Loop on `stream.Next()` until it returns `io.EOF`, then check `stream.Err()`. - Chunks have a `Type`. Text arrives as `provider.ChunkTypeText` with the delta in `chunk.Text`. - To send the stream to a browser, use `ai.PipeUIMessageStreamToResponse` or `PipeTextStreamToResponse`. ## Go deeper - [Guide: Streaming](https://goaisdk.com/docs/foundations/streaming.md) - [Reference: Stream text](https://goaisdk.com/docs/reference/ai/stream-text.md) - [Reference: Stream transport helpers](https://goaisdk.com/docs/reference/ai/stream-transport-helpers.md) The full program is at [`examples/recipes/stream-text/main.go`](https://github.com/digitallysavvy/go-ai/blob/main/examples/recipes/stream-text/main.go). --- # Recipes > Short, complete programs for common tasks with the Go AI SDK, each with a runnable example in the repository. Canonical URL: https://goaisdk.com/docs/recipes/ Documentation index: https://goaisdk.com/llms.txt Each recipe is one task, one complete program and links to the guide and reference for more depth. Every program lives in [`examples/recipes`](https://github.com/digitallysavvy/go-ai/tree/main/examples/recipes), is built by CI, and compiles without API keys. Set the keys a recipe names to run it. | Recipe | Task | | --- | --- | | [Serve a useChat frontend](https://goaisdk.com/docs/recipes/usechat-backend.md) | Serve an endpoint that the `useChat` hook from `@ai-sdk/react` can talk to. | | [Ask before running a tool](https://goaisdk.com/docs/recipes/tool-approval.md) | Let the model request a tool, but run it only after the user says yes. | | [Use tools from an MCP server](https://goaisdk.com/docs/recipes/mcp-tools.md) | Call the tools of a Model Context Protocol (MCP) server from `GenerateText`. | | [Get typed data from a model](https://goaisdk.com/docs/recipes/structured-output.md) | Get a Go struct back from a model, not free text. | | [Switch providers](https://goaisdk.com/docs/recipes/switch-providers.md) | Run the same code on different providers, chosen by configuration. | | [Use Claude with reasoning](https://goaisdk.com/docs/recipes/claude-reasoning.md) | Have Claude think before it answers, and read what it thought. | | [Run Claude Code or Codex](https://goaisdk.com/docs/recipes/harness-coding-agent.md) | Hand a coding task to Claude Code or Codex and read what it changed. | | [Retrieve context with embeddings](https://goaisdk.com/docs/recipes/rag-embeddings.md) | Answer questions from your own documents. | | [Handle rate limits and retries](https://goaisdk.com/docs/recipes/rate-limits-retries.md) | Make calls resilient to 429 and 5xx responses, and fail clearly when they are not. | | [Test code that calls a model](https://goaisdk.com/docs/recipes/mock-model-tests.md) | Write fast, deterministic tests for code that calls a language model. | | [Run a tool loop with an agent](https://goaisdk.com/docs/recipes/agent-with-tools.md) | Let a model call your Go functions until it has an answer. | | [Stream text as it arrives](https://goaisdk.com/docs/recipes/stream-text.md) | Print a model's answer while it is still being written. | ## Reference app For a full application that combines several of these recipes, see the [Build a chat app](https://goaisdk.com/docs/build-a-chat-app.md) guides and the [Shipyard demo](https://github.com/digitallysavvy/go-ai-demo). --- # Changelog > Release history of the Go AI SDK, generated from CHANGELOG.md in the repository. Canonical URL: https://goaisdk.com/docs/changelog/ Documentation index: https://goaisdk.com/llms.txt # Changelog All notable changes to the Go AI SDK will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). ## [Unreleased] ## [0.5.1] - 2026-10-03 Fixes found while building the [Shipyard demo](https://github.com/digitallysavvy/go-ai-demo), a `useChat` frontend on a Go backend, plus a documentation overhaul for human and agent developers. ### Added - `ai.PipeUIMessageChunksToResponse` and `ai.CreateUIMessageChunksResponse` serve any UI message chunk stream over HTTP, such as one built with `CreateUIMessageStreamWithOptions` that writes data parts and merges an agent stream (TS `pipeUIMessageStreamToResponse({ stream })` / `createUIMessageStreamResponse({ stream })`). - `agent.PipeAgentUIStreamFromUIMessagesToResponse` and `agent.CreateAgentUIStreamResponseFromUIMessages` take `useChat`'s UI messages directly (TS `pipeAgentUIStreamToResponse` / `createAgentUIStreamResponse` with `uiMessages`). - `ai.UIMessageStreamHeaders()` (TS `UI_MESSAGE_STREAM_HEADERS`). - `cmd/goai-docs-mcp`: an MCP server over stdio that lets coding agents search and read the docs (`search_docs`, `read_doc`, `list_docs`). It embeds the docs, so it works offline: `go install github.com/digitallysavvy/go-ai/cmd/goai-docs-mcp@latest`. - `skills/go-ai`: an agent skill for projects that use the SDK. - Runnable `Example` functions for `GenerateText`, `StreamText`, structured output, tools, `ToolLoopAgent`, `PipeUIMessageStreamToResponse` and the MCP client, plus package overviews for the core and provider packages, on pkg.go.dev. ### Changed - On an `http.ResponseWriter`, `PipeUIMessageStreamToResponse`, `PipeUIMessageChunksToResponse`, `PipeTextStreamToResponse` and the agent `Pipe*` helpers now set the response headers and write the status before the body, like their TS counterparts on a Node `ServerResponse`. Before, they wrote only the body and the caller had to set the headers. Remove any manual header setup or `WriteHeader` call made before these helpers. Other `io.Writer`s still get only the body. New `PipeTextStreamToResponseWithInit` takes a status and headers. ### Security - Tool approval fails closed. A `ToolApproval` (or `NeedsApproval`) set to a bare function literal, such as `func(ctx context.Context, input map[string]interface{}, opts types.ToolNeedsApprovalOptions) bool`, matched no case and the tool ran without approval. Unnamed literals with the approval function signatures are now treated as their named types, and any other unrecognized value, or an unknown status string such as `"user_approval"`, is an error: the call is reported to the model as a tool error and the tool does not run. This applies to tool-level, per-tool map and call-level `ToolApproval` settings. - The harness credential setup no longer echoes an invalid base URL in its error. A base URL can carry credentials (`https://user:token@host`); the message is now `Invalid URL`, as in TS. ### Fixed - `mcp.MCPClient` now matches JSON-RPC responses to pending requests when the server echoes the request id as a JSON number. Before, responses decoded as `float64` never matched the client's `uint64` request ids, so `Connect` timed out against HTTP servers. - A step that pauses for tool approval keeps the model's finish reason (normally `tool-calls`), as in TS. It was reported as `user-approval`, which TS `useChat` rejects, so the approval step failed with a type validation error in the browser. `types.FinishReasonUserApproval` is deprecated and no longer reported. - `harness.Agent.CreateSession` names the sandbox and its work dir after the generated session ID. Sessions created without a `SessionID` all shared the work dir `-%`. - The default UI message stream headers are canonical, so `Header.Get("X-Vercel-AI-UI-Message-Stream")` finds the protocol header on responses from `CreateUIMessageStreamResponse`. - The harness examples (Claude Code, Codex, Cursor, fx, GitHub Copilot, Grok Build, workflow) give the local sandbox a port; they failed at startup with "needs a TCP port exposed by the sandbox". ### Documentation - New "Build a chat app" guides: serve `useChat` from Go, tool approval end to end, and Claude Code or Codex through the harness. 12 recipes, each backed by a program in `examples/recipes/`. - Rewritten quick start; current model IDs throughout; reference pages for MCP, the harness, workflow and UI message chunks; option tables generated from the source (`docs/scripts/gen-reference.go`); changelog page on the site. - For coding agents: `/agents.md`, `/sitemap.md`, a restructured `llms.txt`, per-section `llms.txt` files and `llms-core.txt`. - CI compiles every Go snippet in the docs (`docs/scripts/compile-snippets.go`) and checks the generated reference for drift. ## [0.5.0] - 2026-10-02 TS SDK parity target: `ai@7.0.127` (was `ai@6.0.137` in v0.4.0). Ships everything merged since the v0.4.0 tag — about 1,000 commits across the May, June, and September 2026 parity cycles. Condensed from and superseded in detail by [`release_notes/RELEASE_NOTES_V0.5.0.md`](https://github.com/digitallysavvy/go-ai/blob/main/release_notes/RELEASE_NOTES_V0.5.0.md); step-by-step upgrade instructions are in `docs/08-migration-guides/from-v0.4-to-v0.5.mdx`. ### Added - **New providers**: Voyage AI (embedding/rerank), Fish Audio, Cartesia (plus Ink 2 realtime transcription), Rev.ai, Hume, Luma, GMI Cloud, Z.AI, MiniMax, TypeSafe AI, QuiverAI, `anthropicaws` (Claude Platform on AWS), and Topaz Labs (image enhance/generation, async video). - **Catch-up to `ai@7.0.127`**: - ToolSearch `MaxResults` and custom `Search` ranking - UI message stream keepalive (`KeepAliveMs`) and `ConvertDataPart` - image-model file/mask input capabilities - speech and transcription telemetry (including streaming transcription) - GPT-6.1 Sol; Claude Sonnet 5.5 with between-tools thinking - Azure MAI-Transcribe / MAI-Voice (including streaming transcription) - Bedrock `requestMetadata` - MCP `AuthorizationServerMismatchError` and conditional token invalidation - harness `ReadHistory`, sub-agent activity events, `WorkDir: "."` and runtime-context forwarding - **Experimental surfaces**: Batch API, Evaluation, Files API v4, async video, streaming transcription/translation, and speech translation, each implemented by two or more providers. - **`pkg/codemode`** (experimental): runs model-written JavaScript in a QuickJS-on-WebAssembly sandbox, with signed continuations, interrupts, and approval flows; TypeScript annotations are stripped with Node `stripTypeScriptTypes` semantics; concurrent tool calls (`Promise.all`) that need approval are batched into one interrupt; the tool catalog lists tools in declaration order (`ToolCallerDefinition.Bind` / `PrepareModelMessage` take an ordered `[]types.Tool`). - **`pkg/harness`** (Go port of `@ai-sdk/harness`): Agent/AgentSession, `StopWhen`, tool approvals, telemetry; adapters for Claude Code, Codex, OpenCode, Deep Agents, ACP, Cursor, fx, GitHub Copilot, Grok Build; a Vercel Sandbox harness provider; `pkg/workflow` harness integration helpers. - **MCP**: full OAuth `Auth()` flow, 2026 protocol support, elicitation requests, resource-template listing, a standing inbound SSE listener for legacy servers. - **Telemetry**: `telemetry.NewOpenTelemetry`, a GenAI-semantic-convention integration; `GenerateObject`/`StreamObject` telemetry spans. - **Core**: stable lifecycle callbacks, `RepairToolCall`, `LogWarnings`, `FingerprintTools`/`DetectToolDrift`, `PrepareStep`, stream retries, `ToolSearch`/`DeferLoading`, `ExperimentalToolCallers`, `UploadFile`/`UploadSkill`, workflow model serialization for every model kind, streaming request bodies (`Request.Body` on `StreamText` steps via `provider.StreamRequestBody`), optional realtime capability interfaces (`RealtimeClientSecretCreator`, `RealtimeWebSocketConfigProvider`), and more (full list in the release notes). - **Providers**: substantial Anthropic, OpenAI Responses, xAI, Google/ Vertex, Bedrock, Gateway, and Cohere feature additions; `openai.Config.TransformRequestBody`; Groq model ID constants; see the release notes' New Features section for the per-provider breakdown. - **Docs for agents and search**: every docs page is published as markdown (append `.md` to its URL), with `llms.txt` and `llms-full.txt` indexes, Copy page / Open in ChatGPT / Open in Claude actions, per-page Open Graph images, `robots.txt`, sitemap dates and schema.org structured data; `AGENTS.md` for coding agents; the docs validator now checks YAML frontmatter; new logo. ### Changed - **`GenerateText` / `StreamText` run a single step unless you set `StopWhen`** (same as v0.4.0 and the TypeScript SDK's default `stopWhen: isStepCount(1)`). If the model calls a tool, the tool runs and its result is returned, but the model is not called again. To keep calling tools until the model answers, set a stop condition, for example `StopWhen: []ai.StopCondition{ai.IsStepCount(5)}`, or use `agent.NewToolLoopAgent` (default 20 steps). Pre-release builds of v0.5.0 briefly looped with no default limit (up to a 1,000-step safety ceiling); that regression is fixed, and the docs and examples now set `StopWhen` wherever they expect a final answer after a tool call. - **Minimum Go version is now 1.26** (`go.mod` declares `go 1.26.0`). Go 1.25 is end-of-life, and the current `golang.org/x/*` modules require Go 1.26. CI tests Go 1.26 and 1.27. - **Dependencies updated to latest**: OpenTelemetry v1.46.0, echo v4.16.0, chi v5.3.2, fiber v2.52.15, grpc v1.84.0, quic-go v0.63.0, and the `golang.org/x/*` modules. See the release notes for the full table. wazero stays at v1.9.0: v1.10+ cannot reuse a compiled module across runtimes, which the code-mode sandbox relies on, and the workaround makes each sandbox start about 5x slower. - **`StreamText` is now asynchronous**, matching TS: it returns before the first model request, and only option-validation errors return from the call itself. - **Full-stream chunk lifecycle redesigned**: `ChunkTypeFinish` now fires once per call; steps are bracketed by new `ChunkTypeStartStep` / `ChunkTypeFinishStep`. - **Runtime/tool context split**: `ExperimentalContext` → `RuntimeContext` / `ToolsContext`; tool approval is now call-level (`ToolApproval`). - **File data is a tagged union** (`types.FileData`/`FileDataType*`); system messages in `Messages` are rejected by default (`AllowSystemMessages` opts back in). - **Bedrock rebuilt on the Converse API**; **Bedrock-Anthropic rebuilt on `anthropic.LanguageModel`**; **xAI Chat Completions API removed** (use the Responses API); **Gateway `xai/*` model IDs renamed `spacexai/*`**. - **Anthropic**: `system`/user content are arrays, `BaseURL` includes `/v1`, `DisableParallelToolUse` is `*bool`, `ModelOptions` JSON tags are camelCase. - **OpenAI/Azure**: reasoning-model parameter handling changed (`max_completion_tokens`, dropped unsupported params); Azure classifies unrecognized base URLs as custom gateways. - **Google/Vertex**: provider-option key order, thinking-budget formula, Imagen removal, Interactions API wire format. - **Embed/EmbedMany/Rerank**: `MaxRetries` is now `*int`. **Schema**: `additionalProperties: false` now enforced. **Perplexity**: migrated to the Agent API. **Telemetry**: tracers belong to registered integrations; `LegacyOpenTelemetry` span shape overhauled to match TS. - **Outgoing requests now carry a `User-Agent` header** (`ai-sdk-/ go/`, the standards-compliant form TS uses, plus `ai/` from the non-streaming `pkg/ai` calls; `StreamText`, `StreamObject` and `Rerank` add no `ai/` tag), matching the TypeScript SDK. - Vendored `qjs.wasm` rebuilt from pinned upstream sources with job-queue quiescence and module-disabling patches; reproducible via `pkg/internal/third_party/qjs/build/build.sh`. - Internal refactor: the WebSocket transcription and translation streams share a session core in `pkg/providerutils/websocket` (`Session[T]`, `ReportError`, `PumpAudio`, `PumpAudioAfterReady`). No behavior change. - Full list of breaking and behavior changes: release notes' Breaking Changes and Behavior Changes sections. ### Deprecated - `ExperimentalContext`, per-tool `NeedsApproval`, `ExperimentalFilterActiveTools`, `MaxSteps`, `ExperimentalTelemetry`, `OTelTelemetryIntegration`, `ToGoogleMessages`, `StreamTextResult.FullStream()`, and several tool-execution event fields (excluded from JSON). All have stable replacements; see the release notes' Deprecations section. ### Removed - Bedrock-Anthropic's Go-only `CacheConfig` API and `PrepareTools`/`UpgradeToolVersion`/`MapToolName`/`GetBetaHeaders`/ `IsComputerUseTool`; xAI `ChatCompletionsLanguageModel()`/ `NewLanguageModel`/`SearchParameters`; Anthropic `ToAnthropicFormatWithCache`; HuggingFace's `EmbeddingModel`/`ImageModel`; non-Gemini Imagen models on Google/Vertex; `telemetry.Options.Tracer`/ `WithTracer`; Cerebras's retired model constants. ### Fixed - Anthropic multi-step tool use, streaming web tool results, mid-stream errors, and finish-reason mapping; OpenAI Responses previously-dropped output items now decoded; Bedrock rerank key, forced tool choice, and retryable stream errors; Cohere tool calling now works end-to-end; MCP HTTP transport connection leaks and content-type handling; realtime/ WebSocket connections now fail instead of finishing silently on a dropped connection; telemetry spans no longer leak on error/abort; harness Codex/host-tool/turn-release fixes; a stray leading "L" in ~88 error strings; tool-caller messages (e.g. code-mode's tool catalog) persist across steps; a stack-overflow crash in streaming providers on long runs of events with no output; Mistral thinking-mode deltas dropping their text; a concurrent map crash in the shared HTTP client when setting headers during in-flight requests; SSE lines over 64 KiB no longer abort streams (32 MiB limit); a concurrent map write crash and a late-write panic in `CreateUIMessageStreamWithOptions`; realtime session goroutine leak and double-close panic; harness `AgentSession` concurrent turn-start race and host tool executions leaked on cancel; concurrent map crashes in the agent subagent/skill registries; MCP stdio, TUI and workflow transport races and leaks; Azure system-only prompt panic; poller timeouts and cancellation; JSON numeric provider options; Vercel Sandbox `Wait` ctx handling and stream error causes; ACP now rejects a misconfigured `askUserQuestions` that returns a provider-executed tool call instead of silently accepting it; the LangChain adapter's argument fallback (it never ran); agent and workflow lifecycle events now fill in the new `ToolCall`/`ToolOutput`/`Provider`/`Instructions` fields, not only the deprecated ones; Google speech rejects out-of-range sample rates; Anthropic extended-thinking signatures were dropped from streamed responses, breaking multi-step `StreamText` with thinking and tools; the shared streaming tool-call tracker no longer aborts, corrupts, loses or misorders calls when providers send unreliable tool-call labels; UI message stream and text stream pipes now close their source when the client disconnects (no leaked provider connections); Azure's OpenAI-protocol transcription model supports streaming (`gpt-realtime-whisper`); Google Vertex embeddings honor `providerOptions.googleVertex`; harness tool-execution telemetry spans start when the tool starts. Docs pages that showed APIs that don't exist were corrected against the code. Full list in the release notes' Bug Fixes section. ### Security - `pkg/codemode`: closed a sandbox escape that let model-written JavaScript reach QuickJS's `std` / `os` modules (host files under the working directory, environment variables, `exit`, unbounded stdout). The sandbox now mounts no filesystem, passes no environment, removes the libc globals and caps console output, and the QuickJS engine refuses all module imports (no native `qjs:*` modules; the loader rejects every specifier), with a source-level `import()` check and `eval` / `Function` blocking as extra layers. - Tool approvals verified on resume (HMAC v1, TS-compatible). - Downloads: DNS pinning, synced blocklist, bounded reads, credential stripping across cross-origin redirects. MCP OAuth discovery SSRF-guarded. - Dependency bumps: `echo` v4.16.0 (includes the CVE-2026-55677 fix from v4.15.4), `chi` v5.3.2, OTel v1.46.0, `grpc` v1.84.0, `x/text` v0.42.0, `quic-go` v0.63.0 and the other `golang.org/x` modules — `govulncheck` reports no reachable vulnerabilities. - Resource names, regions and locations that would rewrite the request host are rejected (Azure, Bedrock incl. Mantle, Google Vertex incl. MaaS and Anthropic on Vertex); the TUI escapes untrusted terminal control characters; ACP host-tool execution requires a one-use authorization from a matching observed tool call. - CodeQL clean: allocation sizes are overflow-checked, JSON fragments are built with the encoder instead of string splicing, and the examples no longer log raw errors or URLs that can carry credentials. - BFL poll URLs, OpenAI image-edit URL inputs, and Anthropic batch `results_url` now fetched through the SSRF-safe download path. - Removed unused internal download helpers that skipped the SSRF checks. - Harness bridge dial errors no longer include the bridge token. - MCP OAuth and OPA policy HTTP responses are read with a 1 MiB limit. - Bedrock event-stream decoder overflow panic on crafted frames; `anthropicaws` SigV4 credential race; WebSocket dial errors no longer include query strings or userinfo. - `provider.SerializableConfig` redacts credential headers (`Authorization`, `Proxy-Authorization`, `X-Api-Key`, `Api-Key`, `X-Goog-Api-Key`, `Cookie`, `Set-Cookie`, and any `*-api-key` / `*-token` / `*secret*` name, case-insensitive) from a serialized model's `Config.Headers`, so header-based credentials no longer end up in `SerializedModel.Config`. ## [0.4.0] - 2026-03-29 TS SDK parity — fully compatible with TS AI SDK v6.0.137. Commit range `ed17fe86d..429b88a79` plus v6.0.137 backports (13 PRDs, 244 tasks). Full release notes: `release_notes/RELEASE_NOTES_V0.4.0.md`. ### Breaking Changes - **Streaming tool execution** now deferred to after stream end (not mid-stream) - **XAI** `LanguageModel()` returns Responses API model; use `ChatCompletionsLanguageModel()` for legacy - **XAI** removed `grok-2` and `grok-2-vision-1212` model IDs - **MCP** default redirect mode changed to `MCPRedirectError` (reject) ### Added #### Core SDK - **Top-level `Reasoning *types.ReasoningLevel`** on `GenerateTextOptions` and `StreamTextOptions` with 7 levels (provider-default, none, minimal, low, medium, high, xhigh) - **Deferred provider tool results** — `pendingDeferredToolCalls` tracking in generate.go and stream.go step loops; `SupportsDeferredResults bool` on `types.Tool` - **`CustomContent`** and **`ReasoningFileContent`** content types with stream chunks (`ChunkTypeCustom`, `ChunkTypeReasoningFile`) - **`TelemetryIntegration`** global registry with `sync.RWMutex`, `NoopTelemetryIntegration` default, `OTelTelemetryIntegration` wrapper; `Fire*` fan-out in generate.go/stream.go - **Tool-level timeouts** — `TimeoutConfig.ToolMs`, `TimeoutConfig.Tools`, `GetToolTimeout()` - **Embed/rerank callbacks** — `EmbedOnStartEvent`, `EmbedOnFinishEvent`, `RerankOnStartEvent`, `RerankOnFinishEvent` with provider options threading - **`StreamTextResult.ProviderMetadata()`** — accumulated from stream chunks - **`ToolResult.Input`** always populated with `call.Arguments` #### Anthropic Provider - **`WebSearch20260209(config)`** tool — allowedDomains, blockedDomains, userLocation, maxUses; returns `[]WebSearchResult20260209` with encryptedContent - **`WebFetch20260209(config)`** tool — citations, maxContentTokens; returns discriminated `WebFetchSource` (PDF/text) with `IsPDF()`/`IsPlainText()` helpers - **`EagerInputStreaming *bool`** per-tool option; `tool-input-start/delta/end` stream chunks - **Error code preservation** — `error.type` surfaces as `ProviderError.Code` - Beta header `code-execution-web-tools-2026-02-09` auto-injected for 20260209 tools #### OpenAI Provider - **GPT-5.4 model family** — `gpt-5.4`, `gpt-5.4-pro`, `gpt-5.4-mini`, `gpt-5.4-nano`, dated variants; `gpt-5.3-chat-latest` - **Responses API compaction** — parsed as `CustomContent` with `Kind: "openai-compaction"` - **`response.failed` SSE event** — maps `incomplete_details.reason` to finish reason, falls back to `error` - **`ToolSearch(args)`** factory — server (default) and client execution modes - **`CustomTool.Name` removed** — name supplied via `ToTool(name)` method - **store=false** strips unencrypted reasoning from assistant history #### XAI Provider - **Responses API as default** with `ChatCompletionsLanguageModel()` opt-in - **Grok 4.20 GA models** — multi-agent, reasoning, non-reasoning variants - Multi-image editing, b64_json output, quality/user params - `CostInUsdTicks` in image and video metadata - `ReasoningSummary` option, `ModerationError` type, `Logprobs`/`TopLogprobs` - Reasoning extraction fix (summary + content fallback) #### Google Provider - **VALIDATED** function calling mode when `tool.Strict == true` - Grounding metadata accumulation across stream chunks - Multimodal `functionResponse.parts[]` for Gemini 3+ models - `finishMessage` in Vertex provider metadata - 7 native Vertex tools (GoogleSearch, UrlContext, CodeExecution, etc.) - `gemini-embedding-2-preview` with multimodal embedding support #### Multi-Provider Reasoning - DeepSeek, Alibaba, Fireworks, Groq, Mistral, Perplexity (warning), Cohere (warning), Open Responses, XAI — all wired to top-level `Reasoning` #### KlingAI Provider - **v3.0 motion control** — multi-shot, element control, voice control, motion brush - Model IDs: `kling-v3.0-motion-control`, `kling-v2.6-motion-control` #### New Provider: Prodia - **Language model** (img2img) — multipart form-data, 11 aspect ratios - **Video model** — T2V and I2V with multipart response parsing - Shared `prodia_api.go` infrastructure #### Provider Fixes - Alibaba: single-item content array with cache_control preserves array - Perplexity: `ProviderMetadata` restructured to `{Images, Usage, Cost}` sub-objects - MCP: protocol version `2025-11-25` added; `MCPRedirectMode` typed option - MCP: OAuth state uses `crypto/subtle.ConstantTimeCompare` - HTTP transport: custom `User-Agent` header removed ### Fixed - **SSRF protection** — `validateDownloadURL()` rejects private IP redirects (IPv4, IPv6, IPv4-mapped-IPv6, localhost, link-local, CGNAT) - **Streaming tool calls** — accumulate+flush in OpenAI, Groq, DeepSeek, Alibaba; no mid-stream finalization from isParsableJson - **`isProviderExecutedTool()`** hardcoded name map deleted; replaced with `tool.ProviderExecuted` - **OpenAI wire format** for multi-turn tool calls (assistant `tool_calls` array, tool role `tool_call_id`) --- ## [0.3.0] - 2026-03-01 TypeScript AI SDK parity update — commit range `c123363c0..ed17fe86d` (124 commits, 18 PRDs, 321 tasks). Also includes the `StopWhen` / `MaxSteps` agent-loop alignment work (2026-02-16 – 2026-02-19) merged ahead of the PRD cycle, filled in below from git history. ### ⚠️ Breaking Changes - **Anthropic / Bedrock**: `output_format` renamed to `output_config.format` in request builder — matches the live Anthropic API ### Added #### New Provider - **ByteDance (Volcengine)** video generation provider (`pkg/providers/bytedance/`) with async-polling `DoGenerate`, all model ID constants, and README #### Core SDK - `Output[T]` interface and five factories: `TextOutput`, `ObjectOutput`, `ArrayOutput`, `ChoiceOutput`, `JSONOutput` - `WithOutput` parameter for `GenerateText` and `StreamText` - Six structured callback event types: `OnStartEvent`, `OnStepStartEvent`, `OnToolCallStartEvent`, `OnToolCallFinishEvent`, `OnStepFinishEvent`, `OnFinishEvent` - Panic-safe `Notify[E]` dispatch utility - Agent callback merging support - **Stop conditions**: `StopCondition` type and built-in `StepCountIs`/ `HasToolCall` conditions; `StopWhen`/`StopReason` fields on `GenerateText` and agent (`ToolLoopAgent`) options; `StopWhen` evaluation wired into both the `GenerateText` step loop and the agent loop; `MaxSteps` realigned with the Vercel AI SDK v5 `stopWhen` approach (deprecation comments removed) #### Anthropic Provider - `code-execution-20260120` tool with `programmatic-tool-call`, `bash_code_execution`, `text_editor_code_execution` input types - Automatic caching support - `claude-sonnet-4-6` model ID constant - `DisableParallelToolUse` model option - `fine-grained-tool-streaming` and `cache-control` beta header support - Streaming tool calls via `contentBlocks` state tracker in `anthropicStream` - Native MCP client: `MCPServerConfig`, `MCPToolConfiguration`, `mcp_servers` request field, `mcp_tool_use` / `mcp_tool_result` response blocks - Agent container & skills: `ContainerConfig`, `ContainerSkill`, `Container` / `ContainerID` model options, three beta headers - `StructuredOutputMode` option (`auto` / `outputFormat` / `jsonTool`) with `jsonTool` fallback for older models - `SendReasoning *bool` model option with `filterReasoningContent()` helper - `ReasoningContent.Signature` and `ReasoningContent.RedactedData` fields - `ToAnthropicMessages` handler for `ReasoningContent` (thinking block round-trip) #### OpenAI Provider - `CustomTool` with grammar and text format options - `LocalShell`, `Shell`, `ApplyPatch` container tool types - MCP approval response type - `Phase` field on Responses API message items #### Google Provider - `gemini-3.1-pro-preview` and `gemini-3.1-flash-image-preview` model constants - New image aspect ratios and sizes for Google AI and Vertex Imagen #### KlingAI Provider - `kling-v3.0-t2v` and `kling-v3.0-i2v` model ID constants #### Fireworks Provider - Async image generation for `flux-kontext-*` models with 2s polling loop #### Gateway - SSE video streaming with heartbeat / progress / complete / error event handling - `ProjectID` field and `WithProjectID()` option for request observability #### Model IDs - OpenAI: `gpt-5.3-codex` and additional model constants - XAI: resolution option, image model IDs - Bedrock: complete Anthropic model ID set - TogetherAI: `TOGETHER_API_KEY` environment variable #### Docs & Examples - Provider architecture guide - Memory management guide - Coding agents guide - `gpt-5.3-codex`, Gemini flash-image, Anthropic context editing examples ### Changed - `GenerateObject` and `StreamObject` deprecated in favor of `WithOutput` - Anthropic thinking blocks now strip `temperature`, `topP`, `topK` automatically - Streaming token counts captured from `message_start` in Anthropic provider - Alibaba cache control applied to all messages (not just first) - Cerebras: deprecated model IDs removed ### Fixed - Unknown tool name in model response produces error `ToolResult` instead of duplicate tool part - Tool choice (`required` / `auto` / specific tool) now always forwarded to provider - Stream resumption no longer flashes status to `submitted` when no active stream - `StreamingToolCallDelta.Type` is now `*string` (nullable) for OpenAI streaming - `WebSearchToolCall.Action` is now `*string` (optional) for OpenAI - OpenAI reasoning parts with `EncryptedContent` included even without `ItemID` - Bedrock and Groq now pass `strict: true` in tool definitions when strict mode set - `chatgpt-image` recognized in OpenAI response format prefix detection - `compaction_delta` null content no longer panics in Anthropic provider - `SupportsStructuredOutput()` now correctly returns `false` for pre-4.5 models --- ## [0.2.0] - 2026-02-16 Release v0.2.0: AI SDK v6. Combines the `2025-12-18` v6.0 API synchronization work and the `2026-02-15` telemetry/audio-provider work (both previously tracked under stale `[Unreleased]` headings), plus provider and platform work filled in below from git history that was never written up in either section. ### 🎉 100% Feature Parity Achieved (2026-02-15) Closed the final gap to achieve complete feature parity with the TypeScript AI SDK through telemetry integration and audio provider additions. ### Added #### Telemetry Integration - **`ExperimentalTelemetry` parameter** added to all core API functions - `GenerateText()` - Full span tracking with input/output attributes - `StreamText()` - Streaming operation telemetry - `GenerateObject()` - Object generation telemetry (all modes) - `StreamObject()` - Streaming object telemetry - `Embed()` - Single embedding telemetry - `EmbedMany()` - Batch embedding telemetry - **OpenTelemetry span tracking** with automatic instrumentation - Input/output attributes captured per operation - Usage metrics (tokens, duration) included in spans - Finish reason tracking - **Privacy controls** - `RecordInputs` - Control whether to record input data in spans - `RecordOutputs` - Control whether to record output data in spans - **MLflow integration example** updated to demonstrate telemetry usage - **Comprehensive test suite** - 5 new tests for telemetry functionality #### Audio Providers (Speech & Transcription) **Gladia Provider** (Speech-to-Text): - Full transcription model implementation - **Features:** - Multipart form upload for audio files - Word-level timestamps support - Multi-language transcription (100+ languages) - Automatic language detection - **Supported formats:** MP3, WAV, M4A, FLAC, OGG - **Model:** Whisper v3 - **Implementation:** - `pkg/providers/gladia/provider.go` - `pkg/providers/gladia/transcription_model.go` - **Tests:** 3 unit tests with mocked HTTP server (all passing) - **Documentation:** - Package README (`pkg/providers/gladia/README.md`) - Official docs (`docs/05-providers/31-gladia.mdx`) - **Examples:** - `examples/gladia-transcription/` - Basic transcription - `examples/gladia-transcription-timestamps/` - Timestamps demo **LMNT Provider** (Text-to-Speech): - Full speech synthesis model implementation - **Features:** - High-quality voice synthesis - Multiple voice options (aurora, lily, harper, sage) - Speed control (0.5x - 2.0x playback) - JSON API with clean interface - **Output format:** MP3 (audio/mpeg) - **Implementation:** - `pkg/providers/lmnt/provider.go` - `pkg/providers/lmnt/speech_model.go` - **Tests:** 3 unit tests with mocked HTTP server (all passing) - **Documentation:** - Package README (`pkg/providers/lmnt/README.md`) - Official docs (`docs/05-providers/32-lmnt.mdx`) - **Examples:** - `examples/lmnt-speech/` - Basic speech synthesis - `examples/lmnt-speech-speed/` - Speed control demo ### Documentation #### Provider Documentation - **Gladia documentation** (`docs/05-providers/31-gladia.mdx`) - Comprehensive setup and configuration guide - Available models and features - Usage examples (basic, timestamps, multi-language) - Supported audio formats reference - Error handling and best practices - Complete working examples - **LMNT documentation** (`docs/05-providers/32-lmnt.mdx`) - Comprehensive setup and configuration guide - Available voices and models - Usage examples (basic, speed control, batch generation) - Advanced features and performance tips - Error handling and best practices - Complete working examples #### Package READMEs - `pkg/providers/gladia/README.md` - Quick start guide with examples - `pkg/providers/lmnt/README.md` - Quick start guide with examples #### Example Applications - 4 new runnable example applications with README documentation - All examples compile successfully - Include setup instructions and expected output ### Testing - **11 new unit tests added** (all passing) - 5 telemetry integration tests - 3 Gladia provider tests - 3 LMNT provider tests - **Test infrastructure improvements** - Mock HTTP servers for audio provider testing - OpenTelemetry span recording for telemetry tests - Type-safe attribute comparison helpers - **No regressions** - All existing tests continue to pass ### Changed - **Telemetry system** - Now accessible through core API options - **Audio provider count** - Increased from 3 to 5 providers - Existing: ElevenLabs, Deepgram, AssemblyAI - New: Gladia, LMNT ### Provider Count Update Total provider count: **28 providers** (26 → 28) - Language models: 16 providers - Image generation: 3 providers - Speech synthesis: 3 providers (ElevenLabs, LMNT, OpenAI TTS) - Speech transcription: 4 providers (Deepgram, AssemblyAI, Gladia, OpenAI Whisper) - Embeddings: 4 providers - Reranking: 1 provider ### Migration Notes No breaking changes in this release. All updates are additive: #### Using Telemetry (Optional) ```go import "github.com/digitallysavvy/go-ai/pkg/telemetry" result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Hello", ExperimentalTelemetry: &telemetry.Settings{ IsEnabled: true, RecordInputs: true, RecordOutputs: true, }, }) ``` #### Using New Audio Providers ```go // Gladia - Speech-to-Text import "github.com/digitallysavvy/go-ai/pkg/providers/gladia" provider := gladia.New(gladia.Config{ APIKey: os.Getenv("GLADIA_API_KEY"), }) model, _ := provider.TranscriptionModel("whisper-v3") result, _ := model.DoTranscribe(ctx, &provider.TranscriptionOptions{ Audio: audioData, MimeType: "audio/mpeg", Timestamps: true, }) // LMNT - Text-to-Speech import "github.com/digitallysavvy/go-ai/pkg/providers/lmnt" provider := lmnt.New(lmnt.Config{ APIKey: os.Getenv("LMNT_API_KEY"), }) model, _ := provider.SpeechModel("default") speed := 1.2 result, _ := model.DoGenerate(ctx, &provider.SpeechGenerateOptions{ Text: "Hello world", Voice: "aurora", Speed: &speed, }) ``` ### Performance - Telemetry integration adds minimal overhead (~1-2% when enabled) - Audio providers use efficient streaming where applicable - HTTP connection pooling for audio API requests - All examples and tests complete in <5 seconds ### Quality Metrics - **Implementation time:** ~8 hours (vs 10-18 estimated) - **Test coverage:** 100% for new features - **Documentation completeness:** 100% parity with TypeScript SDK - **Breaking changes:** 0 (fully backward compatible) ### 🚀 v6.0 API Synchronization (2025-12-18) Synchronized with TypeScript AI SDK v6.0 for complete feature parity. ### 💥 Breaking Changes #### Usage Tracking API Changes - **All `Usage` fields now use pointers (`*int64`) instead of `int64`** - `InputTokens`, `OutputTokens`, `TotalTokens` are now `*int64` to properly distinguish "not set" from "zero" - **Migration**: Update comparisons like `if usage.InputTokens != 0` to `if usage.InputTokens != nil && *usage.InputTokens != 0` #### Tool Execution API Changes - **`ToolExecutor` function signature changed** - Old: `func(ctx context.Context, input map[string]interface{}) (interface{}, error)` - New: `func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error)` - Added `ToolExecutionOptions` parameter providing `ToolCallID`, `UserContext`, and `Usage` #### Callback Signature Changes - **`OnStepFinish` callback signature changed** - Old: `func(step types.StepResult)` - New: `func(ctx context.Context, step types.StepResult, userContext interface{})` - **`OnFinish` callback signature changed** (GenerateText, GenerateObject, StreamObject) - Old: `func(result *GenerateTextResult)` or `func(result *GenerateObjectResult)` - New: `func(ctx context.Context, result *GenerateTextResult, userContext interface{})` - New: `func(ctx context.Context, result *GenerateObjectResult, userContext interface{})` #### GenerateObject API Changes - **`GenerateObject` now requires explicit `Schema` parameter** - Old: `Output: &MyStruct{}` - New: `Schema: schema.NewSimpleJSONSchema(...)` - Provides better control over JSON schema validation ### Added #### Detailed Usage Tracking (v6.0) - **`InputTokenDetails`** - Breakdown of input tokens - `NoCacheTokens` - Tokens not from cache - `CacheReadTokens` - Tokens read from prompt cache (Anthropic, OpenAI, Google) - `CacheWriteTokens` - Tokens written to cache (Anthropic, Bedrock) - **`OutputTokenDetails`** - Breakdown of output tokens - `TextTokens` - Regular text generation tokens - `ReasoningTokens` - Reasoning/thinking tokens (OpenAI o1/o3, Google Gemini thinking, DeepSeek R1) - **`Usage.Raw`** - Raw provider-specific usage data for full transparency #### Enhanced Tool System (v6.0) - **New Tool fields** - `Title` - Human-readable title for better UX - `InputExamples` - Example inputs for better LLM guidance - `Strict` - Enable strict schema validation - `NeedsApproval` - Require approval before execution - `ToModelOutput` - Custom tool output formatting - `OnInputStart`, `OnInputDelta`, `OnInputAvailable` - Streaming callbacks - **`ToolExecutionOptions`** - New context for tool execution - `ToolCallID` - Unique identifier for this tool call - `UserContext` - Flow user context through tool execution - `Usage` - Accumulated token usage - `Metadata` - Additional execution metadata #### Output Objects System (v6.0) - **`ai.ObjectOutput[T](https://github.com/digitallysavvy/go-ai/blob/main/opts)`** - Type-safe object generation - **`ai.ArrayOutput[T](https://github.com/digitallysavvy/go-ai/blob/main/opts)`** - Generate arrays of elements - **`ai.ChoiceOutput[T](https://github.com/digitallysavvy/go-ai/blob/main/opts)`** - Generate enum selections - **`ai.JSONOutput(opts)`** - Flexible JSON generation - **`ai.TextOutput()`** - Plain text output (default) #### Context Flow Management (v6.0) - **`ExperimentalContext`** - Flow custom context through generation - Available in callbacks (`OnStepFinish`, `OnFinish`) - Available in tool execution (`ToolExecutionOptions.UserContext`) - Enables request-scoped data like user IDs, session info, etc. #### Provider Updates All 13 language model providers updated with v6.0 usage tracking: - **OpenAI** - Full cache + reasoning token support - **Anthropic** - Cache read + write tokens - **Google** - Cache + thoughts (reasoning) tokens - **Azure** - OpenAI-compatible with full support - **Bedrock** - Unique cache read + write pattern - **Mistral** - Simple format with input/output details - **Together AI** - OpenAI-compatible with full support - **Fireworks** - OpenAI-compatible for OSS models - **Ollama** - OpenAI-compatible for local LLMs - **xAI** - OpenAI-compatible for Grok models - **Perplexity** - OpenAI-compatible with search augmentation - **DeepSeek** - OpenAI-compatible with reasoning support (R1) - **Huggingface** - Basic support (no token counts) - **Groq** - Simple format with token details - **Cohere** - Simple format with input/output tokens - **Replicate** - Basic support (no token counts) ### Changed - **`Usage.Add(other)`** - Now properly handles pointer arithmetic and nil values - **All provider implementations** - Updated to return detailed usage breakdowns - **Test infrastructure** - Updated all tests for new Usage pointer types ### Migration Guide #### Update Usage Comparisons ```go // Before (v5.0) if result.Usage.TotalTokens > 0 { fmt.Printf("Used %d tokens\n", result.Usage.TotalTokens) } // After (v6.0) if result.Usage.TotalTokens != nil && *result.Usage.TotalTokens > 0 { fmt.Printf("Used %d tokens\n", *result.Usage.TotalTokens) } ``` #### Update Tool Definitions ```go // Before (v5.0) Execute: func(ctx context.Context, input map[string]interface{}) (interface{}, error) { return doSomething(input) } // After (v6.0) Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { fmt.Printf("Tool call ID: %s\n", opts.ToolCallID) return doSomething(input) } ``` #### Update Callbacks ```go // Before (v5.0) OnStepFinish: func(step types.StepResult) { fmt.Printf("Step %d done\n", step.StepNumber) } // After (v6.0) OnStepFinish: func(ctx context.Context, step types.StepResult, userContext interface{}) { fmt.Printf("Step %d done\n", step.StepNumber) } ``` #### Update GenerateObject Calls ```go // Before (v5.0) result, _ := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Output: &Recipe{}, }) // After (v6.0) recipeSchema := schema.NewSimpleJSONSchema(map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "name": map[string]interface{}{"type": "string"}, }, }) result, _ := ai.GenerateObject(ctx, ai.GenerateObjectOptions{ Model: model, Schema: recipeSchema, }) ``` ### Examples - Added `examples/v6_features/main.go` - Comprehensive v6.0 feature demonstration - Updated core library tests for v6.0 API - All provider examples remain compatible ### Documentation - Updated main README.md with v6.0 API examples - Added migration guide for v5.0 → v6.0 - Updated tool calling examples - Updated structured output examples ### Added (additional provider & platform work, filled in from git history) These commits shipped as part of the v0.2.0 range (`v0.1.0..v0.2.0`) but were never written up in either of the sections merged above. #### New Providers - **Alibaba** — chat provider with streaming support - **KlingAI** — video generation provider - **Moonshot** — chat provider - **OpenResponses** — chat provider - **xAI** — full provider implementation (previously listed as supported but not fully implemented), with examples and docs #### Provider Updates - **Google** — image generation support added; Google Vertex updates - **FAL / Alibaba** — improved image-to-video generation - **Fireworks AI, xAI** — updated MCP support and examples - **DeepInfra, MLflow integration** — updates #### Anthropic Provider - Advanced features, agent skills, and subagents support - Context condensing ("condense"), with updated docs and examples #### Core SDK - Token usage / tracking API updates - Session retention support - Security fixes #### Gateway - Video generation gateway support #### Integrations - LangChain callbacks now carry run IDs - LangFlow callback updates #### Development Tools - CI workflows and issue/PR templates added --- ## [0.1.0] - 2025-12-15 ### 🎉 Initial Release The first public release of the Go AI SDK - a complete rewrite of the Vercel AI SDK with full server-side feature parity. ### Added #### Core Features - `GenerateText()` - Synchronous text generation - `StreamText()` - Real-time streaming text generation with channels - `GenerateObject()` - Type-safe structured output generation - `StreamObject()` - Streaming structured output - `Embed()` - Single text embedding generation - `EmbedMany()` - Batch embedding generation - `GenerateImage()` - Text-to-image generation - `GenerateSpeech()` - Text-to-speech synthesis - `Transcribe()` - Speech-to-text transcription - `Rerank()` - Document reranking for search - `CosineSimilarity()` - Vector similarity calculations #### Provider Support (26 Providers) - OpenAI - GPT-4, GPT-3.5, O1, DALL-E, TTS, Whisper - Anthropic - Claude 3.5 Sonnet, Claude 3 family - Google - Gemini Pro, Gemini Flash - AWS Bedrock - Multi-provider access - Azure OpenAI - Enterprise deployment - Mistral - Large, Medium, Small models - Cohere - Command R+, Command R, embeddings, reranking - Groq - Ultra-fast inference (Llama, Mixtral) - xAI - Grok models - DeepSeek - DeepSeek Chat, Coder - Perplexity - Sonar models - Together AI - Open source model hosting - Fireworks AI - Fast model serving - Replicate - All hosted models - Hugging Face - Inference API - Ollama - Local model support - Stability AI - Stable Diffusion - Black Forest Labs - FLUX models - Fal.ai - Fast image generation - ElevenLabs - High-quality TTS - Deepgram - Fast STT - AssemblyAI - Advanced STT - Baseten - Model serving - Cerebras - Ultra-fast inference - DeepInfra - Model hosting - Vercel AI - Gateway integration #### Agent Framework - `agent.New()` - Create autonomous agents - `agent.Execute()` - Run multi-step workflows - Tool loop implementation for autonomous reasoning - Configurable max steps and instructions - Step-by-step execution tracking #### Tool Calling - JSON schema-based tool definitions - Function execution with parameter validation - Multi-tool support - Provider-agnostic tool calling interface #### Middleware System - `WrapLanguageModel()` - Middleware wrapper interface - Logging middleware with multiple output formats - Caching middleware with TTL and LRU eviction - Rate limiting middleware (token bucket, sliding window) - Retry middleware with exponential backoff - Telemetry middleware for observability - Composable middleware chains #### Provider Registry - String-based model resolution (e.g., "openai:gpt-4") - Provider auto-discovery - Model ID parsing and validation #### Telemetry - OpenTelemetry integration - Trace and span support - Metrics collection - Custom instrumentation #### Error Handling - `ProviderError` - Provider-specific errors with retry hints - `ValidationError` - Input validation errors - `ToolExecutionError` - Tool calling errors - `StreamError` - Streaming-specific errors - `RateLimitError` - Rate limit handling - Sentinel errors for common conditions - Structured error types with context #### Context Support - Native Go context throughout - Cancellation support - Timeout handling - Deadline propagation - Graceful shutdown ### Documentation #### Comprehensive Guides (40,000+ Lines) - Getting Started guides - Foundation concepts (providers, prompts, tools, streaming) - Complete API reference for all 12 core functions - 29 provider-specific guides with examples - Agent framework documentation - Middleware implementation guides - Telemetry and observability guides - Error handling reference (7 error types) - Migration guides (TypeScript AI SDK → Go, LangChain → Go) - Troubleshooting guides (6,396 lines) - Common errors and solutions - Rate limit handling - Debugging techniques - Context cancellation patterns ### Examples (50+ Complete Examples) #### HTTP Servers (5) - `http-server` - Standard net/http with SSE streaming - `gin-server` - Gin framework integration - `echo-server` - Echo framework patterns - `fiber-server` - Fiber web framework - `chi-server` - Chi router implementation #### Structured Output (4) - `generate-object/basic` - Type-safe generation - `generate-object/validation` - Schema validation - `generate-object/complex` - Deep nesting - `stream-object` - Real-time streaming #### Provider Features (8) - OpenAI reasoning (o1 models) - OpenAI structured outputs - OpenAI vision - Anthropic prompt caching - Anthropic extended thinking - Anthropic PDF support - Google Gemini integration - Azure OpenAI patterns #### Agents (5) - `math-agent` - Multi-tool problem solver - `web-search-agent` - Research and fact-checking - `streaming-agent` - Real-time step visualization - `multi-agent` - Coordinated systems - `supervisor-agent` - Agent orchestration #### Production Middleware (7) - Logging (console, JSON, file) - Caching (in-memory, file-based, LRU) - Rate limiting (token bucket, sliding window) - Retry with exponential backoff - Telemetry and metrics - Unit testing patterns - Integration testing patterns #### Multimodal (5) - Image generation (DALL-E, Stable Diffusion) - Text-to-speech examples - Speech-to-text transcription - Audio analysis - Vision (image understanding) #### Advanced (6) - Document reranking - Semantic routing - Throughput benchmarks - Latency benchmarks - MCP over stdio - MCP over HTTP ### Package Structure ``` pkg/ ├── ai/ Core AI SDK functions ├── agent/ Autonomous agent framework ├── provider/ Provider interfaces and types ├── providers/ 26 provider implementations ├── middleware/ Middleware system ├── registry/ Provider registry ├── schema/ JSON schema utilities ├── telemetry/ OpenTelemetry integration ├── internal/ Internal utilities └── testutil/ Testing utilities ``` ### Testing - Comprehensive unit tests - Integration tests with real providers - All examples compile and pass `go vet` - Mock providers for testing - Test utilities and helpers ### Development Tools - Contributing guidelines (CONTRIBUTING.md) - Code of conduct (CODE_OF_CONDUCT.md) - Issue templates - PR templates - Development scripts ### Quality Assurance - All code follows Go best practices - Comprehensive error handling throughout - Type-safe APIs - Production-ready patterns - Security best practices (no secrets in code) ## Feature Parity This release achieves **complete server-side parity** with the Vercel AI SDK: - ✅ Text generation (streaming and non-streaming) - ✅ Structured output (streaming and non-streaming) - ✅ Tool calling - ✅ Agent framework - ✅ Embeddings (single and batch) - ✅ Image generation - ✅ Speech synthesis and transcription - ✅ Provider registry - ✅ Middleware system - ✅ Telemetry - ✅ Error handling **Not Included:** React/UI components (client-side only in TypeScript SDK) ## Performance - Efficient streaming with automatic backpressure - Low memory overhead - Concurrent processing with goroutines - Automatic connection pooling - HTTP/2 multiplexing support - Provider-specific optimizations ## Requirements - Go 1.25 or higher - Valid API keys for desired providers ## Installation ```bash go get github.com/digitallysavvy/go-ai ``` ## License Apache 2.0 - See LICENSE for details --- [Unreleased]: https://github.com/digitallysavvy/go-ai/compare/v0.5.1...HEAD [0.5.1]: https://github.com/digitallysavvy/go-ai/compare/v0.5.0...v0.5.1 [0.5.0]: https://github.com/digitallysavvy/go-ai/compare/v0.4.0...v0.5.0 [0.4.0]: https://github.com/digitallysavvy/go-ai/compare/v0.3.0...v0.4.0 [0.3.0]: https://github.com/digitallysavvy/go-ai/compare/v0.2.0...v0.3.0 [0.2.0]: https://github.com/digitallysavvy/go-ai/compare/v0.1.0...v0.2.0 [0.1.0]: https://github.com/digitallysavvy/go-ai/releases/tag/v0.1.0 --- # Download Security and Size Limits > Explains built-in protection against memory-exhaustion attacks when the Go AI SDK downloads files from URLs, covering default limits and customization. Canonical URL: https://goaisdk.com/docs/guides/download-security Documentation index: https://goaisdk.com/llms.txt The Go AI SDK includes built-in protection against memory exhaustion attacks when downloading files from URLs. This guide explains how download security works and how to customize size limits for your use case. ## Overview When generating content with images, videos, or audio from URLs, the SDK automatically enforces size limits to prevent: - Memory exhaustion from downloading extremely large files - Out-of-Memory (OOM) process crashes - Resource depletion on your system - Denial of Service (DoS) attacks ## Default Protection **All downloads are automatically protected with a 2 GiB size limit.** This limit applies to: - Image downloads for vision models - Video downloads for video generation - Audio downloads for transcription (future) - Any user-provided URL passed to the SDK ```go // Automatically protected with 2 GiB limit result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: "analyze this image", Files: []provider.ImageFile{ {Type: "url", URL: "https://example.com/large-image.jpg"}, }, }) ``` ## Why 2 GiB? The 2 GiB default limit: - Prevents most DoS attacks - Allows legitimate large files (high-res images, long videos) - Is easy to remember and reason about - Matches the limit used by the TypeScript AI SDK ## Custom Size Limits For specific use cases, you can configure custom download limits: ### Creating a Custom Download Function ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" ) // Create single-URL download with 100 MB limit customDownload := ai.CreateURLDownloadWithMetadata(&ai.DownloadOptions{ MaxBytes: 100 * 1024 * 1024, // 100 MB }) ``` ### Using Custom Downloads with Video Generation ```go result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: ai.VideoPrompt{ Text: "generate a video from this image", Image: &ai.VideoPromptImage{ URL: "https://example.com/image.jpg", }, }, DownloadWithMetadata: customDownload, // Use custom 100 MB limit }) ``` ### Additional Options You can also set custom HTTP headers: ```go customDownload := ai.CreateURLDownloadWithMetadata(&ai.DownloadOptions{ MaxBytes: 500 * 1024 * 1024, // 500 MB Headers: map[string]string{ "Authorization": "Bearer token", "User-Agent": "MyApp/1.0", }, }) ``` ## How It Works The download protection uses multiple layers: ### 1. Content-Length Check Before downloading, the SDK checks the `Content-Length` HTTP header: ``` Content-Length: 3000000000 (3 GB) Limit: 2147483648 (2 GiB) → Rejects immediately without downloading ``` ### 2. Streaming with Size Tracking Even if `Content-Length` is missing or incorrect, the SDK tracks bytes while streaming: ```go import providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" // Reads incrementally, not all at once limitedReader := io.LimitReader(resp.Body, maxBytes+1) data, err := io.ReadAll(limitedReader) // Checks if limit was exceeded if len(data) > maxBytes { return providererrors.NewDownloadError(url, 0, "", "download exceeds the size limit", nil) } ``` ### 3. Context Cancellation Downloads respect context cancellation for timeouts: ```go ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() download := ai.CreateURLDownload(opts) data, err := download(ctx, url) // Automatically cancelled after 30 seconds ``` ## Error Handling When a download exceeds the size limit, you'll get a `DownloadError`: ```go result, err := ai.GenerateImage(ctx, opts) if err != nil { var downloadErr *providererrors.DownloadError if errors.As(err, &downloadErr) { // Download failed due to size limit or HTTP error fmt.Printf("Download failed for %s: %v\n", downloadErr.URL, err) if downloadErr.StatusCode > 0 { fmt.Printf("HTTP %d: %s\n", downloadErr.StatusCode, downloadErr.StatusText) } } } ``` ### Error Information `DownloadError` provides: - `URL`: The URL that failed - `StatusCode`: HTTP status code (if applicable) - `StatusText`: HTTP status text - `Message`: Detailed error message - `Cause`: Underlying error (if any) ## Best Practices ### 1. Use Appropriate Limits for Your Use Case ```go // Small images for thumbnails thumbnailDownload := ai.CreateURLDownloadWithMetadata(&ai.DownloadOptions{ MaxBytes: 10 * 1024 * 1024, // 10 MB }) // High-resolution images for analysis hiresDownload := ai.CreateURLDownloadWithMetadata(&ai.DownloadOptions{ MaxBytes: 500 * 1024 * 1024, // 500 MB }) // Videos (use larger limit) videoDownload := ai.CreateURLDownloadWithMetadata(&ai.DownloadOptions{ MaxBytes: 2 * 1024 * 1024 * 1024, // 2 GiB (default) }) ``` ### 2. Validate URLs Before Downloading ```go func isAllowedDomain(rawURL string) bool { parsed, err := url.Parse(rawURL) if err != nil { return false } allowedDomains := []string{"cdn.example.com", "storage.example.com"} for _, domain := range allowedDomains { if parsed.Host == domain { return true } } return false } if !isAllowedDomain(userProvidedURL) { return errors.New("URL not from allowed domain") } ``` ### 3. Set Appropriate Timeouts ```go ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second) defer cancel() result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: prompt, }) ``` ### 4. Handle Errors Gracefully ```go result, err := ai.GenerateImage(ctx, opts) if err != nil { var downloadErr *providererrors.DownloadError if errors.As(err, &downloadErr) { // Log error details log.Printf("Download failed: %v", err) // Return user-friendly error return fmt.Errorf("failed to process image: file too large or unavailable") } return err } ``` ## Security Considerations ### Public-Facing Applications If your application accepts URLs from untrusted users: 1. ✅ **Always use size limits** (enabled by default) 2. ✅ **Validate URL domains** (allowlist approach) 3. ✅ **Set timeout contexts** (prevent slowloris attacks) 4. ✅ **Rate limit requests** (prevent abuse) 5. ✅ **Monitor resource usage** (detect anomalies) ### Internal Applications Even for internal services: 1. ✅ **Keep default limits** (defense in depth) 2. ✅ **Validate URL schemes** (prevent SSRF) 3. ✅ **Log download failures** (detect issues early) ## Low-Level API `pkg/internal/fileutil` (the package underlying all of this) is a Go `internal/` package and cannot be imported outside this module. For advanced use cases, build a download function with `ai.CreateURLDownloadWithMetadata` or `ai.CreateURLDownload` instead, then call it directly: ```go import ( "github.com/digitallysavvy/go-ai/pkg/ai" ) // Configure download options download := ai.CreateURLDownloadWithMetadata(&ai.DownloadOptions{ MaxBytes: 50 * 1024 * 1024, // 50 MB Headers: map[string]string{ "User-Agent": "MyApp/1.0", }, }) // Download with options result, err := download(ctx, url) if err != nil { var downloadErr *providererrors.DownloadError if errors.As(err, &downloadErr) { fmt.Printf("Download failed: %v\n", err) } return err } data := result.Data ``` ## FAQ ### Q: Can I disable size limits? **A:** No. Size limits cannot be disabled for security reasons. The maximum you can set is the platform's `int64` limit, but this is strongly discouraged. ### Q: What if I need to download files larger than 2 GiB? **A:** Consider: 1. Downloading the file separately with your own size checking 2. Processing the file in chunks/streaming 3. Using a different approach that doesn't require loading the entire file into memory ### Q: Does this affect generated content from providers? **A:** No. Limits only apply to downloads from user-provided URLs. Content generated by AI providers is handled separately. ### Q: Will this break my existing code? **A:** No. The 2 GiB default is very generous and unlikely to affect normal use cases. If you were downloading files larger than 2 GiB before, you should explicitly configure a larger limit. ## See Also - [Security Advisory: Unbounded Download DoS](https://goaisdk.com/docs/security/ADVISORY-Download-DoS.md) - [Error Handling Guide](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) - [API Reference: CreateDownload](https://goaisdk.com/docs/reference/download-api.md#createdownload) --- # Streaming in Go AI SDK > Explains the Go AI SDK's Next()-based streaming pattern, covering chunk types, advanced patterns, error handling, and migrating from io.Reader. Canonical URL: https://goaisdk.com/docs/guides/STREAMING Documentation index: https://goaisdk.com/llms.txt The Go AI SDK uses a `Next()`-based pattern for streaming responses. This provides better type safety, error handling, and follows Go idioms. ## Basic Pattern ```go import ( "context" "fmt" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func streamExample(model provider.LanguageModel) error { // Start streaming stream, err := model.DoStream(context.Background(), &provider.GenerateOptions{ Prompt: types.Prompt{ Text: "Tell me a story", }, }) if err != nil { return err } defer stream.Close() // Process chunks using Next() for { chunk, err := stream.Next() if err != nil { // Check for end of stream if err.Error() == "EOF" { break } return err } switch chunk.Type { case provider.ChunkTypeText: fmt.Print(chunk.Text) case provider.ChunkTypeFinish: fmt.Printf("\nFinish reason: %s\n", chunk.FinishReason) fmt.Printf("Total tokens: %d\n", chunk.Usage.GetTotalTokens()) } } return nil } ``` ## Why Next() Instead of io.Reader? ### Type Safety **Next() provides typed chunks:** ```go for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } // chunk is typed as *provider.StreamChunk // No parsing or type casting needed if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } } ``` **io.Reader requires manual parsing:** ```go buf := make([]byte, 4096) for { n, err := stream.Read(buf) // Now you need to parse buf[:n] to determine chunk type // Error-prone and complex } ``` ### Error Handling **Next() provides clear error semantics:** ```go chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { // Stream completed successfully break } // Actual error occurred return err } ``` **io.Reader mixes errors with data:** ```go n, err := stream.Read(buf) if err != nil { // Did we get partial data? Should we process it? // Error handling is ambiguous } ``` ### Tool Calls and Structured Data **Next() handles complex types naturally:** ```go for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } switch chunk.Type { case provider.ChunkTypeText: fmt.Print(chunk.Text) case provider.ChunkTypeToolCall: result := executeTool(chunk.ToolCall) // Tool calls are typed structs case provider.ChunkTypeReasoning: fmt.Printf("[Thinking: %s]\n", chunk.Text) } } ``` **io.Reader would require custom serialization:** ```go // How do you represent tool calls in []byte? // Would need JSON encoding/decoding, framing, etc. ``` ## Stream Chunk Types The SDK supports several chunk types: ```go const ( ChunkTypeText ChunkType = "text" // Text content ChunkTypeToolCall ChunkType = "tool-call" // Function/tool invocation ChunkTypeToolResult ChunkType = "tool-result" // Tool execution result ChunkTypeFinish ChunkType = "finish" // Stream completion ChunkTypeError ChunkType = "error" // Error occurred ChunkTypeReasoning ChunkType = "reasoning" // Reasoning/thinking content ChunkTypeUsage ChunkType = "usage" // Token usage information ) ``` ## Advanced Patterns ### Accumulating Text ```go var fullText strings.Builder for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } if chunk.Type == provider.ChunkTypeText { fullText.WriteString(chunk.Text) fmt.Print(chunk.Text) // Also print as it arrives } } fmt.Printf("\nFull response: %s\n", fullText.String()) ``` ### Context Cancellation ```go ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() stream, err := model.DoStream(ctx, &provider.GenerateOptions{ Prompt: types.Prompt{ Text: "Long response...", }, }) if err != nil { return err } defer stream.Close() // Stream will automatically stop when context is cancelled for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } fmt.Print(chunk.Text) } ``` ### Processing Tool Calls ```go for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } switch chunk.Type { case provider.ChunkTypeText: fmt.Print(chunk.Text) case provider.ChunkTypeToolCall: // Execute the tool result, err := executeToolCall(chunk.ToolCall) if err != nil { fmt.Printf("Tool error: %v\n", err) continue } fmt.Printf("Tool %s returned: %v\n", chunk.ToolCall.ToolName, result) case provider.ChunkTypeFinish: fmt.Printf("\nCompleted: %s\n", chunk.FinishReason) } } ``` ### Handling Reasoning/Thinking Tokens Some models (like o1) emit reasoning tokens: ```go var reasoning strings.Builder var response strings.Builder for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } switch chunk.Type { case provider.ChunkTypeReasoning: reasoning.WriteString(chunk.Text) fmt.Printf("[Thinking] %s", chunk.Text) case provider.ChunkTypeText: response.WriteString(chunk.Text) fmt.Print(chunk.Text) case provider.ChunkTypeFinish: if chunk.Usage.OutputDetails != nil && chunk.Usage.OutputDetails.ReasoningTokens != nil { fmt.Printf("\nReasoning tokens: %d\n", *chunk.Usage.OutputDetails.ReasoningTokens) } } } fmt.Printf("\n\nThinking: %s\n", reasoning.String()) fmt.Printf("Response: %s\n", response.String()) ``` ### Concurrent Processing ```go stream, err := model.DoStream(ctx, options) if err != nil { return err } defer stream.Close() // Channel for processing chunks concurrently chunks := make(chan *provider.StreamChunk, 10) // Start processor goroutines var wg sync.WaitGroup for i := 0; i < 3; i++ { wg.Add(1) go func() { defer wg.Done() for chunk := range chunks { processChunk(chunk) } }() } // Read from stream and send to channel for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } chunks <- chunk } close(chunks) wg.Wait() ``` ## Error Handling Best Practices ### Always Check for EOF ```go chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { // Normal stream completion break } // Actual error return fmt.Errorf("stream error: %w", err) } ``` ### Always Close Streams ```go stream, err := model.DoStream(ctx, options) if err != nil { return err } defer stream.Close() // Always close, even on error for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err // defer will still call Close() } // Process chunk } ``` ### Check Stream.Err() ```go for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } // Process chunk } // Check for any lingering errors if err := stream.Err(); err != nil { return fmt.Errorf("stream ended with error: %w", err) } ``` ## Migration from io.Reader Pattern If you were expecting an `io.Reader` interface: **DON'T:** ```go // This pattern is not supported buf := make([]byte, 4096) n, err := stream.Read(buf) ``` **DO:** ```go // Use the Next() method instead chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } if chunk.Type == provider.ChunkTypeText { fmt.Print(chunk.Text) } ``` ## Why Not Both Patterns? Supporting both `io.Reader` and `Next()` would require: 1. **Dual maintenance**: Two streaming implementations per provider 2. **Complexity**: Converting typed chunks to/from byte streams 3. **Performance overhead**: Serialization/deserialization 4. **Type information loss**: Can't represent complex types in []byte 5. **Inconsistent experience**: Users would be confused about which to use The `Next()` pattern is superior for this use case, so that's what we support exclusively. ## Performance Considerations ### Buffering The SDK handles buffering internally. You don't need to add additional buffering unless you have specific requirements: ```go // SDK handles buffering internally - this is sufficient for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } process(chunk) } ``` ### Memory Usage Chunks are processed one at a time, keeping memory usage low: ```go // This uses minimal memory - only one chunk at a time for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } // Chunk is garbage collected after this iteration fmt.Print(chunk.Text) } ``` If you need to accumulate the entire response, use a strings.Builder: ```go var builder strings.Builder for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } if chunk.Type == provider.ChunkTypeText { builder.WriteString(chunk.Text) } } fullText := builder.String() ``` ## Provider-Specific Notes ### OpenAI - Supports all chunk types including reasoning tokens (o1 models) - Tool calls are streamed incrementally ### Anthropic - Supports thinking blocks via ChunkTypeReasoning - Provides context management information in finish chunks ### Google (Vertex AI & Generative AI) - Streams via Server-Sent Events (SSE) - Supports tool calls and multi-modal inputs - Vertex AI supports Google Cloud Storage URLs ### Bedrock - Different stream format per model provider (Claude, Titan, etc.) - SDK normalizes to consistent chunk format ## Common Patterns ### Stream to HTTP Response ```go func handleStream(w http.ResponseWriter, r *http.Request) { stream, err := model.DoStream(r.Context(), options) if err != nil { http.Error(w, err.Error(), http.StatusInternalServerError) return } defer stream.Close() // Set headers for streaming w.Header().Set("Content-Type", "text/plain; charset=utf-8") w.Header().Set("X-Content-Type-Options", "nosniff") flusher, ok := w.(http.Flusher) if !ok { http.Error(w, "Streaming not supported", http.StatusInternalServerError) return } for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } fmt.Fprintf(w, "Error: %v\n", err) return } if chunk.Type == provider.ChunkTypeText { fmt.Fprint(w, chunk.Text) flusher.Flush() } } } ``` ### Stream with Progress Updates ```go func streamWithProgress(model provider.LanguageModel) error { stream, err := model.DoStream(ctx, options) if err != nil { return err } defer stream.Close() chunkCount := 0 for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } return err } chunkCount++ switch chunk.Type { case provider.ChunkTypeText: fmt.Print(chunk.Text) case provider.ChunkTypeFinish: fmt.Printf("\n\n[Received %d chunks]\n", chunkCount) } } return nil } ``` ## Testing Streams ### Mock Stream for Testing ```go type mockStream struct { chunks []*provider.StreamChunk index int } func (m *mockStream) Next() (*provider.StreamChunk, error) { if m.index >= len(m.chunks) { return nil, io.EOF } chunk := m.chunks[m.index] m.index++ return chunk, nil } func (m *mockStream) Err() error { return nil } func (m *mockStream) Close() error { return nil } func TestStreamProcessing(t *testing.T) { stream := &mockStream{ chunks: []*provider.StreamChunk{ {Type: provider.ChunkTypeText, Text: "Hello"}, {Type: provider.ChunkTypeText, Text: " World"}, {Type: provider.ChunkTypeFinish, FinishReason: types.FinishReasonStop}, }, } var result strings.Builder for { chunk, err := stream.Next() if err != nil { if err.Error() == "EOF" { break } t.Fatal(err) } if chunk.Type == provider.ChunkTypeText { result.WriteString(chunk.Text) } } if result.String() != "Hello World" { t.Errorf("Expected 'Hello World', got '%s'", result.String()) } } ``` ## Resources - [Provider Documentation](https://github.com/digitallysavvy/go-ai/tree/main/pkg/providers) - [Language Model Interface](https://github.com/digitallysavvy/go-ai/blob/main/pkg/provider/language_model.go) - [Examples](https://github.com/digitallysavvy/go-ai/tree/main/examples) ## Summary The Go AI SDK uses a `Next()`-based streaming pattern that: - ✅ Provides strong type safety - ✅ Handles complex chunk types naturally - ✅ Follows Go idioms and conventions - ✅ Simplifies error handling - ✅ Supports tool calls and structured data - ✅ Enables concurrent processing The `io.Reader` pattern is intentionally not supported to maintain simplicity and type safety. Use `Next()` for all streaming operations. --- # Tool Reference (Anthropic) > Explains Anthropic tool references in the Go AI SDK, which let you reuse previously defined tool definitions to reduce token usage and improve performance. Canonical URL: https://goaisdk.com/docs/guides/tool-reference Documentation index: https://goaisdk.com/llms.txt Tool references are an Anthropic-specific feature that allows you to reference previously defined tools without sending the full tool definition again. This reduces token usage and improves performance in conversations where tools have already been defined. ## Overview When using features like tool search or tool discovery, instead of returning full tool definitions (which can be hundreds of tokens each), you can return tool references that simply point to tools by name. Claude will use the tools that were already defined earlier in the conversation. ## Benefits - **Reduced token usage**: Reference a tool with ~10 tokens instead of 100+ tokens for a full definition - **Better performance**: Faster API calls with smaller payloads - **Cleaner responses**: Tool search results are more concise and easier to read ## Basic Usage ### Creating a Tool Reference ```go import "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" // Create a tool reference ref := anthropic.ToolReference("calculator") ``` ### Using in Tool Results ```go func toolSearchExecute(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := input["query"].(string) // Search for matching tools matchingTools := searchTools(query) // Returns []string of tool names // Build result with tool references blocks := []types.ToolResultContentBlock{ types.TextContentBlock{Text: fmt.Sprintf("Found %d tools:", len(matchingTools))}, } for _, toolName := range matchingTools { blocks = append(blocks, anthropic.ToolReference(toolName)) } return types.ContentResult(opts.ToolCallID, "search_tools", blocks...), nil } ``` ## Complete Example Here's a complete example of a tool search system using tool references: ```go package main import ( "context" "fmt" "strings" goprovider "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) // Define your actual tools func mathTools() []types.Tool { return []types.Tool{ { Name: "add", Description: "Add two numbers", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "a": map[string]interface{}{"type": "number"}, "b": map[string]interface{}{"type": "number"}, }, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { a := input["a"].(float64) b := input["b"].(float64) return types.SimpleTextResult(opts.ToolCallID, "add", fmt.Sprintf("%.2f", a+b)), nil }, }, { Name: "multiply", Description: "Multiply two numbers", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "a": map[string]interface{}{"type": "number"}, "b": map[string]interface{}{"type": "number"}, }, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { a := input["a"].(float64) b := input["b"].(float64) return types.SimpleTextResult(opts.ToolCallID, "multiply", fmt.Sprintf("%.2f", a*b)), nil }, }, { Name: "divide", Description: "Divide two numbers", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "a": map[string]interface{}{"type": "number"}, "b": map[string]interface{}{"type": "number"}, }, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { a := input["a"].(float64) b := input["b"].(float64) if b == 0 { return types.ErrorResult(opts.ToolCallID, "divide", "Division by zero"), nil } return types.SimpleTextResult(opts.ToolCallID, "divide", fmt.Sprintf("%.2f", a/b)), nil }, }, } } // Tool search tool that returns references func searchToolsTool(availableTools []types.Tool) types.Tool { return types.Tool{ Name: "search_tools", Description: "Search for available tools by keyword", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{ "type": "string", "description": "Search query (e.g., 'math', 'calculation')", }, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := strings.ToLower(input["query"].(string)) // Search for tools matching the query var matchingTools []string for _, tool := range availableTools { if strings.Contains(strings.ToLower(tool.Name), query) || strings.Contains(strings.ToLower(tool.Description), query) { matchingTools = append(matchingTools, tool.Name) } } if len(matchingTools) == 0 { return types.ContentResult(opts.ToolCallID, "search_tools", types.TextContentBlock{Text: "No tools found matching your query."}, ), nil } // Build result with tool references blocks := []types.ToolResultContentBlock{ types.TextContentBlock{ Text: fmt.Sprintf("Found %d tool(s) matching '%s':", len(matchingTools), input["query"]), }, } // Add a tool reference for each matching tool for _, toolName := range matchingTools { blocks = append(blocks, anthropic.ToolReference(toolName)) } return types.ContentResult(opts.ToolCallID, "search_tools", blocks...), nil }, } } func main() { ctx := context.Background() // Create Anthropic provider provider := anthropic.New(anthropic.Config{ APIKey: "your-api-key", }) // Create model model, err := provider.LanguageModel("claude-sonnet-5-5") if err != nil { panic(err) } // Define all available tools (including search) mathToolsList := mathTools() allTools := append(mathToolsList, searchToolsTool(mathToolsList)) // Example conversation messages := []types.Message{ { Role: types.RoleUser, Content: []types.ContentPart{ types.TextContent{Text: "Search for math tools"}, }, }, } // Generate response result, err := model.DoGenerate(ctx, &goprovider.GenerateOptions{ Prompt: types.Prompt{Messages: messages}, Tools: allTools, }) if err != nil { panic(err) } fmt.Printf("Response: %+v\n", result) // Claude may call search_tools and see tool references // Then it can call the actual tools (add, multiply, divide) directly } ``` ## Inspecting Tool References You can detect and extract tool references from results: ### Check if a Block is a Tool Reference ```go if customBlock, ok := block.(types.CustomContentBlock); ok { if toolName, isRef := anthropic.IsToolReference(customBlock); isRef { fmt.Printf("Found tool reference: %s\n", toolName) } } ``` ### Extract All Tool References from a Result ```go result := types.ContentResult("call_123", "search_tools", types.TextContentBlock{Text: "Found tools:"}, anthropic.ToolReference("add"), anthropic.ToolReference("multiply"), ) toolNames := anthropic.ExtractToolReferences(result) // toolNames = []string{"add", "multiply"} ``` ## How It Works ### Without Tool References ```json { "role": "tool", "content": "Found tools: add (adds two numbers), multiply (multiplies two numbers), divide (divides two numbers)" } ``` Token count: ~30 tokens (plus more for full definitions if inline) ### With Tool References ```json { "role": "tool", "content": [ {"type": "text", "text": "Found 3 tools:"}, {"type": "tool_reference", "tool_name": "add"}, {"type": "tool_reference", "tool_name": "multiply"}, {"type": "tool_reference", "tool_name": "divide"} ] } ``` Token count: ~15 tokens (much more efficient!) ## Requirements 1. **Tools must be pre-defined**: The referenced tool must have been included in the `tools` parameter earlier in the conversation. 2. **Tool names must match exactly**: The tool name in the reference must exactly match the name of a defined tool. 3. **Anthropic only**: This is an Anthropic-specific feature and only works with Claude models via the Anthropic API. ## Use Cases ### 1. Tool Discovery Allow the model to search through available tools and suggest relevant ones: ```go searchToolsTool() // Returns tool references for matching tools ``` ### 2. Tool Recommendations Create a tool that recommends other tools based on the task: ```go recommendToolsTool() // Suggests tools by returning references ``` ### 3. Dynamic Tool Loading Load tools dynamically based on context and return references to available tools: ```go loadPluginTools() // Returns references to dynamically loaded tools ``` ## Best Practices 1. **Use for search/discovery**: Tool references are ideal for tool search and discovery features. 2. **Include context**: Add a text block before tool references explaining what they are. 3. **Check tool availability**: Ensure referenced tools are actually defined in the conversation. 4. **Fallback handling**: Have a fallback for when tool references aren't supported (e.g., return tool descriptions as text). ## Limitations - **Anthropic only**: Other providers don't support tool references - **Pre-defined tools only**: Can only reference tools that were already provided - **No cross-conversation**: References don't work across separate API calls ## See Also - [Tool Result Content Arrays Guide](https://goaisdk.com/docs/guides/tool-result-content-arrays.md) - General guide on content blocks - [Anthropic Advanced Features](https://goaisdk.com/docs/providers/anthropic-advanced-features.md) - Other Anthropic-specific features - [Tools Guide](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - General tool development guide --- # Tool Result Content Arrays > Explains the content array feature for tool results in the Go AI SDK, letting tools return rich structured content like text, images, and files. Canonical URL: https://goaisdk.com/docs/guides/tool-result-content-arrays Documentation index: https://goaisdk.com/llms.txt This guide explains how to use the new content array feature for tool results, which allows you to return rich, structured content from tools including text, images, files, and provider-specific content like tool references. ## Overview The content array feature provides a more flexible way to return tool results compared to simple string or JSON values. It enables: - **Multiple content blocks** in a single tool result - **Mixed content types** (text, images, files) - **Provider-specific features** like Anthropic's tool-reference - **Backward compatibility** with existing simple results ## Basic Usage ### Simple Text Result (Old Style - Still Supported) ```go // Old style - still works for simple text results result := types.SimpleTextResult("call_123", "search", "Found 3 results") ``` ### Content Blocks (New Style - Recommended) ```go // New style - supports rich content result := types.ContentResult("call_123", "search", types.TextContentBlock{Text: "Search results:"}, types.TextContentBlock{Text: "Result 1: ..."}, types.TextContentBlock{Text: "Result 2: ..."}, ) ``` ## Content Block Types ### TextContentBlock Simple text content: ```go types.TextContentBlock{ Text: "This is a text block", } ``` ### ImageContentBlock Images with base64 data: ```go imageData, _ := os.ReadFile("chart.png") // Image bytes block := types.ImageContentBlock{ Data: imageData, MediaType: "image/png", } ``` ### FileContentBlock Files like PDFs or documents: ```go pdfData, _ := os.ReadFile("report.pdf") // PDF bytes block := types.FileContentBlock{ Data: pdfData, MediaType: "application/pdf", Filename: "report.pdf", } ``` ### CustomContentBlock Provider-specific content (like tool-reference for Anthropic): ```go // Using Anthropic's ToolReference helper anthropic.ToolReference("calculator") // Or manually: types.CustomContentBlock{ ProviderOptions: map[string]interface{}{ "anthropic": map[string]interface{}{ "type": "tool-reference", "toolName": "calculator", }, }, } ``` ## Complete Examples ### Example 1: Simple Text Blocks ```go package main import ( "context" "fmt" "github.com/digitallysavvy/go-ai/pkg/provider/types" ) func searchTool() types.Tool { return types.Tool{ Name: "search", Description: "Search for information", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{ "type": "string", "description": "Search query", }, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := input["query"].(string) // Return result using content blocks return types.ContentResult(opts.ToolCallID, "search", types.TextContentBlock{Text: fmt.Sprintf("Searching for: %s", query)}, types.TextContentBlock{Text: "Found 3 results:"}, types.TextContentBlock{Text: "1. First result"}, types.TextContentBlock{Text: "2. Second result"}, types.TextContentBlock{Text: "3. Third result"}, ), nil }, } } ``` ### Example 2: Mixed Content (Text + Images) ```go func analyzeTool() types.Tool { return types.Tool{ Name: "analyze_image", Description: "Analyze an image and return results", Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { // Perform analysis... chartData := generateChart() // Returns []byte return types.ContentResult(opts.ToolCallID, "analyze_image", types.TextContentBlock{Text: "Analysis complete:"}, types.TextContentBlock{Text: "Found 5 objects in the image"}, types.ImageContentBlock{ Data: chartData, MediaType: "image/png", }, types.TextContentBlock{Text: "Confidence: 95%"}, ), nil }, } } ``` ### Example 3: Tool References (Anthropic) ```go func toolSearchTool() types.Tool { return types.Tool{ Name: "search_tools", Description: "Search for available tools", Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query := input["query"].(string) // Search for tools matching the query matchingTools := []string{"add", "subtract", "multiply", "divide"} // Build content blocks with tool references blocks := []types.ToolResultContentBlock{ types.TextContentBlock{Text: fmt.Sprintf("Found %d math tools:", len(matchingTools))}, } // Add tool reference for each matching tool for _, toolName := range matchingTools { blocks = append(blocks, anthropic.ToolReference(toolName)) } return types.ContentResult(opts.ToolCallID, "search_tools", blocks...), nil }, } } ``` ## Error Handling Use `ErrorResult` for tool execution errors: ```go func riskyTool() types.Tool { return types.Tool{ Name: "risky_operation", Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { if err := performOperation(); err != nil { return types.ErrorResult(opts.ToolCallID, "risky_operation", err.Error()), nil } return types.ContentResult(opts.ToolCallID, "risky_operation", types.TextContentBlock{Text: "Operation completed successfully"}, ), nil }, } } ``` ## Migration from Old Style ### Before (Old Style) ```go Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { result := "Found 3 results" return types.SimpleTextResult(opts.ToolCallID, "search", result), nil } ``` ### After (New Style) ```go Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { return types.ContentResult(opts.ToolCallID, "search", types.TextContentBlock{Text: "Found 3 results"}, ), nil } ``` ## Provider Compatibility | Feature | Anthropic | OpenAI | Google | |---------|-----------|--------|--------| | Text blocks | ✅ | ✅ | ✅ | | Image blocks | ✅ | ✅ | ✅ | | File blocks | ✅ | ❌ | ❌ | | Tool references | ✅ | ❌ | ❌ | - **Anthropic**: Full support for all content block types - **OpenAI**: Converts to text representation (images shown as base64) - **Google**: Converts to text representation ## Best Practices 1. **Use content blocks for rich results**: When your tool returns multiple pieces of information or mixed content types, use content blocks instead of concatenating strings. 2. **Keep text blocks focused**: Each text block should contain one logical piece of information rather than a large wall of text. 3. **Include context**: When returning images or files, include a text block before them explaining what they contain. 4. **Error handling**: Use `ErrorResult()` for errors rather than returning error text as content. 5. **Backward compatibility**: Old-style results (`SimpleTextResult`, `SimpleJSONResult`) still work if you have existing tools. ## See Also - [Tool Reference Guide](https://goaisdk.com/docs/guides/tool-reference.md) - Anthropic-specific tool-reference feature - [Tool Implementation Guide](https://goaisdk.com/docs/ai-sdk-core/tools-and-tool-calling.md) - General tool development guide - [Anthropic Provider Documentation](https://goaisdk.com/docs/providers/anthropic.md) - Anthropic-specific features --- # Alibaba Video Generation > Guide to image-to-video generation using the Alibaba DashScope provider in the Go AI SDK, covering available models, image formats, and limitations. Canonical URL: https://goaisdk.com/docs/providers/alibaba-video-generation Documentation index: https://goaisdk.com/llms.txt This guide covers video generation using the Alibaba DashScope provider in the Go AI SDK. ## Overview Alibaba provides video generation through their Wan models: - **Text-to-video** (t2v): Generate videos from text descriptions - **Image-to-video** (i2v): Animate static images - **Reference-to-video** (r2v): Generate videos using character references ## Setup ```go import ( "github.com/digitallysavvy/go-ai/pkg/providers/alibaba" ) cfg, err := alibaba.NewConfig(os.Getenv("ALIBABA_API_KEY")) if err != nil { log.Fatal(err) } prov := alibaba.New(cfg) ``` ## Available Models | Model | Type | Description | |-------|------|-------------| | `wan2.6-t2v` | Text-to-video | Generate videos from text | | `wan2.6-i2v` | Image-to-video | Animate static images (standard) | | `wan2.6-i2v-flash` | Image-to-video | Animate static images (faster) | | `wan2.6-r2v` | Reference-to-video | Generate with character references | | `wan2.6-r2v-flash` | Reference-to-video | Fast reference-based generation | ## Image-to-Video ### Using URL ```go model, err := prov.VideoModel("wan2.6-i2v-flash") if err != nil { log.Fatal(err) } duration := 6.0 result, err := model.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "Add gentle camera movement and natural motion", AspectRatio: "16:9", Duration: &duration, Image: &provider.VideoModelV3File{ Type: "url", URL: "https://example.com/landscape.jpg", }, }) ``` ### Using Local File **Important**: File uploads are only supported for Image-to-Video (i2v) models. ```go // Read image file imageData, err := os.ReadFile("landscape.png") if err != nil { log.Fatal(err) } model, err := prov.VideoModel("wan2.6-i2v-flash") if err != nil { log.Fatal(err) } duration := 6.0 result, err := model.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "Add gentle camera movement and natural motion", AspectRatio: "16:9", Duration: &duration, Image: &provider.VideoModelV3File{ Type: "file", Data: imageData, MediaType: "image/png", // Required for file uploads }, }) ``` ## Supported Image Formats For file-based uploads (i2v models only): - JPEG (`image/jpeg`) - PNG (`image/png`) - WebP (`image/webp`) ## Implementation Details When using file uploads with i2v models: - Image data is converted to raw base64 encoding - **Not** wrapped in a data URI (unlike FAL) - Base64 string sent directly in the `img_url` field - No separate upload step required ## Model-Specific Limitations | Model Type | URL Support | File Upload Support | |------------|-------------|---------------------| | t2v | N/A | N/A | | i2v | ✅ Yes | ✅ Yes | | i2v-flash | ✅ Yes | ✅ Yes | | r2v | ✅ Yes | ❌ No | | r2v-flash | ✅ Yes | ❌ No | The `Image` field is only read for i2v models; on t2v and r2v models it is silently ignored (use `InputReferences`/`FrameImages` for r2v instead). ## Parameters | Parameter | Type | Description | |-----------|------|-------------| | `Prompt` | string | Text description for video generation/animation | | `AspectRatio` | string | Video aspect ratio (e.g., "16:9", "9:16", "1:1") | | `Duration` | *float64 | Video duration in seconds (typically 5-10s) | | `Image` | *VideoModelV3File | Input image (URL or file for i2v models) | ## Provider-Specific Options Access Alibaba-specific options through `ProviderOptions`: ```go result, err := model.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "A cat walking", ProviderOptions: map[string]interface{}{ "alibaba": map[string]interface{}{ "negativePrompt": "blurry, distorted", "watermark": false, "pollIntervalMs": 2000, // Poll every 2 seconds "pollTimeoutMs": 300000, // 5 minute timeout }, }, }) ``` ## Example See complete example: `examples/providers/alibaba/09-image-to-video-file.go` ## Memory Considerations - Large images (>10MB) are loaded into memory and base64 encoded - Base64 encoding increases size by ~33% - This is acceptable for typical image sizes ## Best Practices 1. **Model Selection**: - Use `i2v-flash` for faster generation with slightly lower quality - Use `i2v` for higher quality (slower generation) 2. **File Uploads**: - Only use with i2v models - Keep images under 10MB - Provide correct `MediaType` 3. **Prompts**: - Be descriptive about desired motion/animation - Use negative prompts to avoid unwanted effects ## Troubleshooting ### `Image` field seems to have no effect - The top-level `Image` field is only applied on i2v models; t2v and r2v models ignore it without an error - Solution: use `InputReferences` (URLs, or data URIs for images) on r2v models, or switch to an i2v model ### "Invalid image data" error - Verify the image file is not corrupted - Check that the file size is reasonable (<10MB) - Ensure the MIME type matches the actual file format ### Request timeout - Reduce image file size - Use `i2v-flash` instead of `i2v` - Increase `pollTimeoutMs` in provider options ## Comparison with FAL | Aspect | Alibaba | FAL | |--------|---------|-----| | **Format** | Raw base64 | Data URI | | **Field Name** | `img_url` | `image_url` | | **Model Support** | i2v models only | All image-to-video models | | **API Efficiency** | Direct base64 | Data URI wrapper | --- # Anthropic Advanced Features > Covers advanced Claude model features in the Go AI SDK, including Fast Mode, adaptive thinking, combining features, and AWS Bedrock support. Canonical URL: https://goaisdk.com/docs/providers/anthropic-advanced-features Documentation index: https://goaisdk.com/llms.txt This guide covers advanced features available for Anthropic Claude models in the Go AI SDK. ## Fast Mode Fast mode enables 2.5x faster output token speeds for Claude Opus 4.6. This is particularly useful for applications requiring low-latency responses. ### Requirements - **Model**: `claude-opus-4-6` only - **API Version**: Latest Anthropic API ### Usage ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { // Create Anthropic provider provider := anthropic.New(anthropic.Config{ APIKey: "your-api-key", }) // Create model with fast mode enabled options := &anthropic.ModelOptions{ Speed: anthropic.SpeedFast, } model, err := provider.LanguageModelWithOptions("claude-opus-4-6", options) if err != nil { log.Fatal(err) } // Generate text with fast mode result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Quick question: What is 2+2?", }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ### Speed Options - `anthropic.SpeedFast`: Enable fast mode (2.5x faster) - `anthropic.SpeedStandard`: Standard speed (default) ### Notes - Fast mode is only available for `claude-opus-4-6` - The API automatically adds the required beta header (`fast-mode-2026-02-01`) - Fast mode may have slightly different response characteristics ## Adaptive Thinking Adaptive thinking enables Claude to show its reasoning process before providing a final answer. This feature is available in two modes: ### 1. Adaptive Thinking (Opus 4.6+) Claude dynamically adjusts reasoning effort based on task complexity. ```go options := &anthropic.ModelOptions{ Thinking: &anthropic.ThinkingConfig{ Type: anthropic.ThinkingTypeAdaptive, }, } model, err := provider.LanguageModelWithOptions("claude-opus-4-6", options) ``` ### 2. Extended Thinking (Pre-Opus 4.6) For older models, you can specify a thinking budget. ```go budget := 5000 // Token budget for thinking options := &anthropic.ModelOptions{ Thinking: &anthropic.ThinkingConfig{ Type: anthropic.ThinkingTypeEnabled, BudgetTokens: &budget, }, } model, err := provider.LanguageModelWithOptions("claude-sonnet-4", options) ``` ### Complete Example ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" ) func main() { provider := anthropic.New(anthropic.Config{ APIKey: "your-api-key", }) // Enable adaptive thinking options := &anthropic.ModelOptions{ Thinking: &anthropic.ThinkingConfig{ Type: anthropic.ThinkingTypeAdaptive, }, } model, err := provider.LanguageModelWithOptions("claude-opus-4-6", options) if err != nil { log.Fatal(err) } // Ask a complex question result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Solve this logic puzzle: Three switches control three light bulbs in another room...", }) if err != nil { log.Fatal(err) } fmt.Println("Answer:", result.Text) // Access thinking content from the result's Reasoning field for _, reasoning := range result.Reasoning { fmt.Println("\nClaude's thinking:") fmt.Println(reasoning.Text) } } ``` ### Thinking Types - `anthropic.ThinkingTypeAdaptive`: Adaptive thinking (Opus 4.6+) - `anthropic.ThinkingTypeEnabled`: Extended thinking with optional budget (older models) - `anthropic.ThinkingTypeDisabled`: Disable thinking (default) ### Budget Tokens For `ThinkingTypeEnabled`, you can optionally specify a token budget: ```go budget := 5000 config := &anthropic.ThinkingConfig{ Type: anthropic.ThinkingTypeEnabled, BudgetTokens: &budget, // Must be at least 1,024 tokens } ``` - Minimum: 1,024 tokens - Counts towards the `max_tokens` limit - Optional for `ThinkingTypeEnabled` - Not used for `ThinkingTypeAdaptive` ### Response Structure When thinking is enabled, `ai.GenerateTextResult` carries the thinking content in its `Reasoning` field (`[]types.ReasoningContent`, deprecated in favor of `FinalStep.Reasoning`): ```go type ReasoningContent struct { Text string // Thinking content Signature string // Cryptographic signature, required to replay the block // ... other fields } ``` ## Combining Features You can combine fast mode and adaptive thinking: ```go options := &anthropic.ModelOptions{ Speed: anthropic.SpeedFast, Thinking: &anthropic.ThinkingConfig{ Type: anthropic.ThinkingTypeAdaptive, }, } model, err := provider.LanguageModelWithOptions("claude-opus-4-6", options) ``` This provides: - 2.5x faster output speeds (fast mode) - Visible reasoning process (adaptive thinking) ## AWS Bedrock Support Adaptive thinking is also supported for Claude models running on AWS Bedrock: ```go import "github.com/digitallysavvy/go-ai/pkg/providers/bedrock" provider := bedrock.New(bedrock.Config{ AWSAccessKeyID: "your-access-key", AWSSecretAccessKey: "your-secret-key", Region: "us-east-1", }) options := &bedrock.ModelOptions{ Thinking: &bedrock.ThinkingConfig{ Type: bedrock.ThinkingTypeAdaptive, }, } model, err := provider.LanguageModelWithOptions( "anthropic.claude-opus-4-6-v1:0", options, ) ``` ## Error Handling The SDK handles various error conditions: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, }) if err != nil { if providererrors.IsRateLimitError(err) { // Handle rate limit } else if providererrors.IsInvalidArgumentError(err) { // Handle invalid request (e.g., wrong model for fast mode) } else { // Handle other errors } } ``` ## Best Practices ### Fast Mode 1. **Use with Opus 4.6**: Fast mode only works with `claude-opus-4-6` 2. **Latency-sensitive applications**: Ideal for real-time chat, live assistance 3. **Monitor quality**: Test that fast mode meets your quality requirements ### Adaptive Thinking 1. **Complex tasks**: Use thinking for logic puzzles, math problems, strategic planning 2. **Debugging**: The thinking content helps understand Claude's reasoning 3. **Token budget**: For enabled mode, set budget based on task complexity 4. **Reasoning field**: Access thinking content through `result.Reasoning` ### Combined Features 1. **Evaluate trade-offs**: Fast mode + thinking may produce different results 2. **Test thoroughly**: Validate that combined features work for your use case 3. **Monitor costs**: Thinking tokens count towards usage ## Troubleshooting ### Fast Mode Not Working - Verify you're using `claude-opus-4-6` - Check API key has access to beta features - Review error messages for specific issues ### Thinking Content Missing - Verify thinking is enabled in model options - Check `result.Reasoning` for thinking content - Ensure the model supports thinking (Claude 3.7 and later) ### Budget Token Issues - Budget must be at least 1,024 tokens - Budget counts towards `max_tokens` limit - Only applicable for `ThinkingTypeEnabled` ## API Reference ### Model Options ```go type ModelOptions struct { Speed Speed // Fast mode configuration Thinking *ThinkingConfig // Thinking configuration // ... other options } ``` ### Thinking Configuration ```go type ThinkingConfig struct { Type ThinkingType // Thinking mode BudgetTokens *int // Optional token budget (enabled mode only) } ``` ### Thinking Types ```go const ( ThinkingTypeAdaptive ThinkingType = "adaptive" // Opus 4.6+ ThinkingTypeEnabled ThinkingType = "enabled" // Older models ThinkingTypeDisabled ThinkingType = "disabled" // Default ) ``` ### Speed Options ```go const ( SpeedFast Speed = "fast" // 2.5x faster (Opus 4.6 only) SpeedStandard Speed = "standard" // Standard speed ) ``` ## Version Compatibility | Feature | Minimum Model | API Version | |---------|--------------|-------------| | Fast Mode | claude-opus-4-6 | 2023-06-01+ | | Adaptive Thinking | claude-opus-4-6 | 2023-06-01+ | | Extended Thinking | Claude 3.7+ | 2023-06-01+ | ## Code Execution Tool (2026-01-20) The code execution tool (`code-execution_20260120`) enables Claude to request client-side code execution across three modes: - **Python programmatic tool calls** — execute Python code that may trigger client-side tools - **Bash command execution** — run shell commands and receive stdout/stderr/exit code - **Text editor file operations** — view, create, or str_replace files The required beta header `code-execution-20260120` is **automatically injected** when this tool is included in the tool list. ### Requirements - **Models**: `claude-opus-4-6`, `claude-sonnet-4-6` - **Beta header**: injected automatically — no manual configuration needed ### Usage ```go import ( "context" "encoding/json" "fmt" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/anthropic" anthropicTools "github.com/digitallysavvy/go-ai/pkg/providers/anthropic/tools" ) // runBash and handleTextEditor are placeholders for your own // sandboxed command/file-operation implementations. func runBash(command string) (stdout, stderr string, returnCode int) { return "", "", 0 } func handleTextEditor(v *anthropicTools.TextEditorInput) interface{} { return nil } func main() { prov := anthropic.New(anthropic.Config{APIKey: "your-api-key"}) model, _ := prov.LanguageModel("claude-sonnet-4-6") // Create the tool with a client-side executor codeExecTool := anthropicTools.CodeExecution20260120() codeExecTool.Execute = func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { inputJSON, _ := json.Marshal(input) parsed, err := anthropicTools.UnmarshalCodeExecutionInput(inputJSON) if err != nil { return nil, err } switch v := parsed.(type) { case *anthropicTools.BashCodeExecutionInput: stdout, stderr, code := runBash(v.Command) return &anthropicTools.BashExecutionResult{ Type: anthropicTools.CodeExecutionResultTypeBash, Stdout: stdout, Stderr: stderr, ReturnCode: code, }, nil case *anthropicTools.TextEditorInput: // handle view/create/str_replace based on v.Command return handleTextEditor(v), nil } return nil, fmt.Errorf("unsupported input type: %T", parsed) } result, _ := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "List the files in /tmp using bash", Tools: []types.Tool{codeExecTool}, }) fmt.Println(result.Text) } ``` ### Input Types | Type constant | JSON `type` value | Fields | |---|---|---| | `CodeExecutionInputTypeProgrammatic` | `programmatic-tool-call` | `code string` | | `CodeExecutionInputTypeBash` | `bash_code_execution` | `command string` | | `CodeExecutionInputTypeTextEditor` | `text_editor_code_execution` | `command`, `path`, `file_text?`, `old_str?`, `new_str?` | Use `anthropicTools.UnmarshalCodeExecutionInput(data)` to decode the discriminated union. ### Result Types | Type constant | JSON `type` value | Description | |---|---|---| | `CodeExecutionResultTypeProgrammatic` | `code_execution_result` | Python execution result | | `CodeExecutionResultTypeBash` | `bash_code_execution_result` | Bash execution result | | `CodeExecutionResultTypeBashError` | `bash_code_execution_tool_result_error` | Bash execution error | | `CodeExecutionResultTypeTextEditorError` | `text_editor_code_execution_tool_result_error` | Text editor error | | `CodeExecutionResultTypeViewResult` | `text_editor_code_execution_view_result` | File view result | | `CodeExecutionResultTypeCreateResult` | `text_editor_code_execution_create_result` | File create/update result | | `CodeExecutionResultTypeStrReplaceResult` | `text_editor_code_execution_str_replace_result` | String replace result | Error codes: `invalid_tool_input`, `unavailable`, `too_many_requests`, `execution_time_exceeded`, `output_file_too_large` (bash), `file_not_found` (text editor). Use `anthropicTools.UnmarshalCodeExecutionResult(data)` to decode result JSON. ### Complete Examples - `examples/anthropic-code-execution-bash/` — bash execution with result handling - `examples/anthropic-code-execution-text-editor/` — text editor view/create/str_replace flow ## Further Reading - [Anthropic API Documentation](https://docs.anthropic.com/) - [Claude Model Overview](https://www.anthropic.com/claude) - [Extended Thinking Guide](https://www.anthropic.com/research/extended-thinking) --- # ByteDance Video Generation > Guide to text-to-video and image-to-video generation with the ByteDance Volcengine provider in the Go AI SDK, covering start/end frames and resolution. Canonical URL: https://goaisdk.com/docs/providers/bytedance-video-generation Documentation index: https://goaisdk.com/llms.txt This guide covers video generation using the ByteDance (Volcengine) provider in the Go AI SDK. ## Overview ByteDance's Volcengine Ark platform offers high-quality video generation models through the Seedance model family: - Text-to-video generation - Image-to-video animation - Start + end frame interpolation - Reference image guided generation ## Setup ```go import ( "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/bytedance" ) prov, err := bytedance.New(bytedance.Config{ APIKey: os.Getenv("BYTEDANCE_API_KEY"), }) model, err := prov.VideoModel(string(bytedance.ModelSeedance10Pro)) ``` ## Text-to-Video ```go dur := 5.0 response, err := model.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "A cherry blossom tree sways gently in the spring breeze", AspectRatio: "16:9", Duration: &dur, }) if err != nil { log.Fatal(err) } fmt.Println(response.Videos[0].URL) ``` ## Image-to-Video Animate a static image with an optional text prompt: ```go response, err := model.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "The cat slowly blinks and turns its head", Image: &provider.VideoModelV3File{ Type: "url", URL: "https://example.com/cat.png", }, }) ``` ## Start + End Frame Generation Generate video that transitions between two specific frames: ```go response, err := model.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "The flower blooms slowly", Image: &provider.VideoModelV3File{ Type: "url", URL: "https://example.com/bud.png", }, ProviderOptions: map[string]interface{}{ "bytedance": map[string]interface{}{ "lastFrameImage": "https://example.com/bloom.png", }, }, }) ``` ## Resolution Pass a `WxH` resolution string to control output quality. The SDK maps these to ByteDance API values: ```go response, err := model.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "Mountain sunrise", Resolution: "1920x1080", // maps to "1080p" }) ``` | Resolution | API Value | |------------|-----------| | 1920×1080, 1080×1920, etc. | `1080p` | | 1280×720, 720×1280, etc. | `720p` | | 864×480, 480×864, etc. | `480p` | ## Available Models | Constant | Model ID | |---|---| | `bytedance.ModelSeedance20` | `dreamina-seedance-2-0-260128` | | `bytedance.ModelSeedance20Fast` | `dreamina-seedance-2-0-fast-260128` | | `bytedance.ModelSeedance15Pro` | `seedance-1-5-pro-251215` | | `bytedance.ModelSeedance10Pro` | `seedance-1-0-pro-250528` | | `bytedance.ModelSeedance10ProFast` | `seedance-1-0-pro-fast-251015` | | `bytedance.ModelSeedance10LiteT2V` | `seedance-1-0-lite-t2v-250428` | | `bytedance.ModelSeedance10LiteI2V` | `seedance-1-0-lite-i2v-250428` | ## Provider Options Pass ByteDance-specific options under the `"bytedance"` key in `ProviderOptions`: ```go response, err := model.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "A futuristic city at night", ProviderOptions: map[string]interface{}{ "bytedance": map[string]interface{}{ "generateAudio": true, "watermark": false, "serviceTier": "flex", "draft": false, }, }, }) ``` | Option | Type | Description | |--------|------|-------------| | `watermark` | `bool` | Add a watermark | | `generateAudio` | `bool` | Generate audio track | | `cameraFixed` | `bool` | Fix camera position | | `returnLastFrame` | `bool` | Include last frame in response | | `serviceTier` | `string` | `"default"` or `"flex"` | | `draft` | `bool` | Draft mode (faster but lower quality) | | `lastFrameImage` | `string` | URL of the desired last frame | | `referenceImages` | `[]string` | Reference image URLs | | `pollIntervalMs` | `int` | Status polling interval when calling `model.DoGenerate` directly (default: 5000ms). Deprecated: emits a warning, and has no effect through `ai.GenerateVideo`/`ai.ExperimentalStartVideo` — use `Poll: &ai.VideoPollOptions{IntervalMs: ...}` instead. | | `pollTimeoutMs` | `int` | Max polling timeout when calling `model.DoGenerate` directly (default: 600000ms). Same deprecation as `pollIntervalMs` — use `ai.VideoPollOptions.TimeoutMs` instead. | ## Async Polling Video generation is asynchronous. The SDK handles the submit-then-poll loop automatically: 1. POST to `/contents/generations/tasks` → receive `task_id` 2. GET `/contents/generations/tasks/{task_id}` every `pollIntervalMs` ms (default 5000ms) 3. When `status == "succeeded"`, return the video URL Context cancellation stops the polling loop immediately: ```go ctx, cancel := context.WithTimeout(context.Background(), 10*time.Minute) defer cancel() response, err := model.DoGenerate(ctx, opts) ``` ## Error Handling ```go response, err := model.DoGenerate(ctx, opts) if err != nil { if bdErr, ok := err.(*bytedance.Error); ok { fmt.Printf("ByteDance error %d: %s\n", bdErr.Code, bdErr.Message) } } ``` Common error codes: | Code | Meaning | |------|---------| | 400 | Invalid request (bad prompt, unsupported option) | | 401 | Unauthorized (invalid API key) | | 504 | Polling timed out | ## Complete Example ```go package main import ( "context" "fmt" "log" "os" "github.com/digitallysavvy/go-ai/pkg/provider" "github.com/digitallysavvy/go-ai/pkg/providers/bytedance" ) func main() { prov, err := bytedance.New(bytedance.Config{ APIKey: os.Getenv("BYTEDANCE_API_KEY"), }) if err != nil { log.Fatal(err) } model, err := prov.VideoModel(string(bytedance.ModelSeedance10Pro)) if err != nil { log.Fatal(err) } ctx := context.Background() dur := 5.0 response, err := model.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "A majestic eagle soaring over snow-capped mountains", AspectRatio: "16:9", Duration: &dur, ProviderOptions: map[string]interface{}{ "bytedance": map[string]interface{}{ "generateAudio": true, }, }, }) if err != nil { log.Fatal(err) } fmt.Printf("Video URL: %s\n", response.Videos[0].URL) meta := response.ProviderMetadata["bytedance"].(map[string]interface{}) fmt.Printf("Task ID: %v\n", meta["taskId"]) } ``` ## References - [ByteDance Ark Platform](https://ark.volces.com) - [Provider package documentation](https://github.com/digitallysavvy/go-ai/blob/main/pkg/providers/bytedance/README.md) - [Example: text-to-video](https://github.com/digitallysavvy/go-ai/blob/main/examples/providers/bytedance/01-text-to-video.go) --- # FAL Video Generation > Guide to image-to-video generation using the FAL provider in the Go AI SDK, covering supported image formats, parameters, and memory considerations. Canonical URL: https://goaisdk.com/docs/providers/fal-video-generation Documentation index: https://goaisdk.com/llms.txt This guide covers video generation using the FAL provider in the Go AI SDK. ## Overview FAL provides high-quality video generation models through their API, including: - Text-to-video generation - Image-to-video animation - Various quality and speed tiers (pro, standard, turbo) ## Setup ```go import ( "github.com/digitallysavvy/go-ai/pkg/providers/fal" ) cfg := fal.Config{ APIKey: os.Getenv("FAL_API_KEY"), } prov := fal.New(cfg) model, err := prov.VideoModel("fal-ai/kling-video/v2.5-turbo/pro/image-to-video") ``` ## Image-to-Video ### Using URL ```go result, err := model.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "The cat slowly turns its head and blinks", AspectRatio: "16:9", Image: &provider.VideoModelV3File{ Type: "url", URL: "https://example.com/cat.jpg", }, }) ``` ### Using Local File ```go // Read image file imageData, err := os.ReadFile("cat.png") if err != nil { log.Fatal(err) } result, err := model.DoGenerate(ctx, &provider.VideoModelV3CallOptions{ Prompt: "The cat slowly turns its head and blinks", AspectRatio: "16:9", Image: &provider.VideoModelV3File{ Type: "file", Data: imageData, MediaType: "image/png", // Required for file uploads }, }) ``` ## Supported Image Formats For file-based uploads, FAL supports: - JPEG (`image/jpeg`) - PNG (`image/png`) - WebP (`image/webp`) - GIF (`image/gif`) ## Implementation Details When using file uploads: - Image data is converted to a base64-encoded data URI - Format: `data:;base64,` - No separate upload step required - Data is sent directly in the API request body ## Parameters | Parameter | Type | Description | |-----------|------|-------------| | `Prompt` | string | Text description of desired video animation | | `AspectRatio` | string | Video aspect ratio (e.g., "16:9", "9:16", "1:1") | | `Resolution` | string | Video resolution (e.g., "1920x1080", "1280x720") | | `Duration` | *float64 | Video duration in seconds | | `FPS` | *int | Frames per second (24, 30, 60) | | `Seed` | *int | Seed for reproducible generation | | `Image` | *VideoModelV3File | Input image (URL or file data) | ## Example See complete example: `examples/providers/fal/image-to-video-file.go` ## Memory Considerations - Large images (>10MB) are loaded into memory and base64 encoded - Base64 encoding increases size by ~33% - This is acceptable for typical image sizes but should be considered for very large images ## Best Practices 1. **File Size**: Keep images under 10MB for optimal performance 2. **Format**: Use JPEG or PNG for best compatibility 3. **MIME Type**: Always provide correct `MediaType` for file uploads 4. **Error Handling**: Handle encoding failures gracefully ## Troubleshooting ### "Invalid image data" error - Verify the image file is not corrupted - Check that the file size is reasonable (<10MB) - Ensure the MIME type matches the actual file format ### Request timeout - Reduce image file size - Try a different model tier (turbo vs pro) - Increase timeout in provider options --- # Fireworks Kimi K2.5: Extended Reasoning > Explains extended reasoning and thinking options for Fireworks AI's Kimi K2.5 model in the Go AI SDK, covering model IDs, token budgets, and history modes. Canonical URL: https://goaisdk.com/docs/providers/fireworks-kimi-k2-5 Documentation index: https://goaisdk.com/llms.txt Fireworks AI's Kimi K2.5 model supports extended reasoning and thinking capabilities, allowing the model to spend additional compute on complex problems before generating a response. ## Overview The Kimi K2.5 model provides two key features: 1. **Thinking Mode**: Enable extended reasoning with configurable token budgets 2. **Reasoning History**: Control how intermediate reasoning steps are included in responses ## Model IDs ```go // Kimi K2 variants "accounts/fireworks/models/kimi-k2-instruct" // Standard instruction following "accounts/fireworks/models/kimi-k2-thinking" // Thinking-optimized variant "accounts/fireworks/models/kimi-k2p5" // Kimi K2.5 (latest) ``` ## Basic Usage ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/fireworks" ) func main() { provider := fireworks.New(fireworks.Config{ APIKey: "your-api-key", }) model, _ := provider.LanguageModel("accounts/fireworks/models/kimi-k2p5") result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Solve this complex math problem step by step: ...", // Enable thinking with default settings ProviderOptions: map[string]interface{}{ "thinking": map[string]interface{}{ "type": "enabled", }, }, }) if err != nil { log.Fatal(err) } fmt.Println(result.Text) } ``` ## Thinking Options ### Enable/Disable Thinking ```go // Enable thinking ProviderOptions: map[string]interface{}{ "thinking": map[string]interface{}{ "type": "enabled", }, } // Disable thinking (default) ProviderOptions: map[string]interface{}{ "thinking": map[string]interface{}{ "type": "disabled", }, } ``` ### Set Token Budget Control how many tokens the model can use for reasoning: ```go ProviderOptions: map[string]interface{}{ "thinking": map[string]interface{}{ "type": "enabled", "budgetTokens": 2048, // Minimum: 1024 }, } ``` **Important**: `budgetTokens` must be at least 1024. Values passed directly through `ProviderOptions` as shown above are sent to Fireworks as-is — the SDK does **not** clamp them. If you want the 1024-token floor enforced client-side, build the option with `fireworks.WithThinkingBudget(n)` instead, which clamps any value below 1024 up to 1024: ```go import "github.com/digitallysavvy/go-ai/pkg/providers/fireworks" thinking := fireworks.WithThinkingBudget(512) // clamped to 1024 ``` ## Reasoning History Modes Control how intermediate reasoning steps appear in the response: ### Disabled (Default) Don't include reasoning history in the response: ```go ProviderOptions: map[string]interface{}{ "reasoningHistory": "disabled", } ``` ### Interleaved Mix reasoning steps with the final response: ```go ProviderOptions: map[string]interface{}{ "reasoningHistory": "interleaved", } ``` The response will contain both the reasoning process and the final answer interwoven. ### Preserved Keep reasoning separate from the final response: ```go ProviderOptions: map[string]interface{}{ "reasoningHistory": "preserved", } ``` The reasoning steps are preserved but clearly separated from the final answer. ## Complete Example ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/fireworks" ) func main() { provider := fireworks.New(fireworks.Config{ APIKey: "your-fireworks-api-key", }) model, err := provider.LanguageModel("accounts/fireworks/models/kimi-k2p5") if err != nil { log.Fatal(err) } // Complex problem requiring deep reasoning prompt := ` A train leaves Station A at 60 mph heading east. Another train leaves Station B (240 miles east of A) at 40 mph heading west. A bird starts at Station A and flies at 100 mph between the trains until they meet. How far does the bird travel? ` result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: prompt, // Enable extended reasoning with large token budget ProviderOptions: map[string]interface{}{ "thinking": map[string]interface{}{ "type": "enabled", "budgetTokens": 3072, // Allow extensive reasoning }, "reasoningHistory": "preserved", // Show the thinking process }, // Optional: Adjust temperature for reasoning Temperature: floatPtr(0.7), }) if err != nil { log.Fatal(err) } fmt.Printf("Problem: %s\n\n", prompt) fmt.Printf("Solution:\n%s\n", result.Text) // Check token usage if result.Usage.InputTokens != nil { fmt.Printf("\nInput tokens: %d\n", *result.Usage.InputTokens) } if result.Usage.OutputTokens != nil { fmt.Printf("Output tokens: %d\n", *result.Usage.OutputTokens) } } func floatPtr(f float64) *float64 { return &f } ``` ## Combined Configuration Use both thinking and reasoning history together: ```go ProviderOptions: map[string]interface{}{ // Extended reasoning "thinking": map[string]interface{}{ "type": "enabled", "budgetTokens": 2048, }, // Show the reasoning process "reasoningHistory": "preserved", } ``` ## Token Budget Guidelines Choose token budgets based on problem complexity: | Problem Type | Recommended Budget | |--------------|-------------------| | Simple calculations | 1024 (minimum) | | Medium complexity | 2048 | | Complex reasoning | 3072 - 4096 | | Very complex/multi-step | 4096+ | **Note**: Higher budgets allow more thorough reasoning but increase costs and latency. ## API Parameter Conversion The SDK automatically converts Go-style camelCase parameters to the API's snake_case format: | Go Parameter | API Parameter | |--------------|---------------| | `budgetTokens` | `budget_tokens` | | `reasoningHistory` | `reasoning_history` | You don't need to worry about this conversion - it's handled automatically. ## Cost Considerations Extended reasoning uses additional tokens: ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: "Complex problem...", ProviderOptions: map[string]interface{}{ "thinking": map[string]interface{}{ "type": "enabled", "budgetTokens": 2048, }, }, }) // Check actual token usage if result.Usage.OutputTokens != nil { fmt.Printf("Output tokens used: %d\n", *result.Usage.OutputTokens) } // The model may use up to budgetTokens additional tokens for reasoning // beyond the normal response length ``` ## Use Cases ### 1. Mathematical Problem Solving ```go ProviderOptions: map[string]interface{}{ "thinking": map[string]interface{}{ "type": "enabled", "budgetTokens": 3072, }, "reasoningHistory": "preserved", // Show work } ``` ### 2. Code Review and Analysis ```go ProviderOptions: map[string]interface{}{ "thinking": map[string]interface{}{ "type": "enabled", "budgetTokens": 4096, // Thorough analysis }, "reasoningHistory": "interleaved", // Mix analysis with findings } ``` ### 3. Strategic Planning ```go ProviderOptions: map[string]interface{}{ "thinking": map[string]interface{}{ "type": "enabled", "budgetTokens": 2048, }, "reasoningHistory": "disabled", // Only final plan } ``` ### 4. Simple Queries (No Thinking) ```go // For simple queries, disable thinking to save tokens ProviderOptions: map[string]interface{}{ "thinking": map[string]interface{}{ "type": "disabled", }, } ``` ## Best Practices 1. **Match budget to complexity**: Don't use large budgets for simple problems 2. **Monitor token usage**: Track costs, especially with large budgets 3. **Choose reasoning history mode carefully**: - Use `"preserved"` for educational/debugging purposes - Use `"disabled"` for production to reduce output tokens - Use `"interleaved"` for detailed explanations 4. **Test with different budgets**: Find the optimal budget for your use case 5. **Combine with temperature**: Lower temperature (0.5-0.7) often works well with thinking mode ## Error Handling ```go result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{ Model: model, Prompt: prompt, ProviderOptions: map[string]interface{}{ "thinking": map[string]interface{}{ "type": "enabled", "budgetTokens": 1024, }, }, }) if err != nil { // Handle errors log.Printf("Generation failed: %v", err) return } // Check for warnings if len(result.Warnings) > 0 { for _, warning := range result.Warnings { log.Printf("Warning: %s", warning.Message) } } ``` ## Comparison with Other Providers | Feature | Fireworks Kimi K2.5 | OpenAI o1 | Anthropic Claude | |---------|---------------------|-----------|------------------| | Extended reasoning | ✅ Configurable | ✅ Automatic | ❌ Not available | | Token budget control | ✅ Yes | ❌ No | N/A | | Reasoning history | ✅ 3 modes | ✅ Limited | N/A | ## See Also - [Fireworks AI Provider](https://goaisdk.com/docs/providers/fireworks.md) - [Provider Options](https://goaisdk.com/docs/foundations/provider-options.md) - Cost Optimization --- # OpenAI Responses API: Custom Tools & Shell Container Tools > Covers Custom Tool, Tool Search, and Shell Container Tool types in the OpenAI Responses API, and how to use PrepareTools with them in the Go AI SDK. Canonical URL: https://goaisdk.com/docs/providers/openai-responses-tools Documentation index: https://goaisdk.com/llms.txt This guide covers the Custom Tool, Tool Search, and Shell Container Tool types available in the OpenAI Responses API, and how to use them with the Go AI SDK. ## Overview The OpenAI Responses API supports several tool types beyond standard function calling: | Tool Type | Wire Type | Package | |--------------|-----------------|----------------------------------------------| | Function | `function` | built-in (`types.Tool`) | | Custom | `custom` | `pkg/providers/openai/tool` | | Tool Search | `tool_search` | `pkg/providers/openai/tool` | | Local Shell | `local_shell` | `pkg/providers/openai/responses` | | Shell | `shell` | `pkg/providers/openai/responses` | | Apply Patch | `apply_patch` | `pkg/providers/openai/responses` | Use `responses.PrepareTools()` to convert any combination of these into the wire format required by the Responses API. --- ## Custom Tools A custom tool constrains the model's output to a specific format — either free-form text or a grammar (regex or Lark). ### Package ```go import openaitool "github.com/digitallysavvy/go-ai/pkg/providers/openai/tool" ``` ### Types ```go type CustomToolFormat struct { Type string // "grammar" or "text" Syntax *string // "regex" or "lark" (grammar type only) Definition *string // the grammar or regex string (grammar type only) } // CustomTool does not store a Name field — the name is supplied to ToTool("name"). type CustomTool struct { Description *string Format *CustomToolFormat // Async, when true, lets the model continue generating after calling // this tool without waiting for its result (GPT-6+ only). Async *bool } ``` ### Factory Function ```go func NewCustomTool(opts ...CustomToolOption) CustomTool ``` Available options: ```go // Package openaitool func WithDescription(desc string) CustomToolOption func WithFormat(format CustomToolFormat) CustomToolOption func WithAsync(async bool) CustomToolOption ``` The tool name is **not** stored in `CustomTool`. Supply it when calling `ToTool("name")` so the name is derived from the caller's context (e.g., the key in a tools map), matching the TypeScript SDK convention. ### Usage ```go // Custom tool with Lark grammar syntax := "lark" definition := `start: OBJECT` ct := openaitool.NewCustomTool( openaitool.WithDescription("Extract JSON from the provided text"), openaitool.WithFormat(openaitool.CustomToolFormat{ Type: "grammar", Syntax: &syntax, Definition: &definition, }), ) // Convert to types.Tool — supply the name here tool := ct.ToTool("json-extractor") ``` ```go // Custom tool with regex grammar syntax := "regex" definition := `^\d{4}-\d{2}-\d{2}$` ct := openaitool.NewCustomTool( openaitool.WithDescription("Extract a date in YYYY-MM-DD format"), openaitool.WithFormat(openaitool.CustomToolFormat{ Type: "grammar", Syntax: &syntax, Definition: &definition, }), ) tool := ct.ToTool("date-extractor") ``` ```go // Custom tool with text format (unconstrained output) ct := openaitool.NewCustomTool( openaitool.WithDescription("Analyze the sentiment of the provided text"), openaitool.WithFormat(openaitool.CustomToolFormat{Type: "text"}), ) tool := ct.ToTool("sentiment-analyzer") ``` ### Wire Format When `ct.ToTool("name")` is passed through `responses.PrepareTools()`, it serializes to: ```json { "type": "custom", "name": "json-extractor", "description": "Extract JSON from the provided text", "format": { "type": "grammar", "syntax": "lark", "definition": "start: OBJECT" } } ``` --- ## Tool Search The tool_search tool enables the model to search across deferred tools. There are two execution modes: - **Server mode** (default): OpenAI resolves tool matches internally. No `tool_search_call` event is emitted. - **Client mode**: The model emits a `tool_search_call` event. The client's `Execute` function is called with search arguments and should return matching tool names. ### Factory Function ```go type ToolSearchArgs struct { Execution string // "server" (default) or "client" Description string // describes the search capability (client mode) Parameters map[string]interface{} // JSON schema for search arguments (client mode) Execute func(ctx, input, opts) (interface{}, error) // called in client mode } func ToolSearch(args ToolSearchArgs) types.Tool ``` ### Usage ```go // Server mode (default): OpenAI handles the search searchTool := openaitool.ToolSearch(openaitool.ToolSearchArgs{}) ``` ```go // Client mode: route tool_search_call events to your Execute function searchTool := openaitool.ToolSearch(openaitool.ToolSearchArgs{ Execution: "client", Description: "Find tools matching a query", Parameters: map[string]interface{}{ "type": "object", "properties": map[string]interface{}{ "query": map[string]interface{}{"type": "string"}, }, "required": []string{"query"}, }, Execute: func(ctx context.Context, input map[string]interface{}, opts types.ToolExecutionOptions) (interface{}, error) { query, _ := input["query"].(string) // Return tool names matching the query return []string{"get_weather", "search_web"}, nil }, }) ``` ### Wire Format Server mode serializes to: ```json {"type": "tool_search"} ``` Client mode serializes to: ```json { "type": "tool_search", "execution": "client", "description": "Find tools matching a query", "parameters": {"type": "object", "properties": {"query": {"type": "string"}}} } ``` --- ## Shell Container Tools Shell tools allow the model to interact with a sandboxed environment. There are three variants: ### Package ```go import "github.com/digitallysavvy/go-ai/pkg/providers/openai/responses" ``` ### Local Shell Tool Executes commands in a local (sandboxed) shell. No container configuration needed. ```go tool := responses.NewLocalShellTool() ``` Wire format: ```json {"type": "local_shell"} ``` ### Shell Tool Runs commands in a managed container environment. Supports auto-provisioned and referenced containers. ```go // Shell tool with auto-provisioned container memLimit := "2g" tool := responses.NewShellTool( responses.WithShellEnvironment(responses.ShellEnvironment{ Type: "container_auto", MemoryLimit: &memLimit, FileIDs: []string{"file-abc123"}, // files to mount NetworkPolicy: &responses.ShellNetworkPolicy{ Type: "allowlist", AllowedDomains: []string{"pypi.org"}, }, }), ) ``` ```go // Shell tool referencing an existing container containerID := "cntr_abc123" tool := responses.NewShellTool( responses.WithShellEnvironment(responses.ShellEnvironment{ Type: "container_reference", ContainerID: &containerID, }), ) ``` ```go // Shell tool with no environment (uses default) tool := responses.NewShellTool() ``` #### ShellEnvironment Fields | Field | Type | Description | |----------------|-----------------------|----------------------------------------------------| | `Type` | `string` | `"container_auto"`, `"container_reference"`, `"local"` | | `FileIDs` | `[]string` | Files to mount (container_auto only) | | `MemoryLimit` | `*string` | Memory limit e.g. `"2g"` (container_auto only) | | `NetworkPolicy`| `*ShellNetworkPolicy` | Network access control (container_auto only) | | `Skills` | `[]ShellSkill` | Executable skills in the container | | `ContainerID` | `*string` | Existing container ID (container_reference only) | #### ShellNetworkPolicy ```go type ShellNetworkPolicy struct { Type string // "allowlist" or "disabled" AllowedDomains []string // allowed hostnames DomainSecrets []ShellDomainSecret } ``` ### Apply Patch Tool Enables the model to create, update, or delete files using unified diffs. ```go tool := responses.NewApplyPatchTool() ``` Wire format: ```json {"type": "apply_patch"} ``` --- ## PrepareTools `PrepareTools` converts a `[]types.Tool` slice into the `[]interface{}` slice expected as the `"tools"` field in a Responses API request body. Function tools with nil or implicit object parameters serialize with an explicit `{"type":"object","properties":{}}` schema when needed, matching the TypeScript SDK request shape. ```go import ( "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/openai/responses" openaitool "github.com/digitallysavvy/go-ai/pkg/providers/openai/tool" ) memLimit := "4g" tools := responses.PrepareTools([]types.Tool{ // Standard function tool {Name: "get_weather", Description: "Get weather"}, // Custom tool — name supplied to ToTool("name") openaitool.NewCustomTool().ToTool("json-extractor"), // Shell tools responses.NewLocalShellTool(), responses.NewShellTool(responses.WithShellEnvironment(responses.ShellEnvironment{ Type: "container_auto", MemoryLimit: &memLimit, })), responses.NewApplyPatchTool(), }) // tools is ready to be marshaled into a Responses API request ``` ### Restricting Callable Tools Use the OpenAI Responses `allowedTools` provider option when you want to keep the full tools list in the request for prompt caching, but restrict which function tools may be called in this step. This overrides request-level `ToolChoice` and serializes `tool_choice` as `allowed_tools`. ```go result, err := model.DoGenerate(ctx, &provider.GenerateOptions{ Prompt: prompt, Tools: []types.Tool{ {Name: "weather", Description: "Get weather"}, {Name: "cityAttractions", Description: "Find city attractions"}, }, ProviderOptions: map[string]interface{}{ "openai": map[string]interface{}{ "allowedTools": map[string]interface{}{ "toolNames": []string{"weather"}, // Optional: "auto" (default) or "required". "mode": "auto", }, }, }, }) ``` Wire shape: ```json { "tools": [ {"type": "function", "name": "weather"}, {"type": "function", "name": "cityAttractions"} ], "tool_choice": { "type": "allowed_tools", "mode": "auto", "tools": [{"type": "function", "name": "weather"}] } } ``` --- ## Response Types When the model uses shell tools, the Responses API returns typed output items. These types are defined in the `responses` package: ### LocalShellCallOutput ```go type LocalShellCallOutput struct { Type string // "local_shell_call_output" CallID string Output string // combined stdout/stderr } ``` ### ShellCallOutput ```go type ShellCallOutput struct { Type string CallID string Status *string Output []ShellCallOutputEntry } type ShellCallOutputEntry struct { Stdout string Stderr string Outcome ShellOutcome } type ShellOutcome struct { Type string // "exit" or "timeout" ExitCode *int } ``` ### ApplyPatchCallOutput ```go type ApplyPatchCallOutput struct { Type string CallID string Status string // "completed" or "failed" Output *string } ``` ### AssistantMessageItem with Phase The Responses API may include a `phase` field on assistant message items to indicate the agentic flow phase: ```go type AssistantMessageItem struct { Type string Role string ID string Phase *string // "commentary", "final_answer", or nil Content []AssistantMessageContent } ``` Values: - `"commentary"` — intermediate reasoning or commentary from the model - `"final_answer"` — the model's conclusive response - `nil` — phase not specified (standard non-agentic response) --- ## Examples Full runnable examples are in: - `examples/providers/openai/responses/custom-tool-grammar/` — Custom tool grammar formats - `examples/providers/openai/responses/shell-tool/` — Shell container tool configurations - `examples/providers/openai/tool-search/` — Tool search (server and client modes) --- ## Server-Side Compaction When the Responses API compacts the conversation context server-side, it emits a `compaction` event in the SSE stream. The Go AI SDK surfaces this as a `ChunkTypeCustom` stream chunk with `CustomContent{Kind: "openai.compaction"}`. The `ProviderMetadata` JSON on the chunk contains: | Field | Description | |--------------------|------------------------------------------------------| | `type` | Always `"compaction"` | | `itemId` | The item ID of the compacted item | | `encryptedContent` | Opaque encrypted context blob (forward in next turn) | ### Converting a compaction event ```go import "github.com/digitallysavvy/go-ai/pkg/providers/openai/responses" // In your streaming parser, when you receive a compaction event: event := responses.CompactionEvent{ Type: "compaction", ItemID: "item_abc", EncryptedContent: "enc_xyz...", } chunk := responses.CompactionEventToChunk(event) // chunk.Type == provider.ChunkTypeCustom // chunk.CustomContent.Kind == "openai.compaction" ``` When you pass that custom content back in a later assistant message, the Responses prompt converter forwards it as a `compaction` item. If `store` is enabled and the item has an OpenAI item ID, the converter sends an `item_reference` instead, matching the TypeScript SDK replay behavior. --- # XAI Advanced Usage Reporting > Explains advanced token usage tracking for xAI models in the Go AI SDK, covering cached tokens, reasoning tokens, and cost tracking. Canonical URL: https://goaisdk.com/docs/providers/xai-usage-reporting Documentation index: https://goaisdk.com/llms.txt The Go AI SDK provides sophisticated token usage tracking for XAI (formerly Twitter/X AI) models, including support for cached tokens and reasoning tokens. ## Overview XAI models report detailed token usage information that goes beyond basic input/output counts: - **Cached tokens**: Tokens read from cache (prompt caching) - **Reasoning tokens**: Tokens used for extended reasoning (Grok models) The `Usage` struct also has `TextTokens`/`ImageTokens` fields for a text/image input split, but xAI's current Responses API usage conversion does not populate them — see [Multi-Modal Token Tracking](#multi-modal-token-tracking). ## Basic Usage ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/providers/xai" ) func main() { provider := xai.New(xai.Config{ APIKey: "your-xai-api-key", }) model, _ := provider.LanguageModel("grok-beta") result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Explain quantum computing", }) if err != nil { log.Fatal(err) } // Access detailed usage information usage := result.Usage fmt.Printf("Total tokens: %d\n", *usage.TotalTokens) fmt.Printf("Input tokens: %d\n", *usage.InputTokens) fmt.Printf("Output tokens: %d\n", *usage.OutputTokens) // Access detailed breakdowns if available if usage.InputDetails != nil { if usage.InputDetails.CacheReadTokens != nil { fmt.Printf("Cache hits: %d tokens\n", *usage.InputDetails.CacheReadTokens) } if usage.InputDetails.NoCacheTokens != nil { fmt.Printf("Non-cached: %d tokens\n", *usage.InputDetails.NoCacheTokens) } } if usage.OutputDetails != nil { if usage.OutputDetails.ReasoningTokens != nil { fmt.Printf("Reasoning: %d tokens\n", *usage.OutputDetails.ReasoningTokens) } } } ``` ## Usage Structure ```go type Usage struct { // Basic counts InputTokens *int64 // Total input tokens OutputTokens *int64 // Total output tokens (including reasoning) TotalTokens *int64 // Sum of input and output // Detailed breakdowns InputDetails *InputTokenDetails OutputDetails *OutputTokenDetails // Raw API response data Raw map[string]interface{} } type InputTokenDetails struct { NoCacheTokens *int64 // Tokens not from cache CacheReadTokens *int64 // Tokens read from cache CacheWriteTokens *int64 // Tokens written to cache TextTokens *int64 // Text-only input tokens ImageTokens *int64 // Image input tokens } type OutputTokenDetails struct { TextTokens *int64 // Regular output tokens ReasoningTokens *int64 // Extended reasoning tokens } ``` ## Cached Token Tracking xAI's Responses API (the only language model path xAI uses — Chat Completions was removed) reports `input_tokens` as the total, with cached tokens broken out separately under `input_tokens_details.cached_tokens`. The SDK derives `NoCacheTokens` by subtracting the cached count from the total: ```text // API Response: // input_tokens: 200 // input_tokens_details: { cached_tokens: 150 } // SDK converts to: InputTokens: 200 // Total, taken directly from input_tokens InputDetails: { NoCacheTokens: 50, // 200 - 150 CacheReadTokens: 150, // From input_tokens_details.cached_tokens } ``` `InputDetails` is only populated when `cached_tokens > 0`; otherwise it stays `nil`. The SDK does not special-case `cached_tokens` exceeding `input_tokens` — if the API ever reports that, `NoCacheTokens` would be negative. ## Reasoning Token Tracking For models with extended reasoning (like Grok with reasoning mode), xAI's `output_tokens` field is already the full total (text plus reasoning). The SDK breaks out the reasoning portion from `output_tokens_details.reasoning_tokens` and derives the text-only count by subtraction — it does not sum two separate fields: ```text // API Response: // output_tokens: 278 // output_tokens_details: { reasoning_tokens: 228 } // SDK converts to: OutputTokens: 278 // Total, taken directly from output_tokens OutputDetails: { TextTokens: 50, // 278 - 228 ReasoningTokens: 228, // From output_tokens_details.reasoning_tokens } ``` `OutputDetails` is only populated when `reasoning_tokens > 0`. ## Multi-Modal Token Tracking xAI's Responses API usage conversion does not currently populate `InputDetails.TextTokens` / `InputDetails.ImageTokens` — those fields stay `nil` for xAI models even on multi-modal requests. If you need a text/image token split, use a provider that supports it (see the [Token usage differentiation guide](https://goaisdk.com/docs/migration-guides/token-usage-differentiation.md) for current per-provider support). ## Complete Example: Cost Tracking ```go package main import ( "context" "fmt" "log" "github.com/digitallysavvy/go-ai/pkg/ai" "github.com/digitallysavvy/go-ai/pkg/provider/types" "github.com/digitallysavvy/go-ai/pkg/providers/xai" ) // XAI pricing (example - check current pricing) const ( COST_PER_1K_INPUT = 0.01 COST_PER_1K_OUTPUT = 0.03 COST_PER_1K_CACHED = 0.001 // Cached tokens are cheaper ) func calculateCost(usage types.Usage) float64 { var cost float64 // Input cost (accounting for cache) if usage.InputDetails != nil { // Non-cached input tokens if usage.InputDetails.NoCacheTokens != nil { tokens := float64(*usage.InputDetails.NoCacheTokens) / 1000.0 cost += tokens * COST_PER_1K_INPUT } // Cached tokens (cheaper) if usage.InputDetails.CacheReadTokens != nil { tokens := float64(*usage.InputDetails.CacheReadTokens) / 1000.0 cost += tokens * COST_PER_1K_CACHED } } else if usage.InputTokens != nil { // Fallback if no detailed breakdown tokens := float64(*usage.InputTokens) / 1000.0 cost += tokens * COST_PER_1K_INPUT } // Output cost if usage.OutputTokens != nil { tokens := float64(*usage.OutputTokens) / 1000.0 cost += tokens * COST_PER_1K_OUTPUT } return cost } func main() { provider := xai.New(xai.Config{ APIKey: "your-xai-api-key", }) model, _ := provider.LanguageModel("grok-beta") result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{ Model: model, Prompt: "Write a detailed analysis of...", }) if err != nil { log.Fatal(err) } // Calculate cost cost := calculateCost(result.Usage) // Print detailed breakdown fmt.Println("=== Token Usage ===") fmt.Printf("Total tokens: %d\n", result.Usage.GetTotalTokens()) if result.Usage.InputDetails != nil { fmt.Println("\nInput breakdown:") if result.Usage.InputDetails.NoCacheTokens != nil { fmt.Printf(" Non-cached: %d tokens\n", *result.Usage.InputDetails.NoCacheTokens) } if result.Usage.InputDetails.CacheReadTokens != nil { fmt.Printf(" Cached: %d tokens\n", *result.Usage.InputDetails.CacheReadTokens) } if result.Usage.InputDetails.TextTokens != nil { fmt.Printf(" Text: %d tokens\n", *result.Usage.InputDetails.TextTokens) } if result.Usage.InputDetails.ImageTokens != nil { fmt.Printf(" Images: %d tokens\n", *result.Usage.InputDetails.ImageTokens) } } if result.Usage.OutputDetails != nil { fmt.Println("\nOutput breakdown:") if result.Usage.OutputDetails.TextTokens != nil { fmt.Printf(" Text: %d tokens\n", *result.Usage.OutputDetails.TextTokens) } if result.Usage.OutputDetails.ReasoningTokens != nil { fmt.Printf(" Reasoning: %d tokens\n", *result.Usage.OutputDetails.ReasoningTokens) } } fmt.Printf("\nEstimated cost: $%.4f\n", cost) } ``` ## Advanced: Budget Enforcement ```go type UsageBudget struct { MaxInputTokens int64 MaxOutputTokens int64 MaxTotalCost float64 } func (b *UsageBudget) Check(usage types.Usage) error { if usage.InputTokens != nil && *usage.InputTokens > b.MaxInputTokens { return fmt.Errorf("input tokens (%d) exceeds budget (%d)", *usage.InputTokens, b.MaxInputTokens) } if usage.OutputTokens != nil && *usage.OutputTokens > b.MaxOutputTokens { return fmt.Errorf("output tokens (%d) exceeds budget (%d)", *usage.OutputTokens, b.MaxOutputTokens) } cost := calculateCost(usage) if cost > b.MaxTotalCost { return fmt.Errorf("cost ($%.4f) exceeds budget ($%.2f)", cost, b.MaxTotalCost) } return nil } // Usage budget := UsageBudget{ MaxInputTokens: 5000, MaxOutputTokens: 2000, MaxTotalCost: 0.50, } result, err := ai.GenerateText(ctx, opts) if err != nil { log.Fatal(err) } if err := budget.Check(result.Usage); err != nil { log.Printf("Budget violation: %v", err) } ``` ## Raw API Data All raw token data from the API is preserved in the `Raw` field: ```go result, err := ai.GenerateText(ctx, opts) if err != nil { log.Fatal(err) } // Access raw API response rawData := result.Usage.Raw // Get direct API fields. Raw is populated via a plain json.Unmarshal into // map[string]interface{}, so JSON numbers decode as float64, not int64. if cachedTokens, ok := rawData["cached_tokens"].(float64); ok { fmt.Printf("Raw cached_tokens: %.0f\n", cachedTokens) } if reasoningTokens, ok := rawData["reasoning_tokens"].(float64); ok { fmt.Printf("Raw reasoning_tokens: %.0f\n", reasoningTokens) } // Nested structures if details, ok := rawData["prompt_tokens_details"].(map[string]interface{}); ok { fmt.Printf("Prompt token details: %+v\n", details) } ``` ## Token Calculation Examples ### Example 1: Basic Request (No Cache, No Reasoning) ``` API Response: input_tokens: 100 output_tokens: 50 SDK Output: InputTokens: 100 OutputTokens: 50 TotalTokens: 150 (100 + 50, computed by the SDK) InputDetails: nil OutputDetails: nil ``` ### Example 2: With Cache ``` API Response: input_tokens: 200 input_tokens_details: { cached_tokens: 150 } SDK Output: InputTokens: 200 InputDetails: { NoCacheTokens: 50 (200 - 150) CacheReadTokens: 150 } ``` ### Example 3: With Cache and Reasoning ``` API Response: input_tokens: 4142 input_tokens_details: { cached_tokens: 2000 } output_tokens: 300 output_tokens_details: { reasoning_tokens: 200 } SDK Output: InputTokens: 4142 OutputTokens: 300 TotalTokens: 4442 (4142 + 300, computed by the SDK) InputDetails: { NoCacheTokens: 2142 (4142 - 2000) CacheReadTokens: 2000 } OutputDetails: { TextTokens: 100 (300 - 200) ReasoningTokens: 200 } ``` ## Best Practices 1. **Always check for nil**: Token details may not always be present ```go if usage.InputDetails != nil && usage.InputDetails.CacheReadTokens != nil { // Use cached token count } ``` 2. **Use detailed breakdowns for cost calculation**: Don't rely on basic counts alone 3. **Monitor cache effectiveness**: ```go if usage.InputDetails != nil { cacheRate := float64(*usage.InputDetails.CacheReadTokens) / float64(*usage.InputTokens) * 100 fmt.Printf("Cache hit rate: %.1f%%\n", cacheRate) } ``` 4. **Track reasoning token usage**: Reasoning tokens can significantly increase costs 5. **Log raw data for debugging**: The `Raw` field contains all original API data ## Comparison with TypeScript SDK The Go implementation matches the TypeScript AI SDK's usage tracking: | TypeScript | Go | Notes | |------------|-----|-------| | `usage.inputTokens.total` | `InputTokens` | Total input | | `usage.inputTokens.noCache` | `InputDetails.NoCacheTokens` | Non-cached | | `usage.inputTokens.cacheRead` | `InputDetails.CacheReadTokens` | Cached | | `usage.outputTokens.total` | `OutputTokens` | Total output | | `usage.outputTokens.text` | `OutputDetails.TextTokens` | Text output | | `usage.outputTokens.reasoning` | `OutputDetails.ReasoningTokens` | Reasoning | | `usage.raw` | `Raw` | Raw API data | The cached token inclusivity logic and reasoning token additive behavior match the TypeScript implementation exactly. ## See Also - [XAI Provider](https://goaisdk.com/docs/providers/xai.md) - Cost Optimization - [Usage Tracking](https://goaisdk.com/docs/reference/types/usage.md) - [Prompt Caching](https://goaisdk.com/docs/advanced/caching.md) --- # Download API Reference > API reference for the Go AI SDK's Download API, which fetches files securely with built-in size limits that prevent memory-exhaustion attacks. Canonical URL: https://goaisdk.com/docs/reference/download-api Documentation index: https://goaisdk.com/llms.txt ## Overview The Download API provides secure file downloading with built-in size limits to prevent memory exhaustion attacks. ## Default Size Limit The default maximum download size is **2 GiB** (2,147,483,648 bytes), used by `ai.CreateDownload(nil)`, `ai.CreateURLDownload(nil)`, and `ai.CreateURLDownloadWithMetadata(nil)` when `options` is `nil` or `MaxBytes` is `0`. This limit prevents memory exhaustion from unbounded downloads while allowing legitimate large files. All download operations use this limit by default. There is no exported constant for this value in `pkg/ai`; pass an explicit `MaxBytes` on `ai.DownloadOptions` to override it. ## Types ### URLDownloadFunction ```go type URLDownloadFunction func(ctx context.Context, url string) ([]byte, error) ``` A single-URL download function that returns downloaded bytes. **Parameters:** - `ctx`: Context for cancellation and timeout - `url`: URL to download from **Returns:** - `[]byte`: Downloaded data - `error`: Error if download fails or exceeds size limit ### URLDownloadWithMetadataFunction ```go type URLDownloadWithMetadataFunction func(ctx context.Context, url string) (*DownloadResult, error) ``` A single-URL download function that returns downloaded bytes plus an optional media type. `GenerateVideoOptions.DownloadWithMetadata` uses this shape to match TypeScript `generateVideo` custom downloads. ### DownloadFunction ```go type DownloadFunction func(ctx context.Context, requests []DownloadRequest) ([]*DownloadResult, error) ``` A batch download function for prompt file URL normalization. A nil result leaves the original URL in place when the model can consume it directly. ### DownloadResult ```go type DownloadResult struct { Data []byte MediaType string } ``` Downloaded bytes plus an optional media type. ### DownloadOptions ```go type DownloadOptions struct { // MaxBytes is the maximum allowed download size in bytes. // Default: 2 GiB (DefaultMaxDownloadSize) MaxBytes int64 // Headers are additional HTTP headers to include in download requests. Headers map[string]string } ``` Configuration options for creating custom download functions. **Fields:** - `MaxBytes`: Maximum file size in bytes (default: 2 GiB) - `Headers`: Custom HTTP headers to send with requests ### DownloadError ```go type DownloadError struct { // URL that was being downloaded URL string // HTTP status code (if applicable) StatusCode int // HTTP status text StatusText string // Error message Message string // Underlying cause Cause error } ``` Error type returned when downloads fail due to size limits or HTTP errors. **Methods:** #### Error ```go func (e *DownloadError) Error() string ``` Returns a formatted error message. #### Unwrap ```go func (e *DownloadError) Unwrap() error ``` Returns the underlying cause of the error. ## Functions ### CreateDownload ```go func CreateDownload(options *DownloadOptions) DownloadFunction ``` Creates a download function with configurable options. The returned function enforces size limits to prevent memory exhaustion. Pass `nil` for default options (2 GiB limit). **Parameters:** - `options`: Download configuration (nil for defaults) **Returns:** - `DownloadFunction`: Configured download function ### CreateURLDownload ```go func CreateURLDownload(options *DownloadOptions) URLDownloadFunction ``` Creates a bytes-only single-URL download function. ### CreateURLDownloadWithMetadata ```go func CreateURLDownloadWithMetadata(options *DownloadOptions) URLDownloadWithMetadataFunction ``` Creates a single-URL download function that preserves response media type metadata for APIs such as `GenerateVideo`. **Example:** ```go // Create a batch prompt download function with default 2 GiB limit defaultDownload := ai.CreateDownload(nil) // Create a single-URL download function with custom 100 MB limit customDownload := ai.CreateURLDownload(&ai.DownloadOptions{ MaxBytes: 100 * 1024 * 1024, }) // Create a single-URL metadata download function with custom headers authDownload := ai.CreateURLDownloadWithMetadata(&ai.DownloadOptions{ MaxBytes: 500 * 1024 * 1024, Headers: map[string]string{ "Authorization": "Bearer token", }, }) ``` ### DefaultDownload ```go var DefaultDownload = CreateDownload(nil) ``` The default batch prompt download function with 2 GiB limit. Use `CreateURLDownload(nil)` when you need a single-URL downloader: ```go download := ai.CreateURLDownload(nil) data, err := download(ctx, url) ``` ## Internal Implementation The actual HTTP download logic (Content-Length checking, incremental reads with an abort once the limit is exceeded, timeouts) lives in `pkg/internal/fileutil`. That package is under an `internal/` path, so it is **not importable from outside this module** — attempting to import `github.com/digitallysavvy/go-ai/pkg/internal/fileutil` from your own code fails to build with `use of internal package ... not allowed`. Use the public `ai.CreateDownload`, `ai.CreateURLDownload`, and `ai.CreateURLDownloadWithMetadata` constructors above instead; they wrap `fileutil`'s behavior (Content-Length pre-check, incremental read limit, `*providererrors.DownloadError` on failure) behind the public API. ## Error Handling ### Checking for DownloadError ```go import ( "errors" providererrors "github.com/digitallysavvy/go-ai/pkg/provider/errors" ) download := ai.CreateURLDownload(nil) data, err := download(ctx, url) if err != nil { var downloadErr *providererrors.DownloadError if errors.As(err, &downloadErr) { // Handle download-specific error fmt.Printf("Failed to download %s\n", downloadErr.URL) if downloadErr.StatusCode > 0 { fmt.Printf("HTTP %d: %s\n", downloadErr.StatusCode, downloadErr.StatusText) } if downloadErr.Message != "" { fmt.Printf("Details: %s\n", downloadErr.Message) } } } ``` ### IsDownloadError ```go func IsDownloadError(err error) bool ``` Helper function to check if an error is a DownloadError: ```go if providererrors.IsDownloadError(err) { // Handle download error } ``` ### Common Error Messages | Error Message | Cause | Solution | |--------------|-------|----------| | `download of {url} exceeded maximum size of {limit} bytes (Content-Length: {size})` | File size in Content-Length header exceeds limit | Increase limit or validate file size before downloading | | `download of {url} exceeded maximum size of {limit} bytes` | File exceeded limit during streaming download | Increase limit or use chunked processing | | `failed to download {url}: {status} {text}` | HTTP error response | Check URL validity and accessibility | | `failed to download {url}: {cause}` | Network or other error | Check network connectivity and URL | ## Usage in Generation Functions ### Video Generation ```go customDownload := ai.CreateURLDownloadWithMetadata(&ai.DownloadOptions{ MaxBytes: 200 * 1024 * 1024, // 200 MB }) result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: ai.VideoPrompt{ Text: "create video from this image", Image: &ai.VideoPromptImage{ URL: "https://example.com/input.jpg", }, }, DownloadWithMetadata: customDownload, }) ``` ### Image Generation Image downloads happen automatically within providers and use the default 2 GiB limit. Provider implementations use `fileutil.Download` internally. ## Size Limit Guidelines | Use Case | Recommended Limit | Rationale | |----------|------------------|-----------| | Thumbnails | 10 MB | Small images | | Standard images | 100 MB | Most images | | High-res images | 500 MB | Professional photography | | Short videos | 500 MB | < 1 minute | | Long videos | 2 GiB (default) | Several minutes | ## Security Best Practices 1. **Always use size limits** - Never disable or set to unlimited 2. **Validate URLs** - Check domain, scheme, and format 3. **Set timeouts** - Use context with timeout for all downloads 4. **Rate limit** - Prevent abuse in public-facing applications 5. **Log failures** - Monitor for attack attempts ## Performance Considerations ### Memory Usage Downloads are read into memory entirely. For large files: - Peak memory = file size + overhead - Consider streaming approaches for very large files - Monitor memory usage in production ### Timeout Guidelines ```go // Quick images (< 10 MB) ctx, cancel := context.WithTimeout(ctx, 10*time.Second) // Standard images/videos (< 500 MB) ctx, cancel := context.WithTimeout(ctx, 60*time.Second) // Large files (> 500 MB) ctx, cancel := context.WithTimeout(ctx, 5*time.Minute) ``` ## See Also - [Download Security Guide](https://goaisdk.com/docs/guides/download-security.md) - [Security Advisory](https://goaisdk.com/docs/security/ADVISORY-Download-DoS.md) - [Error Handling](https://goaisdk.com/docs/ai-sdk-core/error-handling.md) --- # Security Advisory: Unbounded Download DoS Prevention > Security advisory describing an unbounded download denial-of-service vulnerability in the Go AI SDK, its fix, mitigation steps, and impact assessment. Canonical URL: https://goaisdk.com/docs/security/ADVISORY-Download-DoS Documentation index: https://goaisdk.com/llms.txt **Status:** Fixed **Severity:** High **CVE:** None assigned **Affected Versions:** All versions before v0.2.0 **Fixed in:** v0.2.0 ## Summary Prior versions of the Go AI SDK allowed unbounded memory growth when downloading from user-provided URLs (images, videos, audio), enabling Denial of Service (DoS) attacks through memory exhaustion. ## Vulnerability Details ### Attack Vector The SDK's download functions did not limit file sizes when fetching resources from user-provided URLs. An attacker could provide URLs pointing to extremely large files (e.g., 100GB+), causing: - Memory exhaustion - Process crashes from Out-of-Memory (OOM) errors - Resource depletion on the host system - Service unavailability ### Example Attack ```go // Attacker provides URL to a huge file model := xai.NewImageModel(xaiProvider, "grok-2-vision-1212") result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{ Model: model, Prompt: "analyze this image", Files: []provider.ImageFile{ {Type: "url", URL: "https://evil.com/100gb-file.jpg"}, // Causes OOM crash }, }) // Process crashes from memory exhaustion ``` ### Affected Components - Image download functions in providers (XAI, BFL, FAL, Replicate) - Video download functions in providers (Google, FAL, Replicate) - Any code path that downloads from user-provided URLs ## Fix ### Changes The fix implements a **2 GiB default size limit** for all downloads with the following protections: 1. **Early rejection**: Checks `Content-Length` header before downloading 2. **Streaming with limits**: Reads response body incrementally with `io.LimitReader` 3. **Proper error handling**: Throws `DownloadError` when limit exceeded 4. **Context cancellation**: Respects context cancellation for aborting downloads ### Default Protection All download operations now automatically enforce a 2 GiB limit. The default download function (used internally, and available via `ai.CreateURLDownload` for your own code) enforces this limit without any extra configuration: ```go // Automatically protected with 2 GiB limit download := ai.CreateURLDownload(nil) data, err := download(ctx, url) ``` ### Custom Limits Applications can configure custom size limits: ```go // Create single-URL download function with custom 100 MB limit customDownload := ai.CreateURLDownloadWithMetadata(&ai.DownloadOptions{ MaxBytes: 100 * 1024 * 1024, // 100 MB }) // Use in generation functions result, err := ai.GenerateVideo(ctx, ai.GenerateVideoOptions{ Model: model, Prompt: videoPrompt, DownloadWithMetadata: customDownload, }) ``` ## Mitigation ### Immediate Action **Upgrade to v0.2.0 or later immediately.** No workaround is available for earlier versions. ### Version Check ```bash go list -m github.com/digitallysavvy/go-ai ``` If the version is below v0.2.0, upgrade with: ```bash go get github.com/digitallysavvy/go-ai@latest ``` ## Impact Assessment ### Who is Affected? Any application that: - Accepts user-provided image/video/audio URLs - Uses the Go AI SDK for generation with file inputs - Runs in production environments with untrusted input ### Risk Level - **High** for public-facing services with user-generated content - **Medium** for internal services with trusted users - **Low** for applications that only use SDK-generated content ## Timeline - **Discovered:** 2026-02-12 - **Fixed:** 2026-02-15 - **Released:** 2026-02-XX - **Public Disclosure:** 2026-02-XX ## Credits Security fix implemented based on TypeScript AI SDK commit 4024a3af6. ## References - [CWE-770: Allocation of Resources Without Limits](https://cwe.mitre.org/data/definitions/770.html) - [TypeScript AI SDK Fix](https://github.com/vercel/ai/commit/4024a3af6) ## Contact For security concerns, please report to: security@digitallysavvy.com