# AWS Bedrock Provider

AWS Bedrock provides access to multiple foundation models from different providers through a single API. Includes models from Anthropic, Meta, Cohere, AI21, Amazon, and more with enterprise-grade security and compliance.

Cohere embedding models are detected both as direct model IDs such as `cohere.embed-v4` and cross-region inference profile IDs such as `us.cohere.embed-v4:0`. The request body uses the Cohere embedding shape in both cases.

## Setup

### Installation

The AWS Bedrock provider is included in the Go-AI SDK:

```go
import (
    "github.com/digitallysavvy/go-ai/pkg/ai"
    "github.com/digitallysavvy/go-ai/pkg/providers/bedrock"
)
```

### Configuration

```go
provider := bedrock.New(bedrock.Config{
    Region:             "us-east-1",
    AWSAccessKeyID:     os.Getenv("AWS_ACCESS_KEY_ID"),
    AWSSecretAccessKey: os.Getenv("AWS_SECRET_ACCESS_KEY"),
})

model, err := provider.LanguageModel("anthropic.claude-3-5-sonnet-20241022-v2:0")
if err != nil {
    log.Fatal(err)
}
```

### Get Started

1. Sign up for AWS account at [aws.amazon.com](https://aws.amazon.com)
2. Request model access in [Bedrock Console](https://console.aws.amazon.com/bedrock)
3. Create IAM user with Bedrock permissions
4. Set environment variables:

```bash
export AWS_ACCESS_KEY_ID=AKIA...
export AWS_SECRET_ACCESS_KEY=...
export AWS_REGION=us-east-1
```

## Available Models

### Anthropic Claude Models

| Model ID | Context | Input Price | Output Price | Best For |
|----------|---------|-------------|--------------|----------|
| anthropic.claude-3-5-sonnet-20241022-v2:0 | 200K | $3.00/1M | $15.00/1M | Best balance |
| anthropic.claude-3-opus-20240229-v1:0 | 200K | $15.00/1M | $75.00/1M | Complex tasks |
| anthropic.claude-3-haiku-20240307-v1:0 | 200K | $0.25/1M | $1.25/1M | Fast responses |

### Meta Llama Models

| Model ID | Context | Input Price | Output Price | Best For |
|----------|---------|-------------|--------------|----------|
| us.meta.llama3-3-70b-instruct-v1:0 | 128K | $0.99/1M | $0.99/1M | Open source |
| meta.llama3-2-90b-instruct-v1:0 | 128K | $2.65/1M | $2.65/1M | Multimodal |
| meta.llama3-2-11b-instruct-v1:0 | 128K | $0.35/1M | $0.35/1M | Efficient |

### Amazon Titan Models

| Model ID | Context | Input Price | Output Price | Best For |
|----------|---------|-------------|--------------|----------|
| amazon.titan-text-premier-v1:0 | 32K | $0.50/1M | $1.50/1M | General purpose |
| amazon.titan-text-express-v1 | 8K | $0.20/1M | $0.60/1M | Fast, cheap |

### Cohere Command Models

| Model ID | Context | Input Price | Output Price | Best For |
|----------|---------|-------------|--------------|----------|
| cohere.command-r-plus-v1:0 | 128K | $3.00/1M | $15.00/1M | RAG, search |
| cohere.command-r-v1:0 | 128K | $0.50/1M | $1.50/1M | Cost-effective |

### AI21 Jurassic Models

| Model ID | Context | Input Price | Output Price | Best For |
|----------|---------|-------------|--------------|----------|
| ai21.jamba-1-5-large-v1:0 | 256K | $2.00/1M | $8.00/1M | Long context |
| ai21.jamba-1-5-mini-v1:0 | 256K | $0.20/1M | $0.40/1M | Fast tasks |

### Mistral AI Models

| Model ID | Context | Input Price | Output Price | Best For |
|----------|---------|-------------|--------------|----------|
| mistral.mistral-large-2407-v1:0 | 128K | $3.00/1M | $9.00/1M | Complex reasoning |
| mistral.mistral-small-2402-v1:0 | 32K | $1.00/1M | $3.00/1M | Efficient |

### Embedding Models

| Model ID | Dimensions | Price | Best For |
|----------|-----------|-------|----------|
| amazon.titan-embed-text-v2:0 | 1024 | $0.02/1M | General embeddings |
| cohere.embed-english-v3 | 1024 | $0.10/1M | English text |
| cohere.embed-multilingual-v3 | 1024 | $0.10/1M | Multilingual |

### Image Generation

| Model ID | Quality | Price | Best For |
|----------|---------|-------|----------|
| amazon.titan-image-generator-v2:0 | High | $0.04/image | General images |
| stability.stable-diffusion-xl-v1 | High | $0.04/image | Artistic |

## Provider-Specific Features

### Model Access Control

Request access to models in the [Bedrock Console](https://console.aws.amazon.com/bedrock/home#/modelaccess) before calling them; the SDK does not expose an API to list or manage model access.

### Cross-Region Inference

Cross-region inference profiles are plain model IDs with a region prefix (e.g. `us.`, `eu.`, or no prefix for global profiles) — pass the profile ID directly to `LanguageModel`, there is no separate config field:

```go
provider := bedrock.New(bedrock.Config{
    Region:             "us-east-1",
    AWSAccessKeyID:     os.Getenv("AWS_ACCESS_KEY_ID"),
    AWSSecretAccessKey: os.Getenv("AWS_SECRET_ACCESS_KEY"),
})

model, err := provider.LanguageModel("us.anthropic.claude-3-5-sonnet-20241022-v2:0")
```

### Guardrails

Apply AWS Bedrock Guardrails for content filtering:

```go
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:  model,
    Prompt: prompt,
    ProviderOptions: map[string]interface{}{
        "amazonBedrock": map[string]interface{}{
            "guardrailIdentifier": "your-guardrail-id",
            "guardrailVersion":    "1",
        },
    },
})
```

## Examples

### Basic Text Generation

```go
package main

import (
    "context"
    "fmt"
    "log"
    "os"

    "github.com/digitallysavvy/go-ai/pkg/ai"
    "github.com/digitallysavvy/go-ai/pkg/providers/bedrock"
)

func main() {
    ctx := context.Background()
    provider := bedrock.New(bedrock.Config{
        Region:             "us-east-1",
        AWSAccessKeyID:     os.Getenv("AWS_ACCESS_KEY_ID"),
        AWSSecretAccessKey: os.Getenv("AWS_SECRET_ACCESS_KEY"),
    })

    model, err := provider.LanguageModel("anthropic.claude-3-5-sonnet-20241022-v2:0")
    if err != nil {
        log.Fatal(err)
    }

    result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
        Model:  model,
        Prompt: "Explain AWS Bedrock benefits",
    })
    if err != nil {
        log.Fatal(err)
    }

    fmt.Println(result.Text)
}
```

### Multi-Model Comparison

```go
models := []string{
    "anthropic.claude-3-5-sonnet-20241022-v2:0",
    "us.meta.llama3-3-70b-instruct-v1:0",
    "cohere.command-r-plus-v1:0",
}

prompt := "Write a haiku about clouds"

for _, modelID := range models {
    model, err := provider.LanguageModel(modelID)
    if err != nil {
        log.Printf("Failed to load %s: %v", modelID, err)
        continue
    }

    result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
    if err != nil {
        log.Printf("Failed to generate with %s: %v", modelID, err)
        continue
    }

    fmt.Printf("\n=== %s ===\n%s\n", modelID, result.Text)
}
```

### Image Generation with Titan

```go
imageModel, err := provider.ImageModel("amazon.titan-image-generator-v2:0")
if err != nil {
    log.Fatal(err)
}

result, err := ai.GenerateImage(ctx, ai.GenerateImageOptions{
    Model:  imageModel,
    Prompt: "A serene mountain landscape at sunset",
    Size:   "1024x1024",
})
if err != nil {
    log.Fatal(err)
}

// Save image
os.WriteFile("output.png", result.Images[0].Data, 0644)
```

### Embeddings with Cohere

```go
embeddingModel, err := provider.EmbeddingModel("cohere.embed-english-v3")
if err != nil {
    log.Fatal(err)
}

texts := []string{
    "AWS Bedrock provides access to foundation models",
    "Amazon offers cloud computing services",
    "Machine learning powers AI applications",
}

result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{
    Model:  embeddingModel,
    Inputs: texts,
})
if err != nil {
    log.Fatal(err)
}

fmt.Printf("Generated %d embeddings\n", len(result.Embeddings))
```

### Streaming with Claude

```go
model, err := provider.LanguageModel("anthropic.claude-3-5-sonnet-20241022-v2:0")
if err != nil {
    log.Fatal(err)
}

stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a detailed essay about serverless computing"})
if err != nil {
    log.Fatal(err)
}
defer stream.Close()

for chunk := range stream.Chunks() {
    fmt.Print(chunk.Text)
}
```

## Advanced Configuration

### IAM Role Authentication

```go
// Use IAM role credentials (for EC2, ECS, Lambda)
provider := bedrock.New(bedrock.Config{
    Region: "us-east-1",
    // Credentials automatically loaded from IAM role
})
```

### Session Token (STS)

```go
provider := bedrock.New(bedrock.Config{
    Region:             "us-east-1",
    AWSAccessKeyID:     os.Getenv("AWS_ACCESS_KEY_ID"),
    AWSSecretAccessKey: os.Getenv("AWS_SECRET_ACCESS_KEY"),
    SessionToken:       os.Getenv("AWS_SESSION_TOKEN"),
})
```

When explicit credentials are provided, the provider uses the explicit `SessionToken` value only. It does not merge in `AWS_SESSION_TOKEN` from the environment unless credentials are being resolved from the environment, matching the TypeScript credential precedence.

### Custom Endpoint

```go
provider := bedrock.New(bedrock.Config{
    Region:  "us-east-1",
    BaseURL: "https://bedrock-runtime.us-east-1.amazonaws.com",
})
```

### Reasoning and Structured Output

Bedrock drops reasoning blocks that belong to a foreign provider before request signing, so unsigned provider-specific reasoning is not forwarded accidentally. When top-level reasoning is set, custom reasoning configuration properties are merged into the Bedrock request instead of replacing the reasoning config.

Anthropic-on-Bedrock supports native structured output when reasoning is not enabled. The SDK preserves empty text blocks that accompany reasoning content, matching the TypeScript content-shape behavior required by Claude reasoning models.

```go
model, _ := provider.LanguageModel("anthropic.claude-opus-4-7")

// invoiceSchema is a schema.Schema, e.g. schema.NewSimpleJSONSchema(map[string]interface{}{...}).
result, err := ai.GenerateObject(ctx, ai.GenerateObjectOptions{
    Model:  model,
    Prompt: "Extract the invoice total.",
    Schema: invoiceSchema,
})
```

## Error Handling

### Common Errors

Bedrock surfaces errors as `*providererrors.ProviderError`, like every other provider. AWS's exception type is folded into `Message` as `"{type}: {message}"` (there is no separate error-code field), so match on the message prefix:

```go
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
if err != nil {
    var providerErr *providererrors.ProviderError
    if errors.As(err, &providerErr) {
        switch {
        case strings.HasPrefix(providerErr.Message, "AccessDeniedException"):
            log.Fatal("Access denied - check IAM permissions")
        case strings.HasPrefix(providerErr.Message, "ResourceNotFoundException"):
            log.Fatal("Model not found - request access in Bedrock console")
        case strings.HasPrefix(providerErr.Message, "ThrottlingException"):
            log.Println("Throttled, retrying...")
            time.Sleep(time.Second * 5)
        case strings.HasPrefix(providerErr.Message, "ValidationException"):
            log.Printf("Invalid request: %s", providerErr.Message)
        case strings.HasPrefix(providerErr.Message, "ModelNotReadyException"):
            log.Println("Model still loading, retry later")
        case strings.HasPrefix(providerErr.Message, "ServiceQuotaExceededException"):
            log.Fatal("Quota exceeded - request limit increase")
        default:
            log.Printf("Bedrock error (%d): %s", providerErr.StatusCode, providerErr.Message)
        }
    }
}
```

### IAM Permission Requirements

Minimum IAM policy for Bedrock:

```json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "bedrock:InvokeModel",
        "bedrock:InvokeModelWithResponseStream"
      ],
      "Resource": "arn:aws:bedrock:*::foundation-model/*"
    }
  ]
}
```

## Best Practices

1. **Model Selection**
   - Use Claude for complex reasoning and analysis
   - Use Llama for cost-effective open-source alternative
   - Use Titan for AWS-native integration
   - Use Cohere for RAG and search applications

2. **Cost Management**
   - Compare prices across providers
   - Use smaller models for simple tasks
   - Monitor usage with AWS Cost Explorer
   - Set up billing alerts

3. **Security**
   - Use IAM roles instead of access keys
   - Apply least privilege IAM policies
   - Enable CloudTrail logging
   - Use VPC endpoints for private access

4. **Performance**
   - Choose region closest to users
   - Use cross-region inference for high availability
   - Implement caching for repeated queries
   - Use streaming for long responses

5. **Compliance**
   - Review model provider compliance certifications
   - Use Guardrails for content filtering
   - Enable logging for audit trails
   - Configure data retention policies

## Rate Limits & Pricing

### Rate Limits

Varies by model and region:

| Model Family | Default Tokens/Min | Burst |
|--------------|-------------------|-------|
| Claude | 100K - 400K | 200K |
| Llama | 200K - 800K | 400K |
| Titan | 400K - 1.6M | 800K |

Request quota increases through AWS Support.

### Cost Comparison

```go
func compareCosts(promptTokens, completionTokens int) {
    models := map[string][2]float64{
        "Claude Sonnet 3.5": {3.00 / 1_000_000, 15.00 / 1_000_000},
        "Llama 3.3 70B":     {0.99 / 1_000_000, 0.99 / 1_000_000},
        "Command R+":        {3.00 / 1_000_000, 15.00 / 1_000_000},
        "Titan Premier":     {0.50 / 1_000_000, 1.50 / 1_000_000},
    }

    for model, rates := range models {
        cost := float64(promptTokens)*rates[0] + float64(completionTokens)*rates[1]
        fmt.Printf("%s: $%.4f\n", model, cost)
    }
}
```

## See Also

- [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md)
- [Anthropic Provider](https://goaisdk.com/docs/providers/anthropic.md) - Direct Anthropic access
- [AWS Bedrock Documentation](https://docs.aws.amazon.com/bedrock/)
- [Bedrock Model Access](https://console.aws.amazon.com/bedrock/home#/modelaccess)

## May 2026 parity updates

### Credential precedence

Explicit credentials in `bedrock.Config` take precedence over environment and shared credential file values. When `AWSAccessKeyID` and `AWSSecretAccessKey` are set, `AWS_SESSION_TOKEN` is not used unless `SessionToken` is also explicitly configured.

### Reasoning and structured output

`bedrock.ModelOptions.ReasoningConfig` mirrors the TypeScript `amazonBedrock.reasoningConfig` option. Partial provider configs merge with top-level `Reasoning`. Foreign-provider reasoning blocks are dropped rather than sent unsigned.

```go
budget := 4096
model, err := p.LanguageModelWithOptions("anthropic.claude-opus-4-7", &bedrock.ModelOptions{
    ReasoningConfig: &bedrock.ReasoningConfig{
        Type:         "enabled",
        BudgetTokens: &budget,
    },
    ServiceTier: "priority",
})
```

### Transport behavior

The TypeScript lazy fetch resolution change has no direct Go runtime equivalent. The Go provider uses its AWS signing request path and explicit configuration fields; no application migration is required for JavaScript `fetch` patching behavior.

### requestMetadata

Pass `providerOptions.amazonBedrock.requestMetadata` (a `map[string]string`) to attach key-value pairs to a Converse/ConverseStream request for invocation-log cost attribution:

```go
_, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model: model,
    Prompt: "hello",
    ProviderOptions: map[string]interface{}{
        "amazonBedrock": map[string]interface{}{
            "requestMetadata": map[string]interface{}{"team": "search", "environment": "prod"},
        },
    },
})
```

It is forwarded as a top-level `requestMetadata` field on the request body, separate from `additionalModelRequestFields`.

### Nova 2 Lite high reasoning

`amazon.nova-2-lite-v1:0` rejects a request that combines `maxOutputTokens` with `high`/`xhigh`/`max` reasoning effort. The provider now drops `inferenceConfig.maxTokens` and returns an `unsupported` warning for `maxOutputTokens` instead of sending a request Bedrock would reject; `medium` (and lower) reasoning keeps `maxOutputTokens` unchanged.

### Region validation

`Region` (and Bedrock Mantle's `Region`) is interpolated directly into the request host (`https://bedrock-runtime.{region}.amazonaws.com`, `https://bedrock-mantle.{region}.api.aws`). A value that isn't a single DNS label (letters, digits, and hyphens) is rejected with an `InvalidArgumentError` before any request is sent — this only applies when no explicit `BaseURL` or endpoint environment variable is configured, since those make `Region` unused.
