# Alibaba Cloud Provider

Alibaba Cloud provides powerful Qwen language models with strong Chinese language capabilities and Wan video generation models. Known for cost-effective prompt caching, reasoning capabilities, and innovative video generation features.

## Setup

### Installation

```go
import (
    "github.com/digitallysavvy/go-ai/pkg/ai"
    "github.com/digitallysavvy/go-ai/pkg/providers/alibaba"
)
```

### Configuration

```go
provider := alibaba.New(alibaba.Config{
    APIKey: os.Getenv("ALIBABA_API_KEY"),
})

// Chat model
model, err := provider.LanguageModel("qwen-plus")
if err != nil {
    log.Fatal(err)
}

// Video model
videoModel, err := provider.VideoModel("wan2.6-t2v")
if err != nil {
    log.Fatal(err)
}
```

### Get API Key

1. Sign up at [Alibaba Cloud DashScope](https://dashscope.console.aliyun.com/)
2. Navigate to API Keys section
3. Create a new API key
4. Set environment variable:

```bash
export ALIBABA_API_KEY=sk-...
```

## Available Models

### Qwen Language Models

| Model ID | Context | Best For |
|----------|---------|----------|
| qwen-plus | 32K | Balanced performance and cost |
| qwen-turbo | 8K | Fast responses, economical |
| qwen3-max | 32K | Most capable, highest quality |
| qwq-32b | 32K | Complex reasoning with thinking |
| qwen-vl-max | 32K | Vision + language understanding |

### Wan Video Models

`Provider.VideoModel` accepts any model ID (there is no fixed allowlist);
the constants below (`pkg/providers/alibaba/video_model_ids.go`) cover the
documented catalog. An empty model ID defaults to `wan2.6-t2v`.

| Model ID | Go constant | Type | Best For |
|----------|-------------|------|----------|
| wan2.5-t2v-preview | `VideoModelWan2_5T2VPreview` | Text-to-video | Preview quality (720p) |
| wan2.6-t2v | `VideoModelWan2_6T2V` | Text-to-video | High quality generation |
| wan2.7-t2v | `VideoModelWan2_7T2V` | Text-to-video | Latest generation |
| wan2.6-i2v | `VideoModelWan2_6I2V` | Image-to-video | Animate static images |
| wan2.6-i2v-flash | `VideoModelWan2_6I2VFlash` | Image-to-video | Faster generation |
| wan2.6-r2v | `VideoModelWan2_6R2V` | Reference-to-video | Style transfer from reference |
| wan2.6-r2v-flash | `VideoModelWan2_6R2VFlash` | Reference-to-video | Faster style transfer |
| wan2.7-r2v | `VideoModelWan2_7R2V` | Reference-to-video | Latest reference-to-video |
| wan3.0-video | `VideoModelWan3Video` | All-in-one | One model ID serves text-, image-, and reference-to-video, and supports custom aspect ratios |

`wan2.7` and `wan3` are new this cycle — `wan3.0-video` in particular
replaces separate t2v/i2v/r2v model IDs with one model that branches on
which inputs (`Prompt` only, `Image`, or `InputReferences`) you pass.

### Embedding Models

Alibaba DashScope text embeddings are exposed through `Provider.EmbeddingModel` and `Provider.Embedding`.

| Model ID | Go constant | Notes |
|----------|-------------|-------|
| text-embedding-v4 | `alibaba.AlibabaEmbeddingTextV4` | Dense, sparse, or dense+sparse output |
| text-embedding-v3 | `alibaba.AlibabaEmbeddingTextV3` | Dense text embeddings |

```go
embeddingModel, err := provider.EmbeddingModel(alibaba.AlibabaEmbeddingTextV4)
if err != nil {
    log.Fatal(err)
}

dimension := 1024.0
result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{
    Model:  embeddingModel,
    Inputs: []string{"first document", "second document"},
    ProviderOptions: map[string]interface{}{
        "alibaba": alibaba.AlibabaEmbeddingModelOptions{
            TextType:   "document",
            Dimension:  &dimension,
            OutputType: alibaba.AlibabaEmbeddingOutputDense,
        },
    },
})
```

`OutputType` accepts `dense`, `sparse`, or `dense&sparse`. Sparse-only responses are reported as unsupported for the SDK's dense embedding result contract; dense+sparse responses preserve sparse vectors in provider metadata.

## Provider-Specific Features

### Thinking/Reasoning

Enable reasoning for complex problems with token tracking:

```go
result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
    Prompt: types.Prompt{
        Text: "Solve this logic puzzle: ...",
    },
    ProviderOptions: map[string]interface{}{
        "alibaba": map[string]interface{}{
            "enableThinking":  true,
            "thinkingBudget": 1000, // Max reasoning tokens
        },
    },
})

// Access reasoning tokens
if result.Usage.OutputDetails != nil && result.Usage.OutputDetails.ReasoningTokens != nil {
    fmt.Printf("Reasoning tokens used: %d\n", *result.Usage.OutputDetails.ReasoningTokens)
}
```

### Prompt Caching

Automatic caching for repeated content with cost savings:

```go
// System prompts are automatically cached
prompt := types.Prompt{
    System: "You are an expert software architect...", // Cached
    Text:   "Design a microservices system",
}

result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
    Prompt: prompt,
})

// Check cache usage
if result.Usage.InputDetails != nil {
    if result.Usage.InputDetails.CacheReadTokens != nil {
        fmt.Printf("Cache hit: %d tokens saved!\n", *result.Usage.InputDetails.CacheReadTokens)
    }
    if result.Usage.InputDetails.CacheWriteTokens != nil {
        fmt.Printf("Cached: %d tokens for future use\n", *result.Usage.InputDetails.CacheWriteTokens)
    }
}
```

### Vision Capabilities

Image understanding with Qwen VL models:

```go
messages := []types.Message{
    {
        Role: types.RoleUser,
        Content: []types.ContentPart{
            types.FileContent{MediaType: "image/jpeg", URL: "https://example.com/photo.jpg"},
            types.TextContent{Text: "What's in this image?"},
        },
    },
}

visionModel, _ := provider.LanguageModel("qwen-vl-max")
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:    visionModel,
    Messages: messages,
})
```

### Tool Calling

Function calling with single or parallel execution:

```go
tools := []types.Tool{
    {
        Name:        "get_weather",
        Description: "Get current weather for a location",
        Parameters: map[string]interface{}{
            "type": "object",
            "properties": map[string]interface{}{
                "location": map[string]interface{}{
                    "type":        "string",
                    "description": "City name",
                },
            },
            "required": []string{"location"},
        },
    },
}

result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
    Prompt: types.Prompt{Text: "What's the weather in Paris?"},
    Tools:  tools,
    ProviderOptions: map[string]interface{}{
        "alibaba": map[string]interface{}{
            "parallelToolCalls": true, // Enable parallel execution
        },
    },
})

// Handle tool calls
if len(result.ToolCalls) > 0 {
    // Execute tools and send results back
}
```

### Video Generation

#### Text-to-Video

```go
videoModel, _ := provider.VideoModel("wan2.6-t2v")
duration := 5.0

result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
    Prompt:      "A golden retriever playing in a park, sunny day",
    AspectRatio: "16:9",
    Duration:    &duration,
})

if len(result.Videos) > 0 {
    fmt.Printf("Video URL: %s\n", result.Videos[0].URL)
}
```

#### Image-to-Video

Animate static images:

```go
videoModel, _ := provider.VideoModel("wan2.6-i2v-flash")
duration := 6.0

result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
    Prompt:      "Add gentle camera movement and natural motion",
    AspectRatio: "16:9",
    Duration:    &duration,
    Image: &provider.VideoModelV3File{
        Type: "url",
        URL:  "https://example.com/landscape.jpg",
    },
})
```

#### Reference-to-Video (Style Transfer)

Apply style from reference image to new content:

```go
videoModel, _ := provider.VideoModel("wan2.6-r2v")
duration := 5.0

result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
    Prompt:      "A cat walking through a garden",
    AspectRatio: "16:9",
    Duration:    &duration,
    Image: &provider.VideoModelV3File{
        Type: "url",
        URL:  "https://example.com/anime-style.jpg", // Reference style
    },
})
```

#### Async Video

`VideoModel` also implements `provider.VideoModelStarter` /
`VideoModelStatusChecker`, so it works with `ai.ExperimentalStartVideo` /
`ai.ExperimentalGetVideoStatus` instead of the polling `ai.GenerateVideo`.
The default poll interval/timeout are 5s / 10 minutes, and the
provider-returned task ID is path-encoded when polling status.

```go
started, err := ai.ExperimentalStartVideo(ctx, ai.StartVideoOptions{
    Model:  videoModel,
    Prompt: ai.VideoPrompt{Text: "A golden retriever playing in a park"},
})
if err != nil {
    log.Fatal(err)
}

status, err := ai.ExperimentalGetVideoStatus(ctx, videoModel, ai.GetVideoStatusOptions{
    Operation: started.Operation,
})
```

## Examples

### Basic Text Generation

```go
package main

import (
    "context"
    "fmt"
    "log"
    "os"

    "github.com/digitallysavvy/go-ai/pkg/provider"
    "github.com/digitallysavvy/go-ai/pkg/provider/types"
    "github.com/digitallysavvy/go-ai/pkg/providers/alibaba"
)

func main() {
    cfg, err := alibaba.NewConfig(os.Getenv("ALIBABA_API_KEY"))
    if err != nil {
        log.Fatal(err)
    }

    prov := alibaba.New(cfg)
    model, _ := prov.LanguageModel("qwen-plus")

    result, err := model.DoGenerate(context.Background(),
        &provider.GenerateOptions{
            Prompt: types.Prompt{
                Text: "Write a haiku about artificial intelligence.",
            },
        })
    if err != nil {
        log.Fatal(err)
    }

    fmt.Println(result.Text)
}
```

### Reasoning with Thinking

```go
model, _ := provider.LanguageModel("qwq-32b")

result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
    Prompt: types.Prompt{
        Text: `Solve this logic puzzle:
Three friends - Alice, Bob, and Carol - each have different pets.
- Alice doesn't have a dog
- Bob is allergic to cats
- Carol lives in an apartment that doesn't allow fish
Who has which pet?`,
    },
    ProviderOptions: map[string]interface{}{
        "alibaba": map[string]interface{}{
            "enableThinking":  true,
            "thinkingBudget": 1000,
        },
    },
})
```

### Video Generation

```go
videoModel, _ := provider.VideoModel("wan2.6-t2v")
duration := 5.0

result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
    Prompt:      "A person dancing in the rain",
    AspectRatio: "9:16", // Vertical for mobile
    Duration:    &duration,
})
```

## Best Practices

1. **Use Prompt Caching**
   - Cache long system prompts for cost savings
   - Second request with same system prompt uses cached tokens
   - Significant cost reduction for repeated queries

2. **Choose the Right Model**
   - `qwen-turbo`: Simple tasks, fast responses
   - `qwen-plus`: Balanced for most use cases
   - `qwen3-max`: Complex reasoning, highest quality
   - `qwq-32b`: Logic puzzles, math problems

3. **Video Generation**
   - Text-to-video takes 1-2 minutes
   - Use flash variants when speed > quality
   - Aspect ratios: 16:9 (landscape), 9:16 (portrait), 1:1 (square)
   - Reference-to-video for consistent visual style

4. **Thinking Budget**
   - Set limits to control costs
   - Monitor reasoning token usage
   - Higher budgets for complex problems

## Rate Limits & Pricing

### Rate Limits

Check [Alibaba Cloud DashScope documentation](https://help.aliyun.com/zh/dashscope/) for current limits.

### Cost Optimization

- Enable prompt caching for repeated content
- Use `qwen-turbo` for simple tasks
- Set thinking budget limits for reasoning models
- Use flash video variants when appropriate

## API Endpoints

- **Chat API**: `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` (OpenAI-compatible)
- **Video API**: `https://dashscope-intl.aliyuncs.com` (DashScope native)

## Complete Examples

See the [Alibaba provider examples](https://github.com/digitallysavvy/go-ai/tree/main/examples/providers/alibaba) directory for 8 complete examples:

1. Basic chat
2. Thinking/reasoning
3. Prompt caching
4. Text-to-video
5. Image-to-video
6. Vision chat
7. Tool calling
8. Reference-to-video

## Workflow Serialization

Alibaba embedding models can cross a workflow boundary with
`provider.SerializeEmbeddingModel` / `DeserializeEmbeddingModel` (language
models could already be serialized with `providerutils.SerializeModel` /
`DeserializeModel`). See
[Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization)
for the mechanism; video models are not yet serializable.

## See Also

- [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md)
- [API Reference: GenerateVideo](https://goaisdk.com/docs/ai-sdk-core/video-generation.md)
- [Alibaba Cloud DashScope Documentation](https://help.aliyun.com/zh/dashscope/)
- [Qwen Models](https://qwen.readthedocs.io/)
