Skip to main content

Alibaba Cloud Provider

Alibaba Cloud provides powerful Qwen language models with strong Chinese language capabilities and Wan video generation models. Known for cost-effective prompt caching, reasoning capabilities, and innovative video generation features.

Setup​

Installation​

import (
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/alibaba"
)

Configuration​

provider := alibaba.New(alibaba.Config{
APIKey: os.Getenv("ALIBABA_API_KEY"),
})

// Chat model
model, err := provider.LanguageModel("qwen-plus")
if err != nil {
log.Fatal(err)
}

// Video model
videoModel, err := provider.VideoModel("wan2.6-t2v")
if err != nil {
log.Fatal(err)
}

Get API Key​

  1. Sign up at Alibaba Cloud DashScope
  2. Navigate to API Keys section
  3. Create a new API key
  4. Set environment variable:
export ALIBABA_API_KEY=sk-...

Available Models​

Qwen Language Models​

Model IDContextBest For
qwen-plus32KBalanced performance and cost
qwen-turbo8KFast responses, economical
qwen3-max32KMost capable, highest quality
qwq-32b32KComplex reasoning with thinking
qwen-vl-max32KVision + language understanding

Wan Video Models​

Provider.VideoModel accepts any model ID (there is no fixed allowlist); the constants below (pkg/providers/alibaba/video_model_ids.go) cover the documented catalog. An empty model ID defaults to wan2.6-t2v.

Model IDGo constantTypeBest For
wan2.5-t2v-previewVideoModelWan2_5T2VPreviewText-to-videoPreview quality (720p)
wan2.6-t2vVideoModelWan2_6T2VText-to-videoHigh quality generation
wan2.7-t2vVideoModelWan2_7T2VText-to-videoLatest generation
wan2.6-i2vVideoModelWan2_6I2VImage-to-videoAnimate static images
wan2.6-i2v-flashVideoModelWan2_6I2VFlashImage-to-videoFaster generation
wan2.6-r2vVideoModelWan2_6R2VReference-to-videoStyle transfer from reference
wan2.6-r2v-flashVideoModelWan2_6R2VFlashReference-to-videoFaster style transfer
wan2.7-r2vVideoModelWan2_7R2VReference-to-videoLatest reference-to-video
wan3.0-videoVideoModelWan3VideoAll-in-oneOne model ID serves text-, image-, and reference-to-video, and supports custom aspect ratios

wan2.7 and wan3 are new this cycle — wan3.0-video in particular replaces separate t2v/i2v/r2v model IDs with one model that branches on which inputs (Prompt only, Image, or InputReferences) you pass.

Embedding Models​

Alibaba DashScope text embeddings are exposed through Provider.EmbeddingModel and Provider.Embedding.

Model IDGo constantNotes
text-embedding-v4alibaba.AlibabaEmbeddingTextV4Dense, sparse, or dense+sparse output
text-embedding-v3alibaba.AlibabaEmbeddingTextV3Dense text embeddings
embeddingModel, err := provider.EmbeddingModel(alibaba.AlibabaEmbeddingTextV4)
if err != nil {
log.Fatal(err)
}

dimension := 1024.0
result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{
Model: embeddingModel,
Inputs: []string{"first document", "second document"},
ProviderOptions: map[string]interface{}{
"alibaba": alibaba.AlibabaEmbeddingModelOptions{
TextType: "document",
Dimension: &dimension,
OutputType: alibaba.AlibabaEmbeddingOutputDense,
},
},
})

OutputType accepts dense, sparse, or dense&sparse. Sparse-only responses are reported as unsupported for the SDK's dense embedding result contract; dense+sparse responses preserve sparse vectors in provider metadata.

Provider-Specific Features​

Thinking/Reasoning​

Enable reasoning for complex problems with token tracking:

result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: types.Prompt{
Text: "Solve this logic puzzle: ...",
},
ProviderOptions: map[string]interface{}{
"alibaba": map[string]interface{}{
"enableThinking": true,
"thinkingBudget": 1000, // Max reasoning tokens
},
},
})

// Access reasoning tokens
if result.Usage.OutputDetails != nil && result.Usage.OutputDetails.ReasoningTokens != nil {
fmt.Printf("Reasoning tokens used: %d\n", *result.Usage.OutputDetails.ReasoningTokens)
}

Prompt Caching​

Automatic caching for repeated content with cost savings:

// System prompts are automatically cached
prompt := types.Prompt{
System: "You are an expert software architect...", // Cached
Text: "Design a microservices system",
}

result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: prompt,
})

// Check cache usage
if result.Usage.InputDetails != nil {
if result.Usage.InputDetails.CacheReadTokens != nil {
fmt.Printf("Cache hit: %d tokens saved!\n", *result.Usage.InputDetails.CacheReadTokens)
}
if result.Usage.InputDetails.CacheWriteTokens != nil {
fmt.Printf("Cached: %d tokens for future use\n", *result.Usage.InputDetails.CacheWriteTokens)
}
}

Vision Capabilities​

Image understanding with Qwen VL models:

messages := []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.FileContent{MediaType: "image/jpeg", URL: "https://example.com/photo.jpg"},
types.TextContent{Text: "What's in this image?"},
},
},
}

visionModel, _ := provider.LanguageModel("qwen-vl-max")
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: visionModel,
Messages: messages,
})

Tool Calling​

Function calling with single or parallel execution:

tools := []types.Tool{
{
Name: "get_weather",
Description: "Get current weather for a location",
Parameters: map[string]interface{}{
"type": "object",
"properties": map[string]interface{}{
"location": map[string]interface{}{
"type": "string",
"description": "City name",
},
},
"required": []string{"location"},
},
},
}

result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: types.Prompt{Text: "What's the weather in Paris?"},
Tools: tools,
ProviderOptions: map[string]interface{}{
"alibaba": map[string]interface{}{
"parallelToolCalls": true, // Enable parallel execution
},
},
})

// Handle tool calls
if len(result.ToolCalls) > 0 {
// Execute tools and send results back
}

Video Generation​

Text-to-Video​

videoModel, _ := provider.VideoModel("wan2.6-t2v")
duration := 5.0

result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
Prompt: "A golden retriever playing in a park, sunny day",
AspectRatio: "16:9",
Duration: &duration,
})

if len(result.Videos) > 0 {
fmt.Printf("Video URL: %s\n", result.Videos[0].URL)
}

Image-to-Video​

Animate static images:

videoModel, _ := provider.VideoModel("wan2.6-i2v-flash")
duration := 6.0

result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
Prompt: "Add gentle camera movement and natural motion",
AspectRatio: "16:9",
Duration: &duration,
Image: &provider.VideoModelV3File{
Type: "url",
URL: "https://example.com/landscape.jpg",
},
})

Reference-to-Video (Style Transfer)​

Apply style from reference image to new content:

videoModel, _ := provider.VideoModel("wan2.6-r2v")
duration := 5.0

result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
Prompt: "A cat walking through a garden",
AspectRatio: "16:9",
Duration: &duration,
Image: &provider.VideoModelV3File{
Type: "url",
URL: "https://example.com/anime-style.jpg", // Reference style
},
})

Async Video​

VideoModel also implements provider.VideoModelStarter / VideoModelStatusChecker, so it works with ai.ExperimentalStartVideo / ai.ExperimentalGetVideoStatus instead of the polling ai.GenerateVideo. The default poll interval/timeout are 5s / 10 minutes, and the provider-returned task ID is path-encoded when polling status.

started, err := ai.ExperimentalStartVideo(ctx, ai.StartVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{Text: "A golden retriever playing in a park"},
})
if err != nil {
log.Fatal(err)
}

status, err := ai.ExperimentalGetVideoStatus(ctx, videoModel, ai.GetVideoStatusOptions{
Operation: started.Operation,
})

Examples​

Basic Text Generation​

package main

import (
"context"
"fmt"
"log"
"os"

"github.com/digitallysavvy/go-ai/pkg/provider"
"github.com/digitallysavvy/go-ai/pkg/provider/types"
"github.com/digitallysavvy/go-ai/pkg/providers/alibaba"
)

func main() {
cfg, err := alibaba.NewConfig(os.Getenv("ALIBABA_API_KEY"))
if err != nil {
log.Fatal(err)
}

prov := alibaba.New(cfg)
model, _ := prov.LanguageModel("qwen-plus")

result, err := model.DoGenerate(context.Background(),
&provider.GenerateOptions{
Prompt: types.Prompt{
Text: "Write a haiku about artificial intelligence.",
},
})
if err != nil {
log.Fatal(err)
}

fmt.Println(result.Text)
}

Reasoning with Thinking​

model, _ := provider.LanguageModel("qwq-32b")

result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: types.Prompt{
Text: `Solve this logic puzzle:
Three friends - Alice, Bob, and Carol - each have different pets.
- Alice doesn't have a dog
- Bob is allergic to cats
- Carol lives in an apartment that doesn't allow fish
Who has which pet?`,
},
ProviderOptions: map[string]interface{}{
"alibaba": map[string]interface{}{
"enableThinking": true,
"thinkingBudget": 1000,
},
},
})

Video Generation​

videoModel, _ := provider.VideoModel("wan2.6-t2v")
duration := 5.0

result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
Prompt: "A person dancing in the rain",
AspectRatio: "9:16", // Vertical for mobile
Duration: &duration,
})

Best Practices​

  1. Use Prompt Caching

    • Cache long system prompts for cost savings
    • Second request with same system prompt uses cached tokens
    • Significant cost reduction for repeated queries
  2. Choose the Right Model

    • qwen-turbo: Simple tasks, fast responses
    • qwen-plus: Balanced for most use cases
    • qwen3-max: Complex reasoning, highest quality
    • qwq-32b: Logic puzzles, math problems
  3. Video Generation

    • Text-to-video takes 1-2 minutes
    • Use flash variants when speed > quality
    • Aspect ratios: 16:9 (landscape), 9:16 (portrait), 1:1 (square)
    • Reference-to-video for consistent visual style
  4. Thinking Budget

    • Set limits to control costs
    • Monitor reasoning token usage
    • Higher budgets for complex problems

Rate Limits & Pricing​

Rate Limits​

Check Alibaba Cloud DashScope documentation for current limits.

Cost Optimization​

  • Enable prompt caching for repeated content
  • Use qwen-turbo for simple tasks
  • Set thinking budget limits for reasoning models
  • Use flash video variants when appropriate

API Endpoints​

  • Chat API: https://dashscope-intl.aliyuncs.com/compatible-mode/v1 (OpenAI-compatible)
  • Video API: https://dashscope-intl.aliyuncs.com (DashScope native)

Complete Examples​

See the Alibaba provider examples directory for 8 complete examples:

  1. Basic chat
  2. Thinking/reasoning
  3. Prompt caching
  4. Text-to-video
  5. Image-to-video
  6. Vision chat
  7. Tool calling
  8. Reference-to-video

Workflow Serialization​

Alibaba embedding models can cross a workflow boundary with provider.SerializeEmbeddingModel / DeserializeEmbeddingModel (language models could already be serialized with providerutils.SerializeModel / DeserializeModel). See Provider Serialization for the mechanism; video models are not yet serializable.

See Also​