Alibaba Cloud Provider
Alibaba Cloud provides powerful Qwen language models with strong Chinese language capabilities and Wan video generation models. Known for cost-effective prompt caching, reasoning capabilities, and innovative video generation features.
Setup
Installation
import (
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/alibaba"
)
Configuration
provider := alibaba.New(alibaba.Config{
APIKey: os.Getenv("ALIBABA_API_KEY"),
})
// Chat model
model, err := provider.LanguageModel("qwen-plus")
if err != nil {
log.Fatal(err)
}
// Video model
videoModel, err := provider.VideoModel("wan2.6-t2v")
if err != nil {
log.Fatal(err)
}
Get API Key
- Sign up at Alibaba Cloud DashScope
- Navigate to API Keys section
- Create a new API key
- Set environment variable:
export ALIBABA_API_KEY=sk-...
Available Models
Qwen Language Models
| Model ID | Context | Best For |
|---|---|---|
| qwen-plus | 32K | Balanced performance and cost |
| qwen-turbo | 8K | Fast responses, economical |
| qwen3-max | 32K | Most capable, highest quality |
| qwq-32b | 32K | Complex reasoning with thinking |
| qwen-vl-max | 32K | Vision + language understanding |
Wan Video Models
Provider.VideoModel accepts any model ID (there is no fixed allowlist);
the constants below (pkg/providers/alibaba/video_model_ids.go) cover the
documented catalog. An empty model ID defaults to wan2.6-t2v.
| Model ID | Go constant | Type | Best For |
|---|---|---|---|
| wan2.5-t2v-preview | VideoModelWan2_5T2VPreview | Text-to-video | Preview quality (720p) |
| wan2.6-t2v | VideoModelWan2_6T2V | Text-to-video | High quality generation |
| wan2.7-t2v | VideoModelWan2_7T2V | Text-to-video | Latest generation |
| wan2.6-i2v | VideoModelWan2_6I2V | Image-to-video | Animate static images |
| wan2.6-i2v-flash | VideoModelWan2_6I2VFlash | Image-to-video | Faster generation |
| wan2.6-r2v | VideoModelWan2_6R2V | Reference-to-video | Style transfer from reference |
| wan2.6-r2v-flash | VideoModelWan2_6R2VFlash | Reference-to-video | Faster style transfer |
| wan2.7-r2v | VideoModelWan2_7R2V | Reference-to-video | Latest reference-to-video |
| wan3.0-video | VideoModelWan3Video | All-in-one | One model ID serves text-, image-, and reference-to-video, and supports custom aspect ratios |
wan2.7 and wan3 are new this cycle — wan3.0-video in particular
replaces separate t2v/i2v/r2v model IDs with one model that branches on
which inputs (Prompt only, Image, or InputReferences) you pass.
Embedding Models
Alibaba DashScope text embeddings are exposed through Provider.EmbeddingModel and Provider.Embedding.
| Model ID | Go constant | Notes |
|---|---|---|
| text-embedding-v4 | alibaba.AlibabaEmbeddingTextV4 | Dense, sparse, or dense+sparse output |
| text-embedding-v3 | alibaba.AlibabaEmbeddingTextV3 | Dense text embeddings |
embeddingModel, err := provider.EmbeddingModel(alibaba.AlibabaEmbeddingTextV4)
if err != nil {
log.Fatal(err)
}
dimension := 1024.0
result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{
Model: embeddingModel,
Inputs: []string{"first document", "second document"},
ProviderOptions: map[string]interface{}{
"alibaba": alibaba.AlibabaEmbeddingModelOptions{
TextType: "document",
Dimension: &dimension,
OutputType: alibaba.AlibabaEmbeddingOutputDense,
},
},
})
OutputType accepts dense, sparse, or dense&sparse. Sparse-only responses are reported as unsupported for the SDK's dense embedding result contract; dense+sparse responses preserve sparse vectors in provider metadata.
Provider-Specific Features
Thinking/Reasoning
Enable reasoning for complex problems with token tracking:
result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: types.Prompt{
Text: "Solve this logic puzzle: ...",
},
ProviderOptions: map[string]interface{}{
"alibaba": map[string]interface{}{
"enableThinking": true,
"thinkingBudget": 1000, // Max reasoning tokens
},
},
})
// Access reasoning tokens
if result.Usage.OutputDetails != nil && result.Usage.OutputDetails.ReasoningTokens != nil {
fmt.Printf("Reasoning tokens used: %d\n", *result.Usage.OutputDetails.ReasoningTokens)
}
Prompt Caching
Automatic caching for repeated content with cost savings:
// System prompts are automatically cached
prompt := types.Prompt{
System: "You are an expert software architect...", // Cached
Text: "Design a microservices system",
}
result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: prompt,
})
// Check cache usage
if result.Usage.InputDetails != nil {
if result.Usage.InputDetails.CacheReadTokens != nil {
fmt.Printf("Cache hit: %d tokens saved!\n", *result.Usage.InputDetails.CacheReadTokens)
}
if result.Usage.InputDetails.CacheWriteTokens != nil {
fmt.Printf("Cached: %d tokens for future use\n", *result.Usage.InputDetails.CacheWriteTokens)
}
}
Vision Capabilities
Image understanding with Qwen VL models:
messages := []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.FileContent{MediaType: "image/jpeg", URL: "https://example.com/photo.jpg"},
types.TextContent{Text: "What's in this image?"},
},
},
}
visionModel, _ := provider.LanguageModel("qwen-vl-max")
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: visionModel,
Messages: messages,
})
Tool Calling
Function calling with single or parallel execution:
tools := []types.Tool{
{
Name: "get_weather",
Description: "Get current weather for a location",
Parameters: map[string]interface{}{
"type": "object",
"properties": map[string]interface{}{
"location": map[string]interface{}{
"type": "string",
"description": "City name",
},
},
"required": []string{"location"},
},
},
}
result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: types.Prompt{Text: "What's the weather in Paris?"},
Tools: tools,
ProviderOptions: map[string]interface{}{
"alibaba": map[string]interface{}{
"parallelToolCalls": true, // Enable parallel execution
},
},
})
// Handle tool calls
if len(result.ToolCalls) > 0 {
// Execute tools and send results back
}
Video Generation
Text-to-Video
videoModel, _ := provider.VideoModel("wan2.6-t2v")
duration := 5.0
result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
Prompt: "A golden retriever playing in a park, sunny day",
AspectRatio: "16:9",
Duration: &duration,
})
if len(result.Videos) > 0 {
fmt.Printf("Video URL: %s\n", result.Videos[0].URL)
}
Image-to-Video
Animate static images:
videoModel, _ := provider.VideoModel("wan2.6-i2v-flash")
duration := 6.0
result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
Prompt: "Add gentle camera movement and natural motion",
AspectRatio: "16:9",
Duration: &duration,
Image: &provider.VideoModelV3File{
Type: "url",
URL: "https://example.com/landscape.jpg",
},
})
Reference-to-Video (Style Transfer)
Apply style from reference image to new content:
videoModel, _ := provider.VideoModel("wan2.6-r2v")
duration := 5.0
result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
Prompt: "A cat walking through a garden",
AspectRatio: "16:9",
Duration: &duration,
Image: &provider.VideoModelV3File{
Type: "url",
URL: "https://example.com/anime-style.jpg", // Reference style
},
})
Async Video
VideoModel also implements provider.VideoModelStarter /
VideoModelStatusChecker, so it works with ai.ExperimentalStartVideo /
ai.ExperimentalGetVideoStatus instead of the polling ai.GenerateVideo.
The default poll interval/timeout are 5s / 10 minutes, and the
provider-returned task ID is path-encoded when polling status.
started, err := ai.ExperimentalStartVideo(ctx, ai.StartVideoOptions{
Model: videoModel,
Prompt: ai.VideoPrompt{Text: "A golden retriever playing in a park"},
})
if err != nil {
log.Fatal(err)
}
status, err := ai.ExperimentalGetVideoStatus(ctx, videoModel, ai.GetVideoStatusOptions{
Operation: started.Operation,
})
Examples
Basic Text Generation
package main
import (
"context"
"fmt"
"log"
"os"
"github.com/digitallysavvy/go-ai/pkg/provider"
"github.com/digitallysavvy/go-ai/pkg/provider/types"
"github.com/digitallysavvy/go-ai/pkg/providers/alibaba"
)
func main() {
cfg, err := alibaba.NewConfig(os.Getenv("ALIBABA_API_KEY"))
if err != nil {
log.Fatal(err)
}
prov := alibaba.New(cfg)
model, _ := prov.LanguageModel("qwen-plus")
result, err := model.DoGenerate(context.Background(),
&provider.GenerateOptions{
Prompt: types.Prompt{
Text: "Write a haiku about artificial intelligence.",
},
})
if err != nil {
log.Fatal(err)
}
fmt.Println(result.Text)
}
Reasoning with Thinking
model, _ := provider.LanguageModel("qwq-32b")
result, err := model.DoGenerate(ctx, &provider.GenerateOptions{
Prompt: types.Prompt{
Text: `Solve this logic puzzle:
Three friends - Alice, Bob, and Carol - each have different pets.
- Alice doesn't have a dog
- Bob is allergic to cats
- Carol lives in an apartment that doesn't allow fish
Who has which pet?`,
},
ProviderOptions: map[string]interface{}{
"alibaba": map[string]interface{}{
"enableThinking": true,
"thinkingBudget": 1000,
},
},
})
Video Generation
videoModel, _ := provider.VideoModel("wan2.6-t2v")
duration := 5.0
result, err := videoModel.DoGenerate(ctx, &provider.VideoModelV3CallOptions{
Prompt: "A person dancing in the rain",
AspectRatio: "9:16", // Vertical for mobile
Duration: &duration,
})
Best Practices
-
Use Prompt Caching
- Cache long system prompts for cost savings
- Second request with same system prompt uses cached tokens
- Significant cost reduction for repeated queries
-
Choose the Right Model
qwen-turbo: Simple tasks, fast responsesqwen-plus: Balanced for most use casesqwen3-max: Complex reasoning, highest qualityqwq-32b: Logic puzzles, math problems
-
Video Generation
- Text-to-video takes 1-2 minutes
- Use flash variants when speed > quality
- Aspect ratios: 16:9 (landscape), 9:16 (portrait), 1:1 (square)
- Reference-to-video for consistent visual style
-
Thinking Budget
- Set limits to control costs
- Monitor reasoning token usage
- Higher budgets for complex problems
Rate Limits & Pricing
Rate Limits
Check Alibaba Cloud DashScope documentation for current limits.
Cost Optimization
- Enable prompt caching for repeated content
- Use
qwen-turbofor simple tasks - Set thinking budget limits for reasoning models
- Use flash video variants when appropriate
API Endpoints
- Chat API:
https://dashscope-intl.aliyuncs.com/compatible-mode/v1(OpenAI-compatible) - Video API:
https://dashscope-intl.aliyuncs.com(DashScope native)
Complete Examples
See the Alibaba provider examples directory for 8 complete examples:
- Basic chat
- Thinking/reasoning
- Prompt caching
- Text-to-video
- Image-to-video
- Vision chat
- Tool calling
- Reference-to-video
Workflow Serialization
Alibaba embedding models can cross a workflow boundary with
provider.SerializeEmbeddingModel / DeserializeEmbeddingModel (language
models could already be serialized with providerutils.SerializeModel /
DeserializeModel). See
Provider Serialization
for the mechanism; video models are not yet serializable.