DeepInfra Provider
DeepInfra provides serverless GPU inference for open-source models with excellent performance and competitive pricing. Access hundreds of models through a simple API with no infrastructure management.
Setup
Installation
import (
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/deepinfra"
)
Configuration
provider := deepinfra.New(deepinfra.Config{
APIKey: os.Getenv("DEEPINFRA_API_KEY"),
})
model, err := provider.LanguageModel("meta-llama/Meta-Llama-3.1-70B-Instruct")
Note: The DeepInfra provider automatically fixes token counting issues for Gemini/Gemma models. See Token Counting Fix section for details.
Get API Key
- Sign up at deepinfra.com
- Get API key from dashboard
- Set environment variable:
export DEEPINFRA_API_KEY=...
Available Models
Language Models
| Model ID | Parameters | Input Price | Output Price | Best For |
|---|---|---|---|---|
| meta-llama/Meta-Llama-3.1-70B-Instruct | 70B | $0.52/1M | $0.75/1M | General purpose |
| meta-llama/Meta-Llama-3.1-8B-Instruct | 8B | $0.06/1M | $0.06/1M | Fast, cheap |
| mistralai/Mixtral-8x7B-Instruct-v0.1 | 47B | $0.24/1M | $0.24/1M | Balanced |
| Qwen/Qwen2.5-72B-Instruct | 72B | $0.35/1M | $0.40/1M | Multilingual |
| microsoft/WizardLM-2-8x22B | 141B | $0.65/1M | $0.65/1M | Complex tasks |
Embedding Models
| Model ID | Dimensions | Price | Best For |
|---|---|---|---|
| BAAI/bge-large-en-v1.5 | 1024 | $0.005/1M | English embeddings |
| sentence-transformers/all-MiniLM-L6-v2 | 384 | $0.005/1M | Fast embeddings |
Vision Models
| Model ID | Capability | Price | Best For |
|---|---|---|---|
| meta-llama/Llama-3.2-90B-Vision-Instruct | Vision | $0.50/1M | Image understanding |
Provider-Specific Features
Token Counting Fix
The Go-AI SDK automatically corrects token counting issues for Gemini and Gemma models on DeepInfra.
The Issue:
DeepInfra's API has a bug where reasoning_tokens are not included in completion_tokens for Gemini/Gemma models with thinking capabilities. This violates the OpenAI-compatible spec and can result in negative token counts.
Example of Incorrect API Response:
{
"completion_tokens": 84, // Text-only tokens
"completion_tokens_details": {
"reasoning_tokens": 1081 // Not included in completion_tokens!
}
}
This would incorrectly calculate: text_tokens = 84 - 1081 = -997 ❌
The Fix: The Go-AI SDK automatically detects and corrects this:
// When reasoning_tokens > completion_tokens, automatically add them
// corrected_completion_tokens = 84 + 1081 = 1165 ✅
Affected Models:
google/gemini-2.0-flash-thinking-exp-1219google/gemini-2.0-flash-thinking-exp:freegoogle/gemma-2-9b-it- Other Gemini/Gemma variants with reasoning support
Usage:
provider := deepinfra.New(deepinfra.Config{
APIKey: os.Getenv("DEEPINFRA_API_KEY"),
})
model, err := provider.LanguageModel("google/gemini-2.0-flash-thinking-exp-1219")
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain quantum computing"})
// Token usage is automatically corrected
if result.Usage.OutputDetails != nil {
textTokens := result.Usage.OutputDetails.TextTokens // Correct text tokens
reasoningTokens := result.Usage.OutputDetails.ReasoningTokens // Correct reasoning tokens
totalOutput := result.Usage.OutputTokens // Correct total (text + reasoning)
}
No configuration needed - the fix is automatic and transparent. See examples/deepinfra-token-fix for a complete example.
Serverless Infrastructure
No infrastructure management required:
// Automatic scaling and GPU allocation
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
// DeepInfra handles all infrastructure
Pay-Per-Use
Pay only for actual inference time:
// No idle costs, no minimum commitments
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
// Charged only for tokens used
Wide Model Selection
Access hundreds of open-source models:
// Browse available models
models := []string{
"meta-llama/Meta-Llama-3.1-70B-Instruct",
"mistralai/Mixtral-8x7B-Instruct-v0.1",
"google/gemma-2-9b-it",
"microsoft/Phi-3-medium-128k-instruct",
"Qwen/Qwen2.5-72B-Instruct",
}
for _, modelID := range models {
model, err := provider.LanguageModel(modelID)
// Test different models easily
}
Examples
Basic Text Generation
package main
import (
"context"
"fmt"
"log"
"os"
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/deepinfra"
)
func main() {
provider := deepinfra.New(deepinfra.Config{
APIKey: os.Getenv("DEEPINFRA_API_KEY"),
})
model, err := provider.LanguageModel("meta-llama/Meta-Llama-3.1-70B-Instruct")
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{
Model: model,
Prompt: "Explain serverless GPU inference",
})
if err != nil {
log.Fatal(err)
}
fmt.Println(result.Text)
fmt.Printf("Cost: $%.6f\n",
calculateCost(result.Usage.GetInputTokens(),
result.Usage.GetOutputTokens()))
}
func calculateCost(inputTokens, outputTokens int64) float64 {
inputCost := float64(inputTokens) * 0.52 / 1_000_000
outputCost := float64(outputTokens) * 0.75 / 1_000_000
return inputCost + outputCost
}
Model Comparison
func compareModels(prompt string) {
models := map[string]string{
"Llama 3.1 70B": "meta-llama/Meta-Llama-3.1-70B-Instruct",
"Mixtral 8x7B": "mistralai/Mixtral-8x7B-Instruct-v0.1",
"Qwen 2.5 72B": "Qwen/Qwen2.5-72B-Instruct",
}
for name, modelID := range models {
model, err := provider.LanguageModel(modelID)
if err != nil {
log.Printf("Failed to load %s: %v", name, err)
continue
}
start := time.Now()
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
if err != nil {
log.Printf("Failed %s: %v", name, err)
continue
}
elapsed := time.Since(start)
fmt.Printf("\n=== %s ===\n", name)
fmt.Printf("Response: %s\n", result.Text)
fmt.Printf("Time: %v\n", elapsed)
fmt.Printf("Tokens: %d\n", result.Usage.GetTotalTokens())
}
}
Embeddings
embeddingModel, err := provider.EmbeddingModel("BAAI/bge-large-en-v1.5")
if err != nil {
log.Fatal(err)
}
texts := []string{
"DeepInfra provides serverless AI",
"GPU inference without infrastructure",
"Open-source models on demand",
}
result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{
Model: embeddingModel,
Inputs: texts,
})
if err != nil {
log.Fatal(err)
}
fmt.Printf("Generated %d embeddings of dimension %d\n",
len(result.Embeddings), len(result.Embeddings[0]))
// Use embeddings for semantic search, clustering, etc.
Streaming
stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a detailed article"})
if err != nil {
log.Fatal(err)
}
defer stream.Close()
for chunk := range stream.Chunks() {
fmt.Print(chunk.Text)
}
Best Practices
-
Model Selection
- Use Llama 3.1 70B for high quality
- Use Llama 3.1 8B for cost efficiency
- Use Mixtral for balanced performance
- Use Qwen for multilingual tasks
-
Cost Optimization
- Compare prices across models
- Use smaller models for simple tasks
- Monitor token usage
- Cache results when appropriate
-
Performance
- Serverless = no cold starts to worry about
- Automatic scaling for traffic spikes
- Geographic distribution for low latency
-
Reliability
- Implement retry logic
- Handle rate limits gracefully
- Monitor API status
Rate Limits & Pricing
Rate Limits
Varies by plan:
- Free tier: 10 requests/min
- Pay-as-you-go: Higher limits
- Enterprise: Custom limits
Pricing Comparison
func compareProviderCosts(inputTokens, outputTokens int) {
providers := map[string][2]float64{
"DeepInfra Llama 70B": {0.52 / 1_000_000, 0.75 / 1_000_000},
"DeepInfra Llama 8B": {0.06 / 1_000_000, 0.06 / 1_000_000},
"DeepInfra Mixtral": {0.24 / 1_000_000, 0.24 / 1_000_000},
"DeepInfra Qwen 72B": {0.35 / 1_000_000, 0.40 / 1_000_000},
}
fmt.Println("Cost comparison for 1M input + 1M output tokens:")
for provider, rates := range providers {
inputCost := float64(inputTokens) * rates[0]
outputCost := float64(outputTokens) * rates[1]
totalCost := inputCost + outputCost
fmt.Printf("%s: $%.2f\n", provider, totalCost)
}
}
Error Handling
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
if err != nil {
if strings.Contains(err.Error(), "rate_limit") {
log.Println("Rate limited, implement backoff")
time.Sleep(time.Second * 5)
// Retry
} else if strings.Contains(err.Error(), "model_not_found") {
log.Fatal("Model not available on DeepInfra")
} else if strings.Contains(err.Error(), "insufficient_credits") {
log.Fatal("Add credits to account")
}
log.Fatal(err)
}
Advanced Features
Custom Deployments
Deploy custom models:
// Contact DeepInfra for custom model deployments
// Bring your own fine-tuned models
API Compatibility
Works with OpenAI SDK:
// DeepInfra is OpenAI-compatible
// Easy migration from OpenAI
// Use same code, just change base URL
See Also
- API Reference: GenerateText
- DeepInfra Documentation
- DeepInfra Model Library
- Together AI Provider - Alternative
- Fireworks AI Provider - Alternative