Skip to main content

DeepInfra Provider

DeepInfra provides serverless GPU inference for open-source models with excellent performance and competitive pricing. Access hundreds of models through a simple API with no infrastructure management.

Setup​

Installation​

import (
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/deepinfra"
)

Configuration​

provider := deepinfra.New(deepinfra.Config{
APIKey: os.Getenv("DEEPINFRA_API_KEY"),
})

model, err := provider.LanguageModel("meta-llama/Meta-Llama-3.1-70B-Instruct")

Note: The DeepInfra provider automatically fixes token counting issues for Gemini/Gemma models. See Token Counting Fix section for details.

Get API Key​

  1. Sign up at deepinfra.com
  2. Get API key from dashboard
  3. Set environment variable:
export DEEPINFRA_API_KEY=...

Available Models​

Language Models​

Model IDParametersInput PriceOutput PriceBest For
meta-llama/Meta-Llama-3.1-70B-Instruct70B$0.52/1M$0.75/1MGeneral purpose
meta-llama/Meta-Llama-3.1-8B-Instruct8B$0.06/1M$0.06/1MFast, cheap
mistralai/Mixtral-8x7B-Instruct-v0.147B$0.24/1M$0.24/1MBalanced
Qwen/Qwen2.5-72B-Instruct72B$0.35/1M$0.40/1MMultilingual
microsoft/WizardLM-2-8x22B141B$0.65/1M$0.65/1MComplex tasks

Embedding Models​

Model IDDimensionsPriceBest For
BAAI/bge-large-en-v1.51024$0.005/1MEnglish embeddings
sentence-transformers/all-MiniLM-L6-v2384$0.005/1MFast embeddings

Vision Models​

Model IDCapabilityPriceBest For
meta-llama/Llama-3.2-90B-Vision-InstructVision$0.50/1MImage understanding

Provider-Specific Features​

Token Counting Fix​

The Go-AI SDK automatically corrects token counting issues for Gemini and Gemma models on DeepInfra.

The Issue: DeepInfra's API has a bug where reasoning_tokens are not included in completion_tokens for Gemini/Gemma models with thinking capabilities. This violates the OpenAI-compatible spec and can result in negative token counts.

Example of Incorrect API Response:

{
"completion_tokens": 84, // Text-only tokens
"completion_tokens_details": {
"reasoning_tokens": 1081 // Not included in completion_tokens!
}
}

This would incorrectly calculate: text_tokens = 84 - 1081 = -997 ❌

The Fix: The Go-AI SDK automatically detects and corrects this:

// When reasoning_tokens > completion_tokens, automatically add them
// corrected_completion_tokens = 84 + 1081 = 1165 ✅

Affected Models:

  • google/gemini-2.0-flash-thinking-exp-1219
  • google/gemini-2.0-flash-thinking-exp:free
  • google/gemma-2-9b-it
  • Other Gemini/Gemma variants with reasoning support

Usage:

provider := deepinfra.New(deepinfra.Config{
APIKey: os.Getenv("DEEPINFRA_API_KEY"),
})

model, err := provider.LanguageModel("google/gemini-2.0-flash-thinking-exp-1219")

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain quantum computing"})

// Token usage is automatically corrected
if result.Usage.OutputDetails != nil {
textTokens := result.Usage.OutputDetails.TextTokens // Correct text tokens
reasoningTokens := result.Usage.OutputDetails.ReasoningTokens // Correct reasoning tokens
totalOutput := result.Usage.OutputTokens // Correct total (text + reasoning)
}

No configuration needed - the fix is automatic and transparent. See examples/deepinfra-token-fix for a complete example.

Serverless Infrastructure​

No infrastructure management required:

// Automatic scaling and GPU allocation
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
// DeepInfra handles all infrastructure

Pay-Per-Use​

Pay only for actual inference time:

// No idle costs, no minimum commitments
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
// Charged only for tokens used

Wide Model Selection​

Access hundreds of open-source models:

// Browse available models
models := []string{
"meta-llama/Meta-Llama-3.1-70B-Instruct",
"mistralai/Mixtral-8x7B-Instruct-v0.1",
"google/gemma-2-9b-it",
"microsoft/Phi-3-medium-128k-instruct",
"Qwen/Qwen2.5-72B-Instruct",
}

for _, modelID := range models {
model, err := provider.LanguageModel(modelID)
// Test different models easily
}

Examples​

Basic Text Generation​

package main

import (
"context"
"fmt"
"log"
"os"

"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/deepinfra"
)

func main() {
provider := deepinfra.New(deepinfra.Config{
APIKey: os.Getenv("DEEPINFRA_API_KEY"),
})

model, err := provider.LanguageModel("meta-llama/Meta-Llama-3.1-70B-Instruct")
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateText(context.Background(), ai.GenerateTextOptions{
Model: model,
Prompt: "Explain serverless GPU inference",
})
if err != nil {
log.Fatal(err)
}

fmt.Println(result.Text)
fmt.Printf("Cost: $%.6f\n",
calculateCost(result.Usage.GetInputTokens(),
result.Usage.GetOutputTokens()))
}

func calculateCost(inputTokens, outputTokens int64) float64 {
inputCost := float64(inputTokens) * 0.52 / 1_000_000
outputCost := float64(outputTokens) * 0.75 / 1_000_000
return inputCost + outputCost
}

Model Comparison​

func compareModels(prompt string) {
models := map[string]string{
"Llama 3.1 70B": "meta-llama/Meta-Llama-3.1-70B-Instruct",
"Mixtral 8x7B": "mistralai/Mixtral-8x7B-Instruct-v0.1",
"Qwen 2.5 72B": "Qwen/Qwen2.5-72B-Instruct",
}

for name, modelID := range models {
model, err := provider.LanguageModel(modelID)
if err != nil {
log.Printf("Failed to load %s: %v", name, err)
continue
}

start := time.Now()
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
if err != nil {
log.Printf("Failed %s: %v", name, err)
continue
}

elapsed := time.Since(start)
fmt.Printf("\n=== %s ===\n", name)
fmt.Printf("Response: %s\n", result.Text)
fmt.Printf("Time: %v\n", elapsed)
fmt.Printf("Tokens: %d\n", result.Usage.GetTotalTokens())
}
}

Embeddings​

embeddingModel, err := provider.EmbeddingModel("BAAI/bge-large-en-v1.5")
if err != nil {
log.Fatal(err)
}

texts := []string{
"DeepInfra provides serverless AI",
"GPU inference without infrastructure",
"Open-source models on demand",
}

result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{
Model: embeddingModel,
Inputs: texts,
})
if err != nil {
log.Fatal(err)
}

fmt.Printf("Generated %d embeddings of dimension %d\n",
len(result.Embeddings), len(result.Embeddings[0]))

// Use embeddings for semantic search, clustering, etc.

Streaming​

stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a detailed article"})
if err != nil {
log.Fatal(err)
}
defer stream.Close()

for chunk := range stream.Chunks() {
fmt.Print(chunk.Text)
}

Best Practices​

  1. Model Selection

    • Use Llama 3.1 70B for high quality
    • Use Llama 3.1 8B for cost efficiency
    • Use Mixtral for balanced performance
    • Use Qwen for multilingual tasks
  2. Cost Optimization

    • Compare prices across models
    • Use smaller models for simple tasks
    • Monitor token usage
    • Cache results when appropriate
  3. Performance

    • Serverless = no cold starts to worry about
    • Automatic scaling for traffic spikes
    • Geographic distribution for low latency
  4. Reliability

    • Implement retry logic
    • Handle rate limits gracefully
    • Monitor API status

Rate Limits & Pricing​

Rate Limits​

Varies by plan:

  • Free tier: 10 requests/min
  • Pay-as-you-go: Higher limits
  • Enterprise: Custom limits

Pricing Comparison​

func compareProviderCosts(inputTokens, outputTokens int) {
providers := map[string][2]float64{
"DeepInfra Llama 70B": {0.52 / 1_000_000, 0.75 / 1_000_000},
"DeepInfra Llama 8B": {0.06 / 1_000_000, 0.06 / 1_000_000},
"DeepInfra Mixtral": {0.24 / 1_000_000, 0.24 / 1_000_000},
"DeepInfra Qwen 72B": {0.35 / 1_000_000, 0.40 / 1_000_000},
}

fmt.Println("Cost comparison for 1M input + 1M output tokens:")
for provider, rates := range providers {
inputCost := float64(inputTokens) * rates[0]
outputCost := float64(outputTokens) * rates[1]
totalCost := inputCost + outputCost
fmt.Printf("%s: $%.2f\n", provider, totalCost)
}
}

Error Handling​

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
if err != nil {
if strings.Contains(err.Error(), "rate_limit") {
log.Println("Rate limited, implement backoff")
time.Sleep(time.Second * 5)
// Retry
} else if strings.Contains(err.Error(), "model_not_found") {
log.Fatal("Model not available on DeepInfra")
} else if strings.Contains(err.Error(), "insufficient_credits") {
log.Fatal("Add credits to account")
}
log.Fatal(err)
}

Advanced Features​

Custom Deployments​

Deploy custom models:

// Contact DeepInfra for custom model deployments
// Bring your own fine-tuned models

API Compatibility​

Works with OpenAI SDK:

// DeepInfra is OpenAI-compatible
// Easy migration from OpenAI
// Use same code, just change base URL

See Also​