Token usage differentiation
types.InputTokenDetails reports separate text and image input token counts on multimodal requests, so you can cost a vision call accurately instead of billing every input token at the text rate. The fields shipped in v0.2.0 and are optional — code written against v0.1.x keeps compiling and running unchanged. This applies to anyone billing or budgeting multimodal calls per token.
TextTokens and ImageTokens on InputTokenDetails
Impact: additive — existing code is unaffected.
type InputTokenDetails struct {
NoCacheTokens *int64
CacheReadTokens *int64
CacheWriteTokens *int64
TextTokens *int64 // added in v0.2.0
ImageTokens *int64 // added in v0.2.0
}
Support varies by provider:
| Provider | Populates TextTokens / ImageTokens |
|---|---|
| OpenAI, XAI, Azure, DeepSeek, Together, Fireworks, Perplexity, Groq, Mistral | Yes — parsed from the OpenAI-compatible prompt_tokens_details.text_tokens / image_tokens wire fields |
| Yes — computed from Gemini's per-modality token counts | |
| Anthropic | No — both fields stay nil; the API doesn't report a text/image split |
Before (text-rate cost for every input token):
result, err := ai.GenerateText(ctx, opts)
if err != nil {
return err
}
inputCost := float64(result.Usage.GetInputTokens()) * textInputRate
outputCost := float64(result.Usage.GetOutputTokens()) * outputTokenRate
totalCost := inputCost + outputCost
After (accurate multimodal cost, with a fallback for providers that don't report the split):
result, err := ai.GenerateText(ctx, opts)
if err != nil {
return err
}
var inputCost float64
details := result.Usage.InputDetails
if details != nil && details.TextTokens != nil && details.ImageTokens != nil {
inputCost = float64(*details.TextTokens)*textInputRate + float64(*details.ImageTokens)*imageInputRate
} else {
inputCost = float64(result.Usage.GetInputTokens()) * textInputRate
}
outputCost := float64(result.Usage.GetOutputTokens()) * outputTokenRate
totalCost := inputCost + outputCost
GetInputTokens(), GetOutputTokens(), and GetTotalTokens() are unaffected — they keep returning the total (text + image) regardless of whether the detailed breakdown is available, and return 0 instead of panicking when the underlying pointer is nil.
Example
package main
import (
"context"
"fmt"
"log"
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/provider/types"
"github.com/digitallysavvy/go-ai/pkg/providers/openai"
)
const (
textInputRate = 0.0000025 // $2.50 per 1M tokens (GPT-4o)
imageInputRate = 0.0000075 // ~$7.50 per 1M tokens (varies by image size)
outputTokenRate = 0.0000100
)
func main() {
ctx := context.Background()
prov := openai.New(openai.Config{APIKey: "your-api-key"})
model, err := prov.LanguageModel("gpt-4o")
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "What's in this image?"},
types.ImageContent{URL: "https://example.com/image.jpg"},
},
},
},
})
if err != nil {
log.Fatal(err)
}
details := result.Usage.InputDetails
if details != nil && details.TextTokens != nil && details.ImageTokens != nil {
fmt.Printf("text tokens: %d, image tokens: %d\n", *details.TextTokens, *details.ImageTokens)
}
}
Reference
- Full field list: Usage types reference