Skip to main content

Prodia Provider

Prodia is a fast inference platform for generative AI, offering high-speed image generation with FLUX and Stable Diffusion models, plus video generation with Wan 2.2 Lightning models.

Setup​

Installation​

import (
"github.com/digitallysavvy/go-ai/pkg/providers/prodia"
)

Configuration​

provider := prodia.New(prodia.Config{
APIKey: os.Getenv("PRODIA_TOKEN"),
})

You can use the following optional settings to customize the Prodia provider instance:

  • APIKey string — API key sent as a Bearer token in the Authorization header. Defaults to the PRODIA_TOKEN environment variable.
  • BaseURL string — Custom URL prefix for API calls. Defaults to https://inference.prodia.com/v2.
  • Headers map[string]string — Custom headers to include in all requests.

Get API keys​

  1. Sign up at prodia.com
  2. Navigate to your API settings
  3. Generate an API token
  4. Set the environment variable:
export PRODIA_TOKEN=your-api-key

Image model​

Prodia's image model handles text-to-image and image-to-image generation. You create an image model using the ImageModel() factory method.

Basic usage​

package main

import (
"context"
"fmt"
"log"
"github.com/digitallysavvy/go-ai/pkg/provider"
"github.com/digitallysavvy/go-ai/pkg/providers/prodia"
)

func main() {
prov := prodia.New(prodia.Config{})

model, err := prov.ImageModel("inference.flux-fast.schnell.txt2img.v2")
if err != nil {
log.Fatal(err)
}

result, err := model.DoGenerate(context.Background(), &provider.ImageGenerateOptions{
Prompt: "A cat wearing an intricate robe",
Size: "1024x1024",
})
if err != nil {
log.Fatal(err)
}

fmt.Printf("Generated image: %d bytes, type: %s\n", len(result.Image), result.MimeType)
}

Model capabilities​

Model IDDescription
inference.flux-fast.schnell.txt2img.v2Fast FLUX Schnell model for text-to-image generation
inference.flux.schnell.txt2img.v2FLUX Schnell model for text-to-image generation

Note: Prodia supports additional model IDs. Check the Prodia documentation for the full list of available models.

Language Model​

Prodia's language model handles img2img generation via multipart form-data. It accepts a text prompt and an optional input image and returns a transformed image along with an optional text description. Because Prodia does not expose SSE for this endpoint, DoStream wraps DoGenerate and emits the result as a synchronous chunk sequence.

Model IDs​

Model IDDescription
inference.nano-banana.img2img.v2Nano Banana image-to-image (default)

DoGenerate Example​

package main

import (
"context"
"fmt"
"log"

"github.com/digitallysavvy/go-ai/pkg/provider"
"github.com/digitallysavvy/go-ai/pkg/provider/types"
"github.com/digitallysavvy/go-ai/pkg/providers/prodia"
)

func main() {
prov := prodia.New(prodia.Config{})

model, err := prov.LanguageModel("inference.nano-banana.img2img.v2")
if err != nil {
log.Fatal(err)
}

result, err := model.DoGenerate(context.Background(),
&provider.GenerateOptions{
Prompt: types.Prompt{
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Transform this image into a watercolor painting"},
types.ImageContent{Image: imageBytes, MimeType: "image/png"},
},
},
},
},
ProviderOptions: map[string]interface{}{
"prodia": map[string]interface{}{
"aspectRatio": "16:9",
},
},
})
if err != nil {
log.Fatal(err)
}

fmt.Printf("Finish reason: %s\n", result.FinishReason)
fmt.Printf("Content parts: %d\n", len(result.Content))
}

The language model does not support temperature, topP, topK, maxTokens, stopSequences, tools, toolChoice, responseFormat, or reasoning. Unsupported options emit warnings but do not cause failures.

Video Models​

Prodia supports text-to-video (T2V) and image-to-video (I2V) generation with Wan 2.2 Lightning models. Text-to-video requests are sent as JSON; image-to-video requests are sent as multipart/form-data.

Model IDs​

Model IDTypeDescription
inference.wan2-2.lightning.txt2vid.v0T2VWan 2.2 Lightning text-to-video (default)
inference.wan2-2.lightning.img2vid.v0I2VWan 2.2 Lightning image-to-video

Text-to-Video Example​

model, err := prov.VideoModel("inference.wan2-2.lightning.txt2vid.v0")
if err != nil {
log.Fatal(err)
}

result, err := model.DoGenerate(context.Background(),
&provider.VideoModelV3CallOptions{
Prompt: "A sunrise over mountains in 4K",
AspectRatio: "16:9",
})
if err != nil {
log.Fatal(err)
}

fmt.Printf("Video MIME: %s\n", result.Videos[0].MediaType)
fmt.Printf("Video bytes: %d\n", len(result.Videos[0].Binary))

Image-to-Video Example​

model, err := prov.VideoModel("inference.wan2-2.lightning.img2vid.v0")
if err != nil {
log.Fatal(err)
}

result, err := model.DoGenerate(context.Background(),
&provider.VideoModelV3CallOptions{
Prompt: "The cat slowly turns its head",
AspectRatio: "9:16",
Image: &provider.VideoModelV3File{
Type: "file",
Data: imageBytes,
MediaType: "image/png",
},
})
if err != nil {
log.Fatal(err)
}

fmt.Printf("Video bytes: %d\n", len(result.Videos[0].Binary))

Max videos per call is 1.

Image size​

You can specify the image size using the Size field in WIDTHxHEIGHT format. Width and height must be between 256 and 1920 pixels.

result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
Prompt: "A serene mountain landscape at sunset",
Size: "1024x768",
})

Provider options​

Prodia image models support additional options through the ProviderOptions map under the "prodia" key:

steps := 4
stylePreset := "cinematic"

result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
Prompt: "A cat wearing an intricate robe",
Size: "1024x768",
ProviderOptions: map[string]interface{}{
"prodia": map[string]interface{}{
"steps": steps,
"stylePreset": stylePreset,
},
},
})

Image model options​

OptionTypeDescription
stepsintNumber of computational iterations (1-4). More steps produce higher quality results.
widthintOutput width in pixels (256-1920). Overrides width from Size.
heightintOutput height in pixels (256-1920). Overrides height from Size.
stylePresetstringVisual theme for the output image (see style presets below).
loras[]stringUp to 3 LoRA model references for output augmentation.
progressiveboolReturn a progressive JPEG when using JPEG output.

Style presets​

Supported values: 3d-model, analog-film, anime, cinematic, comic-book, digital-art, enhance, fantasy-art, isometric, line-art, low-poly, neon-punk, origami, photographic, pixel-art, texture, craft-clay.

Language and video model options​

OptionTypeApplies toDescription
aspectRatiostringLanguage, VideoOne of: 1:1, 2:3, 3:2, 4:5, 5:4, 4:7, 7:4, 9:16, 16:9, 9:21, 21:9
resolutionstringVideo onlyOutput resolution (e.g. "480p", "720p")

Seed​

You can use the Seed parameter to get reproducible results:

seed := 12345

result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
Prompt: "A serene mountain landscape at sunset",
Seed: &seed,
})

The seed used for generation is available in the provider metadata (see below), which is useful when you do not set a specific seed and want to reproduce a result later.

Provider metadata​

The response includes provider-specific metadata in ProviderMetadata["prodia"]. The nesting depends on the model type:

  • Image models: metadata is nested under "images" as a slice — ProviderMetadata["prodia"]["images"][0]
  • Video models: metadata is nested under "videos" as a slice — ProviderMetadata["prodia"]["videos"][0]
  • Language models: metadata is flat — ProviderMetadata["prodia"] directly contains the fields

Each metadata object may contain:

FieldTypeDescription
jobIdstringUnique identifier for the generation job
seedfloat64Seed used for generation (useful for reproducing results)
elapsedfloat64Generation time in seconds
iterationsPerSecondfloat64Processing speed metric
createdAtstringTimestamp when the job was created
updatedAtstringTimestamp when the job was last updated
dollarsfloat64Cost of the generation in USD
result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
Prompt: "A serene mountain landscape at sunset",
})
if err != nil {
log.Fatal(err)
}

// Access provider metadata
if prodiaMeta, ok := result.ProviderMetadata["prodia"].(map[string]interface{}); ok {
if images, ok := prodiaMeta["images"].([]interface{}); ok && len(images) > 0 {
if img, ok := images[0].(map[string]interface{}); ok {
fmt.Println("Job ID:", img["jobId"])
fmt.Println("Seed:", img["seed"])
fmt.Println("Elapsed:", img["elapsed"])
}
}
}

Workflow Serialization​

Prodia image models can cross a workflow boundary with provider.SerializeImageModel / DeserializeImageModel (language models could already be serialized with providerutils.SerializeModel / DeserializeModel). See Provider Serialization for the mechanism; video models are not yet serializable.

See also​