Prodia Provider
Prodia is a fast inference platform for generative AI, offering high-speed image generation with FLUX and Stable Diffusion models, plus video generation with Wan 2.2 Lightning models.
Setup
Installation
import (
"github.com/digitallysavvy/go-ai/pkg/providers/prodia"
)
Configuration
provider := prodia.New(prodia.Config{
APIKey: os.Getenv("PRODIA_TOKEN"),
})
You can use the following optional settings to customize the Prodia provider instance:
- APIKey
string— API key sent as a Bearer token in theAuthorizationheader. Defaults to thePRODIA_TOKENenvironment variable. - BaseURL
string— Custom URL prefix for API calls. Defaults tohttps://inference.prodia.com/v2. - Headers
map[string]string— Custom headers to include in all requests.
Get API keys
- Sign up at prodia.com
- Navigate to your API settings
- Generate an API token
- Set the environment variable:
export PRODIA_TOKEN=your-api-key
Image model
Prodia's image model handles text-to-image and image-to-image generation. You create an image model using the ImageModel() factory method.
Basic usage
package main
import (
"context"
"fmt"
"log"
"github.com/digitallysavvy/go-ai/pkg/provider"
"github.com/digitallysavvy/go-ai/pkg/providers/prodia"
)
func main() {
prov := prodia.New(prodia.Config{})
model, err := prov.ImageModel("inference.flux-fast.schnell.txt2img.v2")
if err != nil {
log.Fatal(err)
}
result, err := model.DoGenerate(context.Background(), &provider.ImageGenerateOptions{
Prompt: "A cat wearing an intricate robe",
Size: "1024x1024",
})
if err != nil {
log.Fatal(err)
}
fmt.Printf("Generated image: %d bytes, type: %s\n", len(result.Image), result.MimeType)
}
Model capabilities
| Model ID | Description |
|---|---|
inference.flux-fast.schnell.txt2img.v2 | Fast FLUX Schnell model for text-to-image generation |
inference.flux.schnell.txt2img.v2 | FLUX Schnell model for text-to-image generation |
Note: Prodia supports additional model IDs. Check the Prodia documentation for the full list of available models.
Language Model
Prodia's language model handles img2img generation via multipart form-data. It accepts a text prompt and an optional input image and returns a transformed image along with an optional text description. Because Prodia does not expose SSE for this endpoint, DoStream wraps DoGenerate and emits the result as a synchronous chunk sequence.
Model IDs
| Model ID | Description |
|---|---|
inference.nano-banana.img2img.v2 | Nano Banana image-to-image (default) |
DoGenerate Example
package main
import (
"context"
"fmt"
"log"
"github.com/digitallysavvy/go-ai/pkg/provider"
"github.com/digitallysavvy/go-ai/pkg/provider/types"
"github.com/digitallysavvy/go-ai/pkg/providers/prodia"
)
func main() {
prov := prodia.New(prodia.Config{})
model, err := prov.LanguageModel("inference.nano-banana.img2img.v2")
if err != nil {
log.Fatal(err)
}
result, err := model.DoGenerate(context.Background(),
&provider.GenerateOptions{
Prompt: types.Prompt{
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Transform this image into a watercolor painting"},
types.ImageContent{Image: imageBytes, MimeType: "image/png"},
},
},
},
},
ProviderOptions: map[string]interface{}{
"prodia": map[string]interface{}{
"aspectRatio": "16:9",
},
},
})
if err != nil {
log.Fatal(err)
}
fmt.Printf("Finish reason: %s\n", result.FinishReason)
fmt.Printf("Content parts: %d\n", len(result.Content))
}
The language model does not support temperature, topP, topK, maxTokens, stopSequences, tools, toolChoice, responseFormat, or reasoning. Unsupported options emit warnings but do not cause failures.
Video Models
Prodia supports text-to-video (T2V) and image-to-video (I2V) generation with Wan 2.2 Lightning models. Text-to-video requests are sent as JSON; image-to-video requests are sent as multipart/form-data.
Model IDs
| Model ID | Type | Description |
|---|---|---|
inference.wan2-2.lightning.txt2vid.v0 | T2V | Wan 2.2 Lightning text-to-video (default) |
inference.wan2-2.lightning.img2vid.v0 | I2V | Wan 2.2 Lightning image-to-video |
Text-to-Video Example
model, err := prov.VideoModel("inference.wan2-2.lightning.txt2vid.v0")
if err != nil {
log.Fatal(err)
}
result, err := model.DoGenerate(context.Background(),
&provider.VideoModelV3CallOptions{
Prompt: "A sunrise over mountains in 4K",
AspectRatio: "16:9",
})
if err != nil {
log.Fatal(err)
}
fmt.Printf("Video MIME: %s\n", result.Videos[0].MediaType)
fmt.Printf("Video bytes: %d\n", len(result.Videos[0].Binary))
Image-to-Video Example
model, err := prov.VideoModel("inference.wan2-2.lightning.img2vid.v0")
if err != nil {
log.Fatal(err)
}
result, err := model.DoGenerate(context.Background(),
&provider.VideoModelV3CallOptions{
Prompt: "The cat slowly turns its head",
AspectRatio: "9:16",
Image: &provider.VideoModelV3File{
Type: "file",
Data: imageBytes,
MediaType: "image/png",
},
})
if err != nil {
log.Fatal(err)
}
fmt.Printf("Video bytes: %d\n", len(result.Videos[0].Binary))
Max videos per call is 1.
Image size
You can specify the image size using the Size field in WIDTHxHEIGHT format. Width and height must be between 256 and 1920 pixels.
result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
Prompt: "A serene mountain landscape at sunset",
Size: "1024x768",
})
Provider options
Prodia image models support additional options through the ProviderOptions map under the "prodia" key:
steps := 4
stylePreset := "cinematic"
result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
Prompt: "A cat wearing an intricate robe",
Size: "1024x768",
ProviderOptions: map[string]interface{}{
"prodia": map[string]interface{}{
"steps": steps,
"stylePreset": stylePreset,
},
},
})
Image model options
| Option | Type | Description |
|---|---|---|
steps | int | Number of computational iterations (1-4). More steps produce higher quality results. |
width | int | Output width in pixels (256-1920). Overrides width from Size. |
height | int | Output height in pixels (256-1920). Overrides height from Size. |
stylePreset | string | Visual theme for the output image (see style presets below). |
loras | []string | Up to 3 LoRA model references for output augmentation. |
progressive | bool | Return a progressive JPEG when using JPEG output. |
Style presets
Supported values: 3d-model, analog-film, anime, cinematic, comic-book, digital-art, enhance, fantasy-art, isometric, line-art, low-poly, neon-punk, origami, photographic, pixel-art, texture, craft-clay.
Language and video model options
| Option | Type | Applies to | Description |
|---|---|---|---|
aspectRatio | string | Language, Video | One of: 1:1, 2:3, 3:2, 4:5, 5:4, 4:7, 7:4, 9:16, 16:9, 9:21, 21:9 |
resolution | string | Video only | Output resolution (e.g. "480p", "720p") |
Seed
You can use the Seed parameter to get reproducible results:
seed := 12345
result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
Prompt: "A serene mountain landscape at sunset",
Seed: &seed,
})
The seed used for generation is available in the provider metadata (see below), which is useful when you do not set a specific seed and want to reproduce a result later.
Provider metadata
The response includes provider-specific metadata in ProviderMetadata["prodia"]. The nesting depends on the model type:
- Image models: metadata is nested under
"images"as a slice —ProviderMetadata["prodia"]["images"][0] - Video models: metadata is nested under
"videos"as a slice —ProviderMetadata["prodia"]["videos"][0] - Language models: metadata is flat —
ProviderMetadata["prodia"]directly contains the fields
Each metadata object may contain:
| Field | Type | Description |
|---|---|---|
jobId | string | Unique identifier for the generation job |
seed | float64 | Seed used for generation (useful for reproducing results) |
elapsed | float64 | Generation time in seconds |
iterationsPerSecond | float64 | Processing speed metric |
createdAt | string | Timestamp when the job was created |
updatedAt | string | Timestamp when the job was last updated |
dollars | float64 | Cost of the generation in USD |
result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
Prompt: "A serene mountain landscape at sunset",
})
if err != nil {
log.Fatal(err)
}
// Access provider metadata
if prodiaMeta, ok := result.ProviderMetadata["prodia"].(map[string]interface{}); ok {
if images, ok := prodiaMeta["images"].([]interface{}); ok && len(images) > 0 {
if img, ok := images[0].(map[string]interface{}); ok {
fmt.Println("Job ID:", img["jobId"])
fmt.Println("Seed:", img["seed"])
fmt.Println("Elapsed:", img["elapsed"])
}
}
}
Workflow Serialization
Prodia image models can cross a workflow boundary with
provider.SerializeImageModel / DeserializeImageModel (language models
could already be serialized with providerutils.SerializeModel /
DeserializeModel). See
Provider Serialization
for the mechanism; video models are not yet serializable.
See also
- Provider options — how to pass provider-specific configuration
- KlingAI provider — another video generation provider
- Prodia documentation