# Prodia Provider

[Prodia](https://prodia.com/) is a fast inference platform for generative AI, offering high-speed image generation with FLUX and Stable Diffusion models, plus video generation with Wan 2.2 Lightning models.

## Setup

### Installation

```go
import (
    "github.com/digitallysavvy/go-ai/pkg/providers/prodia"
)
```

### Configuration

```go
provider := prodia.New(prodia.Config{
    APIKey: os.Getenv("PRODIA_TOKEN"),
})
```

You can use the following optional settings to customize the Prodia provider instance:

- **APIKey** `string` — API key sent as a Bearer token in the `Authorization` header. Defaults to the `PRODIA_TOKEN` environment variable.
- **BaseURL** `string` — Custom URL prefix for API calls. Defaults to `https://inference.prodia.com/v2`.
- **Headers** `map[string]string` — Custom headers to include in all requests.

### Get API keys

1. Sign up at [prodia.com](https://prodia.com)
2. Navigate to your API settings
3. Generate an API token
4. Set the environment variable:

```bash
export PRODIA_TOKEN=your-api-key
```

## Image model

Prodia's image model handles text-to-image and image-to-image generation. You create an image model using the `ImageModel()` factory method.

### Basic usage

```go
package main

import (
    "context"
    "fmt"
    "log"
    "github.com/digitallysavvy/go-ai/pkg/provider"
    "github.com/digitallysavvy/go-ai/pkg/providers/prodia"
)

func main() {
    prov := prodia.New(prodia.Config{})

    model, err := prov.ImageModel("inference.flux-fast.schnell.txt2img.v2")
    if err != nil {
        log.Fatal(err)
    }

    result, err := model.DoGenerate(context.Background(), &provider.ImageGenerateOptions{
        Prompt: "A cat wearing an intricate robe",
        Size:   "1024x1024",
    })
    if err != nil {
        log.Fatal(err)
    }

    fmt.Printf("Generated image: %d bytes, type: %s\n", len(result.Image), result.MimeType)
}
```

### Model capabilities

| Model ID | Description |
|----------|-------------|
| `inference.flux-fast.schnell.txt2img.v2` | Fast FLUX Schnell model for text-to-image generation |
| `inference.flux.schnell.txt2img.v2` | FLUX Schnell model for text-to-image generation |

> **Note:** Prodia supports additional model IDs. Check the [Prodia documentation](https://docs.prodia.com/) for the full list of available models.

## Language Model

Prodia's language model handles img2img generation via multipart form-data. It accepts a text prompt and an optional input image and returns a transformed image along with an optional text description. Because Prodia does not expose SSE for this endpoint, `DoStream` wraps `DoGenerate` and emits the result as a synchronous chunk sequence.

### Model IDs

| Model ID | Description |
|----------|-------------|
| `inference.nano-banana.img2img.v2` | Nano Banana image-to-image (default) |

### DoGenerate Example

```go
package main

import (
    "context"
    "fmt"
    "log"

    "github.com/digitallysavvy/go-ai/pkg/provider"
    "github.com/digitallysavvy/go-ai/pkg/provider/types"
    "github.com/digitallysavvy/go-ai/pkg/providers/prodia"
)

func main() {
    prov := prodia.New(prodia.Config{})

    model, err := prov.LanguageModel("inference.nano-banana.img2img.v2")
    if err != nil {
        log.Fatal(err)
    }

    result, err := model.DoGenerate(context.Background(),
        &provider.GenerateOptions{
            Prompt: types.Prompt{
                Messages: []types.Message{
                    {
                        Role: types.RoleUser,
                        Content: []types.ContentPart{
                            types.TextContent{Text: "Transform this image into a watercolor painting"},
                            types.ImageContent{Image: imageBytes, MimeType: "image/png"},
                        },
                    },
                },
            },
            ProviderOptions: map[string]interface{}{
                "prodia": map[string]interface{}{
                    "aspectRatio": "16:9",
                },
            },
        })
    if err != nil {
        log.Fatal(err)
    }

    fmt.Printf("Finish reason: %s\n", result.FinishReason)
    fmt.Printf("Content parts: %d\n", len(result.Content))
}
```

The language model does not support temperature, topP, topK, maxTokens, stopSequences, tools, toolChoice, responseFormat, or reasoning. Unsupported options emit warnings but do not cause failures.

## Video Models

Prodia supports text-to-video (T2V) and image-to-video (I2V) generation with Wan 2.2 Lightning models. Text-to-video requests are sent as JSON; image-to-video requests are sent as multipart/form-data.

### Model IDs

| Model ID | Type | Description |
|----------|------|-------------|
| `inference.wan2-2.lightning.txt2vid.v0` | T2V | Wan 2.2 Lightning text-to-video (default) |
| `inference.wan2-2.lightning.img2vid.v0` | I2V | Wan 2.2 Lightning image-to-video |

### Text-to-Video Example

```go
model, err := prov.VideoModel("inference.wan2-2.lightning.txt2vid.v0")
if err != nil {
    log.Fatal(err)
}

result, err := model.DoGenerate(context.Background(),
    &provider.VideoModelV3CallOptions{
        Prompt:      "A sunrise over mountains in 4K",
        AspectRatio: "16:9",
    })
if err != nil {
    log.Fatal(err)
}

fmt.Printf("Video MIME: %s\n", result.Videos[0].MediaType)
fmt.Printf("Video bytes: %d\n", len(result.Videos[0].Binary))
```

### Image-to-Video Example

```go
model, err := prov.VideoModel("inference.wan2-2.lightning.img2vid.v0")
if err != nil {
    log.Fatal(err)
}

result, err := model.DoGenerate(context.Background(),
    &provider.VideoModelV3CallOptions{
        Prompt:      "The cat slowly turns its head",
        AspectRatio: "9:16",
        Image: &provider.VideoModelV3File{
            Type:      "file",
            Data:      imageBytes,
            MediaType: "image/png",
        },
    })
if err != nil {
    log.Fatal(err)
}

fmt.Printf("Video bytes: %d\n", len(result.Videos[0].Binary))
```

Max videos per call is 1.

## Image size

You can specify the image size using the `Size` field in `WIDTHxHEIGHT` format. Width and height must be between 256 and 1920 pixels.

```go
result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
    Prompt: "A serene mountain landscape at sunset",
    Size:   "1024x768",
})
```

## Provider options

Prodia image models support additional options through the `ProviderOptions` map under the `"prodia"` key:

```go
steps := 4
stylePreset := "cinematic"

result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
    Prompt: "A cat wearing an intricate robe",
    Size:   "1024x768",
    ProviderOptions: map[string]interface{}{
        "prodia": map[string]interface{}{
            "steps":       steps,
            "stylePreset": stylePreset,
        },
    },
})
```

### Image model options

| Option | Type | Description |
|--------|------|-------------|
| `steps` | `int` | Number of computational iterations (1-4). More steps produce higher quality results. |
| `width` | `int` | Output width in pixels (256-1920). Overrides width from `Size`. |
| `height` | `int` | Output height in pixels (256-1920). Overrides height from `Size`. |
| `stylePreset` | `string` | Visual theme for the output image (see style presets below). |
| `loras` | `[]string` | Up to 3 LoRA model references for output augmentation. |
| `progressive` | `bool` | Return a progressive JPEG when using JPEG output. |

### Style presets

Supported values: `3d-model`, `analog-film`, `anime`, `cinematic`, `comic-book`, `digital-art`, `enhance`, `fantasy-art`, `isometric`, `line-art`, `low-poly`, `neon-punk`, `origami`, `photographic`, `pixel-art`, `texture`, `craft-clay`.

### Language and video model options

| Option | Type | Applies to | Description |
|--------|------|------------|-------------|
| `aspectRatio` | `string` | Language, Video | One of: `1:1`, `2:3`, `3:2`, `4:5`, `5:4`, `4:7`, `7:4`, `9:16`, `16:9`, `9:21`, `21:9` |
| `resolution` | `string` | Video only | Output resolution (e.g. `"480p"`, `"720p"`) |

## Seed

You can use the `Seed` parameter to get reproducible results:

```go
seed := 12345

result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
    Prompt: "A serene mountain landscape at sunset",
    Seed:   &seed,
})
```

The seed used for generation is available in the provider metadata (see below), which is useful when you do not set a specific seed and want to reproduce a result later.

## Provider metadata

The response includes provider-specific metadata in `ProviderMetadata["prodia"]`. The nesting depends on the model type:

- **Image models**: metadata is nested under `"images"` as a slice — `ProviderMetadata["prodia"]["images"][0]`
- **Video models**: metadata is nested under `"videos"` as a slice — `ProviderMetadata["prodia"]["videos"][0]`
- **Language models**: metadata is flat — `ProviderMetadata["prodia"]` directly contains the fields

Each metadata object may contain:

| Field | Type | Description |
|-------|------|-------------|
| `jobId` | `string` | Unique identifier for the generation job |
| `seed` | `float64` | Seed used for generation (useful for reproducing results) |
| `elapsed` | `float64` | Generation time in seconds |
| `iterationsPerSecond` | `float64` | Processing speed metric |
| `createdAt` | `string` | Timestamp when the job was created |
| `updatedAt` | `string` | Timestamp when the job was last updated |
| `dollars` | `float64` | Cost of the generation in USD |

```go
result, err := model.DoGenerate(ctx, &provider.ImageGenerateOptions{
    Prompt: "A serene mountain landscape at sunset",
})
if err != nil {
    log.Fatal(err)
}

// Access provider metadata
if prodiaMeta, ok := result.ProviderMetadata["prodia"].(map[string]interface{}); ok {
    if images, ok := prodiaMeta["images"].([]interface{}); ok && len(images) > 0 {
        if img, ok := images[0].(map[string]interface{}); ok {
            fmt.Println("Job ID:", img["jobId"])
            fmt.Println("Seed:", img["seed"])
            fmt.Println("Elapsed:", img["elapsed"])
        }
    }
}
```

## Workflow Serialization

Prodia image models can cross a workflow boundary with
`provider.SerializeImageModel` / `DeserializeImageModel` (language models
could already be serialized with `providerutils.SerializeModel` /
`DeserializeModel`). See
[Provider Serialization](https://goaisdk.com/docs/agents/workflow-agent.md#provider-serialization)
for the mechanism; video models are not yet serializable.

## See also

- [Provider options](https://goaisdk.com/docs/foundations/provider-options.md) — how to pass provider-specific configuration
- [KlingAI provider](https://goaisdk.com/docs/providers/klingai.md) — another video generation provider
- [Prodia documentation](https://docs.prodia.com/)
