Skip to main content

OpenAI-compatible providers

Go does not expose a separate openai-compatible provider package. OpenAI-compatible request behavior is shared by provider implementations such as Groq, Together, Moonshot, Alibaba, and DeepSeek; use each provider's package and provider-specific option key.

Setup​

Configure the provider package that matches the target API. Providers that use the OpenAI-compatible option resolver accept camelCase provider options and keep deprecated fallback keys only for compatibility warnings.

May 2026 parity updates​

Provider options should use camelCase keys to match the TypeScript SDK. Prefer the concrete provider key, for example "deepseek" or "groq", rather than a generic "openai-compatible" key.

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Write a short status update.",
ProviderOptions: map[string]interface{}{
"deepseek": map[string]interface{}{
"reasoningEffort": "low",
},
},
})

Assistant Content Serialization​

Assistant messages with tool calls no longer use one shape for every OpenAI-compatible provider — each matches its own API's expectations:

  • Groq sends the text content verbatim, including an empty string (""), never null.
  • Cohere omits the content key entirely on a tool-call assistant turn.
  • Other OpenAI-compatible providers keep the previous null-content shape for tool-call turns.

Video Parts​

video/* file parts can now be sent as OpenAI-compatible video_url content parts. This is controlled internally by prompt.ToOpenAIMessagesOptions.AllowVideo (default false in the plain openai package, since not every OpenAI-compatible API accepts video content): Fireworks, GMI Cloud, Z.AI, Together, Baseten, Cerebras, and DeepInfra all set it to true for their own requests, so passing a video/* types.FileContent to those providers just works —

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model, // a Fireworks, GMI Cloud, Z.AI, Together, Baseten, Cerebras, or DeepInfra model
Messages: []types.Message{{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "What happens in this clip?"},
types.FileContent{MediaType: "video/mp4", URL: "https://example.com/clip.mp4"},
},
}},
})

— there is no per-call or per-provider AllowVideo setting to opt into; the plain openai package used directly against a custom BaseURL still omits video parts unless the endpoint is one of the providers above.

Multipart Tool Result Content​

By default, a tool result's structured content output (text and file blocks, e.g. from toModelOutput) is sent as a single JSON-stringified content field, since not every OpenAI-compatible endpoint accepts structured tool-result content. Endpoints documented to extend Chat Completions with structured tool-result content can opt in internally via prompt.ToOpenAIMessagesOptions.SupportsMultiPartToolContent, which instead sends an array of content parts (text/image_url/video_url/input_audio/file) so a multimodal model can see an image a tool returned, rather than its JSON-stringified representation:

toolMessages := prompt.ToOpenAIMessages(messages, prompt.ToOpenAIMessagesOptions{
SupportsMultiPartToolContent: true,
})

There is no per-call setting; a provider implementation sets this only for an endpoint it has confirmed accepts the structured shape.

Streams Truncated Without A Finish Reason​

If the underlying SSE stream ends without ever sending a finish_reason (a truncated or misbehaving upstream), the shared OpenAICompatStream base (used by Fireworks, Together, Mistral, Ollama, XAI, Azure, Perplexity, and others) now emits an InvalidResponseDataError error chunk ("Response stream ended without a finish reason.") followed by a finish chunk with reason "error", instead of ending the stream silently as in v0.4.0.

Trailing Usage Chunks​

Streamed responses on providers built on the shared base now report token usage: the trailing stream_options.include_usage chunk that v0.4.0 dropped is applied to the final usage totals.

Unreliable Streamed Tool-Call Labels​

Some OpenAI-compatible gateways omit, blank, repeat, or change streamed tool-call id and index values mid-stream, and may strip the type field from individual deltas. The shared tool-call tracker (used by OpenAI, OpenAI-compatible providers, Groq, DeepSeek, Alibaba, Mistral, and MoonshotAI) now correlates fragments using the wire ID, index, function name, incremental argument structure, and explicit-start evidence together, rather than treating any single label as authoritative:

  • Tool calls that reuse the same index (or the same id, or both) are kept distinct once another signal — a new function name, or a fresh JSON object/array start — shows the provider has moved on to a new call.
  • A continuation fragment with a conflicting or unattributable label (no matching id/index, and more than one call it could belong to) is dropped instead of being merged into the wrong call.
  • A tool call whose function name only arrives on a later, correlated delta (rather than on the delta that first introduces its id/index) is still reported correctly, with its buffered arguments included.

This requires no code changes — it is a correctness fix to the existing tool-input-start / tool-input-delta / tool-input-end / tool-call chunk sequence for well-formed streams. A tool call whose function name never arrives by the end of the stream still surfaces as an InvalidResponseDataError error chunk, exactly as before.