Skip to main content

GMI Cloud Provider

GMI Cloud provides GPU inference for open-weight models over an OpenAI-compatible chat API. It only supports language models — EmbeddingModel, ImageModel, SpeechModel, TranscriptionModel, and RerankingModel all return errors.

Setup​

Installation​

import (
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/gmicloud"
)

Configuration​

provider := gmicloud.New(gmicloud.Config{
APIKey: os.Getenv("GMI_CLOUD_APIKEY"),
})

model, err := provider.LanguageModel(gmicloud.ModelDeepSeekV4FlashChat)

gmicloud.Config also accepts BaseURL (default https://api.gmi-serving.com/v1), Headers, and HTTPClient.

Get API Key​

export GMI_CLOUD_APIKEY=...

Available Models​

Model IDGo constant
deepseek-ai/DeepSeek-V4-Flash-0731gmicloud.ModelDeepSeekV4FlashChat
Qwen/Qwen3.8-Maxgmicloud.ModelQwen38Max
moonshotai/kimi-k3gmicloud.ModelKimiK3
zai-org/GLM-5.2-FP8gmicloud.ModelGLM52FP8
MiniMaxAI/MiniMax-M3gmicloud.ModelMiniMaxM3

ChatModel is an alias for LanguageModel.

Basic Usage​

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: "Explain GPU inference in one paragraph",
})

Provider Options​

GMI Cloud-specific options are passed under the "gmicloud" ProviderOptions key, resolved through the same OpenAI-compatible option resolver used by the other OpenAI-compatible providers in this SDK (see OpenAI-compatible providers).

Video Parts​

Chat requests sent through this provider set AllowVideo: true internally, so video/* file parts are sent as video_url content parts — see OpenAI-compatible providers: Video Parts.

Workflow Serialization​

GMI Cloud language models can cross a workflow boundary with providerutils.SerializeModel / DeserializeModel. See Provider Serialization for the mechanism.

See Also​