Skip to main content

Ollama Provider

Ollama enables running large language models locally on your machine. Perfect for development, privacy-sensitive applications, and offline usage.

Setup​

Installation​

  1. Install Ollama from ollama.com
  2. Pull a model:
ollama pull llama3.1

Configuration​

Ollama has its own provider package; no API key is needed:

import (
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/ollama"
)

provider := ollama.New(ollama.Config{
// BaseURL defaults to http://localhost:11434
})

model, err := provider.LanguageModel("llama3.1")

Ollama does not support image generation, speech synthesis, transcription, or reranking — ImageModel, SpeechModel, TranscriptionModel, and RerankingModel all return errors.

Available Models​

ModelParametersSizeBest For
llama3.170B/8B40GB/4.7GBGeneral purpose
llama3.23B/1B2GB/1.3GBFast, efficient
mistral7B4.1GBBalanced
phi33.8B2.3GBSmall, capable
qwen2.572B/7B43GB/4.7GBMultilingual
codellama34B/13B/7B19GB/7.4GB/3.8GBCode generation

Specialized Models​

ModelSizeBest For
nomic-embed-text137MEmbeddings
llava7BVision
stable-code3BCode completion

Examples​

Basic Text Generation​

package main

import (
"context"
"fmt"
"log"

"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/ollama"
)

func main() {
ctx := context.Background()
provider := ollama.New(ollama.Config{})

model, err := provider.LanguageModel("llama3.1")
if err != nil {
log.Fatal(err)
}

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain Ollama"})
if err != nil {
log.Fatal(err)
}

fmt.Println(result.Text)
}

Streaming​

stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a story"})
if err != nil {
log.Fatal(err)
}
defer stream.Close()

for chunk := range stream.Chunks() {
fmt.Print(chunk.Text)
}

Vision with LLaVA​

model, err := provider.LanguageModel("llava")

imageData, _ := os.ReadFile("image.jpg")

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Describe this image"},
types.FileContent{
Data: imageData,
MediaType: "image/jpeg",
},
},
},
},
})

Embeddings​

embeddingModel, err := provider.EmbeddingModel("nomic-embed-text")

texts := []string{
"Ollama runs models locally",
"Local inference is private",
}

result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{
Model: embeddingModel,
Inputs: texts,
})

Best Practices​

  1. Hardware Requirements

    • 8GB RAM minimum for 7B models
    • 16GB RAM for 13B models
    • 32GB+ RAM for 70B models
    • GPU recommended for better performance
  2. Model Selection

    • Start with smaller models (7B)
    • Use quantized versions for less RAM
    • Test different models for your use case
  3. Performance

    • Use GPU acceleration when available
    • Adjust context window based on needs
    • Monitor system resources
  4. Development

    • Perfect for development and testing
    • No API costs
    • Complete privacy
    • Works offline

Managing Models​

Pull Models​

# Pull latest version
ollama pull llama3.1

# Pull specific version
ollama pull llama3.1:70b

# Pull with tag
ollama pull llama3.1:8b-instruct-q4_0

List Models​

ollama list

Remove Models​

ollama rm llama3.1

Create Custom Model​

Create Modelfile:

FROM llama3.1
PARAMETER temperature 0.8
SYSTEM You are a helpful coding assistant

Build model:

ollama create my-coding-assistant -f Modelfile

Configuration Options​

GPU Configuration​

// Ollama automatically uses GPU if available
// Configure via OLLAMA_NUM_GPU environment variable

Context Window​

maxTokens := 4096
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: prompt,
MaxTokens: &maxTokens, // Adjust based on model
})

Remote Ollama​

provider := ollama.New(ollama.Config{
BaseURL: "http://remote-server:11434",
})

Error Handling​

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
if err != nil {
if strings.Contains(err.Error(), "connection refused") {
log.Fatal("Ollama server not running - run: ollama serve")
}
if strings.Contains(err.Error(), "model not found") {
log.Fatal("Model not found - run: ollama pull llama3.1")
}
log.Fatal(err)
}

Performance Tips​

  1. Use Quantized Models

    • q4_0: 4-bit quantization (smallest)
    • q5_0: 5-bit quantization (balanced)
    • q8_0: 8-bit quantization (high quality)
  2. Optimize Context

    • Reduce context window for faster inference
    • Clear conversation history periodically
  3. GPU Acceleration

    • Ensure CUDA/ROCm installed
    • Monitor GPU memory usage
    • Adjust num_gpu parameter

See Also​