Ollama Provider
Ollama enables running large language models locally on your machine. Perfect for development, privacy-sensitive applications, and offline usage.
Setup
Installation
- Install Ollama from ollama.com
- Pull a model:
ollama pull llama3.1
Configuration
Ollama has its own provider package; no API key is needed:
import (
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/ollama"
)
provider := ollama.New(ollama.Config{
// BaseURL defaults to http://localhost:11434
})
model, err := provider.LanguageModel("llama3.1")
Ollama does not support image generation, speech synthesis, transcription, or reranking — ImageModel, SpeechModel, TranscriptionModel, and RerankingModel all return errors.
Available Models
Popular Models
| Model | Parameters | Size | Best For |
|---|---|---|---|
| llama3.1 | 70B/8B | 40GB/4.7GB | General purpose |
| llama3.2 | 3B/1B | 2GB/1.3GB | Fast, efficient |
| mistral | 7B | 4.1GB | Balanced |
| phi3 | 3.8B | 2.3GB | Small, capable |
| qwen2.5 | 72B/7B | 43GB/4.7GB | Multilingual |
| codellama | 34B/13B/7B | 19GB/7.4GB/3.8GB | Code generation |
Specialized Models
| Model | Size | Best For |
|---|---|---|
| nomic-embed-text | 137M | Embeddings |
| llava | 7B | Vision |
| stable-code | 3B | Code completion |
Examples
Basic Text Generation
package main
import (
"context"
"fmt"
"log"
"github.com/digitallysavvy/go-ai/pkg/ai"
"github.com/digitallysavvy/go-ai/pkg/providers/ollama"
)
func main() {
ctx := context.Background()
provider := ollama.New(ollama.Config{})
model, err := provider.LanguageModel("llama3.1")
if err != nil {
log.Fatal(err)
}
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain Ollama"})
if err != nil {
log.Fatal(err)
}
fmt.Println(result.Text)
}
Streaming
stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a story"})
if err != nil {
log.Fatal(err)
}
defer stream.Close()
for chunk := range stream.Chunks() {
fmt.Print(chunk.Text)
}
Vision with LLaVA
model, err := provider.LanguageModel("llava")
imageData, _ := os.ReadFile("image.jpg")
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Messages: []types.Message{
{
Role: types.RoleUser,
Content: []types.ContentPart{
types.TextContent{Text: "Describe this image"},
types.FileContent{
Data: imageData,
MediaType: "image/jpeg",
},
},
},
},
})
Embeddings
embeddingModel, err := provider.EmbeddingModel("nomic-embed-text")
texts := []string{
"Ollama runs models locally",
"Local inference is private",
}
result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{
Model: embeddingModel,
Inputs: texts,
})
Best Practices
-
Hardware Requirements
- 8GB RAM minimum for 7B models
- 16GB RAM for 13B models
- 32GB+ RAM for 70B models
- GPU recommended for better performance
-
Model Selection
- Start with smaller models (7B)
- Use quantized versions for less RAM
- Test different models for your use case
-
Performance
- Use GPU acceleration when available
- Adjust context window based on needs
- Monitor system resources
-
Development
- Perfect for development and testing
- No API costs
- Complete privacy
- Works offline
Managing Models
Pull Models
# Pull latest version
ollama pull llama3.1
# Pull specific version
ollama pull llama3.1:70b
# Pull with tag
ollama pull llama3.1:8b-instruct-q4_0
List Models
ollama list
Remove Models
ollama rm llama3.1
Create Custom Model
Create Modelfile:
FROM llama3.1
PARAMETER temperature 0.8
SYSTEM You are a helpful coding assistant
Build model:
ollama create my-coding-assistant -f Modelfile
Configuration Options
GPU Configuration
// Ollama automatically uses GPU if available
// Configure via OLLAMA_NUM_GPU environment variable
Context Window
maxTokens := 4096
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
Model: model,
Prompt: prompt,
MaxTokens: &maxTokens, // Adjust based on model
})
Remote Ollama
provider := ollama.New(ollama.Config{
BaseURL: "http://remote-server:11434",
})
Error Handling
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
if err != nil {
if strings.Contains(err.Error(), "connection refused") {
log.Fatal("Ollama server not running - run: ollama serve")
}
if strings.Contains(err.Error(), "model not found") {
log.Fatal("Model not found - run: ollama pull llama3.1")
}
log.Fatal(err)
}
Performance Tips
-
Use Quantized Models
- q4_0: 4-bit quantization (smallest)
- q5_0: 5-bit quantization (balanced)
- q8_0: 8-bit quantization (high quality)
-
Optimize Context
- Reduce context window for faster inference
- Clear conversation history periodically
-
GPU Acceleration
- Ensure CUDA/ROCm installed
- Monitor GPU memory usage
- Adjust num_gpu parameter