# Ollama Provider

Ollama enables running large language models locally on your machine. Perfect for development, privacy-sensitive applications, and offline usage.

## Setup

### Installation

1. Install Ollama from [ollama.com](https://ollama.com)
2. Pull a model:

```bash
ollama pull llama3.1
```

### Configuration

Ollama has its own provider package; no API key is needed:

```go
import (
    "github.com/digitallysavvy/go-ai/pkg/ai"
    "github.com/digitallysavvy/go-ai/pkg/providers/ollama"
)

provider := ollama.New(ollama.Config{
    // BaseURL defaults to http://localhost:11434
})

model, err := provider.LanguageModel("llama3.1")
```

Ollama does not support image generation, speech synthesis, transcription, or reranking — `ImageModel`, `SpeechModel`, `TranscriptionModel`, and `RerankingModel` all return errors.

## Available Models

### Popular Models

| Model | Parameters | Size | Best For |
|-------|-----------|------|----------|
| llama3.1 | 70B/8B | 40GB/4.7GB | General purpose |
| llama3.2 | 3B/1B | 2GB/1.3GB | Fast, efficient |
| mistral | 7B | 4.1GB | Balanced |
| phi3 | 3.8B | 2.3GB | Small, capable |
| qwen2.5 | 72B/7B | 43GB/4.7GB | Multilingual |
| codellama | 34B/13B/7B | 19GB/7.4GB/3.8GB | Code generation |

### Specialized Models

| Model | Size | Best For |
|-------|------|----------|
| nomic-embed-text | 137M | Embeddings |
| llava | 7B | Vision |
| stable-code | 3B | Code completion |

## Examples

### Basic Text Generation

```go
package main

import (
    "context"
    "fmt"
    "log"

    "github.com/digitallysavvy/go-ai/pkg/ai"
    "github.com/digitallysavvy/go-ai/pkg/providers/ollama"
)

func main() {
    ctx := context.Background()
    provider := ollama.New(ollama.Config{})

    model, err := provider.LanguageModel("llama3.1")
    if err != nil {
        log.Fatal(err)
    }

    result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: "Explain Ollama"})
    if err != nil {
        log.Fatal(err)
    }

    fmt.Println(result.Text)
}
```

### Streaming

```go
stream, err := ai.StreamText(ctx, ai.StreamTextOptions{Model: model, Prompt: "Write a story"})
if err != nil {
    log.Fatal(err)
}
defer stream.Close()

for chunk := range stream.Chunks() {
    fmt.Print(chunk.Text)
}
```

### Vision with LLaVA

```go
model, err := provider.LanguageModel("llava")

imageData, _ := os.ReadFile("image.jpg")

result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model: model,
    Messages: []types.Message{
        {
            Role: types.RoleUser,
            Content: []types.ContentPart{
                types.TextContent{Text: "Describe this image"},
                types.FileContent{
                    Data:      imageData,
                    MediaType: "image/jpeg",
                },
            },
        },
    },
})
```

### Embeddings

```go
embeddingModel, err := provider.EmbeddingModel("nomic-embed-text")

texts := []string{
    "Ollama runs models locally",
    "Local inference is private",
}

result, err := ai.EmbedMany(ctx, ai.EmbedManyOptions{
    Model:  embeddingModel,
    Inputs: texts,
})
```

## Best Practices

1. **Hardware Requirements**
   - 8GB RAM minimum for 7B models
   - 16GB RAM for 13B models
   - 32GB+ RAM for 70B models
   - GPU recommended for better performance

2. **Model Selection**
   - Start with smaller models (7B)
   - Use quantized versions for less RAM
   - Test different models for your use case

3. **Performance**
   - Use GPU acceleration when available
   - Adjust context window based on needs
   - Monitor system resources

4. **Development**
   - Perfect for development and testing
   - No API costs
   - Complete privacy
   - Works offline

## Managing Models

### Pull Models

```bash
# Pull latest version
ollama pull llama3.1

# Pull specific version
ollama pull llama3.1:70b

# Pull with tag
ollama pull llama3.1:8b-instruct-q4_0
```

### List Models

```bash
ollama list
```

### Remove Models

```bash
ollama rm llama3.1
```

### Create Custom Model

Create Modelfile:

```
FROM llama3.1
PARAMETER temperature 0.8
SYSTEM You are a helpful coding assistant
```

Build model:

```bash
ollama create my-coding-assistant -f Modelfile
```

## Configuration Options

### GPU Configuration

```go
// Ollama automatically uses GPU if available
// Configure via OLLAMA_NUM_GPU environment variable
```

### Context Window

```go
maxTokens := 4096
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{
    Model:     model,
    Prompt:    prompt,
    MaxTokens: &maxTokens, // Adjust based on model
})
```

### Remote Ollama

```go
provider := ollama.New(ollama.Config{
    BaseURL: "http://remote-server:11434",
})
```

## Error Handling

```go
result, err := ai.GenerateText(ctx, ai.GenerateTextOptions{Model: model, Prompt: prompt})
if err != nil {
    if strings.Contains(err.Error(), "connection refused") {
        log.Fatal("Ollama server not running - run: ollama serve")
    }
    if strings.Contains(err.Error(), "model not found") {
        log.Fatal("Model not found - run: ollama pull llama3.1")
    }
    log.Fatal(err)
}
```

## Performance Tips

1. **Use Quantized Models**
   - q4_0: 4-bit quantization (smallest)
   - q5_0: 5-bit quantization (balanced)
   - q8_0: 8-bit quantization (high quality)

2. **Optimize Context**
   - Reduce context window for faster inference
   - Clear conversation history periodically

3. **GPU Acceleration**
   - Ensure CUDA/ROCm installed
   - Monitor GPU memory usage
   - Adjust num_gpu parameter

## See Also

- [API Reference: GenerateText](https://goaisdk.com/docs/reference/ai/generate-text.md)
- [Ollama Documentation](https://ollama.com/docs)
- [Ollama Model Library](https://ollama.com/library)
