baalajimaestro/ollama

Michael Yang 145e060855

Apply suggestions from code review

Co-authored-by: Bruce MacDonald <brucewmacdonald@gmail.com>

2023-11-06 11:32:23 -08:00

813 B

Raw Blame History

Deploy Ollama to Kubernetes

Prerequisites

Ollama: https://ollama.ai/download
Kubernetes cluster. This example will use Google Kubernetes Engine.

Steps

Create the Ollama namespace, daemon set, and service
```
kubectl apply -f cpu.yaml
```
Port forward the Ollama service to connect and use it locally
```
kubectl -n ollama port-forward service/ollama 11434:80
```
Pull and run a model, for example orca-mini:3b
```
ollama run orca-mini:3b
```

(Optional) Hardware Acceleration

Hardware acceleration in Kubernetes requires NVIDIA's k8s-device-plugin. Follow the link for more details.

Once configured, create a GPU enabled Ollama deployment.

kubectl apply -f gpu.yaml