baalajimaestro/ollama

Jeffrey Morgan 1c8435ffa9

Update domain name references in docs and install script (#2435 )

2024-02-09 15:19:30 -08:00

805 B

Raw Blame History

Deploy Ollama to Kubernetes

Prerequisites

Ollama: https://ollama.com/download
Kubernetes cluster. This example will use Google Kubernetes Engine.

Steps

Create the Ollama namespace, daemon set, and service
```
kubectl apply -f cpu.yaml
```
Port forward the Ollama service to connect and use it locally
```
kubectl -n ollama port-forward service/ollama 11434:80
```
Pull and run a model, for example orca-mini:3b
```
ollama run orca-mini:3b
```

(Optional) Hardware Acceleration

Hardware acceleration in Kubernetes requires NVIDIA's k8s-device-plugin. Follow the link for more details.

Once configured, create a GPU enabled Ollama deployment.

kubectl apply -f gpu.yaml