baalajimaestro/ollama

Michael Yang dccac8c8fa k8s example

2023-11-01 14:52:58 -07:00

792 B

Raw Blame History

Deploy Ollama to Kubernetes

Prerequisites

Ollama: https://ollama.ai/download
Kubernetes cluster. This example will use Google Kubernetes Engine.

Steps

Create the Ollama namespace, daemon set, and service
```
kubectl apply -f cpu.yaml
```
Port forward the Ollama service to connect and use it locally
```
kubectl -n ollama port-forward service/ollama 11434:80
```
Pull and run orca-mini:3b
```
ollama run orca-mini:3b
```

(Optional) Hardware Acceleration

Hardware acceleration in Kubernetes requires NVIDIA's k8s-device-plugin. Follow the link for more details.

Once configured, create a GPU enabled Ollama deployment.

kubectl apply -f gpu.yaml