History

Jesse Gross 65973ceb64 runner.go: Make KV entry accounting more robust The structure of the accounting for KV cache shifting was carried over from the old runner but it now doesn't feel natural with the new runner. There are a number of invariants that should hold true but are difficult to reason about. There is at least one bug report that would imply that the invariants are not holding. This reduces the number of implicit assumptions and is more forgiving of unexpected situations. It also improves behavior around which input tokens are kept when truncation occurs. Bug #7545		2024-11-11 20:23:03 -08:00
..
cache.go	runner.go: Make KV entry accounting more robust	2024-11-11 20:23:03 -08:00
cache_test.go	runner.go: Better abstract vision model integration	2024-10-30 14:53:43 -07:00
image.go	runner.go: Check for zero length images	2024-11-08 09:39:32 -08:00
image_test.go	runner.go: Better abstract vision model integration	2024-10-30 14:53:43 -07:00
README.md	Re-introduce the `llama` package (#5034 )	2024-10-08 08:53:54 -07:00
requirements.go	Re-introduce the `llama` package (#5034 )	2024-10-08 08:53:54 -07:00
runner.go	runner.go: Make KV entry accounting more robust	2024-11-11 20:23:03 -08:00
stop.go	runner.go: Handle truncation of tokens for stop sequences	2024-10-09 20:39:04 -07:00
stop_test.go	runner.go: Handle truncation of tokens for stop sequences	2024-10-09 20:39:04 -07:00

`runner`

Note: this is a work in progress

A minimial runner for loading a model and running inference via a http web server.

./runner -model <model binary>

curl -X POST -H "Content-Type: application/json" -d '{"prompt": "hi"}' http://localhost:8080/completion

curl -X POST -H "Content-Type: application/json" -d '{"prompt": "turn me into an embedding"}' http://localhost:8080/embeddings