ollama

baalajimaestro/ollama

Fork 0

Commit graph

Author	SHA1	Message	Date
Daniel Hiltgen	34b9db5afc	Request and model concurrency This change adds support for multiple concurrent requests, as well as loading multiple models by spawning multiple runners. The default settings are currently set at 1 concurrent request per model and only 1 loaded model at a time, but these can be adjusted by setting OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.	2024-04-22 19:29:12 -07:00
Daniel Hiltgen	aeb1fb5192	Add test case for context exhaustion Confirmed this fails on 0.1.30 with known regression but passes on main	2024-04-04 07:42:17 -07:00

Author

SHA1

Message

Date

Daniel Hiltgen

34b9db5afc

Request and model concurrency

This change adds support for multiple concurrent requests, as well as
loading multiple models by spawning multiple runners. The default
settings are currently set at 1 concurrent request per model and only 1
loaded model at a time, but these can be adjusted by setting
OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.

2024-04-22 19:29:12 -07:00

Daniel Hiltgen

aeb1fb5192

Add test case for context exhaustion

Confirmed this fails on 0.1.30 with known regression
but passes on main

2024-04-04 07:42:17 -07:00

2 commits