ollama

baalajimaestro/ollama

Fork 0

Commit graph

Author	SHA1	Message	Date
Bryce Reitano	36a6daccab	Restructure loading conditional chain	2024-04-24 17:37:03 -06:00
Bryce Reitano	284e02bed0	Move ggml loading to when we attempt fitting	2024-04-24 17:17:24 -06:00
Daniel Hiltgen	34b9db5afc	Request and model concurrency This change adds support for multiple concurrent requests, as well as loading multiple models by spawning multiple runners. The default settings are currently set at 1 concurrent request per model and only 1 loaded model at a time, but these can be adjusted by setting OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.	2024-04-22 19:29:12 -07:00

Author

SHA1

Message

Date

Bryce Reitano

36a6daccab

Restructure loading conditional chain

2024-04-24 17:37:03 -06:00

Bryce Reitano

284e02bed0

Move ggml loading to when we attempt fitting

2024-04-24 17:17:24 -06:00

Daniel Hiltgen

34b9db5afc

Request and model concurrency

This change adds support for multiple concurrent requests, as well as
loading multiple models by spawning multiple runners. The default
settings are currently set at 1 concurrent request per model and only 1
loaded model at a time, but these can be adjusted by setting
OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.

2024-04-22 19:29:12 -07:00

3 commits