ollama/llm/patches/01-load-progress.diff

diff --git a/common/common.cpp b/common/common.cpp
index 73ff0e85..6adb1a92 100644
--- a/common/common.cpp
+++ b/common/common.cpp
@@ -2447,6 +2447,8 @@ struct llama_model_params llama_model_params_from_gpt_params(const gpt_params &
     mparams.use_mmap        = params.use_mmap;
     mparams.use_mlock       = params.use_mlock;
     mparams.check_tensors   = params.check_tensors;
+    mparams.progress_callback = params.progress_callback;
+    mparams.progress_callback_user_data = params.progress_callback_user_data;
     if (params.kv_overrides.empty()) {
         mparams.kv_overrides = NULL;
     } else {
diff --git a/common/common.h b/common/common.h
index 58ed72f4..0bb2605e 100644
--- a/common/common.h
+++ b/common/common.h
@@ -180,6 +180,13 @@ struct gpt_params {
     std::string mmproj = "";        // path to multimodal projector
     std::vector<std::string> image; // path to image file(s)
 
+    // Called with a progress value between 0.0 and 1.0. Pass NULL to disable.
+    // If the provided progress_callback returns true, model loading continues.
+    // If it returns false, model loading is immediately aborted.
+    llama_progress_callback progress_callback = NULL;
+    // context pointer passed to the progress callback
+    void * progress_callback_user_data;
+
     // server params
     int32_t port           = 8080;         // server listens on this network port
     int32_t timeout_read   = 600;          // http read timeout in seconds
Wire up load progress This doesn't expose a UX yet, but wires the initial server portion of progress reporting during load 2024-05-20 23:41:43 +00:00			`diff --git a/common/common.cpp b/common/common.cpp`
llm: update llama.cpp commit to `7c26775` (#4896) * llm: update llama.cpp submodule to `7c26775` * disable `LLAMA_BLAS` for now * `-DLLAMA_OPENMP=off` 2024-06-17 19:56:16 +00:00			`index 73ff0e85..6adb1a92 100644`
Wire up load progress This doesn't expose a UX yet, but wires the initial server portion of progress reporting during load 2024-05-20 23:41:43 +00:00			`--- a/common/common.cpp`
			`+++ b/common/common.cpp`
llm: update llama.cpp commit to `7c26775` (#4896) * llm: update llama.cpp submodule to `7c26775` * disable `LLAMA_BLAS` for now * `-DLLAMA_OPENMP=off` 2024-06-17 19:56:16 +00:00			`@@ -2447,6 +2447,8 @@ struct llama_model_params llama_model_params_from_gpt_params(const gpt_params &`
Wire up load progress This doesn't expose a UX yet, but wires the initial server portion of progress reporting during load 2024-05-20 23:41:43 +00:00			`mparams.use_mmap = params.use_mmap;`
			`mparams.use_mlock = params.use_mlock;`
			`mparams.check_tensors = params.check_tensors;`
			`+ mparams.progress_callback = params.progress_callback;`
			`+ mparams.progress_callback_user_data = params.progress_callback_user_data;`
			`if (params.kv_overrides.empty()) {`
			`mparams.kv_overrides = NULL;`
			`} else {`
			`diff --git a/common/common.h b/common/common.h`
llm: update llama.cpp commit to `7c26775` (#4896) * llm: update llama.cpp submodule to `7c26775` * disable `LLAMA_BLAS` for now * `-DLLAMA_OPENMP=off` 2024-06-17 19:56:16 +00:00			`index 58ed72f4..0bb2605e 100644`
Wire up load progress This doesn't expose a UX yet, but wires the initial server portion of progress reporting during load 2024-05-20 23:41:43 +00:00			`--- a/common/common.h`
			`+++ b/common/common.h`
llm: update llama.cpp commit to `7c26775` (#4896) * llm: update llama.cpp submodule to `7c26775` * disable `LLAMA_BLAS` for now * `-DLLAMA_OPENMP=off` 2024-06-17 19:56:16 +00:00			`@@ -180,6 +180,13 @@ struct gpt_params {`
Wire up load progress This doesn't expose a UX yet, but wires the initial server portion of progress reporting during load 2024-05-20 23:41:43 +00:00			`std::string mmproj = ""; // path to multimodal projector`
			`std::vector<std::string> image; // path to image file(s)`
llm: update llama.cpp commit to `7c26775` (#4896) * llm: update llama.cpp submodule to `7c26775` * disable `LLAMA_BLAS` for now * `-DLLAMA_OPENMP=off` 2024-06-17 19:56:16 +00:00
Wire up load progress This doesn't expose a UX yet, but wires the initial server portion of progress reporting during load 2024-05-20 23:41:43 +00:00			`+ // Called with a progress value between 0.0 and 1.0. Pass NULL to disable.`
			`+ // If the provided progress_callback returns true, model loading continues.`
			`+ // If it returns false, model loading is immediately aborted.`
			`+ llama_progress_callback progress_callback = NULL;`
			`+ // context pointer passed to the progress callback`
			`+ void * progress_callback_user_data;`
llm: update llama.cpp commit to `7c26775` (#4896) * llm: update llama.cpp submodule to `7c26775` * disable `LLAMA_BLAS` for now * `-DLLAMA_OPENMP=off` 2024-06-17 19:56:16 +00:00			`+`
			`// server params`
			`int32_t port = 8080; // server listens on this network port`
			`int32_t timeout_read = 600; // http read timeout in seconds`