ollama

Author	SHA1	Message	Date
suncloudsmoon	18237be9b2	readme: add TextCraft to community integrations (#7377 )	2024-11-03 16:53:51 -08:00
Daniel Hiltgen	29ab9fa7d7	nvidia libs have inconsistent ordering (#7473 ) The runtime and management libraries may not always have identical ordering, so use the device UUID to correlate instead of ID.	2024-11-02 16:35:41 -07:00
Daniel Hiltgen	b8d5036e33	CI: omit unused tools for faster release builds (#7432 ) This leverages caching, and some reduced installer scope to try to speed up builds. It also tidies up some windows build logic that was only relevant for the older generate/cmake builds.	2024-11-02 13:56:54 -07:00
Jesse Gross	312d9de1d1	llama: Improve error handling Check for NULL return values from llama.cpp in more places and convert them into Go errors, which should make debugging easier in the future rather than having hidden surprises in our data structures.	2024-11-02 13:37:55 -07:00
Jesse Gross	a103dae01e	runner.go: Only allocate 1 element embedding batches for mllama Mllama has large embeddings (100 MB per image) and each embedding is represented as 1 token when passed to llama.cpp. Batches are pre- allocated for the size of the tokens times the batch size, so this results in allocations of over 50 GB at the default batch size. On some systems, these mallocs will fail. Since an image is represented as a single token and mllama doesn't support more than 1 image per request, we only need to allocate a batch size of 1, which is much more reasonable. In addition, for non-multimodal models, we don't need to allocate the embedding batches at all. Fixes #7464	2024-11-02 13:37:55 -07:00
Michael Yang	d07cf41a97	refactor kv estimation	2024-11-01 16:23:55 -07:00
Michael Yang	8c238e70ab	mllama cross attention	2024-11-01 16:23:55 -07:00
Daniel Hiltgen	8a9bb0d000	Add basic mllama integration tests (#7455 )	2024-10-31 17:25:48 -07:00
Jesse Gross	26acdcf44e	runner.go: Don't set cross attention before sending embeddings Currently if an input has embeddings at any point then we will set cross attention to true from the beginning. This means that any tokens before the embeddings are sent will incorrectly have cross attention layers applied. This only sets cross attention when we have an embedding, either previously in this sequence or in the cache. It also makes cross attention capable of supporting parallelism at the runner level, though the mllama implementation doesn't support that yet.	2024-10-31 13:56:08 -07:00
Daniel Hiltgen	921779bb10	Give unicode test more time to run (#7437 ) * Give unicode test more time to run Some slower GPUs (or partial CPU/GPU loads) can take more than the default 30s to complete this test * Give more time for concurrency test CPU inference can be very slow under stress	2024-10-31 13:35:31 -07:00
Daniel Hiltgen	16f4eabe2d	Refine default thread selection for NUMA systems (#7322 ) Until we have full NUMA support, this adjusts the default thread selection algorithm to count up the number of performance cores across all sockets.	2024-10-30 15:05:45 -07:00
Jesse Gross	c826e57475	runner.go: Better abstract vision model integration -Update mllama to take the cross attention state as embeddings in a batch, more similar to how Llava handles it. This improves integration with the input cache. -Pass locations in a prompt for embeddings using tags similar to Llava. -Abstract interface to vision models so the main runner accesses Clip and Mllama similarly Co-authored-by: Michael Yang <mxyng@pm.me>	2024-10-30 14:53:43 -07:00
Daniel Hiltgen	712e99d477	Soften windows clang requirement (#7428 ) This will no longer error if built with regular gcc on windows. To help triage issues that may come in related to different compilers, the runner now reports the compier used by cgo.	2024-10-30 12:28:36 -07:00
Daniel Hiltgen	b754f5a6a3	Remove submodule and shift to Go server - 0.4.0 (#7157 ) * Remove llama.cpp submodule and shift new build to top * CI: install msys and clang gcc on win Needed for deepseek to work properly on windows	2024-10-30 10:34:28 -07:00
Daniel Hiltgen	a805e5947e	Move windows app out of preview (#7347 )	2024-10-30 09:24:59 -07:00
Daniel Hiltgen	91dfbb1bba	windows: Support alt install paths, fit and finish (#6967 ) * windows: Support alt install paths Advanced users are leveraging innosetup's /DIR switch to target an alternate location, but we get confused by things not existing in the LocalAppData dir. This also hardens the server path lookup code for a future attempt to unify with a ./bin prefix * Fit and finish improvements for windows app Document alternate install location instructions for binaries and model. Pop up progress UI for upgrades (automatic, with cancel button). Expose non-default port in menu to disambiguate mutiple instances. Set minimum Windows version to 10 22H2	2024-10-30 09:24:31 -07:00
Patrick Devine	db1842b9e1	add more tests for getting the optimal tiled canvas (#7411 )	2024-10-29 16:28:02 -07:00
Daniel Hiltgen	c9ca386131	Switch windows to clang (#7407 ) * Switch over to clang for deepseek on windows The patch for deepseek requires clang on windows. gcc on windows has a buggy c++ library and can't handle the unicode characters * Fail fast with wrong compiler on windows Avoid users mistakenly building with GCC when we need clang	2024-10-29 13:15:04 -07:00
Jesse Gross	078f666f73	tests: Add test for Unicode processing	2024-10-28 18:12:29 -07:00
Jesse Gross	de1557a0dc	runner.go: Better handle return NULL values from llama.cpp Llama.cpp sometimes returns NULL as a return value to report an error. We should explicitly check for this and convert it to a Go error rather than putting NULL in our data structures and waiting for it to blow up later.	2024-10-28 18:12:29 -07:00
Patrick Devine	084929c293	add mllama image processing to the generate handler (#7384 )	2024-10-28 13:51:19 -07:00
Daniel Hiltgen	abd5dfd06a	Bump to latest Go 1.22 patch (#7379 )	2024-10-26 17:03:37 -07:00
Daniel Hiltgen	099f7077a1	Fix deepseek deseret regex (#7369 ) On windows compiled with gcc the c++ regex library failed to handle the characters	2024-10-26 14:58:54 -07:00
Daniel Hiltgen	d7c94e0ca6	Better support for AMD multi-GPU on linux (#7212 ) * Better support for AMD multi-GPU This resolves a number of problems related to AMD multi-GPU setups on linux. The numeric IDs used by rocm are not the same as the numeric IDs exposed in sysfs although the ordering is consistent. We have to count up from the first valid gfx (major/minor/patch with non-zero values) we find starting at zero. There are 3 different env vars for selecting GPUs, and only ROCR_VISIBLE_DEVICES supports UUID based identification, so we should favor that one, and try to use UUIDs if detected to avoid potential ordering bugs with numeric IDs * ROCR_VISIBLE_DEVICES only works on linux Use the numeric ID only HIP_VISIBLE_DEVICES on windows	2024-10-26 14:04:14 -07:00
Daniel Hiltgen	35ec7f079f	Fix unicode output on windows with redirect to file (#7358 ) If we're not writing out to a terminal, avoid setting the console mode on windows, which corrupts the output file.	2024-10-25 13:43:16 -07:00
Daniel Hiltgen	5231ae52d9	Fix incremental build file deps (#7361 ) The common src/hdr defs should be in the common definitions, not gpu specific.	2024-10-25 11:50:45 -07:00
Daniel Hiltgen	3085c47bea	Improve dependency gathering logic (#7345 ) This unfies the rocm/cuda dependency logic into the makefile and fixes a missing define which broke windows rocm	2024-10-24 09:51:53 -07:00
Bill Wang	0ccc73251a	fix #7247 - invalid image input (#7249 ) --------- Co-authored-by: Bill Wang <bill.wang@bill.wang>	2024-10-23 10:31:04 -07:00
Daniel Hiltgen	dc6fe82051	integration: harden embedding test (#7306 ) Use cosine similarity to make the embeddings tests more robust	2024-10-22 15:25:22 -07:00
Patrick Devine	d78fb62056	default to "FROM ." if a Modelfile isn't present (#7250 )	2024-10-22 13:32:24 -07:00
Daniel Hiltgen	5c44461ccf	Fix rocm windows build and clean up dependency gathering (#7305 ) On windows ensure windows version define is properly set for rocm. Remove duplicate rocm arch flags. Resolve wildcards in the targets so parallel builds don't race. Use readlink to resolve rocm dependencies since wildcards omit libelf Keep windows rocm deps aligned with unified packaging model	2024-10-22 12:54:15 -07:00
Jesse Gross	03e40efa51	runner.go: Merge partial unicode characters before sending We check for partial unicode characters and accumulate them before sending. However, when we did send, we still sent each individual piece separately, leading to broken output. This combines everything into a single group, which is also more efficient. This also switches to the built-in check for valid unicode characters, which is stricter. After this, we should never send back an invalid sequence. Fixes #7290	2024-10-22 12:07:51 -07:00
Mattt	23f746508d	readme: add Ollama for Swift to the community integrations (#7295 )	2024-10-21 22:29:11 -07:00
Jeffrey Morgan	48708ca0d5	server: allow vscode-webview origin (#7273 )	2024-10-19 14:06:41 -07:00
Patrick Devine	c7cb0f0602	image processing for llama3.2 (#6963 ) Co-authored-by: jmorganca <jmorganca@gmail.com> Co-authored-by: Michael Yang <mxyng@pm.me> Co-authored-by: Jesse Gross <jesse@ollama.com>	2024-10-18 16:12:35 -07:00
Daniel Hiltgen	bf4018b9ec	llama: Decouple patching script from submodule (#7139 ) * Refine llama.cpp vendoring workflow tools Switch from the sync.sh over to make based tooling * Run new make sync and patch flow	2024-10-17 15:03:09 -07:00
Daniel Hiltgen	f86d00cd95	llama: add compiler tags for cpu features (#7137 ) This adds the ability to customize the default runner with user specified flags	2024-10-17 13:43:20 -07:00
Gabe Goodhart	f2890a4494	IBM granite/granitemoe architecture support (#6760 ) * fix(ext_server): Port llama.cpp sampling refactors to ext_server This was a fairly large changeset. I closely followed the changes here: `df270ef745` Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(server.cpp): Refactor server.cpp logging for llama.cpp overhaul Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat: Bump llama.cpp to the latest master with `granite` support This does not yet have granite MoE support, but that can come in a follow up PR Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(patches): Update all patches (except solar-pro) to work with bumped llama.cpp Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(solar): Update solar patch for llama.cpp bump Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat(llama.cpp): Bump llama.cpp for granitemoe support Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat(llama.cpp): Bump llama.cpp for granitemoe support Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(solar): Update the solar-pro patch for latest llama.cpp bump Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat(llama.cpp): Bump to the latest master of llama.cpp Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(patches): Update all patches for latest bump Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat(llama): Always run sync.sh from the right directory Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llama/patches): Update llama patches Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat(llama)!: Rough sync with llama.cpp submodule There are a number of changes that will need to be propagated to llama.go before any of this works! Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llama/patches): Add a patch and update for missing ggml-impl.h include This include is where the ggml_cgraph struct is defined. It is included in many of the .c files to define the forward declartion in ggml.h. It seems that with the subset of code included here, the import was somehow lost (or out-of-order) when building, so adding this include to llama.cpp fixes the missing definition. Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llama/sync): Add missing ggml-cpu-impl.h copy-over in sync.sh Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llama): Add missing log.cpp This was added as part of the logging overhaul done in llama.cpp Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llama): Overhaul use of sampling module for llama.cpp changes The changes here reflect the changes made in the big llama.cpp sampling PR https://github.com/ggerganov/llama.cpp/pull/9294 The sampling functionality is now broken into the base interface (llama_sampler) and the generation implementation (gpt_sampler). The changes here reflect that. Since the sampling.h/sampling.cpp code uses c++ STL headers, the sampling_ext.[h\|cpp] wrapper is maintained to allow go to access a pure-C interface. Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llama): Fix the impl of SampleTokenGreedy for new sampling I don't think this method is currently used, so it could probably just be removed so that all sampling goes through the GPT interface, but in the interest of doing no harm, this should keep the method working as expected. Branch: IBMGraniteArchitectureSupport * fix(llama): Remove unused SampleTokenGreedy Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(sync): Remove bash-specific change to sync.sh Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * chore(gofumpt): Format on llama.go to pass linting Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llm): Fix missing <thread> include in ext_server Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llama): Remove TODO about grammar_first This feature was not used/needed previously so should be fine without plumbing it through now. Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llama): Better naming for sampling wrapper and args Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llama): Fix patch 05 to use new wrapper api and re-sync Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * runner: Flush pending responses before returning If there are any pending reponses (such as from potential stop tokens) then we should send them back before ending the sequence. Otherwise, we can be missing tokens at the end of a response. Fixes #6707 * fix(llama/sampling): Use gpt_sampler with a forward declaration Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llama): Remove unnecessary patch for gguf impl header This was caused by an earlier mistake in the embeddings patch that was dereferencing the pointer instead of using the wrapper API. Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(llm): Remove use of deprecated --log-disable flag Branch: IBMGraniteArchitectureSupport Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> --------- Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>	2024-10-17 11:59:52 -07:00
Daniel Hiltgen	05cd82ef94	Rename gpu package discover (#7143 ) Cleaning up go package naming	2024-10-16 17:45:00 -07:00
Daniel Hiltgen	7d6eb0d4c3	Move macos v11 support flags to build script (#7203 ) Having v11 support hard-coded into the cgo settings causes warnings for newer Xcode versions. This should help keep the build clean for users building from source with the latest tools, while still allow us to target the older OS via our CI processes.	2024-10-16 12:49:46 -07:00
Daniel Hiltgen	24636dfa87	Discovery CPU details for default thread selection (#6264 ) On windows, detect large multi-socket systems and reduce to the number of cores in one socket for best performance	2024-10-15 11:36:08 -07:00
JHubi1	1d7fa3ad2d	Adding 'Ollama App' as community integrations (#6465 )	2024-10-15 09:57:32 -07:00
frob	09035b71cd	Add missing BF16 tensor type. (#7193 ) Co-authored-by: Richard Lyons <frob@cloudstaff.com>	2024-10-14 17:06:35 -07:00
Daniel Hiltgen	f3c8b898cd	Track GPU discovery failure information (#5820 ) * Expose GPU discovery failure information * Remove exposed API for now	2024-10-14 16:26:45 -07:00
Daniel Hiltgen	5dd0477fd4	Fix regression on older macos versions (#7192 ) The new cgo compilation requires a flag to target older macos versions	2024-10-13 10:47:42 -07:00
Daniel Hiltgen	c3d321d405	llm: Remove GGML_CUDA_NO_PEER_COPY for ROCm (#7174 ) This workaround logic in llama.cpp is causing crashes for users with less system memory than VRAM.	2024-10-12 09:56:49 -07:00
Jesse Gross	7fe3902552	cli: Send all images in conversation history Currently the CLI only sends images from the most recent image- containing message. This prevents doing things like sending one message with an image and then a follow message with a second image and asking for comparision based on additional information not present in any text that was output. It's possible that some models have a problem with this but the CLI is not the right place to do this since any adjustments are model-specific and should affect all clients. Both llava:34b and minicpm-v do reasonable things with multiple images in the history.	2024-10-10 11:21:51 -07:00
Jesse Gross	0077e22d52	runner.go: Handle truncation of tokens for stop sequences When a single token contains both text to be return and a stop sequence, this causes an out of bounds error when we update the cache to match our text. This is because we currently assume that the removing the stop sequence will consume at least one token. This also inverts the logic to deal with positive numbers, rather than a value to be subtracted, which is easier to reason about. Fixes #7153	2024-10-09 20:39:04 -07:00
Jesse Gross	03408f3437	server: Don't clear cmd when closing a server Close can be called on an LLM server if the runner subprocess dies. However, the Ollama scheduler code may not know about this yet and still try to access it. In this case, it is important that 'cmd' is still available as it is used to check on the status of the subprocess. If this happens, Kill may be called twice on the subprocess - that is fine. In addition, model unloading may race with new accesses, so we should hold a lock around this. This may result in the model being reloaded after the first close call - this is also fine as close will be called again later.	2024-10-09 20:39:04 -07:00
Daniel Hiltgen	cd7e01e8b9	fix vendoring attribute for metal (#7156 ) Add missing metal files to vendoring list	2024-10-09 15:22:36 -07:00

1 2 3 4 5 ...

3612 commits