ollama

Author	SHA1	Message	Date
Michael Yang	5121b7ac9c	remove double newlines in /set parameter	2024-01-12 11:21:15 -08:00
Michael Yang	a70262c6b2	Update README.md Co-authored-by: Jeffrey Morgan <jmorganca@gmail.com>	2024-01-12 09:43:04 -08:00
Tristram Oaten	40a0a90a88	Add group delete to uninstall instructions (#1924 ) After executing the `userdel ollama` command, I saw this message: ```sh $ sudo userdel ollama userdel: group ollama not removed because it has other members. ``` Which reminded me that I had to remove the dangling group too. For completeness, the uninstall instructions should do this too. Thanks!	2024-01-12 00:07:00 -05:00
Michael Yang	cbe20c4375	update readme	2024-01-11 16:24:37 -08:00
Michael Yang	5ffbbea1d7	remove client.py	2024-01-11 15:53:10 -08:00
Daniel Hiltgen	3773fb6465	Merge pull request #1935 from dhiltgen/cpu_fallback Fix up the CPU fallback selection	2024-01-11 15:52:32 -08:00
Daniel Hiltgen	7427fa1387	Fix up the CPU fallback selection The memory changes and multi-variant change had some merge glitches I missed. This fixes them so we actually get the cpu llm lib and best variant for the given system.	2024-01-11 15:27:06 -08:00
Michael Yang	f84537e0e0	Merge pull request #1934 from jmorganca/mxyng/fix-slices fix build and lint	2024-01-11 14:36:20 -08:00
Michael Yang	d2be6387c9	fix typo	2024-01-11 14:25:21 -08:00
Michael Yang	d7af35d3d0	import fmt	2024-01-11 14:22:32 -08:00
Michael Yang	defc1dbd6e	use x/exp/slices	2024-01-11 14:20:13 -08:00
Daniel Hiltgen	de2fbdec99	Merge pull request #1819 from dhiltgen/multi_variant Support multiple LLM libs; ROCm v5 and v6; Rosetta, AVX, and AVX2 compatible CPU builds	2024-01-11 14:00:48 -08:00
Eduard van Valkenburg	f5faf79aa1	Add semantic kernel to Readme (#1931 )	2024-01-11 14:40:23 -05:00
Michael Yang	f4f939de28	Merge pull request #1552 from jmorganca/mxyng/lint-test add lint and test on pull_request	2024-01-11 09:37:45 -08:00
Daniel Hiltgen	39928a42e8	Always dynamically load the llm server library This switches darwin to dynamic loading, and refactors the code now that no static linking of the library is used on any platform	2024-01-11 08:42:47 -08:00
Daniel Hiltgen	d88c527be3	Build multiple CPU variants and pick the best This reduces the built-in linux version to not use any vector extensions which enables the resulting builds to run under Rosetta on MacOS in Docker. Then at runtime it checks for the actual CPU vector extensions and loads the best CPU library available	2024-01-11 08:42:47 -08:00
Fabian Preiß	3bc8b9832b	fix gpu_test.go Error (same type) uint64->uint32 (#1921 )	2024-01-11 08:22:23 -05:00
Jeffrey Morgan	ab6be852c7	revisit memory allocation to account for full kv cache on main gpu	2024-01-11 01:45:31 -05:00
Daniel Hiltgen	052b33b81b	DRY out the Dockefile.build	2024-01-10 17:27:51 -08:00
Daniel Hiltgen	8da7bef05f	Support multiple variants for a given llm lib type In some cases we may want multiple variants for a given GPU type or CPU. This adds logic to have an optional Variant which we can use to select an optimal library, but also allows us to try multiple variants in case some fail to load. This can be useful for scenarios such as ROCm v5 vs v6 incompatibility or potentially CPU features.	2024-01-10 17:27:51 -08:00
Jeffrey Morgan	b24e8d17b2	Increase minimum CUDA memory allocation overhead and fix minimum overhead for multi-gpu (#1896 ) * increase minimum cuda overhead and fix minimum overhead for multi-gpu * fix multi gpu overhead * limit overhead to 10% of all gpus * better wording * allocate fixed amount before layers * fixed only includes graph alloc	2024-01-10 19:08:51 -05:00
Jeffrey Morgan	f83881390f	revert submodule back to `328b83de23b33240e28f4e74900d1d06726f5eb1`	2024-01-10 18:42:39 -05:00
Daniel Hiltgen	ac70ab6761	Merge pull request #1914 from dhiltgen/smarter_cuda_detection Smarter GPU Management library detection	2024-01-10 15:21:56 -08:00
Daniel Hiltgen	3c49c3ab0d	Harden GPU mgmt library lookup When there are multiple management libraries installed on a system not every one will be compatible with the current driver. This change improves our management library algorithm to build up a set of discovered libraries based on glob patterns, and then try all of them until we're able to load one without error.	2024-01-10 15:06:41 -08:00
Daniel Hiltgen	9754ae4c89	Support optional override of the target archictures This can help speed up incremental builds when you're only testing one archicture, like amd64. E.g. BUILD_ARCH=amd64 ./scripts/build_linux.sh && scp ./dist/ollama-linux-amd64 test-system:	2024-01-10 14:43:24 -08:00
Jeffrey Morgan	224fbf2795	update submodule to commit `1fc2f265ff9377a37fd2c61eae9cd813a3491bea` until its main branch is fixed	2024-01-10 17:03:15 -05:00
Jeffrey Morgan	2c6e8f5248	Update submodule to `6efb8eb30e7025b168f3fda3ff83b9b386428ad6` (#1885 ) * update submodule to `6efb8eb30e7025b168f3fda3ff83b9b386428ad6` * unblock condition variable in `update_slots` when closing server	2024-01-10 16:48:38 -05:00
Jeffrey Morgan	34344d801c	clean up cmake `build` directory when cross compiling macOS builds	2024-01-09 17:13:56 -05:00
Robin Glauser	e868c8a5c7	Update api.md (#1878 ) Fixed assistant in the example response.	2024-01-09 16:21:17 -05:00
Jeffrey Morgan	c336693f07	calculate overhead based number of gpu devices (#1875 )	2024-01-09 15:53:33 -05:00
Daniel Hiltgen	e89dc1d54b	Merge pull request #1874 from dhiltgen/correct_cuda_min Set corret CUDA minimum compute capability version	2024-01-09 11:37:22 -08:00
Daniel Hiltgen	1961a81f03	Set corret CUDA minimum compute capability version If you attempt to run the current CUDA build on compute capability 5.2 cards, you'll hit the following failure: cuBLAS error 15 at ggml-cuda.cu:7956: the requested functionality is not supported	2024-01-09 11:28:24 -08:00
Jeffrey Morgan	8a8c7e7f8d	only build for metal on `arm64`	2024-01-09 13:51:08 -05:00
Jeffrey Morgan	6df83e6daa	update rough cuda overhead estimate to 15% + 384MiB	2024-01-09 13:51:08 -05:00
Michael Yang	f921e2696e	typo	2024-01-09 09:45:42 -08:00
Michael Yang	4a33cede20	remove unused fields and functions	2024-01-09 09:37:40 -08:00
Michael Yang	f95d2f25f3	fix temporary history file permissions	2024-01-09 09:36:58 -08:00
Michael Yang	2b9892a808	fix(windows): modelpath and list	2024-01-09 09:36:58 -08:00
Michael Yang	2bb2bdd5d4	fix lint	2024-01-09 09:36:58 -08:00
Michael Yang	acfc376efd	add .golangci.yaml	2024-01-09 09:36:58 -08:00
Michael Yang	997253143f	add lint and test on pull_request	2024-01-09 09:36:58 -08:00
Michael Yang	62023177f6	Merge pull request #1614 from jmorganca/mxyng/fix-set-template fix: set template without triple quotes	2024-01-09 09:36:24 -08:00
Jeffrey Morgan	6164f378f2	revert cuda overhead to 20%	2024-01-09 00:54:29 -05:00
Jeffrey Morgan	f387e9631b	use runner if cuda alloc won't fit	2024-01-09 00:44:34 -05:00
Jeffrey Morgan	6566387ae3	add `TODO` for cuda overhead	2024-01-09 00:28:03 -05:00
Jeffrey Morgan	37708931fb	update cuda overhead to 20% to fix crashes when switching between models and large context sizes	2024-01-09 00:05:23 -05:00
Jeffrey Morgan	f6cb0a553c	update cuda overhead to 15% or 400MiB	2024-01-08 23:45:45 -05:00
Jeffrey Morgan	2680078c13	fix build on linux	2024-01-08 23:44:13 -05:00
Jeffrey Morgan	f1b7e5f560	update overhead to 15%	2024-01-08 23:37:45 -05:00
Jeffrey Morgan	cb534e6ac2	use 10% vram overhead for cuda	2024-01-08 23:17:44 -05:00

... 24 25 26 27 28 ...

3035 commits