ollama

Author	SHA1	Message	Date
Nikita Ganzikov	fce30f407a	app: typo in wintray messages const (#7705 )	2024-11-20 22:01:58 -08:00
Daniel Hiltgen	d863298210	docs: Link to AMD guide on multi-GPU guidance (#7744 )	2024-11-20 16:00:46 -08:00
Jesse Gross	c4b34f2a2a	runner.go: Truncate inputs that exceed context rather than shifting Previous versions of the runner would truncate inputs to the context window before beginning processing. The main processing loop relied on this behavior if the context needed to be shifted later (due to token generation). If truncation did not occur then invariants would be broken, causing crashes or infinite loops. Later versions attempted to fix these bugs and make the logic less subtle so that all inputs could be handled. Truncation was removed to make things consistent. However, truncation is much faster than processing and shifting, so removing it caused performance problems when the input vastly exceeded the context size. This restores the input truncation as a performance optimization while keeping the more robust processing logic. Fixes #7762	2024-11-20 12:49:24 -08:00
Jesse Gross	c3ff916431	runner.go: Don't add inputs to cache view until actually processed We need to track which tokens are in the cache ourselves. We currently add tokens to the cache tracker when we add them to batch but they are not actually in the cache until we call Decode. This can cause confusion when we are shifting the cache. Avoids "could not find a KV slot for the batch" issues. Bug #7545	2024-11-20 12:49:24 -08:00
Jesse Gross	3fc1dc0e6f	runner.go: Hard fail on errors rather than potentially infinite looping We try to recover from errors by dropping the tokens that caused the problem and re-trying. However, dropping the tokens is not correct and continuing often leads to infinite loops. To avoid, this we end the sequence if such a condition is detected, which is also surprising. At this point, it is better to just report the error. This will make it easier to find problems and the alternatives are perhaps even more surprising to users. This is not a very satisfactory solution either - we should isolate the error and return it to the user without killing the whole process. However, this is an incremental step and consistent with most other failures (which either manifest as abort() or panic).	2024-11-20 12:49:24 -08:00
Jesse Gross	7121dfa309	runner.go: Retry decoding after defragmentation if needed Fragmentation of the KV cache can occur due to cache shifting or different sequences getting processed. Decode uses a heuristic to decide if it should defrag. However, this heuristic isn't 100% accurate, so decoding can sometimes fail by surprise. For these cases, if decode indicates that there is no KV cache space, we should defrag and then try again.	2024-11-20 12:49:24 -08:00
Jesse Gross	5f68fcab12	runner.go: Use correct index when retrieving embedding results This doesn't have any impact currently because NUM_PARALLEL is forced to 1 for embeddings, so both indicies will always be 0.	2024-11-20 12:49:24 -08:00
Emir Sahin	ecf41eed05	readme: add llm-axe to community integrations (#5931 )	2024-11-20 10:53:14 -08:00
Marcus Ziadé	b8c66d3307	readme: add a swift community integration (#7383 )	2024-11-20 10:49:15 -08:00
thewh1teagle	303f4bc79e	readme: add vibe app to community integrations (#7607 )	2024-11-20 10:45:10 -08:00
Adarsh Mishra	d2a25206b1	readme: add opentalkgpt to community integrations (#7707 )	2024-11-20 10:42:55 -08:00
rohitanshu	2f0a8c8778	docs: fix minor typo in import.md (#7764 ) change 'containg' to 'containing'	2024-11-20 09:57:32 -08:00
Gordon Kamer	bfd30f4286	readme: add Abbey to community integrations (#7746 )	2024-11-19 21:37:15 -08:00
Jonathan Hecl	0ef17ede89	readme: add Gollama to community integrations (#7756 )	2024-11-19 21:31:43 -08:00
Daniel Hiltgen	909a88c5c0	Improve crash reporting (#7728 ) Many model crashes are masked behind "An existing connection was forcibly closed by the remote host" This captures that common error message and wires in any detected errors from the log. This also adds the deepseek context shift error to the known errors we capture.	2024-11-19 16:26:57 -08:00
Daniel Hiltgen	f602ab4de4	expose underlying error on embedding failure (#7743 ) Avoid a round-trip asking users for logs to see what went wrong.	2024-11-19 16:26:05 -08:00
Gabe Goodhart	807ace5b1f	fix(runner): Set logits to 0 if false on Batch.Add https://github.com/ollama/ollama/issues/7656 Branch: Granite3StoppingBug-7656 Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>	2024-11-19 15:45:37 -08:00
Blake Mizerany	4b8a2e341a	server: allow mixed-case model names on push, pull, cp, and create (#7676 ) This change allows for mixed-case model names to be pushed, pulled, copied, and created, which was previously disallowed because the Ollama registry was backed by a Docker registry that enforced a naming convention that disallowed mixed-case names, which is no longer the case. This does not break existing, intended, behaviors. Also, make TestCase test a story of creating, updating, pulling, and copying a model with case variations, ensuring the model's manifest is updated correctly, and not duplicated across different files with different case variations.	2024-11-19 15:05:57 -08:00
frob	e66c29261a	Better error suppresion when getting terminal colours (#7739 ) Co-authored-by: Richard Lyons <frob@cloudstaff.com>	2024-11-19 08:33:52 -08:00
Patrick Devine	712d63c3f0	update the docs (#7731 )	2024-11-18 21:17:38 -08:00
Patrick Sy	6cdf27d154	readme: add Alfred Ollama to community integrations (#7724 )	2024-11-18 19:33:23 -08:00
frob	5c18e66384	Notify the user if systemd is not running (#6693 ) Co-authored-by: Richard Lyons <frob@cloudstaff.com>	2024-11-18 15:02:41 -08:00
Daniel Hiltgen	35096a7eff	win: add right click menu support (#7727 ) Enable both left and right click on the pop-up menu	2024-11-18 14:39:52 -08:00
Daniel Hiltgen	81d55d3e4d	fix index out of range on zero layer metal load (#7696 ) If the model doesn't fit any layers on metal, and we load zero layers we would panic trying to look up the GPU size during scheduling ops	2024-11-18 11:48:13 -08:00
Vinh Nguyen	a14f76491d	readme: improve Community Integrations section (#7718 )	2024-11-17 19:30:22 -08:00
Nicolas Bonamy	760cfa27e5	readme: add Witsy and multi-llm-ts to community integrations (#7713 )	2024-11-17 16:33:10 -08:00
Darius Kocar	c9a5aca3da	readme: add Perfect Memory AI to community integrations (#7431 )	2024-11-17 15:19:26 -08:00
Tushar Adhatrao	d5da2ab7e8	readme: add ollama-haskell library to community integrations (#7451 )	2024-11-17 15:18:04 -08:00
Vinh Nguyen	1c04117114	readme: add the VT app to the community integrations section (#7706 )	2024-11-17 14:35:41 -08:00
Jeffrey Morgan	8b4b243f5f	server: fix warnings in prompt_test.go (#7710 )	2024-11-17 13:01:04 -08:00
Jeffrey Morgan	b42a596425	docs: add customization section in linux.md (#7709 )	2024-11-17 11:48:12 -08:00
Daniel Hiltgen	4759d879f2	Install support for jetpacks (#7632 ) Follow up to #7217 - merge after release	2024-11-15 16:47:54 -08:00
Jesse Gross	d875e99e46	runner.go: Propagate panics back to the user. This is a partial revert of `8a35bb92` "runner.go: Increase survivability of main processing loop", removing the panic handler. Although we want to avoid errors taking down the runner, we also should make the user aware of problems when they happen. In the future, we can restructure things so both parts are true.	2024-11-15 11:52:25 -08:00
Jesse Gross	8a35bb926e	runner.go: Increase survivability of main processing loop Currently, if an error occurs during the prep stages (such as tokenizing) of a single request, it will only affect that request. However, if an error happens during decoding, it can take down the entire runner. Instead, it's better to drop the tokens that triggered the error and try to keep going. However, we also need to stop when we run out of tokens, otherwise, this just causes an infinite loop. This is likely the cause of at least some of the hanging issues that have been reported. Bug #7573	2024-11-14 17:18:41 -08:00
Daniel Hiltgen	a0ea067b63	build: fix arm container image (#7674 ) Fix a rebase glitch from the old C++ runner build model	2024-11-14 16:02:01 -08:00
Patrick Devine	4efb98cb4f	add line numbers for parser errors (#7326 )	2024-11-14 13:59:44 -08:00
Bruce MacDonald	0679d491fe	chore(deps): bump golang.org/x dependencies (#7655 ) - golang.org/x/sync v0.3.0 -> v0.9.0 - golang.org/x/image v0.14.0 -> v0.22.0 - golang.org/x/text v0.15.0 -> v0.20.0	2024-11-14 13:58:25 -08:00
Jesse Gross	c25ffde91d	runner.go: Don't trim whitespace from inputs It's possible to get prompts that consist entirely of whitespace - this is most likely to happen when generating embeddings. Currently, we will trim this away, leaving an empty prompt, which will then generate an error. Generating embeddings from whitespace should not trigger an error, as this may break pipelines. It's better to just leave the whitespace in place and process what we are given. This is consistent with past versions of Ollama. Bug #7578	2024-11-14 11:23:06 -08:00
Jesse Gross	17b386a891	runner.go: Enforce NUM_PARALLEL directly in the runner NUM_PARALEL is currently enforced by the Ollama server process - it will only issue requests to the runner if the maximum number of concurrent requests has not been exceeded. Although this should be sufficient, it is good for the runner to protect its own data structures. Currently, if too many requests get through to the runner, they will just get stuck and never return. This may help with reports of Ollama hanging, though it is unclear how it would actually occur. Bug #7573	2024-11-14 11:21:59 -08:00
Michael Yang	549c2bdfcf	Merge pull request #7657 from ollama/mxyng/sync fix(mllama): sync backend between batches	2024-11-14 09:40:04 -08:00
Blake Mizerany	67691e410d	cmd: preserve exact bytes when displaying template/system layers (#7586 )	2024-11-13 23:53:30 -08:00
Michael Yang	5b3393b6a2	fix(mllama): sync backend between batches	2024-11-13 16:37:21 -08:00
Jesse Gross	d7eb05b936	runner.go: Fix off-by-one for num predicted	2024-11-12 11:35:57 -08:00
Daniel Hiltgen	636a743c2b	CI: give windows lint more time (#7635 ) It looks like 8 minutes isn't quite enough and we're seeing sporadic timeouts	2024-11-12 11:22:39 -08:00
Daniel Hiltgen	df011054fa	Jetpack support for Go server (#7217 ) This adds support for the Jetson JetPack variants into the Go runner	2024-11-12 10:31:52 -08:00
Daniel Hiltgen	ac07160c8d	doc: capture numeric group requirement (#6941 ) Docker uses the container filesystem for name resolution, so we can't guide users to use the name of the host group. Instead they must specify the numeric ID.	2024-11-12 09:13:23 -08:00
Daniel Hiltgen	6606e4243c	docs: Capture docker cgroup workaround (#7519 ) GPU support can break on some systems after a while. This captures a known workaround to solve the problem.	2024-11-12 09:12:50 -08:00
Jesse Gross	65973ceb64	runner.go: Make KV entry accounting more robust The structure of the accounting for KV cache shifting was carried over from the old runner but it now doesn't feel natural with the new runner. There are a number of invariants that should hold true but are difficult to reason about. There is at least one bug report that would imply that the invariants are not holding. This reduces the number of implicit assumptions and is more forgiving of unexpected situations. It also improves behavior around which input tokens are kept when truncation occurs. Bug #7545	2024-11-11 20:23:03 -08:00
Joey Zheng	bebef1e50d	readme: add aichat terminal app to community integrations (#7418 )	2024-11-11 16:44:46 -08:00
Evan	d48c1c5a44	api: fix typos in Go Doc comments (#7620 )	2024-11-11 16:21:58 -08:00

1 2 3 4 5 ...

3687 commits