ollama

Author	SHA1	Message	Date
Patrick Sy	6cdf27d154	readme: add Alfred Ollama to community integrations (#7724 )	2024-11-18 19:33:23 -08:00
frob	5c18e66384	Notify the user if systemd is not running (#6693 ) Co-authored-by: Richard Lyons <frob@cloudstaff.com>	2024-11-18 15:02:41 -08:00
Daniel Hiltgen	35096a7eff	win: add right click menu support (#7727 ) Enable both left and right click on the pop-up menu	2024-11-18 14:39:52 -08:00
Daniel Hiltgen	81d55d3e4d	fix index out of range on zero layer metal load (#7696 ) If the model doesn't fit any layers on metal, and we load zero layers we would panic trying to look up the GPU size during scheduling ops	2024-11-18 11:48:13 -08:00
Vinh Nguyen	a14f76491d	readme: improve Community Integrations section (#7718 )	2024-11-17 19:30:22 -08:00
Nicolas Bonamy	760cfa27e5	readme: add Witsy and multi-llm-ts to community integrations (#7713 )	2024-11-17 16:33:10 -08:00
Darius Kocar	c9a5aca3da	readme: add Perfect Memory AI to community integrations (#7431 )	2024-11-17 15:19:26 -08:00
Tushar Adhatrao	d5da2ab7e8	readme: add ollama-haskell library to community integrations (#7451 )	2024-11-17 15:18:04 -08:00
Vinh Nguyen	1c04117114	readme: add the VT app to the community integrations section (#7706 )	2024-11-17 14:35:41 -08:00
Jeffrey Morgan	8b4b243f5f	server: fix warnings in prompt_test.go (#7710 )	2024-11-17 13:01:04 -08:00
Jeffrey Morgan	b42a596425	docs: add customization section in linux.md (#7709 )	2024-11-17 11:48:12 -08:00
Daniel Hiltgen	4759d879f2	Install support for jetpacks (#7632 ) Follow up to #7217 - merge after release	2024-11-15 16:47:54 -08:00
Jesse Gross	d875e99e46	runner.go: Propagate panics back to the user. This is a partial revert of `8a35bb92` "runner.go: Increase survivability of main processing loop", removing the panic handler. Although we want to avoid errors taking down the runner, we also should make the user aware of problems when they happen. In the future, we can restructure things so both parts are true.	2024-11-15 11:52:25 -08:00
Jesse Gross	8a35bb926e	runner.go: Increase survivability of main processing loop Currently, if an error occurs during the prep stages (such as tokenizing) of a single request, it will only affect that request. However, if an error happens during decoding, it can take down the entire runner. Instead, it's better to drop the tokens that triggered the error and try to keep going. However, we also need to stop when we run out of tokens, otherwise, this just causes an infinite loop. This is likely the cause of at least some of the hanging issues that have been reported. Bug #7573	2024-11-14 17:18:41 -08:00
Daniel Hiltgen	a0ea067b63	build: fix arm container image (#7674 ) Fix a rebase glitch from the old C++ runner build model	2024-11-14 16:02:01 -08:00
Patrick Devine	4efb98cb4f	add line numbers for parser errors (#7326 )	2024-11-14 13:59:44 -08:00
Bruce MacDonald	0679d491fe	chore(deps): bump golang.org/x dependencies (#7655 ) - golang.org/x/sync v0.3.0 -> v0.9.0 - golang.org/x/image v0.14.0 -> v0.22.0 - golang.org/x/text v0.15.0 -> v0.20.0	2024-11-14 13:58:25 -08:00
Jesse Gross	c25ffde91d	runner.go: Don't trim whitespace from inputs It's possible to get prompts that consist entirely of whitespace - this is most likely to happen when generating embeddings. Currently, we will trim this away, leaving an empty prompt, which will then generate an error. Generating embeddings from whitespace should not trigger an error, as this may break pipelines. It's better to just leave the whitespace in place and process what we are given. This is consistent with past versions of Ollama. Bug #7578	2024-11-14 11:23:06 -08:00
Jesse Gross	17b386a891	runner.go: Enforce NUM_PARALLEL directly in the runner NUM_PARALEL is currently enforced by the Ollama server process - it will only issue requests to the runner if the maximum number of concurrent requests has not been exceeded. Although this should be sufficient, it is good for the runner to protect its own data structures. Currently, if too many requests get through to the runner, they will just get stuck and never return. This may help with reports of Ollama hanging, though it is unclear how it would actually occur. Bug #7573	2024-11-14 11:21:59 -08:00
Michael Yang	549c2bdfcf	Merge pull request #7657 from ollama/mxyng/sync fix(mllama): sync backend between batches	2024-11-14 09:40:04 -08:00
Blake Mizerany	67691e410d	cmd: preserve exact bytes when displaying template/system layers (#7586 )	2024-11-13 23:53:30 -08:00
Michael Yang	5b3393b6a2	fix(mllama): sync backend between batches	2024-11-13 16:37:21 -08:00
Jesse Gross	d7eb05b936	runner.go: Fix off-by-one for num predicted	2024-11-12 11:35:57 -08:00
Daniel Hiltgen	636a743c2b	CI: give windows lint more time (#7635 ) It looks like 8 minutes isn't quite enough and we're seeing sporadic timeouts	2024-11-12 11:22:39 -08:00
Daniel Hiltgen	df011054fa	Jetpack support for Go server (#7217 ) This adds support for the Jetson JetPack variants into the Go runner	2024-11-12 10:31:52 -08:00
Daniel Hiltgen	ac07160c8d	doc: capture numeric group requirement (#6941 ) Docker uses the container filesystem for name resolution, so we can't guide users to use the name of the host group. Instead they must specify the numeric ID.	2024-11-12 09:13:23 -08:00
Daniel Hiltgen	6606e4243c	docs: Capture docker cgroup workaround (#7519 ) GPU support can break on some systems after a while. This captures a known workaround to solve the problem.	2024-11-12 09:12:50 -08:00
Jesse Gross	65973ceb64	runner.go: Make KV entry accounting more robust The structure of the accounting for KV cache shifting was carried over from the old runner but it now doesn't feel natural with the new runner. There are a number of invariants that should hold true but are difficult to reason about. There is at least one bug report that would imply that the invariants are not holding. This reduces the number of implicit assumptions and is more forgiving of unexpected situations. It also improves behavior around which input tokens are kept when truncation occurs. Bug #7545	2024-11-11 20:23:03 -08:00
Joey Zheng	bebef1e50d	readme: add aichat terminal app to community integrations (#7418 )	2024-11-11 16:44:46 -08:00
Evan	d48c1c5a44	api: fix typos in Go Doc comments (#7620 )	2024-11-11 16:21:58 -08:00
Prasad Bhalerao	36a8372b28	readme: add GoLamify to community integrations (#7521 )	2024-11-10 22:38:18 -08:00
Ivo Stoykov	4e94227b5d	readme: add browser extension that enables using Ollama for interacting with web pages (#5827 )	2024-11-10 22:14:22 -08:00
frances720	479d551766	docs: add mentions of Llama 3.2 (#7517 )	2024-11-10 19:04:23 -08:00
Evan	76b2b723b2	api: fix typo in python ClientFromEnvironment docs (#7604 )	2024-11-10 17:30:27 -08:00
Arhan Busam	b8d77cdeab	readme: add llama3.2-vision to model list (#7580 )	2024-11-10 13:36:25 -08:00
Jesse Gross	c2e8cbaa14	runner.go: Check for zero length images If we get a request with a zero length image, it will result in an out-of-bounds error when we pass the data to the image encoder.	2024-11-08 09:39:32 -08:00
Edward J. Schwartz	771fab1dd8	docs: update langchainpy.md with proper model name (#7527 )	2024-11-08 09:36:17 -08:00
Daniel Hiltgen	3a5239e6bf	Set macos min version for all architectures (#7579 )	2024-11-08 09:27:04 -08:00
Daniel Hiltgen	3d25e7bf8c	win: remove preview title from installer (#7529 ) This should have been in #7347 but was overlooked.	2024-11-07 14:26:47 -08:00
Daniel Hiltgen	1618700c5a	Workaround buggy P2P ROCm copy on windows (#7466 ) This enables the workaround code only for windows which should help windows users with muliple AMD GPUs	2024-11-07 14:26:31 -08:00
Daniel Hiltgen	b111aa5a91	Debug logging for nvcuda init (#7532 ) Some users are reporting crashes during nvcuda.dll initialization on windows. This should help narrow down where things are going bad.	2024-11-07 14:25:53 -08:00
Daniel Hiltgen	9e83e550e1	Align rocm compiler flags (#7467 ) Bring consistency with the old generate script behavior	2024-11-07 10:20:50 -08:00
Daniel Hiltgen	fc2a0715df	Be explicit for gpu library link dir (#7560 ) On linux nvcc isn't automatically linking to the same cuda version.	2024-11-07 09:20:40 -08:00
Jesse Gross	3020d2dc58	docs: OLLAMA_NEW_RUNNERS no longer exists	2024-11-06 14:39:02 -08:00
Jesse Gross	a909417602	runner.go: Remove unused arguments Now that server.cpp is gone, we don't need to keep passing arguments that were only ignored and only kept for compatibility.	2024-11-06 13:32:18 -08:00
Jesse Gross	6cd566872b	sched: Lift parallel restriction for multimodal models except mllama The Go runner does not have a problem with supporting parallel requests for most multimodal models. Now that we won't be potentially falling back to server.cpp, this restriction can be lifted. However, the new mllama model can't support parallel requests, so we will need to keep a restriction for that.	2024-11-06 13:32:18 -08:00
RAPID ARCHITECT	9d71bcc3e2	Update README.md (#7516 ) added reddit rate below hexabot, ollama powered reddit search and analysis with streamlit for the intervace	2024-11-05 15:07:25 -08:00
Daniel Hiltgen	a4c70fe157	One corrupt manifest should not wedge model operations (#7515 ) One potential failure mode is an empty file which bubbles up as an EOF error, leading to all pulls and listing operations failing. Instead, continue and warn about the corrupt manifest. This also allows re-pulling the corrupt manifest to repair the system.	2024-11-05 14:21:45 -08:00
Jesse Gross	34a75102f7	prompt: Use a single token when estimating mllama context size Currently we assume that images take 768 tokens of context size for the purposes of clipping old messages that exceed the context window. However, our mllama implementation stores the full image embedding in a single token. As a result, there is significant waste of context space. Ideally, we would handle this more generically and have the implementation report the number of tokens. However, at the moment this would just result in a similar set of 'if' conditions in the runner plus APIs to report it back. So for now, we just keep this simple.	2024-11-05 10:11:50 -08:00
Med Marrouchi	4157d1f7b6	readme: add Hexabot to the list of community integrations	2024-11-05 09:06:38 -08:00

1 2 3 4 5 ...

3667 commits