ollama

Author	SHA1	Message	Date
saman-amd	ead27aa9fe	Add gfx1200 & gfx1201 support on linux (#9878 )	2025-03-27 07:35:19 -07:00
Vadim Grinco	45dbd14645	Merged latest ollama 0.6.2 and nasrally's Flash Attention patches (#5 ) * readme: add Ellama to list of community integrations (#9800) * readme: add screenpipe to community integrations (#9786) * Add support for ROCm gfx1151 (#9773) * conditionally enable parallel pipelines * sample: make mutations in transforms explicit (#9743) * updated minP to use early exit making use of sorted tokens * ml/backend/ggml: allocate memory with malloc when loading model (#9822) * runner: remove cache prompt flag from ollama runner (#9826) We do not need to bypass the prompt caching in the ollama runner yet, as only embedding models needed to bypass the prompt caching. When embedding models are implemented they can skip initializing this cache completely. * ollamarunner: Check for minBatch of context space when shifting Models can specify that a group of inputs need to be handled a single batch. However, context shifting didn't respect this and could trigger a break anyways. In this case, we should instead trigger a context shift earlier so that it occurs before the grouped batch. Note that there still some corner cases: - A long prompt that exceeds the context window can get truncated in the middle of an image. With the current models, this will result in the model not recognizing the image at all, which is pretty much the expected result with truncation. - The context window is set less than the minimum batch size. The only solution to this is to refuse to load the model with these settings. However, this can never occur with current models and default settings. Since users are unlikely to run into these scenarios, fixing them is left as a follow up. * Applied latest patches from McBane87 See this for details: https://github.com/whyvl/ollama-vulkan/issues/7#issuecomment-2708820861 Signed-off-by: Vadim Grinco <vadim@grinco.eu> * Add ability to enable flash attention on vulkan (#4) * discover: add flash attention handling for vulkan * envconfig: fix typo in config.go As part of the process some code was refactored and I added a new field FlashAttention to GpuInfo since the previous solution didn't allow for a granular check via vulkan extensions. As a side effect, this now allows for granular per-device FA support checking in other places --------- Signed-off-by: Vadim Grinco <vadim@grinco.eu> Co-authored-by: zeo <108888572+zeozeozeo@users.noreply.github.com> Co-authored-by: Louis Beaumont <louis.beaumont@gmail.com> Co-authored-by: Daniel Hiltgen <dhiltgen@users.noreply.github.com> Co-authored-by: Michael Yang <mxyng@pm.me> Co-authored-by: Parth Sareen <parth.sareen@ollama.com> Co-authored-by: Jeffrey Morgan <jmorganca@gmail.com> Co-authored-by: Bruce MacDonald <brucewmacdonald@gmail.com> Co-authored-by: Jesse Gross <jesse@ollama.com> Co-authored-by: Nikita <50599445+nasrally@users.noreply.github.com>	2025-03-23 12:27:37 +01:00
Michael Yang	74bd09652d	ml/backend/ggml: load tensors in 32KiB chunks	2025-03-21 14:43:52 -07:00
Jesse Gross	0ff28758b3	ollamarunner: Provide mechanism for backends to report loading progress This enables the runner to report progress back to the Ollama server, both for showing status to the user and also to prevent the server from killing the runner if it thinks things have stalled. Most of the infrastructure was already there, this extends it to be available to the backends.	2025-03-21 10:44:26 -07:00
Bruce MacDonald	df94175a0f	ggml: return error on failure to read tensor data (#9872 ) When converting a ggml model if there is a failure to read tensor data a nil error value was being returned. It should be assigned to the actual error from reading.	2025-03-18 16:51:33 -07:00
Michael Yang	021dcf089d	Merge pull request #9824 from ollama/mxyng/sched conditionally enable parallel pipelines	2025-03-17 15:41:37 -07:00
Jeffrey Morgan	364629b8d6	ml/backend/ggml: allocate memory with malloc when loading model (#9822 )	2025-03-17 13:32:40 -07:00
Michael Yang	4561fff36e	conditionally enable parallel pipelines	2025-03-17 09:46:07 -07:00
Vadim Grinco	640f0bb250	Pulled new upstream code for ggml-bulkan backend Signed-off-by: Vadim Grinco <vadim@grinco.eu>	2025-03-16 12:22:22 +01:00
Vadim Grinco	f77b9b99cd	Merge branch 'ollama_vanilla_stable' into vulkan	2025-03-15 21:27:18 +01:00
Vadim Grinco	d1939aa1c6	Fixes SIGSEGV: segmentation violation running gemma3 models on ollama 0.6.0 #21 Patch provided by McBane87 on https://github.com/whyvl/ollama-vulkan/issues/21 Signed-off-by: Vadim Grinco <vadim@grinco.eu>	2025-03-15 20:28:57 +01:00
shane.xb.qian	30d7a59ba8	ollama-debug.c: change 'ld' to 'PRIi64' * macOS has different definition per info from @mxyng	2025-03-13 17:10:37 +08:00
shane.xb.qian	85ab552028	ollama-debug.c: correct mistype Signed-off-by: shane.xb.qian <shane.qian@foxmail.com>	2025-03-12 22:32:30 +08:00
Vadim Grinco	d0afc677db	Merge branch 'vulkan' into ollama_vanilla_stable	2025-03-12 13:33:05 +01:00
Michael Yang	63a394068c	use 2d pooling	2025-03-11 14:49:20 -07:00
Michael Yang	c5cbe4fc2a	fallback to cpu	2025-03-11 14:49:19 -07:00
Michael Yang	9e4642e9b3	ollama debug tensor	2025-03-11 14:49:19 -07:00
Michael Yang	6b0486c216	duplicate token_embd to output	2025-03-11 14:49:19 -07:00
Michael Yang	8934324b72	use fast attention	2025-03-11 14:49:18 -07:00
Michael Yang	0df1800436	set non-causal attention	2025-03-11 14:49:18 -07:00
Michael Yang	4b037a97dc	add gemma vision encoder	2025-03-11 14:49:17 -07:00
Patrick Devine	5f74d1fd47	gemma2 impl	2025-03-11 14:35:08 -07:00
Michael Yang	9926eae015	fix: pad tensor item if ge zero this produces a nicer output since both positive and negative values produces the same width	2025-03-10 16:18:12 -07:00
Vadim Grinco	747898df04	Merge pull request #1 from ollama/main Merged from ollama/main	2025-03-08 08:56:12 +01:00
Jesse Gross	4100ed7bdd	ml: Add support for quantized KV cache Similar to the llama engine, quantizing the KV cache requires flash attention to be enabled through the Ollama server.	2025-03-07 18:43:39 -08:00
Jesse Gross	25f9b152f9	ggml-backend: Ensure allocation meet backend requirements Backends can impose additional alignment requirements on buffer sizes. We should ensure that we meet these or allocations can fail.	2025-03-07 18:43:39 -08:00
Jesse Gross	98272fbd58	additional review comments	2025-03-07 14:08:21 -08:00
Michael Yang	b27e8f3f10	ml/backend/ggml: use backend buffer type this ensures the tensor is created on the right buffer type for backends such as cpu	2025-03-07 14:08:21 -08:00
Michael Yang	45df786f09	comments	2025-03-07 14:08:21 -08:00
Michael Yang	daaf42e4a4	ml/backend/ggml: clean up	2025-03-07 14:08:21 -08:00
Michael Yang	2dc60d4620	ml/backend/ggml: offload vision to cpu temporary until tensor loading can accurately account for vision models	2025-03-07 14:08:21 -08:00
Michael Yang	b5312f30e8	ml/backend/ggml: handle tensor split	2025-03-07 14:08:21 -08:00
Michael Yang	26c2e0bd35	ml/backend/ggml: handle user specified cpu offloading	2025-03-07 14:08:21 -08:00
Michael Yang	bf920883d5	ml/backend/ggml: set cpu n_threads	2025-03-07 14:08:21 -08:00
Michael Yang	7bae7fa5ce	ml/backend/ggml: create tensor on specific backend some tensors should be created on specific backends to reduce number of copies and improve performance	2025-03-07 14:08:21 -08:00
Michael Yang	764e199d67	kvcache: create cache ctx per layer each cache layer creates and maintains its own context instead of using a large context for all layers	2025-03-07 14:08:21 -08:00
Michael Yang	bfce55db3d	model: load non-repeated tensors into multiple backends some tensors are expected to be used in repeating layers but are not themselves repeated. this change copies these tensors into the same backends as their repeating counterparts to minimize copying tensors between backends	2025-03-07 14:08:21 -08:00
Michael Yang	bab6f34dc0	ml/backend/ggml: update model loading for hybrid/multi backends use a similar strategy as llama.cpp for deciding where tensors should be allocated. this will be improved later to be aware of usable memory before assigning the tensor	2025-03-07 14:08:21 -08:00
Jeffrey Morgan	4289c74359	llama: fix kv loading on snowflake-arctic-embed models (#9536 )	2025-03-07 09:25:34 -08:00
Michael Yang	05a01fdecb	ml/backend/ggml: consolidate system info logging - output backend system info when initializing the backend. this ensures this information is always present without needing to be called explicitly - convert to structured logging - enumerate devices rather than backends since devices are ordered - track device indices grouped by device name	2025-03-04 15:14:31 -08:00
Michael Yang	ba7d31240e	fix: own lib/ollama directory expand backend loading error handling to catch more problems and log them instead of panicing	2025-03-03 13:01:18 -08:00
Jesse Gross	21aa666a1e	ml: Enable support for flash attention The GGML flash attention kernel has specific requirements for padding and permutation. This adds support to the KV cache for conforming to these requirements so that flash attention can be enabled. Flash attention can be used in the same situations as the llama engine and is enabled by the user in the same way.	2025-03-01 20:53:23 -08:00
Jesse Gross	ee141cc821	ml: Empty tensor constructor for tensors In cases where we allocate a tensor and then fully overwrite it with copied data, it is wasteful to first zero out the memory.	2025-03-01 20:53:23 -08:00
Jesse Gross	55e5776c44	ggml-backend: Store parent backend as part of tensor It can be important for a tensor to know what backend it came from - for example, to know if flash attention is enabled.	2025-03-01 20:53:23 -08:00
Jesse Gross	854a9195f3	attention: Remove unnecessary contiguous operations Prior to performing attention, we need to permute query, key and value. Currently we call Contiguous after each of these permutations, which is correct but expensive. Avoiding the 3 calls to Contiguous increases performance by over 20%. The permutations of query and key do not violate the continuity rules for mulmat and the Contiguous call can be simply removed. Value requires a different permutation and does require Contiguous. However, we can use the copy into the cache as a way to perform this without further overhead. To support this and avoid unexpected tensor shapes that are seen by models, we need tighter integration between attention, cache and backend. Future optimization will also likely need this structure - for example, flash attention has special padding requirements in the cache and other backends may have their own needs. This further contains the operations that go into attention so that these and other optimizations can be handled transparently. Models that have special requirements for attention can still implement their own version of it.	2025-03-01 20:53:23 -08:00
Michael Yang	3e8b8a1933	ml: update Context.Forward interface update Context.Forward to accept multiple tensors to match Context.Compute signature update Context.Forward to return Context such that it can be chained with Context.Compute	2025-02-27 22:27:16 +00:00
Michael Yang	53d2990d9b	model: add bos token if configured	2025-02-27 21:04:59 +00:00
Michael Yang	a59f665235	ml/backend/ggml: fix debug logging	2025-02-27 18:30:57 +00:00
Jeffrey Morgan	a5272130c4	ml/backend/ggml: follow on fixes after updating vendored code (#9388 ) Fixes sync filters and lowers CUDA version to 11.3 in test.yaml	2025-02-26 22:33:53 -08:00
Jeffrey Morgan	d7d7e99662	llama: update llama.cpp vendor code to commit d7cfe1ff (#9356 )	2025-02-26 20:34:44 -08:00

1 2 3 4

180 Commits