baalajimaestro/llama.cpp

Author	SHA1	Message	Date
Andrei Betlen	9e5b6d675a	Improve logging messages	2023-05-03 10:28:10 -04:00
Andrei Betlen	43f2907e3a	Support smaller state sizes	2023-05-03 09:33:50 -04:00
Andrei Betlen	1d47cce222	Update llama.cpp	2023-05-03 09:33:30 -04:00
Lucas Doyle	b9098b0ef7	llama_cpp server: prompt is a string Not sure why this union type was here but taking a look at llama.py, prompt is only ever processed as a string for completion This was breaking types when generating an openapi client	2023-05-02 14:47:07 -07:00
Matt Hoffner	f97ff3c5bb	Update llama_cpp.py	2023-05-01 20:40:06 -07:00
Andrei	7ab08b8d10	Merge branch 'main' into better-server-params-and-fields	2023-05-01 22:45:57 -04:00
Andrei Betlen	9eafc4c49a	Refactor server to use factory	2023-05-01 22:38:46 -04:00
Andrei Betlen	dd9ad1c759	Formatting	2023-05-01 21:51:16 -04:00
Lucas Doyle	dbbfc4ba2f	llama_cpp server: fix to ChatCompletionRequestMessage When I generate a client, it breaks because it fails to process the schema of ChatCompletionRequestMessage These fix that: - I think `Union[Literal["user"], Literal["channel"], ...]` is the same as Literal["user", "channel", ...] - Turns out default value `Literal["user"]` isn't JSON serializable, so replace with "user"	2023-05-01 15:38:19 -07:00
Lucas Doyle	fa2a61e065	llama_cpp server: fields for the embedding endpoint	2023-05-01 15:38:19 -07:00
Lucas Doyle	8dcbf65a45	llama_cpp server: define fields for chat completions Slight refactor for common fields shared between completion and chat completion	2023-05-01 15:38:19 -07:00
Lucas Doyle	978b6daf93	llama_cpp server: add some more information to fields for completions	2023-05-01 15:38:19 -07:00
Lucas Doyle	a5aa6c1478	llama_cpp server: add missing top_k param to CreateChatCompletionRequest `llama.create_chat_completion` definitely has a `top_k` argument, but its missing from `CreateChatCompletionRequest`. decision: add it	2023-05-01 15:38:19 -07:00
Lucas Doyle	1e42913599	llama_cpp server: move logprobs to supported I think this is actually supported (its in the arguments of `LLama.__call__`, which is how the completion is invoked). decision: mark as supported	2023-05-01 15:38:19 -07:00
Lucas Doyle	b47b9549d5	llama_cpp server: delete some ignored / unused parameters `n`, `presence_penalty`, `frequency_penalty`, `best_of`, `logit_bias`, `user`: not supported, excluded from the calls into llama. decision: delete it	2023-05-01 15:38:19 -07:00
Lucas Doyle	e40fcb0575	llama_cpp server: mark model as required `model` is ignored, but currently marked "optional"... on the one hand could mark "required" to make it explicit in case the server supports multiple llama's at the same time, but also could delete it since its ignored. decision: mark it required for the sake of openai api compatibility. I think out of all parameters, `model` is probably the most important one for people to keep using even if its ignored for now.	2023-05-01 15:38:19 -07:00
Andrei Betlen	b6747f722e	Fix logprob calculation. Fixes #134	2023-05-01 17:45:08 -04:00
Andrei Betlen	9ff9cdd7fc	Fix import error	2023-05-01 15:11:15 -04:00
Andrei Betlen	350a1769e1	Update sampling api	2023-05-01 14:47:55 -04:00
Andrei Betlen	7837c3fdc7	Fix return types and import comments	2023-05-01 14:02:06 -04:00
Andrei Betlen	ccf1ed54ae	Merge branch 'main' of github.com:abetlen/llama_cpp_python into main	2023-05-01 11:35:14 -04:00
Andrei Betlen	80184a286c	Update llama.cpp	2023-05-01 10:44:28 -04:00
Lucas Doyle	efe8e6f879	llama_cpp server: slight refactor to init_llama function Define an init_llama function that starts llama with supplied settings instead of just doing it in the global context of app.py This allows the test to be less brittle by not needing to mess with os.environ, then importing the app	2023-04-29 11:42:23 -07:00
Lucas Doyle	6d8db9d017	tests: simple test for server module	2023-04-29 11:42:20 -07:00
Lucas Doyle	468377b0e2	llama_cpp server: app is now importable, still runnable as a module	2023-04-29 11:41:25 -07:00
Andrei	755f9fa455	Merge pull request #118 from SagsMug/main Fix UnicodeDecodeError permanently	2023-04-29 07:19:01 -04:00
Mug	18a0c10032	Remove excessive errors="ignore" and add utf8 test	2023-04-29 12:19:22 +02:00
Andrei Betlen	ea0faabae1	Update llama.cpp	2023-04-28 15:32:43 -04:00
Mug	b7d14efc8b	Python weirdness	2023-04-28 13:20:31 +02:00
Mug	eed61289b6	Dont detect off tokens, detect off detokenized utf8	2023-04-28 13:16:18 +02:00
Mug	3a98747026	One day, i'll fix off by 1 errors permanently too	2023-04-28 12:54:28 +02:00
Mug	c39547a986	Detect multi-byte responses and wait	2023-04-28 12:50:30 +02:00
Andrei Betlen	9339929f56	Update llama.cpp	2023-04-26 20:00:54 -04:00
Mug	5f81400fcb	Also ignore errors on input prompts	2023-04-26 14:45:51 +02:00
Mug	be2c961bc9	Merge branch 'main' of https://github.com/abetlen/llama-cpp-python	2023-04-26 14:38:09 +02:00
Mug	c4a8491d42	Fix decode errors permanently	2023-04-26 14:37:06 +02:00
Andrei Betlen	cbd26fdcc1	Update llama.cpp	2023-04-25 19:03:41 -04:00
Andrei Betlen	3cab3ef4cb	Update n_batch for server	2023-04-25 09:11:32 -04:00
Andrei Betlen	cc706fb944	Add ctx check and re-order __init__. Closes #112	2023-04-25 09:00:53 -04:00
Andrei Betlen	d484c5634e	Bugfix: Check cache keys as prefix to prompt tokens	2023-04-24 22:18:54 -04:00
Andrei Betlen	cbe95bbb75	Add cache implementation using llama state	2023-04-24 19:54:41 -04:00
Andrei Betlen	2c359a28ff	Merge branch 'main' of github.com:abetlen/llama_cpp_python into main	2023-04-24 17:51:27 -04:00
Andrei Betlen	197cf80601	Add save/load state api for Llama class	2023-04-24 17:51:25 -04:00
Andrei Betlen	86f8e5ad91	Refactor internal state for Llama class	2023-04-24 15:47:54 -04:00
Andrei	f37456133a	Merge pull request #108 from eiery/main Update n_batch default to 512 to match upstream llama.cpp	2023-04-24 13:48:09 -04:00
Andrei Betlen	02cf881317	Update llama.cpp	2023-04-24 09:30:10 -04:00
eiery	aa12d8a81f	Update llama.py update n_batch default to 512 to match upstream llama.cpp	2023-04-23 20:56:40 -04:00
Andrei Betlen	7230599593	Disable mmap when applying lora weights. Closes #107	2023-04-23 14:53:17 -04:00
Andrei Betlen	e99caedbbd	Update llama.cpp	2023-04-22 19:50:28 -04:00
Andrei Betlen	1eb130a6b2	Update llama.cpp	2023-04-21 17:40:27 -04:00
Andrei Betlen	e4647c75ec	Add use_mmap flag to server	2023-04-19 15:57:46 -04:00
Andrei Betlen	0df4d69c20	If lora base is not set avoid re-loading the model by passing NULL	2023-04-18 23:45:25 -04:00
Andrei Betlen	95c0dc134e	Update type signature to allow for null pointer to be passed.	2023-04-18 23:44:46 -04:00
Andrei Betlen	453e517fd5	Add seperate lora_base path for applying LoRA to quantized models using original unquantized model weights.	2023-04-18 10:20:46 -04:00
Andrei Betlen	eb7f278cc6	Add lora_path parameter to Llama model	2023-04-18 01:43:44 -04:00
Andrei Betlen	35abf89552	Add bindings for LoRA adapters. Closes #88	2023-04-18 01:30:04 -04:00
Andrei Betlen	89856ef00d	Bugfix: only eval new tokens	2023-04-15 17:32:53 -04:00
Andrei Betlen	92c077136d	Add experimental cache	2023-04-15 12:03:09 -04:00
Andrei Betlen	a6372a7ae5	Update stop sequences for chat	2023-04-15 12:02:48 -04:00
Andrei Betlen	83b2be6dc4	Update chat parameters	2023-04-15 11:58:43 -04:00
Andrei Betlen	62087514c6	Update chat prompt	2023-04-15 11:58:19 -04:00
Andrei Betlen	02f9fb82fb	Bugfix	2023-04-15 11:39:52 -04:00
Andrei Betlen	3cd67c7bd7	Add type annotations	2023-04-15 11:39:21 -04:00
Andrei Betlen	d7de0e8014	Bugfix	2023-04-15 00:08:04 -04:00
Andrei Betlen	e90e122f2a	Use clear	2023-04-14 23:33:18 -04:00
Andrei Betlen	ac7068a469	Track generated tokens internally	2023-04-14 23:33:00 -04:00
Andrei Betlen	6e298d8fca	Set kv cache size to f16 by default	2023-04-14 22:21:19 -04:00
Andrei Betlen	6c7cec0c65	Fix completion request	2023-04-14 10:01:15 -04:00
Andrei Betlen	6153baab2d	Clean up logprobs implementation	2023-04-14 09:59:33 -04:00
Andrei Betlen	26cc4ee029	Fix signature for stop parameter	2023-04-14 09:59:08 -04:00
Andrei Betlen	6595ad84bf	Add field to disable reseting between generations	2023-04-13 00:28:00 -04:00
Andrei Betlen	22fa5a621f	Revert "Deprecate generate method" This reverts commit `6cf5876538`.	2023-04-13 00:19:55 -04:00
Andrei Betlen	4f5f99ef2a	Formatting	2023-04-12 22:40:12 -04:00
Andrei Betlen	0daf16defc	Enable logprobs on completion endpoint	2023-04-12 19:08:11 -04:00
Andrei Betlen	19598ac4e8	Fix threading bug. Closes #62	2023-04-12 19:07:53 -04:00
Andrei Betlen	005c78d26c	Update llama.cpp	2023-04-12 14:29:00 -04:00
Andrei Betlen	c854c2564b	Don't serialize stateful parameters	2023-04-12 14:07:14 -04:00
Andrei Betlen	2f9b649005	Style fix	2023-04-12 14:06:22 -04:00
Andrei Betlen	6cf5876538	Deprecate generate method	2023-04-12 14:06:04 -04:00
Andrei Betlen	b3805bb9cc	Implement logprobs parameter for text completion. Closes #2	2023-04-12 14:05:11 -04:00
Andrei Betlen	9f1e565594	Update llama.cpp	2023-04-11 11:59:03 -04:00
Andrei Betlen	213cc5c340	Remove async from function signature to avoid blocking the server	2023-04-11 11:54:31 -04:00
Mug	2559e5af9b	Changed the environment variable name into "LLAMA_CPP_LIB"	2023-04-10 17:27:17 +02:00
Mug	ee71ce8ab7	Make windows users happy (hopefully)	2023-04-10 17:12:25 +02:00
Mug	cf339c9b3c	Better custom library debugging	2023-04-10 17:06:58 +02:00
Mug	4132293d2d	Merge branch 'main' of https://github.com/abetlen/llama-cpp-python into local-lib	2023-04-10 17:00:42 +02:00
Mug	76131d5bb8	Use environment variable for library override	2023-04-10 17:00:35 +02:00
Andrei Betlen	1f67ad2a0b	Add use_mmap option	2023-04-10 02:11:35 -04:00
Andrei Betlen	c3c2623e8b	Update llama.cpp	2023-04-09 22:01:33 -04:00
Andrei Betlen	314ce7d1cc	Fix cpu count default	2023-04-08 19:54:04 -04:00
Andrei Betlen	3fbc06361f	Formatting	2023-04-08 16:01:45 -04:00
Andrei Betlen	0067c1a588	Formatting	2023-04-08 16:01:18 -04:00
Andrei Betlen	38f442deb0	Bugfix: Wrong size of embeddings. Closes #47	2023-04-08 15:05:33 -04:00
Andrei Betlen	ae3e9c3d6f	Update shared library extension for macos	2023-04-08 02:45:21 -04:00
Andrei Betlen	da539cc2ee	Safer calculation of default n_threads	2023-04-06 21:22:19 -04:00
Andrei Betlen	930db37dd2	Merge branch 'main' of github.com:abetlen/llama_cpp_python into main	2023-04-06 21:07:38 -04:00
Andrei Betlen	55279b679d	Handle prompt list	2023-04-06 21:07:35 -04:00
MillionthOdin16	c283edd7f2	Set n_batch to default values and reduce thread count: Change batch size to the llama.cpp default of 8. I've seen issues in llama.cpp where batch size affects quality of generations. (It shouldn't) But in case that's still an issue I changed to default. Set auto-determined num of threads to 1/2 system count. ggml will sometimes lock cores at 100% while doing nothing. This is being addressed, but can cause bad experience for user if pegged at 100%	2023-04-05 18:17:29 -04:00
MillionthOdin16	76a82babef	Set n_batch to the default value of 8. I think this is leftover from when n_ctx was missing and n_batch was 2048.	2023-04-05 17:44:53 -04:00
Andrei Betlen	44448fb3a8	Add server as a subpackage	2023-04-05 16:23:25 -04:00

1 2 3 4

200 commits