Xuan Son Nguyen and GitHub
ad3a0505e3
Server: clean up OAI params parsing function ( #6284 )
...
* server: clean up oai parsing function
* fix response_format
* fix empty response_format
* minor fixes
* add TODO for logprobs
* update docs
2024-03-25 09:42:17 +01:00
Xuan Son Nguyen and GitHub
91f8ad167d
Server: version bump for httplib and json ( #6169 )
...
* server: version bump for httplib and json
* fix build
* bring back content_length
2024-03-20 13:30:36 +01:00
Xuan Son Nguyen and GitHub
dfbfdd60f9
readme : add wllama as a wasm binding ( #6100 )
2024-03-16 17:42:08 +02:00
Xuan Son Nguyen and GitHub
aab606a11f
llama : add Orion chat template ( #6066 )
2024-03-15 10:44:57 +02:00
Xuan Son Nguyen and GitHub
99b71c068f
Server: Use multi-task for embeddings endpoint ( #6001 )
...
* use multitask for embd endpoint
* specify types
* remove redundant {"n_predict", 0}
2024-03-13 11:39:11 +01:00
Xuan Son Nguyen and GitHub
caa106d4e0
Server: format error to json ( #5961 )
...
* server: format error to json
* server: do not crash on grammar error
* fix api key test case
* revert limit max n_predict
* small fix
* correct coding style
* update completion.js
* launch_slot_with_task
* update docs
* update_slots
* update webui
* update readme
2024-03-11 10:56:41 +01:00
Xuan Son Nguyen and GitHub
950ba1ab84
Server: reorganize some http logic ( #5939 )
...
* refactor static file handler
* use set_pre_routing_handler for validate_api_key
* merge embedding handlers
* correct http verb for endpoints
* fix embedding response
* fix test case CORS Options
* fix code style
2024-03-09 11:27:53 +01:00
Xuan Son Nguyen and GitHub
4ffcdce2ff
add alias for chat template ( #5858 )
2024-03-04 12:22:08 +01:00
Xuan Son Nguyen and GitHub
6c32d8c7ad
llama : refactor internal quantization functions ( #5830 )
2024-03-02 16:19:09 +02:00
Xuan Son Nguyen and GitHub
052051d8ae
Server: normalize naming ( #5779 )
...
* server: normalize naming
* fix spacing
2024-02-29 21:42:11 +01:00
a693bea1e6
server : hit Ctrl+C twice to exit ( #5734 )
...
* server: twice ctrl+C to exit
* std::atomic_flag
* sigint: message
* sigint: stderr
* Update examples/server/server.cpp
Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com >
---------
Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com >
2024-02-28 10:55:37 +02:00
Xuan Son Nguyen and GitHub
b11a93df41
fix server hangs on empty prompt ( #5733 )
2024-02-26 23:15:48 +01:00
Xuan Son Nguyen and GitHub
373ee3fbba
Add Gemma chat template ( #5665 )
...
* add gemma chat template
* gemma: only apply system_prompt on non-model message
2024-02-22 19:10:21 +01:00
Xuan Son Nguyen and GitHub
a46f50747b
server : fallback to chatml, add AlphaMonarch chat template ( #5628 )
...
* server: fallback to chatml
* add new chat template
* server: add AlphaMonarch to test chat template
* server: only check model template if there is no custom tmpl
* remove TODO
2024-02-22 10:33:24 +02:00
Xuan Son Nguyen and GitHub
7c8bcc11dc
Add docs for llama_chat_apply_template ( #5645 )
...
* add docs for llama_chat_apply_template
* fix typo
2024-02-22 00:31:00 +01:00
9c405c9f9a
Server: use llama_chat_apply_template ( #5593 )
...
* server: use llama_chat_apply_template
* server: remove trailing space
* server: fix format_chat
* server: fix help message
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
* server: fix formatted_chat
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-02-20 15:58:27 +01:00
Xuan Son Nguyen and GitHub
11b12de39b
llama : add llama_chat_apply_template() ( #5538 )
...
* llama: add llama_chat_apply_template
* test-chat-template: remove dedundant vector
* chat_template: do not use std::string for buffer
* add clarification for llama_chat_apply_template
* llama_chat_apply_template: add zephyr template
* llama_chat_apply_template: correct docs
* llama_chat_apply_template: use term "chat" everywhere
* llama_chat_apply_template: change variable name to "tmpl"
2024-02-19 10:23:37 +02:00
907e08c110
server : add llama2 chat template ( #5425 )
...
* server: add mistral chat template
* server: fix typo
* server: rename template mistral to llama2
* server: format_llama2: remove BOS
* server: validate "--chat-template" argument
* server: clean up using_chatml variable
Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com >
---------
Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com >
2024-02-11 12:16:22 +02:00
Xuan Son Nguyen and GitHub
6b91b1e0a9
docker : add build for SYCL, Vulkan + update readme ( #5228 )
...
* add vulkan dockerfile
* intel dockerfile: compile sycl by default
* fix vulkan dockerfile
* add docs for vulkan
* docs: sycl build in docker
* docs: remove trailing spaces
* docs: sycl: add docker section
* docs: clarify install vulkan SDK outside docker
* sycl: use intel/oneapi-basekit docker image
* docs: correct TOC
* docs: correct docker image for Intel oneMKL
2024-02-02 09:56:31 +02:00
Xuan Son Nguyen and GitHub
48c857aa10
server : refactored the task processing logic ( #5065 )
...
* server: add llama_server_queue struct
* server: add llama_server_response_event
* server: add comments
* server: move all mutexes away from server.cpp
* server: correct multitask response
* server: only add back deferred tasks when one slot is available
* server: fix a race condition cause by "request_completion"
2024-01-26 14:42:20 +02:00
2bed4aa3f3
devops : add intel oneapi dockerfile ( #5068 )
...
Co-authored-by: Xuan Son Nguyen <xuanson.nguyen@snowpack.eu >
2024-01-23 09:11:39 +02:00
821f0a271e
server : defer tasks when "slot unavailable" ( #5018 )
...
* server: defer task when no slot is available
* remove unnecessary log
---------
Co-authored-by: Xuan Son Nguyen <xuanson.nguyen@snowpack.eu >
2024-01-18 22:33:05 +02:00