a693bea1e6
server : hit Ctrl+C twice to exit ( #5734 )
...
* server: twice ctrl+C to exit
* std::atomic_flag
* sigint: message
* sigint: stderr
* Update examples/server/server.cpp
Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com >
---------
Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com >
2024-02-28 10:55:37 +02:00
Xuan Son Nguyen and GitHub
b11a93df41
fix server hangs on empty prompt ( #5733 )
2024-02-26 23:15:48 +01:00
Xuan Son Nguyen and GitHub
373ee3fbba
Add Gemma chat template ( #5665 )
...
* add gemma chat template
* gemma: only apply system_prompt on non-model message
2024-02-22 19:10:21 +01:00
Xuan Son Nguyen and GitHub
a46f50747b
server : fallback to chatml, add AlphaMonarch chat template ( #5628 )
...
* server: fallback to chatml
* add new chat template
* server: add AlphaMonarch to test chat template
* server: only check model template if there is no custom tmpl
* remove TODO
2024-02-22 10:33:24 +02:00
Xuan Son Nguyen and GitHub
7c8bcc11dc
Add docs for llama_chat_apply_template ( #5645 )
...
* add docs for llama_chat_apply_template
* fix typo
2024-02-22 00:31:00 +01:00
9c405c9f9a
Server: use llama_chat_apply_template ( #5593 )
...
* server: use llama_chat_apply_template
* server: remove trailing space
* server: fix format_chat
* server: fix help message
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
* server: fix formatted_chat
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-02-20 15:58:27 +01:00
Xuan Son Nguyen and GitHub
11b12de39b
llama : add llama_chat_apply_template() ( #5538 )
...
* llama: add llama_chat_apply_template
* test-chat-template: remove dedundant vector
* chat_template: do not use std::string for buffer
* add clarification for llama_chat_apply_template
* llama_chat_apply_template: add zephyr template
* llama_chat_apply_template: correct docs
* llama_chat_apply_template: use term "chat" everywhere
* llama_chat_apply_template: change variable name to "tmpl"
2024-02-19 10:23:37 +02:00
907e08c110
server : add llama2 chat template ( #5425 )
...
* server: add mistral chat template
* server: fix typo
* server: rename template mistral to llama2
* server: format_llama2: remove BOS
* server: validate "--chat-template" argument
* server: clean up using_chatml variable
Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com >
---------
Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com >
2024-02-11 12:16:22 +02:00
Xuan Son Nguyen and GitHub
6b91b1e0a9
docker : add build for SYCL, Vulkan + update readme ( #5228 )
...
* add vulkan dockerfile
* intel dockerfile: compile sycl by default
* fix vulkan dockerfile
* add docs for vulkan
* docs: sycl build in docker
* docs: remove trailing spaces
* docs: sycl: add docker section
* docs: clarify install vulkan SDK outside docker
* sycl: use intel/oneapi-basekit docker image
* docs: correct TOC
* docs: correct docker image for Intel oneMKL
2024-02-02 09:56:31 +02:00
Xuan Son Nguyen and GitHub
48c857aa10
server : refactored the task processing logic ( #5065 )
...
* server: add llama_server_queue struct
* server: add llama_server_response_event
* server: add comments
* server: move all mutexes away from server.cpp
* server: correct multitask response
* server: only add back deferred tasks when one slot is available
* server: fix a race condition cause by "request_completion"
2024-01-26 14:42:20 +02:00
2bed4aa3f3
devops : add intel oneapi dockerfile ( #5068 )
...
Co-authored-by: Xuan Son Nguyen <xuanson.nguyen@snowpack.eu >
2024-01-23 09:11:39 +02:00
821f0a271e
server : defer tasks when "slot unavailable" ( #5018 )
...
* server: defer task when no slot is available
* remove unnecessary log
---------
Co-authored-by: Xuan Son Nguyen <xuanson.nguyen@snowpack.eu >
2024-01-18 22:33:05 +02:00