Georgi Gerganov
c4a3a4ff47
sync : ggml
2024-04-09 20:29:06 +03:00
Georgi Gerganov and GitHub
e11a8999b5
license : update copyright notice + add AUTHORS ( #6405 )
...
* license : add AUTHORS
* authors : update
* scipts : add LICENSE and gen-authors.sh to sync
2024-04-09 09:23:19 +03:00
cc4a95426d
llama : fix attention layer count sanity check ( #6550 )
...
* llama : fix attention layer count sanity check
* llama : fix parentheses in attention layer count sanity check
There was otherwise a warning when compiling.
---------
Co-authored-by: Francis Couture-Harpin <git@compilade.net >
2024-04-08 22:25:49 +03:00
Georgi Gerganov and GitHub
b73e564b16
quantize : fix precedence of cli args ( #6541 )
2024-04-08 16:23:01 +03:00
b909236c0b
flake.lock: Update ( #6517 )
...
Flake lock file updates:
• Updated input 'flake-parts':
'github:hercules-ci/flake-parts/f7b3c975cf067e56e7cda6cb098ebe3fb4d74ca2' (2024-03-01)
→ 'github:hercules-ci/flake-parts/9126214d0a59633752a136528f5f3b9aa8565b7d' (2024-04-01)
• Updated input 'flake-parts/nixpkgs-lib':
'github:NixOS/nixpkgs/1536926ef5621b09bba54035ae2bb6d806d72ac8?dir=lib' (2024-02-29)
→ 'github:NixOS/nixpkgs/d8fe5e6c92d0d190646fb9f1056741a229980089?dir=lib' (2024-03-29)
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/d8fe5e6c92d0d190646fb9f1056741a229980089' (2024-03-29)
→ 'github:NixOS/nixpkgs/fd281bd6b7d3e32ddfa399853946f782553163b5' (2024-04-03)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-04-07 11:25:30 -07:00
Georgi Gerganov
c37247796b
sync : ggml
2024-04-07 17:05:51 +03:00
Georgi Gerganov
43e8995e75
scripts : sync ggml-cuda folder
2024-04-07 16:08:12 +03:00
Georgi Gerganov
54ea0698fb
sync : ggml
2024-04-06 18:27:46 +03:00
Georgi Gerganov
4399f13fb9
server : remove obsolete --memory-f32 option
2024-04-04 09:34:58 +03:00
Georgi Gerganov and GitHub
076b08649e
readme : update hot topics
2024-04-03 16:11:15 +03:00
f87f7b8986
flake.lock: Update ( #6402 )
...
Flake lock file updates:
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/44d0940ea560dee511026a53f0e2e2cde489b4d4' (2024-03-23)
→ 'github:NixOS/nixpkgs/d8fe5e6c92d0d190646fb9f1056741a229980089' (2024-03-29)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-04-01 09:05:57 -07:00
Georgi Gerganov and GitHub
c50a82ce0f
readme : update hot topics
2024-03-31 11:56:30 +03:00
d48ccf3ad4
sync : ggml ( #6351 )
...
* sync : ggml
ggml-ci
* cuda : move GGML_CUDA_DMMV constants to dmmv.cuh
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-03-29 17:45:46 +02:00
Georgi Gerganov and GitHub
cfde806eb9
ci : fix BGE wget ( #6383 )
...
ggml-ci
2024-03-29 14:34:28 +02:00
Georgi Gerganov and GitHub
bfe7dafc9c
readme : add notice for UI list
2024-03-28 22:56:03 +02:00
Georgi Gerganov
3a0345970e
make : whitespace
2024-03-27 15:02:49 +02:00
Georgi Gerganov
2ab4f00d25
llama2c : open file as binary ( #6332 )
2024-03-27 09:16:02 +02:00
43139cc528
flake.lock: Update ( #6266 )
...
Flake lock file updates:
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/d691274a972b3165335d261cc4671335f5c67de9' (2024-03-14)
→ 'github:NixOS/nixpkgs/44d0940ea560dee511026a53f0e2e2cde489b4d4' (2024-03-23)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-03-25 08:22:27 -07:00
a0e584defd
imatrix : fix wname for mul_mat_id ops ( #6271 )
...
* imatrix : fix wname for mul_mat_id ops
* also filter tensor names in mul_mat_id ops
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-03-24 16:18:45 +02:00
Georgi Gerganov
95562175f8
gitignore : gguf-split
2024-03-23 21:35:23 +02:00
Georgi Gerganov
56a00f0a2f
common : default --hf-file to --model ( #6234 )
2024-03-22 21:10:39 +02:00
Georgi Gerganov and GitHub
80bd33bc2c
common : add HF arg helpers ( #6234 )
...
* common : add HF arg helpers
* common : remove defaults
2024-03-22 15:33:38 +02:00
Georgi Gerganov and GitHub
68e210b354
server : enable continuous batching by default ( #6231 )
2024-03-22 13:08:28 +02:00
Georgi Gerganov and GitHub
b3e94f26ba
metal : proper assert for mat-mat memory alignment ( #6225 )
...
* metal : proper assert for mat-mat memory alignment
ggml-ci
* readme : add notice about the bug fix
* metal : fix the fix
ggml-ci
2024-03-22 11:35:53 +02:00
Georgi Gerganov and GitHub
95d576b48e
metal : pad n_ctx by 32 ( #6177 )
...
* metal : require ne00 >= 128 for mat-mat kernels
ggml-ci
* llama : pad n_ctx by 32
ggml-ci
2024-03-22 09:36:03 +02:00
Georgi Gerganov and GitHub
924ce1dce7
tests : disable system() calls ( #6198 )
...
ggml-ci
2024-03-21 16:20:05 +02:00
Georgi Gerganov
6b7e76d28c
gitignore : ignore curl-related files
2024-03-20 14:17:34 +02:00
Georgi Gerganov
bc0baab2ea
server : allow to override -ngl in tests ( #6170 )
2024-03-20 14:14:32 +02:00
Georgi Gerganov
d795988d9e
Revert "llava : add a MobileVLM_V2-1.7B backup ( #6152 )"
...
This reverts commit f8c4e745e1 .
2024-03-20 13:29:49 +02:00
Georgi Gerganov and GitHub
b80cf3b2d1
common : disable repeat penalties by default ( #6127 )
2024-03-19 10:21:54 +02:00
Georgi Gerganov and GitHub
ac9ee6a4ad
ci : disable stale issue messages ( #6126 )
2024-03-18 13:45:38 +02:00
Georgi Gerganov and GitHub
4f6d1337ca
ci : temporary disable sanitizer builds ( #6128 )
2024-03-18 13:45:27 +02:00
Georgi Gerganov and GitHub
cd776c37c9
ci : close all stale issues at once ( #6115 )
2024-03-17 18:51:57 +01:00
Georgi Gerganov
131b058409
make : ggml-metal.o depends on ggml.h
2024-03-15 11:38:40 +02:00
Georgi Gerganov and GitHub
4755afd1cb
llama : fix integer overflow during quantization ( #6063 )
2024-03-14 22:58:41 +02:00
Georgi Gerganov
044ec4b2a5
embedding : add EOS token if not present ( #899 )
2024-03-14 15:14:14 +02:00
Georgi Gerganov
77178eedc8
gguf-py : fix dtype check ( #6045 )
2024-03-14 13:32:14 +02:00
Georgi Gerganov
a44bc969e4
llama : fix typo
2024-03-14 13:13:06 +02:00
Georgi Gerganov and GitHub
3fe8d7a17f
ggml : designate enum vals for integer types ( #6050 )
2024-03-14 12:38:37 +02:00
Georgi Gerganov
68265ebfc6
embedding : print all resulting embeddings ( #899 )
2024-03-14 12:37:20 +02:00
Georgi Gerganov and GitHub
381da2d9f0
metal : build metallib + fix embed path ( #6015 )
...
* metal : build metallib + fix embed path
ggml-ci
* metal : fix embed build + update library load logic
ggml-ci
* metal : fix embeded library build
ggml-ci
* ci : fix iOS builds to use embedded library
2024-03-14 11:55:23 +02:00
Georgi Gerganov
0fd6c1f015
embedding : print cosine similarity ( #899 )
2024-03-14 10:12:29 +02:00
Georgi Gerganov and GitHub
76a936c893
readme : update API changes and hot topics
2024-03-13 20:33:56 +02:00
Georgi Gerganov and GitHub
8030da7afe
ggml : reuse quantum structs across backends ( #5943 )
...
* ggml : reuse quant blocks across backends
ggml-ci
* ggml : define helper constants only for CUDA and SYCL
ggml-ci
* ggml : define helper quantum constants for SYCL
ggml-ci
2024-03-12 14:27:20 +02:00
Georgi Gerganov and GitHub
184215e783
ggml : fix UB in IQ2_S and IQ3_S ( #6012 )
2024-03-12 13:49:55 +02:00
Georgi Gerganov and GitHub
48358b2e5b
sycl : update IQ1_S kernels (WIP - not working!) ( #5995 )
...
* sycl : try to fix after IQ1_S changes
* sycl : iq1s_grid -> iq1s_grid_gpu
* sycl : fix grid type
2024-03-12 11:15:05 +02:00
Georgi Gerganov and GitHub
05b06210c9
llama : more consistent names of count variables ( #5994 )
...
* llama : more consistent names of count variables
ggml-ci
* llama : n_parallel -> n_seq_max
* common : fix param name
* examples : fix param name
2024-03-11 17:49:47 +02:00
Georgi Gerganov and GitHub
83796e62bc
llama : refactor unicode stuff ( #5992 )
...
* llama : refactor unicode stuff
ggml-ci
* unicode : names
* make : fix c++ compiler
* unicode : names
* unicode : straighten tables
* zig : fix build
* unicode : put nfd normalization behind API
ggml-ci
* swift : fix build
* unicode : add BOM
* unicode : add <cstdint>
ggml-ci
* unicode : pass as cpts as const ref
2024-03-11 17:47:47 +02:00
Georgi Gerganov
ee35600b90
llama : fix F16/F32 downcast + improve names ( #5980 )
2024-03-11 09:56:47 +02:00
Georgi Gerganov and GitHub
bb6d00bbf9
metal : move mm_id indices to shared mem ( #5982 )
2024-03-10 23:12:48 +02:00
Georgi Gerganov and GitHub
d9f65c97c3
readme : update hot topics
2024-03-10 20:58:26 +02:00
Georgi Gerganov
b838b53ad6
sync : ggml
2024-03-10 20:10:46 +02:00
Georgi Gerganov
df4dc3e7cb
ggml : try fix 32-bit arm compat (whisper/1938)
...
* ggml : try fix 32-bit arm compat
* ggml : fix cont
2024-03-10 20:10:39 +02:00
Georgi Gerganov
bf47a5eefc
ggml : remove __constant__ specifier for CUDA tables ( #5940 )
2024-03-10 20:09:24 +02:00
c78541479c
nix: update flake.lock ( #5969 )
...
Flake lock file updates:
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/1536926ef5621b09bba54035ae2bb6d806d72ac8' (2024-02-29)
→ 'github:NixOS/nixpkgs/9df3e30ce24fd28c7b3e2de0d986769db5d6225d' (2024-03-06)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-03-10 16:43:08 +02:00
Georgi Gerganov
77d1ac7e00
server : print chat template info
2024-03-09 22:04:00 +02:00
Georgi Gerganov and GitHub
098dbaab44
readme : update hot topics
2024-03-09 18:14:13 +02:00
Georgi Gerganov and GitHub
8380ecfb21
ggml : fix unnecessary f32 -> f16 -> f32 casts (mmla) ( #5951 )
2024-03-09 17:36:20 +02:00
Georgi Gerganov and GitHub
58308a0ecc
server : fix metrics init ( #5964 )
2024-03-09 17:34:15 +02:00
Georgi Gerganov and GitHub
5b09797321
ggml : remove old quantization functions ( #5942 )
...
* ggml : remove old quantization functions
ggml-ci
* ggml : simplify ggml_quantize_chunk
ggml-ci
* ggml : restrict correctness
ggml-ci
* ggml : remove hist data from the quantization API
ggml-ci
* tests : remove hist usage in test-backend-ops
ggml-ci
* vulkan : remove hist and fix typo
2024-03-09 15:53:59 +02:00
Georgi Gerganov and GitHub
97c09585d6
server : clarify some items in the readme ( #5957 )
...
* server : clarify some items in the readme
* server : fix typo
2024-03-09 15:47:47 +02:00
Georgi Gerganov
2c4f566c88
tests : gitignore ggml-common.h
2024-03-09 14:17:11 +02:00
Georgi Gerganov and GitHub
8a3012a4ad
ggml : add ggml-common.h to deduplicate shared code ( #5940 )
...
* ggml : add ggml-common.h to shared code
ggml-ci
* scripts : update sync scripts
* sycl : reuse quantum tables
ggml-ci
* ggml : minor
* ggml : minor
* sycl : try to fix build
2024-03-09 12:47:57 +02:00
Georgi Gerganov and GitHub
9674aaf35c
server : simplify logic for empty prompts ( #5953 )
2024-03-09 12:34:18 +02:00
Georgi Gerganov and GitHub
af37fd8b30
server : fix EOS token detection with disabled cache ( #5938 )
2024-03-08 12:40:02 +02:00
6cdabe6526
llama-bench : add embeddings option ( #5924 )
...
* llama-bench : add embeddings option
* llama-bench : do not hard code embd default value
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-03-07 16:32:38 +02:00
2002bc96bf
server : refactor ( #5882 )
...
* server : refactoring (wip)
* server : remove llava/clip objects from build
* server : fix empty prompt handling + all slots idle logic
* server : normalize id vars
* server : code style
* server : simplify model chat template validation
* server : code style
* server : minor
* llama : llama_chat_apply_template support null buf
* server : do not process embedding requests when disabled
* server : reorganize structs and enums + naming fixes
* server : merge oai.hpp in utils.hpp
* server : refactor system prompt update at start
* server : disable cached prompts with self-extend
* server : do not process more than n_batch tokens per iter
* server: tests: embeddings use a real embeddings model (#5908 )
* server, tests : bump batch to fit 1 embedding prompt
* server: tests: embeddings fix build type Debug is randomly failing (#5911 )
* server: tests: embeddings, use different KV Cache size
* server: tests: embeddings, fixed prompt do not exceed n_batch, increase embedding timeout, reduce number of concurrent embeddings
* server: tests: embeddings, no need to wait for server idle as it can timout
* server: refactor: clean up http code (#5912 )
* server : avoid n_available var
ggml-ci
* server: refactor: better http codes
* server : simplify json parsing + add comment about t_last
* server : rename server structs
* server : allow to override FQDN in tests
ggml-ci
* server : add comments
---------
Co-authored-by: Pierrick Hymbert <pierrick.hymbert@gmail.com >
2024-03-07 11:41:53 +02:00
Georgi Gerganov
1e35d619a6
convert : remove AWQ remnants ( #5768 )
2024-03-06 09:13:42 +02:00
Georgi Gerganov
82cb31eb93
Revert "grammars : don't allow to output unescaped new line in string ( #5885 )"
...
This reverts commit b1a4e994fd .
2024-03-05 15:56:24 +02:00
Georgi Gerganov and GitHub
29ae62d2ae
llama : fix embeddings ( #5796 )
...
* llama : fix embeddings
ggml-ci
* llama : do not use KV cache for non-causal models
ggml-ci
* embeddings : fix llama_batch_init arg
* llama : add pooling switch
* llama : distinguish token vs sequence embeddings
ggml-ci
* llama : assert pooling tensor
* llama : simplify causal mask condition
ggml-ci
* llama : assert input batch with pooling enabled
* readme : update API changes list
2024-03-04 22:31:20 +02:00
Georgi Gerganov
e0843afe1b
flake : fix
2024-03-04 21:50:50 +02:00
Georgi Gerganov
a1c6d96ed8
ggml : fix unknown status ( #0 )
2024-03-04 20:54:23 +02:00
Georgi Gerganov
efd8533ef8
sync : ggml
...
ggml-ci
2024-03-04 20:54:23 +02:00
Georgi Gerganov
a0fc62661f
sync : ggml
2024-03-04 10:40:04 +02:00
Georgi Gerganov and GitHub
231ae28f07
readme : add API changes section
2024-03-03 12:44:03 +02:00
fa974646e1
flake.lock: Update ( #5842 )
...
Flake lock file updates:
• Updated input 'flake-parts':
'github:hercules-ci/flake-parts/b253292d9c0a5ead9bc98c4e9a26c6312e27d69f' (2024-02-01)
→ 'github:hercules-ci/flake-parts/f7b3c975cf067e56e7cda6cb098ebe3fb4d74ca2' (2024-03-01)
• Updated input 'flake-parts/nixpkgs-lib':
'github:NixOS/nixpkgs/97b17f32362e475016f942bbdfda4a4a72a8a652?dir=lib' (2024-01-29)
→ 'github:NixOS/nixpkgs/1536926ef5621b09bba54035ae2bb6d806d72ac8?dir=lib' (2024-02-29)
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/cbc4211f0afffe6dfd2478a62615dd5175a13f9a' (2024-02-23)
→ 'github:NixOS/nixpkgs/1536926ef5621b09bba54035ae2bb6d806d72ac8' (2024-02-29)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-03-02 20:11:31 -08:00
Georgi Gerganov and GitHub
494c870326
ggml : fix IQ3_S AVX implementation ( #5834 )
...
ggml-ci
2024-03-02 20:00:49 +02:00
Georgi Gerganov
ef2cd694c4
scripts : add pod-llama.sh
2024-03-02 16:54:20 +02:00
Georgi Gerganov and GitHub
38d16b1426
server : remove api_like_OAI.py proxy script ( #5808 )
2024-03-01 20:00:58 +02:00
Georgi Gerganov
f105471ef6
server : fix newlines in help ( #5785 )
2024-03-01 09:59:43 +02:00
Georgi Gerganov and GitHub
87c91c0766
ci : reduce 3b ppl chunks to 1 to avoid timeout ( #5771 )
...
ggml-ci
2024-02-28 21:44:21 +02:00
Georgi Gerganov and GitHub
08c5ee87e4
llama : remove deprecated API ( #5770 )
...
ggml-ci
2024-02-28 18:43:38 +02:00
Georgi Gerganov and GitHub
78aacf3634
awq-py : remove ( #5768 )
2024-02-28 17:36:53 +02:00
Georgi Gerganov
8c0e8f4e73
sync : ggml
2024-02-28 11:17:32 +02:00
Georgi Gerganov and GitHub
9d533a77d0
llama : fix defrag bugs + add parameter ( #5735 )
...
* llama : fix defrag bugs + enable by default
ggml-ci
* llama : add defrag_thold parameter
ggml-ci
* llama : cont
* llama : disable log message
ggml-ci
* llama : fix graph size check during defrag
2024-02-27 14:35:51 +02:00
Georgi Gerganov and GitHub
67fd33132f
unicode : reuse iterator ( #5726 )
2024-02-26 14:02:12 +02:00
Georgi Gerganov
269de86ba0
llama : fix Gemma rope type ( #5691 )
2024-02-26 08:30:17 +02:00
Georgi Gerganov and GitHub
bf08e00643
llama : refactor k-shift implementation + KV defragmentation ( #5691 )
...
* llama : refactor k-shift implementation
ggml-ci
* llama : rename llama_kv_cache_seq_shift to llama_kv_cache_seq_add
* llama : cont k-shift refactoring + normalize type names
ggml-ci
* minor : fix MPI builds
* llama : reuse n_rot from the build context
ggml-ci
* llama : revert enum name changes from this PR
ggml-ci
* llama : update llama_rope_type
* llama : add comment about rope values
* llama : fix build
* passkey : apply kv cache updates explicitly
ggml-ci
* llama : change name to llama_kv_cache_update()
* llama : add llama_kv_cache_seq_pos_max()
* passkey : fix llama_kv_cache_seq_pos_max() usage
* llama : some llama_kv_cell simplifications
* llama : add llama_kv_cache_compress (EXPERIMENTAL)
* llama : add alternative KV cache merging (EXPERIMENTAL)
* llama : add llama_kv_cache_defrag
* llama : comments
* llama : remove llama_kv_cache_compress
will add in a separate PR
ggml-ci
* llama : defragment via non-overlapping moves
* llama : ggml_graph based defrag implementation
ggml-ci
* llama : switch the loop order in build_defrag
* llama : add comments
2024-02-25 22:12:24 +02:00
Georgi Gerganov and GitHub
ab336a9d5e
code : normalize enum names ( #5697 )
...
* coda : normalize enum names
ggml-ci
* code : cont
* code : cont
2024-02-25 12:09:09 +02:00
Georgi Gerganov and GitHub
96633eeca1
gemma : use more bits for the token_embd.weight tensor ( #5650 )
...
* gemma : use Q8_0 for the token_embd.weight tensor
* llama : quantize token_embd.weight using output type
2024-02-22 23:23:46 +02:00
847eedbdb2
py : add Gemma conversion from HF models ( #5647 )
...
* py : add gemma conversion from HF models
* Update convert-hf-to-gguf.py
Co-authored-by: Aarni Koskela <akx@iki.fi >
* Update convert-hf-to-gguf.py
Co-authored-by: Aarni Koskela <akx@iki.fi >
* Update convert-hf-to-gguf.py
Co-authored-by: Jared Van Bortel <jared@nomic.ai >
---------
Co-authored-by: Aarni Koskela <akx@iki.fi >
Co-authored-by: Jared Van Bortel <jared@nomic.ai >
2024-02-22 23:22:48 +02:00
Georgi Gerganov and GitHub
7e4f339c40
ggml : always define ggml_fp16_t as uint16_t ( #5666 )
...
* ggml : always define ggml_fp16_t as uint16_t
ggml-ci
* ggml : cont
ggml-ci
* ggml : cont
* ggml : cont
ggml-ci
* ggml : cont
ggml-ci
* cuda : no longer ggml headers last
ggml-ci
* ggml : fix q6_K FP16 -> FP32 conversion
ggml-ci
* ggml : more FP16 -> FP32 conversion fixes
ggml-ci
2024-02-22 23:21:39 +02:00
Georgi Gerganov
334f76fa38
sync : ggml
2024-02-22 23:21:05 +02:00
Georgi Gerganov
efd56b1c21
ggml : 32-bit arm compat (whisper/1891)
...
* ggml : 32-bit arm compat
* ggml : add ggml_vqtbl1q_s8 impl
* ggml : cont
2024-02-22 23:20:50 +02:00
Georgi Gerganov and GitHub
5a9e2f60ba
py : minor fixes ( #5668 )
2024-02-22 20:13:25 +02:00
Georgi Gerganov
3a03541ced
minor : fix trailing whitespace ( #5638 )
2024-02-22 13:54:03 +02:00
Georgi Gerganov and GitHub
56d03d92be
readme : update hot topics
2024-02-22 10:35:54 +02:00
Georgi Gerganov
5022cf242d
sync : ggml
2024-02-21 16:52:52 +02:00
eccd7a26dd
sync : ggml ( #5633 )
...
* ggml : fix conv_2d batch mode (ggml/737)
Co-authored-by: bssrdf <bssrdf@gmail.com >
* ggml : compute forward no longer pass src tensors (ggml/729)
* sync : ggml
ggml-ci
---------
Co-authored-by: bssrdf <merlintiger@hotmail.com >
Co-authored-by: bssrdf <bssrdf@gmail.com >
2024-02-21 16:17:10 +02:00
Georgi Gerganov and GitHub
c14f72db9c
readme : update hot topics
2024-02-21 15:39:54 +02:00