43139cc528
flake.lock: Update ( #6266 )
...
Flake lock file updates:
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/d691274a972b3165335d261cc4671335f5c67de9' (2024-03-14)
→ 'github:NixOS/nixpkgs/44d0940ea560dee511026a53f0e2e2cde489b4d4' (2024-03-23)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-03-25 08:22:27 -07:00
a0e584defd
imatrix : fix wname for mul_mat_id ops ( #6271 )
...
* imatrix : fix wname for mul_mat_id ops
* also filter tensor names in mul_mat_id ops
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-03-24 16:18:45 +02:00
Georgi Gerganov
95562175f8
gitignore : gguf-split
2024-03-23 21:35:23 +02:00
Georgi Gerganov
56a00f0a2f
common : default --hf-file to --model ( #6234 )
2024-03-22 21:10:39 +02:00
Georgi Gerganov and GitHub
80bd33bc2c
common : add HF arg helpers ( #6234 )
...
* common : add HF arg helpers
* common : remove defaults
2024-03-22 15:33:38 +02:00
Georgi Gerganov and GitHub
68e210b354
server : enable continuous batching by default ( #6231 )
2024-03-22 13:08:28 +02:00
Georgi Gerganov and GitHub
b3e94f26ba
metal : proper assert for mat-mat memory alignment ( #6225 )
...
* metal : proper assert for mat-mat memory alignment
ggml-ci
* readme : add notice about the bug fix
* metal : fix the fix
ggml-ci
2024-03-22 11:35:53 +02:00
Georgi Gerganov and GitHub
95d576b48e
metal : pad n_ctx by 32 ( #6177 )
...
* metal : require ne00 >= 128 for mat-mat kernels
ggml-ci
* llama : pad n_ctx by 32
ggml-ci
2024-03-22 09:36:03 +02:00
Georgi Gerganov and GitHub
924ce1dce7
tests : disable system() calls ( #6198 )
...
ggml-ci
2024-03-21 16:20:05 +02:00
Georgi Gerganov
6b7e76d28c
gitignore : ignore curl-related files
2024-03-20 14:17:34 +02:00
Georgi Gerganov
bc0baab2ea
server : allow to override -ngl in tests ( #6170 )
2024-03-20 14:14:32 +02:00
Georgi Gerganov
d795988d9e
Revert "llava : add a MobileVLM_V2-1.7B backup ( #6152 )"
...
This reverts commit f8c4e745e1 .
2024-03-20 13:29:49 +02:00
Georgi Gerganov and GitHub
b80cf3b2d1
common : disable repeat penalties by default ( #6127 )
2024-03-19 10:21:54 +02:00
Georgi Gerganov and GitHub
ac9ee6a4ad
ci : disable stale issue messages ( #6126 )
2024-03-18 13:45:38 +02:00
Georgi Gerganov and GitHub
4f6d1337ca
ci : temporary disable sanitizer builds ( #6128 )
2024-03-18 13:45:27 +02:00
Georgi Gerganov and GitHub
cd776c37c9
ci : close all stale issues at once ( #6115 )
2024-03-17 18:51:57 +01:00
Georgi Gerganov
131b058409
make : ggml-metal.o depends on ggml.h
2024-03-15 11:38:40 +02:00
Georgi Gerganov and GitHub
4755afd1cb
llama : fix integer overflow during quantization ( #6063 )
2024-03-14 22:58:41 +02:00
Georgi Gerganov
044ec4b2a5
embedding : add EOS token if not present ( #899 )
2024-03-14 15:14:14 +02:00
Georgi Gerganov
77178eedc8
gguf-py : fix dtype check ( #6045 )
2024-03-14 13:32:14 +02:00
Georgi Gerganov
a44bc969e4
llama : fix typo
2024-03-14 13:13:06 +02:00
Georgi Gerganov and GitHub
3fe8d7a17f
ggml : designate enum vals for integer types ( #6050 )
2024-03-14 12:38:37 +02:00
Georgi Gerganov
68265ebfc6
embedding : print all resulting embeddings ( #899 )
2024-03-14 12:37:20 +02:00
Georgi Gerganov and GitHub
381da2d9f0
metal : build metallib + fix embed path ( #6015 )
...
* metal : build metallib + fix embed path
ggml-ci
* metal : fix embed build + update library load logic
ggml-ci
* metal : fix embeded library build
ggml-ci
* ci : fix iOS builds to use embedded library
2024-03-14 11:55:23 +02:00
Georgi Gerganov
0fd6c1f015
embedding : print cosine similarity ( #899 )
2024-03-14 10:12:29 +02:00
Georgi Gerganov and GitHub
76a936c893
readme : update API changes and hot topics
2024-03-13 20:33:56 +02:00
Georgi Gerganov and GitHub
8030da7afe
ggml : reuse quantum structs across backends ( #5943 )
...
* ggml : reuse quant blocks across backends
ggml-ci
* ggml : define helper constants only for CUDA and SYCL
ggml-ci
* ggml : define helper quantum constants for SYCL
ggml-ci
2024-03-12 14:27:20 +02:00
Georgi Gerganov and GitHub
184215e783
ggml : fix UB in IQ2_S and IQ3_S ( #6012 )
2024-03-12 13:49:55 +02:00
Georgi Gerganov and GitHub
48358b2e5b
sycl : update IQ1_S kernels (WIP - not working!) ( #5995 )
...
* sycl : try to fix after IQ1_S changes
* sycl : iq1s_grid -> iq1s_grid_gpu
* sycl : fix grid type
2024-03-12 11:15:05 +02:00
Georgi Gerganov and GitHub
05b06210c9
llama : more consistent names of count variables ( #5994 )
...
* llama : more consistent names of count variables
ggml-ci
* llama : n_parallel -> n_seq_max
* common : fix param name
* examples : fix param name
2024-03-11 17:49:47 +02:00
Georgi Gerganov and GitHub
83796e62bc
llama : refactor unicode stuff ( #5992 )
...
* llama : refactor unicode stuff
ggml-ci
* unicode : names
* make : fix c++ compiler
* unicode : names
* unicode : straighten tables
* zig : fix build
* unicode : put nfd normalization behind API
ggml-ci
* swift : fix build
* unicode : add BOM
* unicode : add <cstdint>
ggml-ci
* unicode : pass as cpts as const ref
2024-03-11 17:47:47 +02:00
Georgi Gerganov
ee35600b90
llama : fix F16/F32 downcast + improve names ( #5980 )
2024-03-11 09:56:47 +02:00
Georgi Gerganov and GitHub
bb6d00bbf9
metal : move mm_id indices to shared mem ( #5982 )
2024-03-10 23:12:48 +02:00
Georgi Gerganov and GitHub
d9f65c97c3
readme : update hot topics
2024-03-10 20:58:26 +02:00
Georgi Gerganov
b838b53ad6
sync : ggml
2024-03-10 20:10:46 +02:00
Georgi Gerganov
df4dc3e7cb
ggml : try fix 32-bit arm compat (whisper/1938)
...
* ggml : try fix 32-bit arm compat
* ggml : fix cont
2024-03-10 20:10:39 +02:00
Georgi Gerganov
bf47a5eefc
ggml : remove __constant__ specifier for CUDA tables ( #5940 )
2024-03-10 20:09:24 +02:00
c78541479c
nix: update flake.lock ( #5969 )
...
Flake lock file updates:
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/1536926ef5621b09bba54035ae2bb6d806d72ac8' (2024-02-29)
→ 'github:NixOS/nixpkgs/9df3e30ce24fd28c7b3e2de0d986769db5d6225d' (2024-03-06)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-03-10 16:43:08 +02:00
Georgi Gerganov
77d1ac7e00
server : print chat template info
2024-03-09 22:04:00 +02:00
Georgi Gerganov and GitHub
098dbaab44
readme : update hot topics
2024-03-09 18:14:13 +02:00
Georgi Gerganov and GitHub
8380ecfb21
ggml : fix unnecessary f32 -> f16 -> f32 casts (mmla) ( #5951 )
2024-03-09 17:36:20 +02:00
Georgi Gerganov and GitHub
58308a0ecc
server : fix metrics init ( #5964 )
2024-03-09 17:34:15 +02:00
Georgi Gerganov and GitHub
5b09797321
ggml : remove old quantization functions ( #5942 )
...
* ggml : remove old quantization functions
ggml-ci
* ggml : simplify ggml_quantize_chunk
ggml-ci
* ggml : restrict correctness
ggml-ci
* ggml : remove hist data from the quantization API
ggml-ci
* tests : remove hist usage in test-backend-ops
ggml-ci
* vulkan : remove hist and fix typo
2024-03-09 15:53:59 +02:00
Georgi Gerganov and GitHub
97c09585d6
server : clarify some items in the readme ( #5957 )
...
* server : clarify some items in the readme
* server : fix typo
2024-03-09 15:47:47 +02:00
Georgi Gerganov
2c4f566c88
tests : gitignore ggml-common.h
2024-03-09 14:17:11 +02:00
Georgi Gerganov and GitHub
8a3012a4ad
ggml : add ggml-common.h to deduplicate shared code ( #5940 )
...
* ggml : add ggml-common.h to shared code
ggml-ci
* scripts : update sync scripts
* sycl : reuse quantum tables
ggml-ci
* ggml : minor
* ggml : minor
* sycl : try to fix build
2024-03-09 12:47:57 +02:00
Georgi Gerganov and GitHub
9674aaf35c
server : simplify logic for empty prompts ( #5953 )
2024-03-09 12:34:18 +02:00
Georgi Gerganov and GitHub
af37fd8b30
server : fix EOS token detection with disabled cache ( #5938 )
2024-03-08 12:40:02 +02:00
6cdabe6526
llama-bench : add embeddings option ( #5924 )
...
* llama-bench : add embeddings option
* llama-bench : do not hard code embd default value
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-03-07 16:32:38 +02:00
2002bc96bf
server : refactor ( #5882 )
...
* server : refactoring (wip)
* server : remove llava/clip objects from build
* server : fix empty prompt handling + all slots idle logic
* server : normalize id vars
* server : code style
* server : simplify model chat template validation
* server : code style
* server : minor
* llama : llama_chat_apply_template support null buf
* server : do not process embedding requests when disabled
* server : reorganize structs and enums + naming fixes
* server : merge oai.hpp in utils.hpp
* server : refactor system prompt update at start
* server : disable cached prompts with self-extend
* server : do not process more than n_batch tokens per iter
* server: tests: embeddings use a real embeddings model (#5908 )
* server, tests : bump batch to fit 1 embedding prompt
* server: tests: embeddings fix build type Debug is randomly failing (#5911 )
* server: tests: embeddings, use different KV Cache size
* server: tests: embeddings, fixed prompt do not exceed n_batch, increase embedding timeout, reduce number of concurrent embeddings
* server: tests: embeddings, no need to wait for server idle as it can timout
* server: refactor: clean up http code (#5912 )
* server : avoid n_available var
ggml-ci
* server: refactor: better http codes
* server : simplify json parsing + add comment about t_last
* server : rename server structs
* server : allow to override FQDN in tests
ggml-ci
* server : add comments
---------
Co-authored-by: Pierrick Hymbert <pierrick.hymbert@gmail.com >
2024-03-07 11:41:53 +02:00
Georgi Gerganov
1e35d619a6
convert : remove AWQ remnants ( #5768 )
2024-03-06 09:13:42 +02:00
Georgi Gerganov
82cb31eb93
Revert "grammars : don't allow to output unescaped new line in string ( #5885 )"
...
This reverts commit b1a4e994fd .
2024-03-05 15:56:24 +02:00
Georgi Gerganov and GitHub
29ae62d2ae
llama : fix embeddings ( #5796 )
...
* llama : fix embeddings
ggml-ci
* llama : do not use KV cache for non-causal models
ggml-ci
* embeddings : fix llama_batch_init arg
* llama : add pooling switch
* llama : distinguish token vs sequence embeddings
ggml-ci
* llama : assert pooling tensor
* llama : simplify causal mask condition
ggml-ci
* llama : assert input batch with pooling enabled
* readme : update API changes list
2024-03-04 22:31:20 +02:00
Georgi Gerganov
e0843afe1b
flake : fix
2024-03-04 21:50:50 +02:00
Georgi Gerganov
a1c6d96ed8
ggml : fix unknown status ( #0 )
2024-03-04 20:54:23 +02:00
Georgi Gerganov
efd8533ef8
sync : ggml
...
ggml-ci
2024-03-04 20:54:23 +02:00
Georgi Gerganov
a0fc62661f
sync : ggml
2024-03-04 10:40:04 +02:00
Georgi Gerganov and GitHub
231ae28f07
readme : add API changes section
2024-03-03 12:44:03 +02:00
fa974646e1
flake.lock: Update ( #5842 )
...
Flake lock file updates:
• Updated input 'flake-parts':
'github:hercules-ci/flake-parts/b253292d9c0a5ead9bc98c4e9a26c6312e27d69f' (2024-02-01)
→ 'github:hercules-ci/flake-parts/f7b3c975cf067e56e7cda6cb098ebe3fb4d74ca2' (2024-03-01)
• Updated input 'flake-parts/nixpkgs-lib':
'github:NixOS/nixpkgs/97b17f32362e475016f942bbdfda4a4a72a8a652?dir=lib' (2024-01-29)
→ 'github:NixOS/nixpkgs/1536926ef5621b09bba54035ae2bb6d806d72ac8?dir=lib' (2024-02-29)
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/cbc4211f0afffe6dfd2478a62615dd5175a13f9a' (2024-02-23)
→ 'github:NixOS/nixpkgs/1536926ef5621b09bba54035ae2bb6d806d72ac8' (2024-02-29)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-03-02 20:11:31 -08:00
Georgi Gerganov and GitHub
494c870326
ggml : fix IQ3_S AVX implementation ( #5834 )
...
ggml-ci
2024-03-02 20:00:49 +02:00
Georgi Gerganov
ef2cd694c4
scripts : add pod-llama.sh
2024-03-02 16:54:20 +02:00
Georgi Gerganov and GitHub
38d16b1426
server : remove api_like_OAI.py proxy script ( #5808 )
2024-03-01 20:00:58 +02:00
Georgi Gerganov
f105471ef6
server : fix newlines in help ( #5785 )
2024-03-01 09:59:43 +02:00
Georgi Gerganov and GitHub
87c91c0766
ci : reduce 3b ppl chunks to 1 to avoid timeout ( #5771 )
...
ggml-ci
2024-02-28 21:44:21 +02:00
Georgi Gerganov and GitHub
08c5ee87e4
llama : remove deprecated API ( #5770 )
...
ggml-ci
2024-02-28 18:43:38 +02:00
Georgi Gerganov and GitHub
78aacf3634
awq-py : remove ( #5768 )
2024-02-28 17:36:53 +02:00
Georgi Gerganov
8c0e8f4e73
sync : ggml
2024-02-28 11:17:32 +02:00
Georgi Gerganov and GitHub
9d533a77d0
llama : fix defrag bugs + add parameter ( #5735 )
...
* llama : fix defrag bugs + enable by default
ggml-ci
* llama : add defrag_thold parameter
ggml-ci
* llama : cont
* llama : disable log message
ggml-ci
* llama : fix graph size check during defrag
2024-02-27 14:35:51 +02:00
Georgi Gerganov and GitHub
67fd33132f
unicode : reuse iterator ( #5726 )
2024-02-26 14:02:12 +02:00
Georgi Gerganov
269de86ba0
llama : fix Gemma rope type ( #5691 )
2024-02-26 08:30:17 +02:00
Georgi Gerganov and GitHub
bf08e00643
llama : refactor k-shift implementation + KV defragmentation ( #5691 )
...
* llama : refactor k-shift implementation
ggml-ci
* llama : rename llama_kv_cache_seq_shift to llama_kv_cache_seq_add
* llama : cont k-shift refactoring + normalize type names
ggml-ci
* minor : fix MPI builds
* llama : reuse n_rot from the build context
ggml-ci
* llama : revert enum name changes from this PR
ggml-ci
* llama : update llama_rope_type
* llama : add comment about rope values
* llama : fix build
* passkey : apply kv cache updates explicitly
ggml-ci
* llama : change name to llama_kv_cache_update()
* llama : add llama_kv_cache_seq_pos_max()
* passkey : fix llama_kv_cache_seq_pos_max() usage
* llama : some llama_kv_cell simplifications
* llama : add llama_kv_cache_compress (EXPERIMENTAL)
* llama : add alternative KV cache merging (EXPERIMENTAL)
* llama : add llama_kv_cache_defrag
* llama : comments
* llama : remove llama_kv_cache_compress
will add in a separate PR
ggml-ci
* llama : defragment via non-overlapping moves
* llama : ggml_graph based defrag implementation
ggml-ci
* llama : switch the loop order in build_defrag
* llama : add comments
2024-02-25 22:12:24 +02:00
Georgi Gerganov and GitHub
ab336a9d5e
code : normalize enum names ( #5697 )
...
* coda : normalize enum names
ggml-ci
* code : cont
* code : cont
2024-02-25 12:09:09 +02:00
Georgi Gerganov and GitHub
96633eeca1
gemma : use more bits for the token_embd.weight tensor ( #5650 )
...
* gemma : use Q8_0 for the token_embd.weight tensor
* llama : quantize token_embd.weight using output type
2024-02-22 23:23:46 +02:00
847eedbdb2
py : add Gemma conversion from HF models ( #5647 )
...
* py : add gemma conversion from HF models
* Update convert-hf-to-gguf.py
Co-authored-by: Aarni Koskela <akx@iki.fi >
* Update convert-hf-to-gguf.py
Co-authored-by: Aarni Koskela <akx@iki.fi >
* Update convert-hf-to-gguf.py
Co-authored-by: Jared Van Bortel <jared@nomic.ai >
---------
Co-authored-by: Aarni Koskela <akx@iki.fi >
Co-authored-by: Jared Van Bortel <jared@nomic.ai >
2024-02-22 23:22:48 +02:00
Georgi Gerganov and GitHub
7e4f339c40
ggml : always define ggml_fp16_t as uint16_t ( #5666 )
...
* ggml : always define ggml_fp16_t as uint16_t
ggml-ci
* ggml : cont
ggml-ci
* ggml : cont
* ggml : cont
ggml-ci
* ggml : cont
ggml-ci
* cuda : no longer ggml headers last
ggml-ci
* ggml : fix q6_K FP16 -> FP32 conversion
ggml-ci
* ggml : more FP16 -> FP32 conversion fixes
ggml-ci
2024-02-22 23:21:39 +02:00
Georgi Gerganov
334f76fa38
sync : ggml
2024-02-22 23:21:05 +02:00
Georgi Gerganov
efd56b1c21
ggml : 32-bit arm compat (whisper/1891)
...
* ggml : 32-bit arm compat
* ggml : add ggml_vqtbl1q_s8 impl
* ggml : cont
2024-02-22 23:20:50 +02:00
Georgi Gerganov and GitHub
5a9e2f60ba
py : minor fixes ( #5668 )
2024-02-22 20:13:25 +02:00
Georgi Gerganov
3a03541ced
minor : fix trailing whitespace ( #5638 )
2024-02-22 13:54:03 +02:00
Georgi Gerganov and GitHub
56d03d92be
readme : update hot topics
2024-02-22 10:35:54 +02:00
Georgi Gerganov
5022cf242d
sync : ggml
2024-02-21 16:52:52 +02:00
eccd7a26dd
sync : ggml ( #5633 )
...
* ggml : fix conv_2d batch mode (ggml/737)
Co-authored-by: bssrdf <bssrdf@gmail.com >
* ggml : compute forward no longer pass src tensors (ggml/729)
* sync : ggml
ggml-ci
---------
Co-authored-by: bssrdf <merlintiger@hotmail.com >
Co-authored-by: bssrdf <bssrdf@gmail.com >
2024-02-21 16:17:10 +02:00
Georgi Gerganov and GitHub
c14f72db9c
readme : update hot topics
2024-02-21 15:39:54 +02:00
Georgi Gerganov
1387cf60f7
llava : remove extra cont ( #5587 )
2024-02-19 15:23:17 +02:00
Georgi Gerganov
337c9cbd52
sync : ggml
...
ggml-ci
2024-02-19 15:09:43 +02:00
Georgi Gerganov
a3145bdc30
ggml-alloc : apply ggml/731
2024-02-19 15:09:43 +02:00
Georgi Gerganov and GitHub
d0e3ce51f4
ci : enable -Werror for CUDA builds ( #5579 )
...
* cmake : pass -Werror through -Xcompiler
ggml-ci
* make, cmake : enable CUDA errors on warnings
ggml-ci
2024-02-19 14:45:41 +02:00
Georgi Gerganov
68a6b98b3c
make : fix CUDA build ( #5580 )
2024-02-19 13:41:51 +02:00
Georgi Gerganov
f53119cec4
minor : fix trailing whitespace ( #5538 )
2024-02-19 10:34:10 +02:00
Georgi Gerganov
14278f55d2
ggml : restore vec dot stride arg names ( #5453 )
2024-02-18 22:58:57 +02:00
Georgi Gerganov and GitHub
b1de96824b
ci : fix wikitext url + compile warnings ( #5569 )
...
ggml-ci
2024-02-18 22:39:30 +02:00
Georgi Gerganov
7ad554f90e
metal : fix unused warnings ( #0 )
2024-02-18 21:39:58 +02:00
Georgi Gerganov
689a091bbe
sampling : do not set min_keep to n_probs ( #5564 )
2024-02-18 19:38:06 +02:00
Georgi Gerganov
f3f28c5395
cmake : fix GGML_USE_SYCL typo ( #5555 )
2024-02-18 19:17:00 +02:00
Georgi Gerganov
1dcc3fde00
common : fix ub ( #5530 )
2024-02-18 18:21:52 +02:00
8f1be0d42f
ggml : add ALiBi support for ggml_soft_max_ext ( #5488 )
...
* ggml : avoid recomputing alibi slopes (CPU)
* llama : reuse hparams.f_max_alibi_bias in all cases
ggml-ci
* ggml : support alibi bias in ggml_soft_max_ext (CPU + Metal)
ggml-ci
* ggml : handle all SRCs (do not break on first null)
ggml-ci
* tests : do not use slope for large soft_max
accumulates too much error
ggml-ci
* ggml : alternative ALiBi without extra tensor
We compute the slopes in the kernel
ggml-ci
* cuda : add ALiBi support in ggml_soft_max_ext
ggml-ci
* ggml : deprecate ggml_alibi
* ggml : support multi-sequence ALiBi (Metal)
ggml-ci
* cuda : add multi-seq ALiBi + remote F16 soft_max
ggml-ci
* ggml : update deprecation message
* ggml : fix pos ptr when no ALiBi
ggml-ci
* cuda : fix performance (pow -> powf)
* cuda : precompute ALiBi constants
* metal : pre-compute ALiBi slopes
ggml-ci
* llama : init kq_pos only if needed
ggml-ci
* test-backend-ops : add null pos test to soft_max
test-backend-ops : replace soft_max tests
ggml-ci
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-02-17 23:04:16 +02:00
Georgi Gerganov and GitHub
5bf2b94dd4
cmake : fix VULKAN and ROCm builds ( #5525 )
...
* cmake : fix VULKAN and ROCm builds
* cmake : fix (cont)
* vulkan : fix compile warnings
ggml-ci
* cmake : fix
ggml-ci
* cmake : minor
ggml-ci
2024-02-16 19:05:56 +02:00
d2819d5577
scripts : add helpers script for bench comparing commits ( #5521 )
...
* scripts : add helpers script for bench comparing commits
* scripts : detect CUDA
* set flags after checking the command line
* fix make flags
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-02-16 15:14:40 +02:00
Georgi Gerganov
594845aab1
ci : fix BERT model download and convert
2024-02-16 09:57:55 +02:00
Georgi Gerganov
c06e45d729
clip : fix wrong loop condition
2024-02-15 18:49:08 +02:00