Diego Devesa and GitHub
59f4db1088
ggml : add predefined list of CPU backend variants to build ( #10626 )
...
* ggml : add predefined list of CPU backend variants to build
* update CPU dockerfiles
2024-12-04 14:45:40 +01:00
Diego Devesa and GitHub
2803540814
ggml-cpu : fix HWCAP2_I8MM value ( #10646 )
2024-12-04 14:40:44 +01:00
Diego Devesa and GitHub
3420909dff
ggml : automatic selection of best CPU backend ( #10606 )
...
* ggml : automatic selection of best CPU backend
* amx : minor opt
* add GGML_AVX_VNNI to enable avx-vnni, fix checks
2024-12-01 16:12:41 +01:00
7cc2d2c889
ggml : move AMX to the CPU backend ( #10570 )
...
* ggml : move AMX to the CPU backend
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-11-29 21:54:58 +01:00
Diego Devesa and GitHub
a3a3048e7a
cleanup UI link list ( #10577 )
...
* cleanup UI link list
* sort list alphabetically
* add missing licenses
2024-11-29 17:45:08 +01:00
Diego Devesa and GitHub
e90688edd0
ci : fix tag name in cuda and hip releases ( #10566 )
2024-11-28 15:58:54 +01:00
Diego Devesa and GitHub
46c69e0e75
ci : faster CUDA toolkit installation method and use ccache ( #10537 )
...
* ci : faster CUDA toolkit installation method and use ccache
* remove fetch-depth
* only pack CUDA runtime on master
2024-11-27 11:03:25 +01:00
Diego Devesa and GitHub
c9b00a70b0
ci : fix cuda releases ( #10532 )
2024-11-26 22:12:10 +01:00
Diego Devesa and GitHub
5a349f2809
ci : remove nix workflows ( #10526 )
2024-11-26 21:13:54 +01:00
Diego Devesa and GitHub
30ec398321
llama : disable warnings for 3rd party sha1 dependency ( #10527 )
2024-11-26 21:01:47 +01:00
Diego Devesa and GitHub
7db3846a94
ci : publish the docker images created during scheduled runs ( #10515 )
2024-11-26 13:05:20 +01:00
Diego Devesa and GitHub
c6807b3f28
ci : add ubuntu cuda build, build with one arch on windows ( #10456 )
2024-11-26 13:05:07 +01:00
Diego Devesa and GitHub
50d5cecbda
ci : build docker images only once daily ( #10503 )
2024-11-25 22:05:39 +01:00
Diego Devesa and GitHub
10bce0450f
llama : accept a list of devices to use to offload a model ( #10497 )
...
* llama : accept a list of devices to use to offload a model
* accept `--dev none` to completely disable offloading
* fix dev list with dl backends
* rename env parameter to LLAMA_ARG_DEVICE for consistency
2024-11-25 19:30:06 +01:00
5931c1f233
ggml : add support for dynamic loading of backends ( #10469 )
...
* ggml : add support for dynamic loading of backends
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-11-25 15:13:39 +01:00
Diego Devesa and GitHub
dc39012cba
llama : fix op mul check with command-r-plus ( #10476 )
2024-11-24 16:10:26 +01:00
Diego Devesa and GitHub
55ed008b2d
ggml : do not use ARM features not included in the build ( #10457 )
2024-11-23 14:41:12 +01:00
a5e47592b6
cuda : optimize argmax ( #10441 )
...
* cuda : optimize argmax
* remove unused parameter
ggml-ci
* fixup : use full warps
ggml-ci
* Apply suggestions from code review
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
* fix ub
* ggml : check ne00 <= INT32_MAX in argmax and argsort
---------
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
2024-11-21 18:18:50 +01:00
slaren and Georgi Gerganov
59b9172822
ggml/sched : do not skip views in pre-assignments
2024-11-21 09:22:05 +02:00
Diego Devesa and GitHub
fab5d30ff6
llama : add .clang-format file ( #10415 )
2024-11-20 12:57:53 +01:00
Diego Devesa and GitHub
3ee6382d48
cuda : fix CUDA_FLAGS not being applied ( #10403 )
2024-11-19 14:29:38 +01:00
Diego Devesa and GitHub
d3481e6316
cuda : only use native when supported by cmake ( #10389 )
2024-11-18 18:43:40 +01:00
Diego Devesa and GitHub
be5caccef9
llama : only use default buffer types for the KV cache ( #10358 )
2024-11-17 12:25:45 +01:00
Diego Devesa and GitHub
eda7e1d4f5
ggml : fix possible buffer use after free in sched reserve ( #9930 )
2024-11-17 08:31:17 +02:00
slaren and Georgi Gerganov
883d206fbd
ggml : fix some build issues
2024-11-15 21:45:32 +02:00
ae8de6d50a
ggml : build backends as libraries ( #10256 )
...
* ggml : build backends as libraries
---------
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com >
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com >
2024-11-14 18:04:35 +01:00
Diego Devesa and GitHub
97404c4a03
ggml : add ggml-cpu.h to the public headers ( #10204 )
2024-11-07 18:16:08 +01:00
Diego Devesa and GitHub
94d8cb8be1
metal : fix from ptr buffer name ( #10189 )
2024-11-06 12:10:07 +01:00
Diego Devesa and GitHub
a9e8a9a030
ggml : fix arch check in bf16_to_fp32 ( #10164 )
2024-11-04 23:17:01 +01:00
Diego Devesa and GitHub
d5a409e57f
ggml : fix gelu tables initialization ( #10172 )
2024-11-04 20:06:58 +01:00
Diego Devesa and GitHub
401558b7ba
ggml : fix q4xx mat mul, increase ggml_aligned_malloc alignment ( #10167 )
2024-11-04 17:34:08 +01:00
Diego Devesa and GitHub
ea02c753eb
cuda : clear error after changing peer access ( #10153 )
2024-11-04 13:10:23 +01:00
Diego Devesa and GitHub
9f40989351
ggml : move CPU backend to a separate file ( #10144 )
2024-11-03 19:34:08 +01:00
Diego Devesa and GitHub
b634f8a26f
simple-chat : only add bos on first prompt ( #10129 )
2024-11-02 13:08:53 +01:00
a6744e43e8
llama : add simple-chat example ( #10124 )
...
* llama : add simple-chat example
---------
Co-authored-by: Xuan Son Nguyen <thichthat@gmail.com >
2024-11-01 23:50:59 +01:00
Diego Devesa and GitHub
e991e3127f
llama : use smart pointers for ggml resources ( #10117 )
2024-11-01 23:48:26 +01:00
Diego Devesa and GitHub
85679d37f3
llama : improve output buffer type selection ( #10098 )
2024-11-01 00:49:53 +01:00
Diego Devesa and GitHub
1e9f94994e
quantize : fix --keep-split ( #10114 )
2024-11-01 00:45:34 +01:00
Diego Devesa and GitHub
c02e5ab2a6
llama : fix buffer checks for mamba and rwk ( #10111 )
...
* llama : fix buffer checks for mamba and rwk
* llama : fix missing worst case flag during reserve
* cuda : fix supports_op for norm
* disable sched SET_CAUSE
2024-10-31 22:54:23 +01:00
Diego Devesa and GitHub
dea5e86051
ggml : check tensor name lengths in gguf files ( #10100 )
2024-10-31 11:40:59 +01:00
Diego Devesa and GitHub
b9e02e8184
ggml : fix memory leaks when loading invalid gguf files ( #10094 )
...
* ggml : fix gguf string leak when reading kv pairs fails
* ggml : avoid crashing with GGML_ABORT when the KV has an invalid type
* ggml : avoid crashing on failed memory allocations when loading a gguf file
2024-10-30 14:51:21 +01:00
Diego Devesa and GitHub
c5b0f4b5d9
llama : refactor model loader with backend registry ( #10026 )
2024-10-30 02:01:23 +01:00
Diego Devesa and GitHub
f010b77a37
vulkan : add backend registry / device interfaces ( #9721 )
...
* vulkan : add backend registry / device interfaces
* llama : print devices used on model load
2024-10-17 02:46:58 +02:00
Diego Devesa and GitHub
96776405a1
ggml : move more prints to the ggml log system ( #9839 )
...
* ggml : move more prints to the ggml log system
* show BLAS OpenMP warnings in all builds using debug print
2024-10-11 15:34:45 +02:00
7eee341bee
common : use common_ prefix for common library functions ( #9805 )
...
* common : use common_ prefix for common library functions
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-10-10 22:57:42 +02:00
Diego Devesa and GitHub
0e9f760eb1
rpc : add backend registry / device interfaces ( #9812 )
...
* rpc : add backend registry / device interfaces
* llama : add llama_supports_rpc API
* ggml_backend_rpc_start_rpc_server -> ggml_backend_rpc_start_server
2024-10-10 20:14:55 +02:00
Diego Devesa and GitHub
c7499c557c
examples : do not use common library in simple example ( #9803 )
...
* examples : do not use common library in simple example
* add command line parser, simplify code
2024-10-10 19:50:49 +02:00
Diego Devesa and GitHub
c81f3bbb05
cmake : do not build common library by default when standalone ( #9804 )
2024-10-09 18:49:52 +02:00
Diego Devesa and GitHub
dca1d4b58a
ggml : fix BLAS with unsupported types ( #9775 )
...
* ggml : do not use BLAS with types without to_float
* ggml : return pointer from ggml_internal_get_type_traits to avoid unnecessary copies
* ggml : rename ggml_internal_get_type_traits -> ggml_get_type_traits
it's not really internal if everybody uses it
2024-10-08 14:21:43 +02:00
Diego Devesa and GitHub
6374743747
ggml : add backend registry / device interfaces to BLAS backend ( #9752 )
...
* ggml : add backend registry / device interfaces to BLAS backend
* fix mmap usage when using host buffers
2024-10-07 21:55:08 +02:00
Diego Devesa and Georgi Gerganov
ff565769f2
ggml : fixes after sync (ggml/983)
...
ggml : remove test-backend-buffer
ggml : fix CUDA build warnings
2024-10-04 18:50:04 +03:00
Diego Devesa and GitHub
a7ad553513
ggml-backend : add device description to CPU backend ( #9720 )
2024-10-03 17:39:18 +02:00
c83ad6d01e
ggml-backend : add device and backend reg interfaces ( #9707 )
...
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
2024-10-03 01:49:47 +02:00
slaren and GitHub
1b2f992cd2
test-backend-ops : use flops for some performance tests ( #9657 )
...
* test-backend-ops : use flops for some performance tests
- parallelize tensor quantization
- use a different set of cases for performance and correctness tests
- run each test for at least one second
2024-09-28 14:32:46 +02:00
slaren and GitHub
d09770cae7
ggml-alloc : fix list of allocated tensors with GGML_ALLOCATOR_DEBUG ( #9573 )
2024-09-21 14:24:23 +02:00
slaren and GitHub
63351143b2
quantize : improve type name parsing ( #9570 )
...
quantize : do not ignore invalid types in arg parsing
quantize : ignore case of type and ftype arguments
2024-09-20 20:55:36 +02:00
64c6af3195
ggml : fix n_threads_cur initialization with one thread ( #9538 )
...
* ggml : fix n_threads_cur initialization with one thread
* Update ggml/src/ggml.c
---------
Co-authored-by: Max Krasnyansky <quic_maxk@quicinc.com >
2024-09-18 10:13:08 -07:00
slaren and GitHub
23e0d70bac
ggml : move common CPU backend impl to new header ( #9509 )
2024-09-16 16:22:07 +02:00
slaren and GitHub
e6deac31f7
gguf-split : add basic checks ( #9499 )
...
* gguf-split : do not overwrite existing files when merging
* gguf-split : error when too many arguments are passed
2024-09-15 19:02:27 +02:00
slaren and GitHub
1b28061400
llama : skip token bounds check when evaluating embeddings ( #9437 )
2024-09-11 17:52:13 +02:00
slaren and GitHub
49006c67b4
llama : move random seed generation to the samplers ( #9398 )
...
* llama_sampler_penalties : clamp penalty_last_n to zero
2024-09-10 18:04:25 +02:00
slaren and GitHub
fb3f249815
make : do not run llama-gen-docs when building ( #9399 )
2024-09-10 09:23:33 +03:00
slaren and GitHub
5fb5e24811
llama : minor sampling refactor (2) ( #9386 )
2024-09-09 17:10:46 +02:00
slaren and GitHub
a249843d89
common : restore --n-gpu-layers ( #9371 )
2024-09-08 16:44:42 +02:00
slaren and GitHub
19f4a7b296
llama : refactor samplers internal implementation ( #9370 )
2024-09-08 15:52:07 +02:00
slaren and GitHub
eae597182c
llama : sanitize tokens in the upper bound ( #9359 )
2024-09-08 12:41:51 +02:00
slaren and GitHub
e32d0816ed
ggml : always check bounds on get_rows operations ( #9354 )
2024-09-07 20:23:07 +02:00
slaren and GitHub
6c89eb0b47
ci : disable rocm image creation ( #9340 )
2024-09-07 10:48:54 +03:00
slaren and GitHub
4db04784f9
cuda : fix defrag with quantized KV ( #9319 )
2024-09-05 11:13:11 +02:00
slaren and GitHub
bdf314f38a
llama-bench : fix NUL terminators in CPU name ( #9313 )
2024-09-05 02:19:39 +02:00
slaren and GitHub
048de848ee
docker : fix missing binaries in full-cuda image ( #9278 )
2024-09-02 18:11:13 +02:00
slaren and GitHub
9fe94ccac9
docker : build images only once ( #9225 )
2024-08-28 17:28:00 +02:00
slaren and GitHub
66b039a501
docker : update CUDA images ( #9213 )
2024-08-28 13:20:36 +02:00
slaren and GitHub
7d787ed96c
ggml : do not crash when quantizing q4_x_x with an imatrix ( #9192 )
2024-08-26 19:44:43 +02:00
slaren and GitHub
0c41e03ceb
metal : gemma2 flash attention support ( #9159 )
2024-08-26 11:08:59 +02:00
slaren and GitHub
f12ceaca0c
ggml-ci : try to improve build time ( #9160 )
2024-08-26 11:03:30 +02:00
slaren and GitHub
6e02327e8b
metal : fix uninitialized abort_callback ( #8968 )
2024-08-10 15:42:10 +02:00
slaren and GitHub
15fa07a5c5
make : use C compiler to build metal embed object ( #8899 )
...
* make : use C compiler to build metal embed object
* use rm + rmdir to avoid -r flag in rm
2024-08-07 18:24:05 +02:00
slaren and GitHub
be55695eff
ggml-backend : fix async copy from CPU ( #8897 )
...
* ggml-backend : fix async copy from CPU
* cuda : more reliable async copy, fix stream used when the devices are the same
2024-08-07 13:29:02 +02:00
slaren and GitHub
7a11eb3a26
cuda : fix dmmv cols requirement to 2*GGML_CUDA_DMMV_X ( #8800 )
...
* cuda : fix dmmv cols requirement to 2*GGML_CUDA_DMMV_X
* update asserts
* only use dmmv for supported types
* add test
2024-08-01 15:26:22 +02:00
slaren and GitHub
2b1f616b20
ggml : reduce hash table reset cost ( #8698 )
...
* ggml : reduce hash table reset cost
* fix unreachable code warnings after GGML_ASSERT(false)
* GGML_ASSERT(false) -> GGML_ABORT("fatal error")
* GGML_ABORT use format string
2024-07-27 04:41:55 +02:00
87e397d00b
ggml : fix quant dot product with odd number of blocks ( #8549 )
...
* ggml : fix iq4_nl dot product with odd number of blocks
* ggml : fix odd blocks for ARM_NEON (#8556 )
* ggml : fix iq4_nl dot product with odd number of blocks
* ggml : fix q4_1
* ggml : fix q5_0
* ggml : fix q5_1
* ggml : fix iq4_nl metal
ggml-ci
* ggml : fix q4_0
* ggml : fix q8_0
ggml-ci
* ggml : remove special Q4_0 code for first 2 blocks
* ggml : fix sumf redefinition
---------
Co-authored-by: slaren <slarengh@gmail.com >
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-07-19 17:17:27 +02:00
slaren and GitHub
5f2d4e60e2
ppl : fix n_seq_max for perplexity ( #8277 )
...
* ppl : fix n_seq_max for perplexity
* use 1 seq for kl_divergence
2024-07-03 20:33:31 +03:00
slaren and GitHub
0e0590adab
cuda : update supports_op for matrix multiplication ( #8245 )
2024-07-02 09:39:38 +03:00
slaren and GitHub
b851b3fba0
cmake : allow user to override default options ( #8178 )
2024-06-28 12:37:45 +02:00
slaren and GitHub
8172ee9da9
cmake : fix deprecated option names not working ( #8171 )
...
* cmake : fix deprecated option names not working
* remove LlAMA_OPENMP
2024-06-27 20:04:39 +02:00
slaren and GitHub
ae5d0f4b89
ci : publish new docker images only when the files change ( #8142 )
2024-06-26 21:59:28 +02:00
slaren and GitHub
31ec3993f6
ggml : add GGML_CUDA_USE_GRAPHS option, restore GGML_CUDA_FORCE_CUBLAS (cmake) ( #8140 )
2024-06-26 21:34:14 +02:00
slaren and GitHub
c7ab7b612c
make : fix missing -O3 ( #8143 )
2024-06-26 21:20:22 +03:00
slaren and GitHub
dd047b476c
disable docker CI on pull requests ( #8110 )
2024-06-25 19:20:06 +02:00
slaren and GitHub
8cb508d0d5
disable publishing the full-rocm docker image ( #8083 )
2024-06-24 08:36:11 +03:00
slaren and GitHub
95f57bb5d5
ggml : remove ggml_task_type and GGML_PERF ( #8017 )
...
* ggml : remove ggml_task_type and GGML_PERF
* check abort_callback on main thread only
* vulkan : remove usage of ggml_compute_params
* remove LLAMA_PERF
2024-06-24 03:07:59 +02:00
slaren and GitHub
b6b9a8e606
fix CI failures ( #8066 )
...
* test-backend-ops : increase cpy max nmse
* server ci : disable thread sanitizer
2024-06-23 13:14:45 +02:00
slaren and GitHub
9c77ec1d74
ggml : synchronize threads using barriers ( #7993 )
2024-06-19 15:04:15 +02:00
slaren and GitHub
99052cd227
sched : offload_op also requires supports_op ( #7977 )
2024-06-17 16:51:42 +02:00
f578b86b21
move BLAS to a separate backend ( #6210 )
...
* move BLAS to a separate backend
* rename GGML_USE_OPENBLAS to GGML_USE_BLAS
* alloc : reuse same buffer when the same buffer type if used multiple times
* set number of threads automatically for openblas and blis
* sched : print assignments when GGML_SCHED_DEBUG env variable is set
* sched : allow ops with weights on an incompatible buffer type
This will cause the weight to be copied to a backend that supports the
op, which is very costly. The weight should have been stored in a buffer
of a backend that can run the op, but llama.cpp cannot do this
automatically at the moment.
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-06-13 03:11:35 +02:00
slaren and GitHub
c2ce6c47e4
fix CUDA CI by using a windows-2019 image ( #7861 )
...
* try to fix CUDA ci with --allow-unsupported-compiler
* trigger when build.yml changes
* another test
* try exllama/bdashore3 method
* install vs build tools before cuda toolkit
* try win-2019
2024-06-11 08:59:20 +03:00
slaren and GitHub
fd5ea0f897
ci : try win-2019 on server windows test ( #7854 )
2024-06-10 15:18:41 +03:00
slaren and GitHub
fe1e3917cf
Revert "[SYCL] Update rpc-server.cpp to include SYCL backend ( #7682 )" ( #7808 )
...
This reverts commit 9422c5e34b .
2024-06-09 01:43:39 +02:00
da799b4189
vulkan : reuse parent extra for views ( #7806 )
...
* vulkan : reuse parent extra for views
* Fix validation error when multiple compute contexts are used in a graph
---------
Co-authored-by: 0cc4m <picard12@live.de >
2024-06-07 19:47:49 +02:00