a6744e43e8
llama : add simple-chat example ( #10124 )
...
* llama : add simple-chat example
---------
Co-authored-by: Xuan Son Nguyen <thichthat@gmail.com >
2024-11-01 23:50:59 +01:00
Diego Devesa and GitHub
e991e3127f
llama : use smart pointers for ggml resources ( #10117 )
2024-11-01 23:48:26 +01:00
Diego Devesa and GitHub
85679d37f3
llama : improve output buffer type selection ( #10098 )
2024-11-01 00:49:53 +01:00
Diego Devesa and GitHub
1e9f94994e
quantize : fix --keep-split ( #10114 )
2024-11-01 00:45:34 +01:00
Diego Devesa and GitHub
c02e5ab2a6
llama : fix buffer checks for mamba and rwk ( #10111 )
...
* llama : fix buffer checks for mamba and rwk
* llama : fix missing worst case flag during reserve
* cuda : fix supports_op for norm
* disable sched SET_CAUSE
2024-10-31 22:54:23 +01:00
Diego Devesa and GitHub
dea5e86051
ggml : check tensor name lengths in gguf files ( #10100 )
2024-10-31 11:40:59 +01:00
Diego Devesa and GitHub
b9e02e8184
ggml : fix memory leaks when loading invalid gguf files ( #10094 )
...
* ggml : fix gguf string leak when reading kv pairs fails
* ggml : avoid crashing with GGML_ABORT when the KV has an invalid type
* ggml : avoid crashing on failed memory allocations when loading a gguf file
2024-10-30 14:51:21 +01:00
Diego Devesa and GitHub
c5b0f4b5d9
llama : refactor model loader with backend registry ( #10026 )
2024-10-30 02:01:23 +01:00
Diego Devesa and GitHub
f010b77a37
vulkan : add backend registry / device interfaces ( #9721 )
...
* vulkan : add backend registry / device interfaces
* llama : print devices used on model load
2024-10-17 02:46:58 +02:00
Diego Devesa and GitHub
96776405a1
ggml : move more prints to the ggml log system ( #9839 )
...
* ggml : move more prints to the ggml log system
* show BLAS OpenMP warnings in all builds using debug print
2024-10-11 15:34:45 +02:00
7eee341bee
common : use common_ prefix for common library functions ( #9805 )
...
* common : use common_ prefix for common library functions
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-10-10 22:57:42 +02:00
Diego Devesa and GitHub
0e9f760eb1
rpc : add backend registry / device interfaces ( #9812 )
...
* rpc : add backend registry / device interfaces
* llama : add llama_supports_rpc API
* ggml_backend_rpc_start_rpc_server -> ggml_backend_rpc_start_server
2024-10-10 20:14:55 +02:00
Diego Devesa and GitHub
c7499c557c
examples : do not use common library in simple example ( #9803 )
...
* examples : do not use common library in simple example
* add command line parser, simplify code
2024-10-10 19:50:49 +02:00
Diego Devesa and GitHub
c81f3bbb05
cmake : do not build common library by default when standalone ( #9804 )
2024-10-09 18:49:52 +02:00
Diego Devesa and GitHub
dca1d4b58a
ggml : fix BLAS with unsupported types ( #9775 )
...
* ggml : do not use BLAS with types without to_float
* ggml : return pointer from ggml_internal_get_type_traits to avoid unnecessary copies
* ggml : rename ggml_internal_get_type_traits -> ggml_get_type_traits
it's not really internal if everybody uses it
2024-10-08 14:21:43 +02:00
Diego Devesa and GitHub
6374743747
ggml : add backend registry / device interfaces to BLAS backend ( #9752 )
...
* ggml : add backend registry / device interfaces to BLAS backend
* fix mmap usage when using host buffers
2024-10-07 21:55:08 +02:00
Diego Devesa and Georgi Gerganov
ff565769f2
ggml : fixes after sync (ggml/983)
...
ggml : remove test-backend-buffer
ggml : fix CUDA build warnings
2024-10-04 18:50:04 +03:00
Diego Devesa and GitHub
a7ad553513
ggml-backend : add device description to CPU backend ( #9720 )
2024-10-03 17:39:18 +02:00
c83ad6d01e
ggml-backend : add device and backend reg interfaces ( #9707 )
...
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
2024-10-03 01:49:47 +02:00
slaren and GitHub
1b2f992cd2
test-backend-ops : use flops for some performance tests ( #9657 )
...
* test-backend-ops : use flops for some performance tests
- parallelize tensor quantization
- use a different set of cases for performance and correctness tests
- run each test for at least one second
2024-09-28 14:32:46 +02:00
slaren and GitHub
d09770cae7
ggml-alloc : fix list of allocated tensors with GGML_ALLOCATOR_DEBUG ( #9573 )
2024-09-21 14:24:23 +02:00
slaren and GitHub
63351143b2
quantize : improve type name parsing ( #9570 )
...
quantize : do not ignore invalid types in arg parsing
quantize : ignore case of type and ftype arguments
2024-09-20 20:55:36 +02:00
64c6af3195
ggml : fix n_threads_cur initialization with one thread ( #9538 )
...
* ggml : fix n_threads_cur initialization with one thread
* Update ggml/src/ggml.c
---------
Co-authored-by: Max Krasnyansky <quic_maxk@quicinc.com >
2024-09-18 10:13:08 -07:00
slaren and GitHub
23e0d70bac
ggml : move common CPU backend impl to new header ( #9509 )
2024-09-16 16:22:07 +02:00
slaren and GitHub
e6deac31f7
gguf-split : add basic checks ( #9499 )
...
* gguf-split : do not overwrite existing files when merging
* gguf-split : error when too many arguments are passed
2024-09-15 19:02:27 +02:00
slaren and GitHub
1b28061400
llama : skip token bounds check when evaluating embeddings ( #9437 )
2024-09-11 17:52:13 +02:00
slaren and GitHub
49006c67b4
llama : move random seed generation to the samplers ( #9398 )
...
* llama_sampler_penalties : clamp penalty_last_n to zero
2024-09-10 18:04:25 +02:00
slaren and GitHub
fb3f249815
make : do not run llama-gen-docs when building ( #9399 )
2024-09-10 09:23:33 +03:00
slaren and GitHub
5fb5e24811
llama : minor sampling refactor (2) ( #9386 )
2024-09-09 17:10:46 +02:00
slaren and GitHub
a249843d89
common : restore --n-gpu-layers ( #9371 )
2024-09-08 16:44:42 +02:00
slaren and GitHub
19f4a7b296
llama : refactor samplers internal implementation ( #9370 )
2024-09-08 15:52:07 +02:00
slaren and GitHub
eae597182c
llama : sanitize tokens in the upper bound ( #9359 )
2024-09-08 12:41:51 +02:00
slaren and GitHub
e32d0816ed
ggml : always check bounds on get_rows operations ( #9354 )
2024-09-07 20:23:07 +02:00
slaren and GitHub
6c89eb0b47
ci : disable rocm image creation ( #9340 )
2024-09-07 10:48:54 +03:00
slaren and GitHub
4db04784f9
cuda : fix defrag with quantized KV ( #9319 )
2024-09-05 11:13:11 +02:00
slaren and GitHub
bdf314f38a
llama-bench : fix NUL terminators in CPU name ( #9313 )
2024-09-05 02:19:39 +02:00
slaren and GitHub
048de848ee
docker : fix missing binaries in full-cuda image ( #9278 )
2024-09-02 18:11:13 +02:00
slaren and GitHub
9fe94ccac9
docker : build images only once ( #9225 )
2024-08-28 17:28:00 +02:00
slaren and GitHub
66b039a501
docker : update CUDA images ( #9213 )
2024-08-28 13:20:36 +02:00
slaren and GitHub
7d787ed96c
ggml : do not crash when quantizing q4_x_x with an imatrix ( #9192 )
2024-08-26 19:44:43 +02:00
slaren and GitHub
0c41e03ceb
metal : gemma2 flash attention support ( #9159 )
2024-08-26 11:08:59 +02:00
slaren and GitHub
f12ceaca0c
ggml-ci : try to improve build time ( #9160 )
2024-08-26 11:03:30 +02:00
slaren and GitHub
6e02327e8b
metal : fix uninitialized abort_callback ( #8968 )
2024-08-10 15:42:10 +02:00
slaren and GitHub
15fa07a5c5
make : use C compiler to build metal embed object ( #8899 )
...
* make : use C compiler to build metal embed object
* use rm + rmdir to avoid -r flag in rm
2024-08-07 18:24:05 +02:00
slaren and GitHub
be55695eff
ggml-backend : fix async copy from CPU ( #8897 )
...
* ggml-backend : fix async copy from CPU
* cuda : more reliable async copy, fix stream used when the devices are the same
2024-08-07 13:29:02 +02:00
slaren and GitHub
7a11eb3a26
cuda : fix dmmv cols requirement to 2*GGML_CUDA_DMMV_X ( #8800 )
...
* cuda : fix dmmv cols requirement to 2*GGML_CUDA_DMMV_X
* update asserts
* only use dmmv for supported types
* add test
2024-08-01 15:26:22 +02:00
slaren and GitHub
2b1f616b20
ggml : reduce hash table reset cost ( #8698 )
...
* ggml : reduce hash table reset cost
* fix unreachable code warnings after GGML_ASSERT(false)
* GGML_ASSERT(false) -> GGML_ABORT("fatal error")
* GGML_ABORT use format string
2024-07-27 04:41:55 +02:00
87e397d00b
ggml : fix quant dot product with odd number of blocks ( #8549 )
...
* ggml : fix iq4_nl dot product with odd number of blocks
* ggml : fix odd blocks for ARM_NEON (#8556 )
* ggml : fix iq4_nl dot product with odd number of blocks
* ggml : fix q4_1
* ggml : fix q5_0
* ggml : fix q5_1
* ggml : fix iq4_nl metal
ggml-ci
* ggml : fix q4_0
* ggml : fix q8_0
ggml-ci
* ggml : remove special Q4_0 code for first 2 blocks
* ggml : fix sumf redefinition
---------
Co-authored-by: slaren <slarengh@gmail.com >
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-07-19 17:17:27 +02:00
slaren and GitHub
5f2d4e60e2
ppl : fix n_seq_max for perplexity ( #8277 )
...
* ppl : fix n_seq_max for perplexity
* use 1 seq for kl_divergence
2024-07-03 20:33:31 +03:00
slaren and GitHub
0e0590adab
cuda : update supports_op for matrix multiplication ( #8245 )
2024-07-02 09:39:38 +03:00
slaren and GitHub
b851b3fba0
cmake : allow user to override default options ( #8178 )
2024-06-28 12:37:45 +02:00
slaren and GitHub
8172ee9da9
cmake : fix deprecated option names not working ( #8171 )
...
* cmake : fix deprecated option names not working
* remove LlAMA_OPENMP
2024-06-27 20:04:39 +02:00
slaren and GitHub
ae5d0f4b89
ci : publish new docker images only when the files change ( #8142 )
2024-06-26 21:59:28 +02:00
slaren and GitHub
31ec3993f6
ggml : add GGML_CUDA_USE_GRAPHS option, restore GGML_CUDA_FORCE_CUBLAS (cmake) ( #8140 )
2024-06-26 21:34:14 +02:00
slaren and GitHub
c7ab7b612c
make : fix missing -O3 ( #8143 )
2024-06-26 21:20:22 +03:00
slaren and GitHub
dd047b476c
disable docker CI on pull requests ( #8110 )
2024-06-25 19:20:06 +02:00
slaren and GitHub
8cb508d0d5
disable publishing the full-rocm docker image ( #8083 )
2024-06-24 08:36:11 +03:00
slaren and GitHub
95f57bb5d5
ggml : remove ggml_task_type and GGML_PERF ( #8017 )
...
* ggml : remove ggml_task_type and GGML_PERF
* check abort_callback on main thread only
* vulkan : remove usage of ggml_compute_params
* remove LLAMA_PERF
2024-06-24 03:07:59 +02:00
slaren and GitHub
b6b9a8e606
fix CI failures ( #8066 )
...
* test-backend-ops : increase cpy max nmse
* server ci : disable thread sanitizer
2024-06-23 13:14:45 +02:00
slaren and GitHub
9c77ec1d74
ggml : synchronize threads using barriers ( #7993 )
2024-06-19 15:04:15 +02:00
slaren and GitHub
99052cd227
sched : offload_op also requires supports_op ( #7977 )
2024-06-17 16:51:42 +02:00
f578b86b21
move BLAS to a separate backend ( #6210 )
...
* move BLAS to a separate backend
* rename GGML_USE_OPENBLAS to GGML_USE_BLAS
* alloc : reuse same buffer when the same buffer type if used multiple times
* set number of threads automatically for openblas and blis
* sched : print assignments when GGML_SCHED_DEBUG env variable is set
* sched : allow ops with weights on an incompatible buffer type
This will cause the weight to be copied to a backend that supports the
op, which is very costly. The weight should have been stored in a buffer
of a backend that can run the op, but llama.cpp cannot do this
automatically at the moment.
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-06-13 03:11:35 +02:00
slaren and GitHub
c2ce6c47e4
fix CUDA CI by using a windows-2019 image ( #7861 )
...
* try to fix CUDA ci with --allow-unsupported-compiler
* trigger when build.yml changes
* another test
* try exllama/bdashore3 method
* install vs build tools before cuda toolkit
* try win-2019
2024-06-11 08:59:20 +03:00
slaren and GitHub
fd5ea0f897
ci : try win-2019 on server windows test ( #7854 )
2024-06-10 15:18:41 +03:00
slaren and GitHub
fe1e3917cf
Revert "[SYCL] Update rpc-server.cpp to include SYCL backend ( #7682 )" ( #7808 )
...
This reverts commit 9422c5e34b .
2024-06-09 01:43:39 +02:00
da799b4189
vulkan : reuse parent extra for views ( #7806 )
...
* vulkan : reuse parent extra for views
* Fix validation error when multiple compute contexts are used in a graph
---------
Co-authored-by: 0cc4m <picard12@live.de >
2024-06-07 19:47:49 +02:00
slaren and GitHub
c9ee7118d5
check for nans in imatrix and quantize ( #7807 )
...
* imatrix : detect nan/inf values
* quantize : check imatrix for nan/inf values
2024-06-07 09:01:29 +03:00
slaren and GitHub
2d08b7fbb4
docker : build only main and server in their images ( #7782 )
...
* add openmp lib to dockerfiles
* build only main and server in their docker images
2024-06-06 08:19:49 +03:00
slaren and GitHub
d67caea0d6
docker : add openmp lib ( #7780 )
2024-06-06 08:17:21 +03:00
slaren and GitHub
adc9ff3841
llama-bench : allow using a different printer for stderr with -oe ( #7722 )
...
compare-commits.sh : hide stdout, use -oe to print markdown
2024-06-04 14:32:42 +02:00
slaren and GitHub
87bdf2a199
ggml : use atomic_flag for critical section ( #7598 )
...
* ggml : use atomic_flag for critical section
* add windows shims
2024-05-29 13:36:39 +02:00
slaren and GitHub
b18532a4ef
phi3 : duplicate rope factors in each layer ( #7447 )
...
* phi3 : duplicate rope factors in each layer
phi3 : set phi-3 model type as 14B
model loader : simplify the process for duplicating model tensors
llama-bench : remove default pg test
* replace bool parameters in llama_model_loader with named flags
2024-05-22 16:10:46 +02:00
slaren and GitHub
d359f30921
llama : remove MPI backend ( #7395 )
2024-05-20 01:17:03 +02:00
slaren and GitHub
e4e6f67be6
ggml : fix another case of quants nans ( #7387 )
2024-05-19 17:08:46 +02:00
slaren and GitHub
ab33f7a338
cuda : clear error after buffer allocation failure ( #7376 )
2024-05-19 14:19:37 +02:00
slaren and GitHub
05834841dc
ggml : fix quants nans when all the group weights are very close to zero ( #7313 )
2024-05-18 02:39:54 +02:00
slaren and GitHub
344f9126cc
ggml : tag ggml_tensor::backend as deprecated ( #7290 )
2024-05-15 15:08:48 +02:00
slaren and GitHub
541600201e
llama : disable pipeline parallelism with nkvo ( #7265 )
2024-05-14 17:33:42 +10:00
slaren and GitHub
b228aba91a
remove convert-lora-to-ggml.py ( #7204 )
2024-05-12 02:29:33 +02:00
slaren and GitHub
e849648888
llama-bench : add pp+tg test type ( #7199 )
2024-05-10 18:03:54 +02:00
slaren and GitHub
25c6e82e7a
llama : use n_vocab to differentiate between mistral 7B and llama3 8B ( #7200 )
2024-05-10 14:28:01 +02:00
slaren and GitHub
eaf4bd8b39
eval-callback : fix conversion to float ( #7184 )
2024-05-10 01:04:12 +02:00
slaren and GitHub
c4ec9c0d3d
ci : exempt confirmed bugs from being tagged as stale ( #7014 )
2024-05-01 08:13:59 +03:00
slaren and GitHub
017e6999b5
add basic tensor data validation function ( #6884 )
...
* add basic tensor data validation function
* add --check-tensors command line argument
tensor validation is disabled by default and can be enabled by adding
`--check-tensors` to the command line arguments.
quantize always validates tensors.
2024-04-26 18:39:58 +02:00
slaren and GitHub
e2764cd7ca
gguf : fix mismatch between alloc and free functions ( #6929 )
2024-04-26 18:07:42 +03:00
slaren and GitHub
d6e1d44f16
llama : synchronize before get/set session data ( #6911 )
2024-04-25 17:59:03 +02:00
slaren and GitHub
0ead1f1072
llama : check that all the tensor data is in the model file ( #6885 )
...
* llama : check that all the tensor data is in the model file
* also check for unsigned overflow
2024-04-25 15:23:47 +02:00
0d56246f4b
ggml : group all experts in a single ggml_mul_mat_id ( #6505 )
...
* ggml : group all experts in a single ggml_mul_mat_id
cuda : improve mmid row copy
* cuda : fix bin bcast with non-cont src0
* test-backend-ops : only run all mul mat tests for base types
* llama : disable moe offloading with SYCL
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-04-18 15:18:48 +02:00
slaren and GitHub
c71bfd736e
llama : fix compatibility with old 2 expert models ( #6735 )
2024-04-18 10:04:47 +03:00
slaren and GitHub
fbbc030ba9
metal : unify mul_mv_id kernels ( #6556 )
2024-04-12 18:13:20 +02:00
slaren and GitHub
4f407a0a35
llama : add model types for mixtral ( #6589 )
2024-04-10 17:24:14 +02:00
slaren and GitHub
65c64dc36f
convert.py : add consolidated.safetensors for mixtral 8x22b ( #6587 )
2024-04-10 15:23:12 +02:00
08a0c02060
ggml : mul_mat_id use the same tensor for all the experts ( #6387 )
...
* ggml : update mul_mat_id to use the same tensor for all the experts
* update cuda
* minor
* update metal
* update test-backend-ops
* fix cuda
* Update ggml-metal.m
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
* update convert.py
* update convert-hf-to-gguf.py
* update convert.py for mixtral hf models
* Update convert-hf-to-gguf.py
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
* cuda : support non-pow-2 number of experts
* allow quantize to work for split and merged experts models in the same way
* cleanup + disable mmap automatically with split tensors models
* update imatrix
* test-backend-ops : test qwen argsort
* update grok model loading
* llama : add merged experts tensors to the grok tensor map
* minor
* gguf : bump version
* fix quantizing of merged experts
* convert-hf-to-gguf.py : update grok (untested)
* make linter happy
* cuda/argsort : use shared memory instead of pool memory
* convert : fix grok tensor names
* metal : add support for non-pow-2 argsort
* llama : more loader cleanup, better error checking
* cuda : fix warning
* llama : still use mmap for loading old models, but copy the data to a host buffer
* add review note
* llama : remove ffn tensor counting + add sanity check
ggml-ci
* convert : fix handling of n_experts == None
ggml-ci
* imatrix : fix ncall counters
* llama : produce error if imatrix size does not match
* quantize : terminate on errors + trace logs
ggml-ci
* metal : pad shared memory to 16 bytes
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2024-04-03 16:07:05 +03:00
slaren and GitHub
e5b89a441a
ggml : fix bounds checking of zero size views ( #6347 )
2024-03-27 15:07:50 +01:00
slaren and GitHub
280345968d
cuda : rename build flag to LLAMA_CUDA ( #6299 )
2024-03-26 01:16:01 +01:00
slaren and GitHub
2f34b865b6
cuda : fix LLAMA_CUDA_F16 build ( #6298 )
2024-03-25 16:43:22 +02:00
slaren and GitHub
ae1f211ce2
cuda : refactor into multiple files ( #6269 )
2024-03-25 13:50:23 +01:00
slaren and GitHub
2f0e81e053
cuda : add LLAMA_CUDA_NO_PEER_COPY to workaround broken ROCm p2p copy ( #6208 )
...
* cuda : add LLAMA_CUDA_NO_PEER_COPY to workaround broken ROCm p2p copy
* add LLAMA_CUDA_NO_PEER_COPY to HIP build
2024-03-22 14:05:31 +01:00
slaren and GitHub
d0a71233fb
cuda : disable host register by default ( #6206 )
2024-03-21 20:54:28 +02:00
slaren and GitHub
03a8f8fafe
cuda : fix LLAMA_CUDA_F16 build ( #6197 )
2024-03-21 14:59:53 +02:00