slaren and GitHub
58b367c2d7
cuBLAS: refactor and optimize f16 mat mul performance ( #1259 )
...
* cuBLAS: refactor, convert fp16 to fp32 on device
* cuBLAS: use multiple streams, choose smartly between mul_mat_q and mul_mat_f16
* fix build
* cuBLAS: update block_q5_1
2023-05-01 18:11:07 +02:00
slaren and GitHub
b925f1f1b0
cuBLAS: fall back to pageable memory if pinned alloc fails ( #1233 )
...
* cuBLAS: fall back to pageable memory if pinned alloc fails
* cuBLAS: do not use pinned memory if env variable GGML_CUDA_NO_PINNED is set
2023-05-01 13:32:22 +02:00
slaren and GitHub
7fc50c051a
cuBLAS: use host pinned memory and dequantize while copying ( #1207 )
...
* cuBLAS: dequantize simultaneously while copying memory
* cuBLAS: use host pinned memory
* cuBLAS: improve ggml_compute_forward_mul_mat_f16_f32 with pinned memory
* cuBLAS: also pin kv cache
* fix rebase
2023-04-29 02:04:18 +02:00
e4cf982e0d
Fix cuda compilation ( #1128 )
...
* Fix: Issue with CUBLAS compilation error due to missing -fPIC flag
---------
Co-authored-by: B1gM8c <89020353+B1gM8c@users.noreply.github.com >
2023-04-24 17:29:58 +02:00
slaren and GitHub
1d78fecdab
Fix LoRA acronym ( #1145 )
2023-04-23 23:03:44 +02:00
slaren and GitHub
50cb666b8a
Improve cuBLAS performance by using a memory pool ( #1094 )
...
* Improve cuBLAS performance by using a memory pool
* Move cuda specific definitions to ggml-cuda.h/cu
* Add CXX flags to nvcc
* Change memory pool synchronization mechanism to a spin lock
General code cleanup
2023-04-21 21:59:17 +02:00
slaren and GitHub
3d59769c3b
Show perplexity ETA in hours and minutes ( #1096 )
2023-04-21 14:57:57 +02:00
slaren and GitHub
2005469ea1
Add Q4_3 support to cuBLAS ( #1086 )
2023-04-20 20:49:53 +02:00
slaren and GitHub
02d6988121
Improve cuBLAS performance by dequantizing on the GPU ( #1065 )
2023-04-20 03:14:14 +02:00
slaren and GitHub
8944a13296
Add NVIDIA cuBLAS support ( #1044 )
2023-04-19 11:22:45 +02:00
6667401238
Multi-threaded ggml_cpy ( #1035 )
...
* Multi-threaded ggml_cpy
* Update ggml.c
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
* Also fix wdata offset in ggml_compute_forward_add_q_f32
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2023-04-19 00:53:24 +02:00
slaren and GitHub
315a95a4d3
Add LoRA support ( #820 )
2023-04-17 17:28:55 +02:00
slaren and GitHub
47f61aaa5f
Fix: do not close file on mmap ( #1017 )
2023-04-16 21:27:38 +02:00
Slaren and Pavol Rusnak
0d054e292e
Show error message when -f fails
2023-04-01 16:08:40 +02:00
slaren and GitHub
1d08882afa
Optimize AVX2 ggml_vec_dot_q4_0 ( #642 )
2023-03-31 15:55:52 +00:00
Slaren and Justine Tunney
a017390358
Initial windows support (untested)
2023-03-30 12:28:25 -07:00
Slaren and Justine Tunney
ac184d5147
Always initialize mm_addr and mm_length in llama_model
2023-03-30 12:28:25 -07:00
Slaren and Justine Tunney
276e5b7811
Unmap the file in llama_free
2023-03-30 12:28:25 -07:00
Slaren and Justine Tunney
d68c5dc435
Make mmap_file static
2023-03-30 12:28:25 -07:00
Slaren and Justine Tunney
64bde3ffd4
Fix ggml_init_params in quantize
2023-03-30 12:28:25 -07:00
Slaren and Justine Tunney
c03ae8dca1
Add mmap support for model files
2023-03-30 12:28:25 -07:00
slaren and GitHub
ed3c680bcd
Fix GGML_F32Cx8_STORE in AVX without F16C path ( #619 )
2023-03-30 11:16:30 +02:00
2a98bc18ea
ggml : add AVX2 implementation of quantize_row_q4_1 ( #515 )
...
* Add AVX2 implementation of quantize_row_q4_1
* Actually use AVX2
* Make quantize_row_q4_1 static
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2023-03-28 21:06:03 +03:00
slaren and GitHub
a6bdc47cba
Fix usage of F16C intrinsics in AVX code ( #563 )
...
* Fix usage of F16C intrinsics in AVX code when F16C is not defined
2023-03-28 17:26:55 +03:00
slaren and GitHub
459e93cce0
Add AVX2 implementation of dequantize_row_q4_1 ( #505 )
2023-03-25 20:31:48 +02:00
slaren and GitHub
09aecbf628
Add AVX2 implementation of dequantize_row_q4_0 ( #467 )
2023-03-25 17:06:49 +02:00
slaren and GitHub
29b7baab67
Add timings for the prompt evaluation ( #478 )
2023-03-25 16:34:23 +02:00
50fae10d03
Add --ignore-eos parameter ( #181 )
...
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2023-03-19 20:22:48 +02:00