Georgi Gerganov
ec2e10c444
llama : add llama_init_backend() API ( close #1527 )
2023-05-20 11:06:37 +03:00
Georgi Gerganov
8a203f9fa1
llama : fix compile warnings in llama_set_state_data()
2023-05-20 10:14:43 +03:00
Georgi Gerganov
4fd3e29297
ggml : fix scalar implementation of Q4_1 dot
2023-05-20 10:13:19 +03:00
Georgi Gerganov and GitHub
2d5db48371
ggml : use F16 instead of F32 in Q4_0, Q4_1, Q8_0 ( #1508 )
...
* ggml : use F16 instead of F32 in Q4_0, Q4_1 and Q8_0
* llama : bump LLAMA_FILE_VERSION to 3
* cuda : update Q4 and Q8 dequantize kernels
* ggml : fix AVX dot products
* readme : update performance table + hot topics
2023-05-19 22:17:18 +03:00
Georgi Gerganov
6986c7835a
tests : add missing header
2023-05-19 21:17:28 +03:00
Georgi Gerganov
4b7e245adf
minor : fix compile warnings
2023-05-19 20:14:51 +03:00
Georgi Gerganov and GitHub
13c351ad72
ggml : various fixes ( #1450 )
...
- `ggml_rope()`
- `ggml_diag_mask_inf()` multi-threaded
- compatibility with scratch buffers
2023-05-14 18:22:50 +03:00
Georgi Gerganov
601a033475
ggml : add GGML_QNT_VERSION to track quantization format changes
...
https://github.com/ggerganov/ggml/issues/150#issuecomment-1546625668
2023-05-14 10:20:19 +03:00
Georgi Gerganov
08737ef720
cuda : fix convert function ( #1412 )
2023-05-13 17:40:58 +03:00
Georgi Gerganov
bda4d7c215
make : fix PERF build with cuBLAS
2023-05-13 17:25:09 +03:00
Georgi Gerganov
5a5aeb1e91
llama : fix unused warning
2023-05-13 16:55:14 +03:00
Georgi Gerganov and GitHub
66841fdb0e
ggml : multi-thread mul and diag_mask ops ( #1428 )
2023-05-13 16:48:03 +03:00
Georgi Gerganov
f048af0230
ggml : sync alibi fix from ggml repo
2023-05-13 11:54:33 +03:00
Georgi Gerganov
0cd22e190a
llama : fix various warnings
2023-05-13 11:23:15 +03:00
Georgi Gerganov and GitHub
cdd5350892
readme : update Q4_0 perplexities
...
I think these were affected by the removal of the `round` during quantization
2023-05-13 09:12:44 +03:00
Georgi Gerganov
738ace394a
llama : free ggml context in set / copy state data ( close #1425 )
2023-05-13 09:08:52 +03:00
Georgi Gerganov
fb62f92433
llama : fix --mtest option ( close #1414 )
2023-05-12 21:44:20 +03:00
b9fd7eee57
ggml : remove bit shuffling ( #1405 )
...
* ggml : remove Q4_0 bit shufling (ARM NEON)
* ggml : remove Q4_1 bit shuffling (ARM NEON + reference)
* ggml : nibbles_from_floats() + bytes_from_nibbles() (ARM NEON)
* ggml : remove Q4_2 bit shuffling (WIP, BROKEN)
* ggml : remove Q5_0 bit shuffling (ARM NEON)
* ggml : 2x faster scalar implementations
* ggml : remove Q5_1 bit shuffling (ARM NEON + scalar)
* ggml : simplify scalar dot
* ggml : remove WASM SIMD bit shuffling + remove vzip for ARM 32-bit
* ggml : fix Q4_1 quantization
* ggml : update cuBLAS + normalize variable names
* ggml : remove Q4_2 mode
* ggml : minor formatting
* ggml : fix Q5_0 quantization
* scripts : add script for measuring the time per token
* AVX implementations (#1370 )
* ggml : uniform 5th bit extraction
* llama : produce error upon loading old model files
* llama : fix model magic/version write
* ggml : speed-up Q5_0 + Q5_1 at 4 threads
* ggml : preserve old Q4 and Q5 formats
* ggml : simplify Q8_1 - no need for low / high sums anymore
* ggml : fix Q8_0 and Q8_1 rounding
* Revert "AVX implementations (#1370 )"
This reverts commit 948d124837f9d287d8490f41338e0e4cceb0814f.
* ggml : fix AVX2 implementation
* sha : update hashes for 7B and 13B
* readme : update timings + remove warning banner
* llama : update v2 PR number to 1405
* ggml : fix WASM comments
* ggml : back to original bit order
* readme : add note that Q4 and Q5 have been changed
* llama : fix return for unknown version
---------
Co-authored-by: Stephan Walter <stephan@walter.name >
2023-05-12 00:23:08 +03:00
Georgi Gerganov and GitHub
56551bc11f
readme : add notice about upcoming breaking change
2023-05-08 22:52:18 +03:00
Georgi Gerganov and GitHub
f9a6364912
llama : require first token to be BOS ( #1303 )
...
* llama : require first token to be BOS
* scripts : add ppl-run-all.sh
* perplexity : add BOS for each chunk
* readme : update perplexity values after BOS fix
* perplexity : add clarifying comments
2023-05-08 17:41:54 +03:00
Georgi Gerganov
799fdc1b5d
ggml : vectorize Q8_0 quantization
...
https://github.com/ggerganov/ggml/pull/127#issuecomment-1533648531
2023-05-03 23:24:20 +03:00
Georgi Gerganov and GitHub
bca9ad938a
minor : fix whitespaces ( #1302 )
2023-05-03 20:09:42 +03:00
Georgi Gerganov
e2a937ca6a
minor : fix trailing whitespaces
2023-05-03 18:43:23 +03:00
Georgi Gerganov
0e6cbff1b7
llama : fix compile warnings
2023-05-02 23:09:08 +03:00
Georgi Gerganov
5d5817ca60
ggml : fix 32-bit ARM
2023-05-02 22:14:50 +03:00
Georgi Gerganov and GitHub
70269cae37
llama : fix session load / save ( #1263 )
2023-05-01 14:54:59 +03:00
Georgi Gerganov
7ff0dcd320
ggml : fix UB (int << 31)
2023-04-30 22:28:51 +03:00
Georgi Gerganov
6bc4400e67
ggml : add Q5 WASM SIMD + GGML_FTYPE
2023-04-30 19:07:43 +03:00
Georgi Gerganov
3e5aa8a1c4
ggml : fix labels for GGML_OP_ALIBI
2023-04-30 10:25:46 +03:00
Georgi Gerganov
c3ca7a5f05
ggml : fix 32-bit ARM NEON
2023-04-29 21:34:23 +03:00
Georgi Gerganov
e8c051611a
ggml : use vzip instead of vuzp for consistency
2023-04-29 21:12:56 +03:00
Georgi Gerganov
0b5a935099
ggml : fix visibility and unused warnings
2023-04-29 19:28:36 +03:00
Georgi Gerganov and GitHub
ec728e44d7
ggml : fix #if for f32_f32 mul_mat (CLBlast) ( #1229 )
2023-04-29 18:43:42 +03:00
Georgi Gerganov and GitHub
214b6a3570
ggml : adjust mul_mat_f16 work memory ( #1226 )
...
* llama : minor - remove explicity int64_t cast
* ggml : reduce memory buffer for F16 mul_mat when not using cuBLAS
* ggml : add asserts to guard for incorrect wsize
2023-04-29 18:43:28 +03:00
Georgi Gerganov
305eb5afd5
build : fix reference to old llama_util.h
2023-04-29 13:53:12 +03:00
Georgi Gerganov
84ca9c2ecf
examples : fix save-load-state + rename llama-util.h
2023-04-29 13:48:11 +03:00
Georgi Gerganov and GitHub
334637e43e
common : change default parameters to pre-#1126 ( #1223 )
2023-04-29 09:51:06 +03:00
Georgi Gerganov and GitHub
7f15c5c477
readme : update hot topics
2023-04-28 21:32:52 +03:00
Georgi Gerganov
55390bcaf2
ggml : sync ggml (ggml_alibi)
2023-04-28 20:51:05 +03:00
Georgi Gerganov
11d902364b
ggml : add helper debug printf in soft_max
2023-04-28 17:59:08 +03:00
Georgi Gerganov and GitHub
f9be42add0
readme : add quantization info
2023-04-26 23:24:42 +03:00
574406dc7e
ggml : add Q5_0 and Q5_1 quantization ( #1187 )
...
* ggml : add Q5_0 quantization (cuBLAS only)
* ggml : fix Q5_0 qh -> uint32_t
* ggml : fix q5_0 histogram stats
* ggml : q5_0 scalar dot product
* ggml : q5_0 ARM NEON dot
* ggml : q5_0 more efficient ARM NEON using uint64_t masks
* ggml : rename Q5_0 -> Q5_1
* ggml : adding Q5_0 mode
* quantize : add Q5_0 and Q5_1 to map
* ggml : AVX2 optimizations for Q5_0, Q5_1 (#1195 )
---------
Co-authored-by: Stephan Walter <stephan@walter.name >
2023-04-26 23:14:13 +03:00
Georgi Gerganov and GitHub
7a32fcb3b2
ggml : add Q8_0 quantization format (rename the old one to Q8_1) (ARM NEON) ( #1179 )
...
* ggml : add Q8_0 quantization format (rename the old one to Q8_1)
* tests : fix test-quantize-fns
* ggml : finalize Q8_0 implementation
* ggml : use q4_0_q8_0 and q4_2_q8_0
* ggml : fix Q8_0 dot product bug (ARM)
* ggml : Q8_0 unroll x2
* ggml : fix bug - using wrong block type
* ggml : extend quantize_fns_t with "vec_dot_type"
* ggml : fix Q8_0 to use 255 values out of 256
* ggml : fix assert using wrong QK4_2 instead of QK4_3
2023-04-25 23:40:51 +03:00
Georgi Gerganov and GitHub
8a0f8673ba
ggml : export symbols ( #1155 )
2023-04-24 22:18:25 +03:00
Georgi Gerganov
957c8ae21d
llama : increase scratch buffer size for 65B (ref #1152 )
...
Temporary solution
2023-04-24 18:47:30 +03:00
Georgi Gerganov and GitHub
c4fe84fb0d
llama : refactor get / set state + remove redundant kv cache API ( #1143 )
2023-04-24 07:40:02 +03:00
Georgi Gerganov
284685f169
scripts : add helper scripts to synch ggml repo
2023-04-23 19:57:09 +03:00
Georgi Gerganov
ec9cdb6752
ggml : do not print perf ops that have not been used at all
2023-04-23 18:32:52 +03:00
Georgi Gerganov
e4422e299c
ggml : better PERF prints + support "LLAMA_PERF=1 make"
2023-04-23 18:15:39 +03:00
Georgi Gerganov
0e018fe008
ggml : fix Q4_3 cuBLAS
2023-04-22 16:32:07 +03:00
Georgi Gerganov
872c365a91
ggml : fix AVX build + update to new Q8_0 format
2023-04-22 11:08:12 +03:00
Georgi Gerganov and GitHub
955ef9a5d5
ggml : alternative Q4_3 implementation using modified Q8_0 ( #1109 )
...
* ggml : prefer vzip to vuzp
This way we always use the same type of instruction across all quantizations
* ggml : alternative Q4_3 implementation using modified Q8_0
* ggml : fix Q4_3 scalar imlpementation
* ggml : slight improvement of Q4_3 - no need for loop unrolling
* ggml : fix AVX paths for Q8_0 quantization
2023-04-22 10:55:35 +03:00
Georgi Gerganov
d40fded93e
llama : fix comment for "output.weight" tensor
2023-04-21 10:24:02 +03:00
Georgi Gerganov
12b5900dbc
ggml : sync ggml (add GPT-NeoX RoPE implementation)
2023-04-20 23:32:59 +03:00
Georgi Gerganov
9ff334f3c9
ggml : fix bug in ggml_compute_forward_dup_f32()
2023-04-20 21:58:38 +03:00
Georgi Gerganov
8a1756abdf
ggml : do not break cuBLAS build (Q4_3 is not yet implemented)
2023-04-20 21:43:50 +03:00
Georgi Gerganov
66aab46079
ggml : fix Q4_3 quantization
...
Broke it during conflict resolution in last PR
2023-04-20 20:44:05 +03:00
Georgi Gerganov and GitHub
e0305ead3a
ggml : add Q4_3 quantization ( #1082 )
2023-04-20 20:35:53 +03:00
884e7d7a2b
ggml : use 8-bit precision for Q4_1 intermediate results ( #1047 )
...
* ggml : use 8-bit precision for Q4_1 intermediate results (ARM)
* ggml : optimize ggml_vec_dot_q4_1_q8_0() via vmalq_n_f32
56 ms/token with Q4_1 !
* ggml : AVX2 implementation of ggml_vec_dot_q4_1_q8_0 (#1051 )
* gitignore : ignore ppl-*.txt files
---------
Co-authored-by: slaren <2141330+slaren@users.noreply.github.com >
2023-04-19 20:10:08 +03:00
Georgi Gerganov and GitHub
7cd5c4a3e9
readme : add warning about Q4_2 and Q4_3
2023-04-19 19:07:54 +03:00
Georgi Gerganov and GitHub
77a73403ca
ggml : add new Q4_2 quantization (ARM only) ( #1046 )
...
* ggml : Q4_2 ARM
* ggml : add ggml_is_quantized()
* llama : update llama_type_name() with Q4_2 entry
* ggml : speed-up q4_2
- 4 threads: ~100ms -> ~90ms
- 8 threads: ~55ms -> ~50ms
* ggml : optimize q4_2 using vmlaq_n_f32 + vmulq_n_f32
2023-04-18 23:54:57 +03:00
Georgi Gerganov
50a8a2af97
ggml : scratch that - vmlaq_n_f32 is always better
...
Had a background process that was messing with the timings
2023-04-18 23:11:23 +03:00
Georgi Gerganov
4caebf6d40
gitignore : vdot
2023-04-18 23:00:08 +03:00
Georgi Gerganov
dcdd65e296
ggml : optimize ggml_vec_dot_q4_0_q8_0() using vectorized accumulators
2023-04-18 22:59:17 +03:00
Georgi Gerganov and GitHub
7faa7460f0
readme : update hot topics about new LoRA functionality
2023-04-18 20:10:26 +03:00
Georgi Gerganov
5af8e32238
ci : do not run on drafts
2023-04-18 19:57:06 +03:00
Georgi Gerganov
eb17a026fd
quantize-stats : fix bug in --type argument
2023-04-17 17:31:06 +03:00
Georgi Gerganov
69b740289f
ggml : avoid using ggml_fp16_to_fp32() and ggml_fp32_to_fp16() in ggml.c
2023-04-17 16:16:23 +03:00
Georgi Gerganov
3173a62eb9
stdout : vertical align outputs for better readibility
2023-04-16 13:59:27 +03:00
e95b6554b4
ggml : add Q8_0 quantization for intermediate results ( #951 )
...
* ggml : add Q8_0 quantization for intermediate results
* quantize-stats : fix test + add it to Makefile default
* Q8: use int8_t, AVX/AVX2 optimizations
* ggml : fix quantize_row_q8_0() ARM_NEON rounding
* minor : updates after rebase to latest master
* quantize-stats : delete obsolete strings
* ggml : fix q4_1 dot func
---------
Co-authored-by: Stephan Walter <stephan@walter.name >
2023-04-15 17:53:22 +03:00
Georgi Gerganov
aa485cee33
ggml : use posix_memalign on non-Windows env
2023-04-15 14:25:45 +03:00
Georgi Gerganov
1623a6e9b4
ggml : minor
2023-04-14 13:31:29 +03:00
Georgi Gerganov
c14e0d2f23
ggml : always allocate buffers with size multiple of GGML_MEM_ALIGN
2023-04-14 13:31:15 +03:00
Georgi Gerganov
0f07cacb05
ggml : fix q4_1 dot product types
2023-04-14 09:45:42 +03:00
Georgi Gerganov
a3a2a0eda8
ggml : add GGML_DEFAULT_N_THREADS
2023-04-13 18:36:48 +03:00
Georgi Gerganov and GitHub
d990e3fffc
ggml : speed-up ggml_vec_dot_q4_1() ARM_NEON + 32-bit ARM support ( #900 )
...
* ggml : speed-up q4_1 ARM_NEON by ~5%
* ggml : implement vaddvq when missing
* ggml : implement vminvq and vmaxvq when missing
* ggml : implement vzip when missing
* ggml : fix comment
* ggml : try to use correct ifdef
2023-04-13 18:32:36 +03:00
Georgi Gerganov
9190e8eac8
llama : merge llama_internal.h into llama.h
...
Hide it behind an #ifdef
2023-04-13 18:04:45 +03:00
Georgi Gerganov
c85980acd0
gitignore : benchmark
2023-04-13 18:01:33 +03:00
Georgi Gerganov and GitHub
f76cb3a34d
readme : change "GPU support" link to discussion
2023-04-12 14:48:57 +03:00
Georgi Gerganov and GitHub
782438070f
readme : update hot topics with link to "GPU support" issue
2023-04-12 14:31:12 +03:00
Georgi Gerganov
461ba9e66e
ggml : fix WASM build
2023-04-10 23:20:01 +03:00
Georgi Gerganov
c3ac702e5e
ggml : add ggml_cont() + optimize ggml_cpy() for contiguous dst
2023-04-10 22:42:28 +03:00
Georgi Gerganov
9d634ef452
ggml : remove trailing whitespaces
2023-04-10 22:42:28 +03:00
Georgi Gerganov
684da25926
ggml : fix quantize_row_q4_1() ARM_NEON ( close #876 )
2023-04-10 19:29:48 +03:00
Georgi Gerganov and GitHub
eeaa7b0492
ggml : multi-thread ggml_rope() (~3-4 times faster on M1) ( #781 )
2023-04-05 22:11:03 +03:00
Georgi Gerganov and GitHub
986b6ce9f9
ggml, llama : avoid heavy V transpose + improvements ( #775 )
...
ggml :
- added ggml_view_3d()
- ggml_view_tensor() now inherits the stride too
- reimplement ggml_cpy() to account for dst stride
- no longer require tensor->data to be memory aligned
llama :
- compute RoPE on 32-bit tensors (should be more accurate)
- store RoPE-ed K in the KV cache
- store transposed V in the KV cache (significant speed-up)
- avoid unnecessary Q copy
2023-04-05 22:07:33 +03:00
Georgi Gerganov and GitHub
3416298929
Update README.md
2023-04-05 19:54:30 +03:00
Georgi Gerganov
62b3e81aae
media : add logos and banners
2023-04-05 18:58:31 +03:00
Georgi Gerganov and GitHub
8d10406d6e
readme : change logo + add bindings + add uis + add wiki
2023-04-05 18:56:20 +03:00
Georgi Gerganov and GitHub
3df890aef4
readme : update supported models
2023-03-30 22:31:54 +03:00
Georgi Gerganov
77efdf5a50
ggml : fix NEON signs ( close #620 , #622 )
2023-03-30 20:27:32 +03:00
Georgi Gerganov
b51c717d5c
ggml : init time on first ggml_init() call
2023-03-29 22:15:34 +03:00
Georgi Gerganov
0ba76c1e73
llama : fix compile warnings when reading the vocab
2023-03-29 22:13:12 +03:00
Georgi Gerganov
cea1c85948
ggml : add ARM_NEON dequantize_row_q4_1()
2023-03-29 22:10:01 +03:00
Georgi Gerganov
f202ada131
ggml : add ARM_NEON quantize_row_q4_1()
2023-03-29 22:03:07 +03:00
Georgi Gerganov
3b44d30d9b
ggml : add ARM_NEON ggml_vec_dot_q4_1()
2023-03-29 22:03:07 +03:00
Georgi Gerganov and GitHub
b467702b87
readme : fix typos
2023-03-29 19:38:31 +03:00
Georgi Gerganov and GitHub
516d88e75c
readme : add GPT4All instructions ( close #588 )
2023-03-29 19:37:20 +03:00
Georgi Gerganov
53635c081c
py : add GPT4All conversion script
...
For now: copy-paste
Too much time for me to deduplicate the python code
2023-03-29 19:29:52 +03:00
Georgi Gerganov
96f9c0506f
ci : make ctest verbose, hopefully we see what is wrong with the sanitizer
2023-03-28 20:01:09 +03:00