Georgi Gerganov and GitHub
c4fe84fb0d
llama : refactor get / set state + remove redundant kv cache API ( #1143 )
2023-04-24 07:40:02 +03:00
Georgi Gerganov
284685f169
scripts : add helper scripts to synch ggml repo
2023-04-23 19:57:09 +03:00
Georgi Gerganov
ec9cdb6752
ggml : do not print perf ops that have not been used at all
2023-04-23 18:32:52 +03:00
Georgi Gerganov
e4422e299c
ggml : better PERF prints + support "LLAMA_PERF=1 make"
2023-04-23 18:15:39 +03:00
Georgi Gerganov
0e018fe008
ggml : fix Q4_3 cuBLAS
2023-04-22 16:32:07 +03:00
Georgi Gerganov
872c365a91
ggml : fix AVX build + update to new Q8_0 format
2023-04-22 11:08:12 +03:00
Georgi Gerganov and GitHub
955ef9a5d5
ggml : alternative Q4_3 implementation using modified Q8_0 ( #1109 )
...
* ggml : prefer vzip to vuzp
This way we always use the same type of instruction across all quantizations
* ggml : alternative Q4_3 implementation using modified Q8_0
* ggml : fix Q4_3 scalar imlpementation
* ggml : slight improvement of Q4_3 - no need for loop unrolling
* ggml : fix AVX paths for Q8_0 quantization
2023-04-22 10:55:35 +03:00
Georgi Gerganov
d40fded93e
llama : fix comment for "output.weight" tensor
2023-04-21 10:24:02 +03:00
Georgi Gerganov
12b5900dbc
ggml : sync ggml (add GPT-NeoX RoPE implementation)
2023-04-20 23:32:59 +03:00
Georgi Gerganov
9ff334f3c9
ggml : fix bug in ggml_compute_forward_dup_f32()
2023-04-20 21:58:38 +03:00
Georgi Gerganov
8a1756abdf
ggml : do not break cuBLAS build (Q4_3 is not yet implemented)
2023-04-20 21:43:50 +03:00
Georgi Gerganov
66aab46079
ggml : fix Q4_3 quantization
...
Broke it during conflict resolution in last PR
2023-04-20 20:44:05 +03:00
Georgi Gerganov and GitHub
e0305ead3a
ggml : add Q4_3 quantization ( #1082 )
2023-04-20 20:35:53 +03:00
884e7d7a2b
ggml : use 8-bit precision for Q4_1 intermediate results ( #1047 )
...
* ggml : use 8-bit precision for Q4_1 intermediate results (ARM)
* ggml : optimize ggml_vec_dot_q4_1_q8_0() via vmalq_n_f32
56 ms/token with Q4_1 !
* ggml : AVX2 implementation of ggml_vec_dot_q4_1_q8_0 (#1051 )
* gitignore : ignore ppl-*.txt files
---------
Co-authored-by: slaren <2141330+slaren@users.noreply.github.com >
2023-04-19 20:10:08 +03:00
Georgi Gerganov and GitHub
7cd5c4a3e9
readme : add warning about Q4_2 and Q4_3
2023-04-19 19:07:54 +03:00
Georgi Gerganov and GitHub
77a73403ca
ggml : add new Q4_2 quantization (ARM only) ( #1046 )
...
* ggml : Q4_2 ARM
* ggml : add ggml_is_quantized()
* llama : update llama_type_name() with Q4_2 entry
* ggml : speed-up q4_2
- 4 threads: ~100ms -> ~90ms
- 8 threads: ~55ms -> ~50ms
* ggml : optimize q4_2 using vmlaq_n_f32 + vmulq_n_f32
2023-04-18 23:54:57 +03:00
Georgi Gerganov
50a8a2af97
ggml : scratch that - vmlaq_n_f32 is always better
...
Had a background process that was messing with the timings
2023-04-18 23:11:23 +03:00
Georgi Gerganov
4caebf6d40
gitignore : vdot
2023-04-18 23:00:08 +03:00
Georgi Gerganov
dcdd65e296
ggml : optimize ggml_vec_dot_q4_0_q8_0() using vectorized accumulators
2023-04-18 22:59:17 +03:00
Georgi Gerganov and GitHub
7faa7460f0
readme : update hot topics about new LoRA functionality
2023-04-18 20:10:26 +03:00
Georgi Gerganov
5af8e32238
ci : do not run on drafts
2023-04-18 19:57:06 +03:00
Georgi Gerganov
eb17a026fd
quantize-stats : fix bug in --type argument
2023-04-17 17:31:06 +03:00
Georgi Gerganov
69b740289f
ggml : avoid using ggml_fp16_to_fp32() and ggml_fp32_to_fp16() in ggml.c
2023-04-17 16:16:23 +03:00
Georgi Gerganov
3173a62eb9
stdout : vertical align outputs for better readibility
2023-04-16 13:59:27 +03:00
e95b6554b4
ggml : add Q8_0 quantization for intermediate results ( #951 )
...
* ggml : add Q8_0 quantization for intermediate results
* quantize-stats : fix test + add it to Makefile default
* Q8: use int8_t, AVX/AVX2 optimizations
* ggml : fix quantize_row_q8_0() ARM_NEON rounding
* minor : updates after rebase to latest master
* quantize-stats : delete obsolete strings
* ggml : fix q4_1 dot func
---------
Co-authored-by: Stephan Walter <stephan@walter.name >
2023-04-15 17:53:22 +03:00
Georgi Gerganov
aa485cee33
ggml : use posix_memalign on non-Windows env
2023-04-15 14:25:45 +03:00
Georgi Gerganov
1623a6e9b4
ggml : minor
2023-04-14 13:31:29 +03:00
Georgi Gerganov
c14e0d2f23
ggml : always allocate buffers with size multiple of GGML_MEM_ALIGN
2023-04-14 13:31:15 +03:00
Georgi Gerganov
0f07cacb05
ggml : fix q4_1 dot product types
2023-04-14 09:45:42 +03:00
Georgi Gerganov
a3a2a0eda8
ggml : add GGML_DEFAULT_N_THREADS
2023-04-13 18:36:48 +03:00
Georgi Gerganov and GitHub
d990e3fffc
ggml : speed-up ggml_vec_dot_q4_1() ARM_NEON + 32-bit ARM support ( #900 )
...
* ggml : speed-up q4_1 ARM_NEON by ~5%
* ggml : implement vaddvq when missing
* ggml : implement vminvq and vmaxvq when missing
* ggml : implement vzip when missing
* ggml : fix comment
* ggml : try to use correct ifdef
2023-04-13 18:32:36 +03:00
Georgi Gerganov
9190e8eac8
llama : merge llama_internal.h into llama.h
...
Hide it behind an #ifdef
2023-04-13 18:04:45 +03:00
Georgi Gerganov
c85980acd0
gitignore : benchmark
2023-04-13 18:01:33 +03:00
Georgi Gerganov and GitHub
f76cb3a34d
readme : change "GPU support" link to discussion
2023-04-12 14:48:57 +03:00
Georgi Gerganov and GitHub
782438070f
readme : update hot topics with link to "GPU support" issue
2023-04-12 14:31:12 +03:00
Georgi Gerganov
461ba9e66e
ggml : fix WASM build
2023-04-10 23:20:01 +03:00
Georgi Gerganov
c3ac702e5e
ggml : add ggml_cont() + optimize ggml_cpy() for contiguous dst
2023-04-10 22:42:28 +03:00
Georgi Gerganov
9d634ef452
ggml : remove trailing whitespaces
2023-04-10 22:42:28 +03:00
Georgi Gerganov
684da25926
ggml : fix quantize_row_q4_1() ARM_NEON ( close #876 )
2023-04-10 19:29:48 +03:00
Georgi Gerganov and GitHub
eeaa7b0492
ggml : multi-thread ggml_rope() (~3-4 times faster on M1) ( #781 )
2023-04-05 22:11:03 +03:00
Georgi Gerganov and GitHub
986b6ce9f9
ggml, llama : avoid heavy V transpose + improvements ( #775 )
...
ggml :
- added ggml_view_3d()
- ggml_view_tensor() now inherits the stride too
- reimplement ggml_cpy() to account for dst stride
- no longer require tensor->data to be memory aligned
llama :
- compute RoPE on 32-bit tensors (should be more accurate)
- store RoPE-ed K in the KV cache
- store transposed V in the KV cache (significant speed-up)
- avoid unnecessary Q copy
2023-04-05 22:07:33 +03:00
Georgi Gerganov and GitHub
3416298929
Update README.md
2023-04-05 19:54:30 +03:00
Georgi Gerganov
62b3e81aae
media : add logos and banners
2023-04-05 18:58:31 +03:00
Georgi Gerganov and GitHub
8d10406d6e
readme : change logo + add bindings + add uis + add wiki
2023-04-05 18:56:20 +03:00
Georgi Gerganov and GitHub
3df890aef4
readme : update supported models
2023-03-30 22:31:54 +03:00
Georgi Gerganov
77efdf5a50
ggml : fix NEON signs ( close #620 , #622 )
2023-03-30 20:27:32 +03:00
Georgi Gerganov
b51c717d5c
ggml : init time on first ggml_init() call
2023-03-29 22:15:34 +03:00
Georgi Gerganov
0ba76c1e73
llama : fix compile warnings when reading the vocab
2023-03-29 22:13:12 +03:00
Georgi Gerganov
cea1c85948
ggml : add ARM_NEON dequantize_row_q4_1()
2023-03-29 22:10:01 +03:00
Georgi Gerganov
f202ada131
ggml : add ARM_NEON quantize_row_q4_1()
2023-03-29 22:03:07 +03:00
Georgi Gerganov
3b44d30d9b
ggml : add ARM_NEON ggml_vec_dot_q4_1()
2023-03-29 22:03:07 +03:00
Georgi Gerganov and GitHub
b467702b87
readme : fix typos
2023-03-29 19:38:31 +03:00
Georgi Gerganov and GitHub
516d88e75c
readme : add GPT4All instructions ( close #588 )
2023-03-29 19:37:20 +03:00
Georgi Gerganov
53635c081c
py : add GPT4All conversion script
...
For now: copy-paste
Too much time for me to deduplicate the python code
2023-03-29 19:29:52 +03:00
Georgi Gerganov
96f9c0506f
ci : make ctest verbose, hopefully we see what is wrong with the sanitizer
2023-03-28 20:01:09 +03:00
Georgi Gerganov
d502bc7c9d
tests : free llama context at the end of the test
2023-03-28 19:51:55 +03:00
Georgi Gerganov
e0670260fb
gitignore : add "embedding"
2023-03-28 18:34:35 +03:00
Georgi Gerganov and GitHub
348d6926ee
Add logo to README.md
2023-03-26 10:20:49 +03:00
Georgi Gerganov
c2b25b6912
Fix colors enabling on WIN32
2023-03-25 21:53:39 +02:00
Georgi Gerganov
79b2b266db
If n_predict == -1, generate forever
2023-03-25 21:51:41 +02:00
Georgi Gerganov
e2d490dafd
Inifinite generation via context swapping ( #71 )
2023-03-25 21:36:22 +02:00
Georgi Gerganov
03f7e33560
Cleanup STL headers + fix embedding examples + minor stuff
2023-03-25 20:51:14 +02:00
Georgi Gerganov
55ad42af84
Move chat scripts into "./examples"
2023-03-25 20:37:09 +02:00
Georgi Gerganov
a316a425d0
Overhaul the examples structure
...
- main -> examples
- utils -> examples (renamed to "common")
- quantize -> examples
- separate tools for "perplexity" and "embedding"
Hope I didn't break something !
2023-03-25 20:26:40 +02:00
Georgi Gerganov and GitHub
ecbe466a36
Retire the ggml_mul_mat() branch for transposed src0 ( #500 )
...
* Retire the ggml_mul_mat() for transposed src0
- It can always be made contiguous with ggml_cpy()
- The code is now simplified
- The results are deterministic in respect to num threads
* SIMD-ify dequantize_row_q4_0() for ARM_NEON (#502 )
* Attempt to SIMD-ify dequantize_row_q4_0() for ARM_NEON
* Fix dequantization - forgot to interleave the quants
2023-03-25 19:47:21 +02:00
Georgi Gerganov
502a400192
Disable prompt verbosity by default and add option to enable ( #480 )
2023-03-25 17:17:16 +02:00
Georgi Gerganov
4640eff23d
Don't interefe with BLAS for large prompts by running only 1 thread
2023-03-25 17:03:10 +02:00
Georgi Gerganov
ab77d76312
Add longer DAN prompt for testing big batch numbers
2023-03-25 16:49:09 +02:00
Georgi Gerganov and GitHub
4a7129acd2
Remove obsolete information from README
2023-03-25 16:30:32 +02:00
Georgi Gerganov
6b6dbc8910
Remove obsolete assert and fix compiler warning
2023-03-25 16:22:05 +02:00
Georgi Gerganov
2a2e63ce05
Fix nasty bug in ggml_compute_forward_mul_mat_f32() and reenable BLAS
2023-03-25 16:10:14 +02:00
Georgi Gerganov
8520fc310e
Disable BLAS altogether - the bug is not just for qunatized mat mul
2023-03-24 23:47:06 +02:00
Georgi Gerganov
b3f460e941
Disable BLAS branch in mul_mat - seems there is a bug
2023-03-24 23:39:17 +02:00
Georgi Gerganov and GitHub
04c6f5ed6f
Immediately start processing the prompt before user input has been provided ( #476 )
2023-03-24 23:17:58 +02:00
Georgi Gerganov and GitHub
7a9b6c3a8b
Reduce memory usage and allocate enough memory for largest context ( #473 )
...
* Reduce memory usage and allocate enough memory for large contexts
* Simpler scratch buffer usage
* Reenable BLAS for quantized mul_mat
* Fix number of layers in 30B and 65B
* Fix KV cache size for F32
2023-03-24 23:17:37 +02:00
Georgi Gerganov
31572d9665
Temporary bump the memory buffer size - hopefully fix issues from 483bab2e
2023-03-24 18:23:56 +02:00
Georgi Gerganov
afd220d9c6
Properly free llama_context on failure
2023-03-24 17:21:01 +02:00
Georgi Gerganov and GitHub
b6b268d441
Add link to Roadmap discussion
2023-03-24 09:13:35 +02:00
Georgi Gerganov
3cd8dde0d1
Revert "Fix memory allocation issues and seg faults"
...
This reverts commit 4870e455b3 .
Will provide the correct fix later
2023-03-24 06:22:28 +02:00
Georgi Gerganov
4870e455b3
Fix memory allocation issues and seg faults
2023-03-24 00:11:53 +02:00
Georgi Gerganov and GitHub
483bab2e3d
Avoid the transposed X branch in the Z = X * Y matrix multiplication ( #439 )
...
Should make results reproducible for different number of threads and batch sizes
2023-03-23 23:22:01 +02:00
Georgi Gerganov
4cc053b6d5
Remove oboslete command from Docker script
2023-03-23 22:39:44 +02:00
Georgi Gerganov
0ba5a3a9a5
Obsolete
2023-03-23 22:32:21 +02:00
Georgi Gerganov and GitHub
93208cfb92
Adjust repetition penalty ..
2023-03-23 10:46:58 +02:00
Georgi Gerganov and GitHub
03ace14cfd
Add link to recent podcast about whisper.cpp and llama.cpp
2023-03-23 09:48:51 +02:00
Georgi Gerganov
ae44e23ee3
When seed <= 0 - use the clock to generate one
2023-03-22 07:47:15 +02:00
Georgi Gerganov
928480ef5b
Init llama_context_params properly from CLI ( #370 )
2023-03-22 07:45:14 +02:00
Georgi Gerganov and GitHub
56817b1f88
Remove temporary notice and update hot topics
2023-03-22 07:34:02 +02:00
Georgi Gerganov and GitHub
f5a77a629b
Introduce C-style API ( #370 )
...
* Major refactoring - introduce C-style API
* Clean up
* Add <cassert>
* Add <iterator>
* Add <algorithm> ....
* Fix timing reporting and accumulation
* Measure eval time only for single-token calls
* Change llama_tokenize return meaning
2023-03-22 07:32:36 +02:00
Georgi Gerganov and GitHub
3366853e41
Add notice about pending change
2023-03-21 22:57:35 +02:00
Georgi Gerganov and GitHub
0f61352708
Update issue templates
2023-03-21 19:47:27 +02:00
Georgi Gerganov and GitHub
1daf4dd712
Minor style changes
2023-03-21 18:10:32 +02:00
Georgi Gerganov
dc6a845b85
Add chat.sh script
2023-03-21 18:09:46 +02:00
Georgi Gerganov
3bfa3b43b7
Fix convert script, warnings alpaca instructions, default params
2023-03-21 17:59:16 +02:00
Georgi Gerganov
8f644a0a85
Change default repeat_penalty to 1.0
...
I feel this penalty is not really helping.
Especially for the example from the README it makes results pretty bad
2023-03-21 17:32:14 +02:00
Georgi Gerganov and GitHub
eb34620aec
Add tokenizer test + revert to C++11 ( #355 )
...
* Add test-tokenizer-0 to do a few tokenizations - feel free to expand
* Added option to convert-pth-to-ggml.py script to dump just the vocabulary
* Added ./models/ggml-vocab.bin containing just LLaMA vocab data (used for tests)
* Added utility to load vocabulary file from previous point (temporary implementation)
* Avoid using std::string_view and drop back to C++11 (hope I didn't break something)
* Rename gpt_vocab -> llama_vocab
* All CMake binaries go into ./bin/ now
2023-03-21 17:29:41 +02:00
Georgi Gerganov
4545539d71
Rename script
2023-03-19 21:58:51 +02:00
Georgi Gerganov
edeba28366
Add temporary helper script for Alpaca chat
2023-03-19 21:57:48 +02:00
Georgi Gerganov and GitHub
160bfb217d
Update hot topics to mention Alpaca support
2023-03-19 19:51:55 +02:00
Georgi Gerganov
c494ed5b94
Fix off-by-one bug ( #115 )
2023-03-19 19:46:32 +02:00