Georgi Gerganov
1a941869cb
metal : disable graph concurrency optimization due to bug ( #2413 )
2023-07-27 11:00:54 +03:00
Georgi Gerganov and GitHub
5b2b2dc6ae
ggml : sync (unary ops refactor, static-correctness) ( #2370 )
...
* ggml : sync (unary ops, tests)
ggml-ci
* tests : remove unnecessary funcs
2023-07-24 14:46:21 +03:00
e76d630df1
llama : grouped-query attention + LLaMAv2 70B support ( #2276 )
...
* CUDA: GQA implementation
* llama : support for GQA and LLaMAv2 70B
ggml-ci
* py : fix hparams parsing (if-else blocks)
ggml-ci
* py : oh boy ..
ggml-ci
* help : fix gqa value for 70B
ggml-ci
---------
Co-authored-by: JohannesGaessler <johannesg@5d6.de >
2023-07-23 15:09:47 +03:00
Georgi Gerganov and GitHub
b47b8a9cfe
llama : optimize memory buffers ( #2325 )
2023-07-22 21:17:57 +03:00
Georgi Gerganov
dd6c67d3cb
ci : fix args
2023-07-22 12:00:56 +03:00
Georgi Gerganov and GitHub
5d500e8ccf
ci : add 7B CUDA tests ( #2319 )
...
* ci : add 7B CUDA tests
ggml-ci
* ci : add Q2_K to the tests
* ci : bump CUDA ppl chunks
ggml-ci
* ci : increase CUDA TG len + add --ignore-eos
* ci : reduce CUDA ppl cunks down to 4 to save time
2023-07-22 11:48:22 +03:00
Georgi Gerganov
0db14fef06
ggml : fix the rope fix ( 513f861953)
2023-07-21 15:16:55 +03:00
Georgi Gerganov
513f861953
ggml : fix rope args order + assert ( #2054 )
2023-07-21 14:51:34 +03:00
Georgi Gerganov
3973b25a64
gitignore : fix final newline
2023-07-21 14:42:41 +03:00
Georgi Gerganov
a814d04f81
make : fix indentation
2023-07-21 13:50:55 +03:00
Georgi Gerganov
4c013bb738
ci : fix MNT realpath usage ( #2250 )
2023-07-21 13:49:18 +03:00
Georgi Gerganov and GitHub
ae178ab46b
llama : make tensor_split ptr instead of array ( #2272 )
2023-07-21 13:10:51 +03:00
Georgi Gerganov
fff0e0eafe
llama : fix regression from #2000 - could not load no-mmap models
2023-07-20 13:47:26 +03:00
Georgi Gerganov and GitHub
d01bccde9f
ci : integrate with ggml-org/ci ( #2250 )
...
* ci : run ctest
ggml-ci
* ci : add open llama 3B-v2 tests
ggml-ci
* ci : disable wget progress output
ggml-ci
* ci : add open llama 3B-v2 tg tests for q4 and q5 quantizations
ggml-ci
* tests : try to fix tail free sampling test
ggml-ci
* ci : add K-quants
ggml-ci
* ci : add short perplexity tests
ggml-ci
* ci : add README.md
* ppl : add --chunks argument to limit max number of chunks
ggml-ci
* ci : update README
2023-07-18 14:24:43 +03:00
Georgi Gerganov
6cbf9dfb32
llama : shorten quantization descriptions
2023-07-18 11:50:49 +03:00
Georgi Gerganov
697966680b
ggml : sync (ggml_conv_2d, fix mul_mat bug, CUDA GLM rope)
2023-07-14 16:36:41 +03:00
Georgi Gerganov and GitHub
975221e954
ggml : broadcast mul_mat + conv batch support ( #2199 )
...
* ggml : broadcast mul_mat + conv batch support
* ggml : apply mul_mat broadcast fix by @jploski
2023-07-12 20:51:29 +03:00
Georgi Gerganov
4523d10d0c
ggml : add ggml_pool_1d and ggml_pool_2d
2023-07-12 20:32:15 +03:00
Georgi Gerganov
680e6f9177
cuda : add gelu support
2023-07-12 20:32:15 +03:00
Georgi Gerganov and GitHub
f7d278faf3
ggml : revert CUDA broadcast changes from #2183 ( #2191 )
2023-07-12 10:54:19 +03:00
Georgi Gerganov and GitHub
20d7740a9b
ggml : sync (abort callback, mul / add broadcast, fix alibi) ( #2183 )
2023-07-11 22:53:34 +03:00
Georgi Gerganov and GitHub
a7e20edf22
ci : switch threads to 1 ( #2138 )
2023-07-07 21:23:57 +03:00
Georgi Gerganov
7242140283
ggml : remove sched_yield() call in ggml_graph_compute_thread() ( #2134 )
2023-07-07 18:37:10 +03:00
Georgi Gerganov
dfd9fce6d6
ggml : fix restrict usage
2023-07-06 19:41:31 +03:00
Georgi Gerganov
ec326d350c
ggml : fix bug introduced in #1237
2023-07-05 20:44:11 +03:00
Georgi Gerganov
1b6efeab82
tests : fix test-grad0
2023-07-05 20:20:25 +03:00
Georgi Gerganov and GitHub
b472f3fca5
readme : add link web chat PR
2023-07-04 22:25:22 +03:00
Georgi Gerganov and GitHub
ed9a54e512
ggml : sync latest (new ops, macros, refactoring) ( #2106 )
...
- add ggml_argmax()
- add ggml_tanh()
- add ggml_elu()
- refactor ggml_conv_1d() and variants
- refactor ggml_conv_2d() and variants
- add helper macros to reduce code duplication in ggml.c
2023-07-04 21:54:11 +03:00
Georgi Gerganov
46088f7231
ggml : fix build with OpenBLAS ( close #2066 )
2023-07-02 09:46:46 +03:00
Georgi Gerganov
463f2f4c4f
llama : fix return value of llama_load_session_file_internal ( #2022 )
2023-07-01 19:05:09 +03:00
Georgi Gerganov
79f634a19d
embd-input : fix returning ptr to temporary
2023-07-01 18:46:00 +03:00
Georgi Gerganov
04606a1599
train : fix compile warning
2023-07-01 18:45:44 +03:00
Georgi Gerganov
181e8d9755
llama : fix rope usage after ChatGLM change
2023-06-27 00:37:33 +03:00
Georgi Gerganov
d9779021bd
ggml : add support for ChatGLM RoPE
2023-06-27 00:06:51 +03:00
Georgi Gerganov
c824d2e368
ggml : avoid conv 2d kernel round up
2023-06-26 21:03:59 +03:00
Georgi Gerganov
9225baef71
k-quants : fix indentation
2023-06-26 20:10:52 +03:00
Georgi Gerganov and GitHub
412c60e473
readme : add link to new k-quants for visibility
2023-06-26 19:45:09 +03:00
Georgi Gerganov and GitHub
447ccbe8c3
readme : add new roadmap + manifesto
2023-06-25 16:08:12 +03:00
Georgi Gerganov
bd34cdde38
ggml : sync latest ggml (custom operators)
2023-06-25 14:25:08 +03:00
Georgi Gerganov and GitHub
66a2555ba6
readme : add Azure CI discussion link
2023-06-25 09:07:03 +03:00
Georgi Gerganov
65bdd52a86
tests : sync test-grad0 from ggml
2023-06-24 19:40:18 +03:00
Georgi Gerganov
11da1a85cd
readme : fix whitespaces
2023-06-24 13:38:18 +03:00
Georgi Gerganov and GitHub
049aa16b8c
readme : add link to p1
2023-06-20 19:05:54 +03:00
Georgi Gerganov
18b35625c3
ggml : fix bug in LBFGS optimizer (found by ggml tests)
2023-06-19 20:43:30 +03:00
Georgi Gerganov
23fc5c219a
cmake : fix trailing whitespaces
2023-06-19 18:18:34 +03:00
Georgi Gerganov and GitHub
b97ca431db
ggml : sync latest ggml repo ( #1924 )
...
* ggml : sync latest ggml repo
* ggml : remove unused comments
* ggml : asserts
2023-06-19 18:12:33 +03:00
Georgi Gerganov and GitHub
ce2c7d72e2
metal : handle buffers larger than device's maxBufferLength ( #1826 )
...
* metal : handle buffers larger than device's maxBufferLength
* metal : print more verbose device info + handle errors
* metal : fix prints for overlapping views
* metal : minimize view overlap to try to utilize device memory better
2023-06-18 09:09:47 +03:00
Georgi Gerganov
b2416493ab
make : do not print help for simple example
2023-06-17 20:55:03 +03:00
Georgi Gerganov
4f9c43e3bd
minor : warning fixes
2023-06-17 20:24:11 +03:00
Georgi Gerganov
051e1b0e6a
llama : fix kv_cache n init ( close #1903 )
2023-06-17 19:31:20 +03:00
Georgi Gerganov
bed9275617
cmake : remove whitespaces
2023-06-15 21:56:50 +03:00
Georgi Gerganov and GitHub
4bfcc855ab
metal : parallel command buffer encoding ( #1860 )
...
* metal : parallel command buffer encoding
* metal : determine number of command buffers based on gf->n_threads
2023-06-15 20:29:48 +03:00
Georgi Gerganov and GitHub
2347e45e7b
llama : do a warm-up eval at start for better timings ( #1824 )
2023-06-13 20:20:07 +03:00
Georgi Gerganov
4de0334f5c
cmake : fix Metal build ( close #1791 )
2023-06-10 22:56:53 +03:00
Georgi Gerganov
17c10acfb4
ggml : force no_alloc == false when creating opt tensors ( close #1699 )
...
This is needed to make operators like ggml_view() be able to store their
parameters in the ggml context's memory and not get discarded when
no_alloc is true
2023-06-10 12:08:15 +03:00
Georgi Gerganov
b33dee282f
metal : fix build "tanhf" -> "tanh"
2023-06-09 11:11:04 +03:00
Georgi Gerganov
0bf7cf1b29
Revert "ggml : load data into int8x16x4_t using vld4q_s8 on arm64 ( #1738 )"
...
This reverts commit 8432d4d9f7 .
2023-06-08 20:48:14 +03:00
Georgi Gerganov
53aba3f393
clang-tidy : restore dot file from accidental deletion
2023-06-08 10:09:08 +03:00
Georgi Gerganov and GitHub
5c64a0952e
k-quants : allow to optionally disable at compile time ( #1734 )
...
* k-quants : put behind optional compile flag LLAMA_K_QUANTS
* build : enable k-quants by default
2023-06-07 10:59:52 +03:00
Georgi Gerganov and GitHub
4dc62c545d
readme : add June roadmap
2023-06-07 07:15:08 +03:00
Georgi Gerganov
2d7bf110ed
llama : fix vram_scratch var
2023-06-06 22:54:39 +03:00
Georgi Gerganov
2a4e41a086
llama : fix compile warnings
2023-06-06 22:41:53 +03:00
Georgi Gerganov
44f906e853
metal : add f16 support
2023-06-06 20:21:56 +03:00
Georgi Gerganov
2d43387daf
ggml : fix builds, add ggml-quants-k.o ( close #1712 , close #1710 )
2023-06-06 10:18:03 +03:00
Georgi Gerganov
7ad7750c5c
gitignore : add .clang-tidy
2023-06-06 09:55:25 +03:00
Georgi Gerganov
7a74dee6b4
llama : temporary disable Q6_K output quantization ( #1711 )
2023-06-06 09:39:38 +03:00
Georgi Gerganov and GitHub
e7fe66e670
ci : disable auto tidy ( #1705 )
2023-06-05 23:05:05 +03:00
Georgi Gerganov
d1f563a743
llama : fix Metal KV cache sync ( close #1695 )
2023-06-05 10:19:03 +03:00
Georgi Gerganov and GitHub
827f5eda91
readme : update hot topics
2023-06-04 23:38:19 +03:00
Georgi Gerganov and GitHub
ecb217db4f
llama : Metal inference ( #1642 )
...
* mtl : export the LLaMA computation graph
* ci : disable temporary
* mtl : adapt the MNIST example as starter
* mtl : no need for mtl-export tool, add cli arg for main instead
* mtl : export just a small part of the graph for now to make it easier
* mtl : move MSL code into separate file for easy editing
* mtl : initial get_rows_q4_0 kernel
* mtl : confirmed get_rows_q4_0 is working correctly
* mtl : add rms_norm kernel + confirm working
* mtl : add mul kernel + confirm working
* mtl : initial mul_mat Q4 kernel (wrong results)
* mtl : mul_mat fixes (still wrong)
* mtl : another mul_mat Q4 (still does not work)
* mtl : working mul_mat q4
* ggml : fix handling of "view" ops in ggml_graph_import()
* mtl : add rope kernel
* mtl : add reshape and transpose handling
* ggml : store offset as opt arg for ggml_view_xd() operators
* mtl : add cpy kernel + handle view ops
* mtl : confirm f16 x f32 attention mul mat
* mtl : add scale kernel
* mtl : add diag_mask_inf kernel
* mtl : fix soft_max kernel
* ggml : update ggml_nbytes() to handle non-contiguous tensors
* mtl : verify V tensor contents
* mtl : add f32 -> f32 cpy kernel
* mtl : add silu kernel
* mtl : add non-broadcast mul kernel
* mtl : full GPU inference of the computation graph
* mtl : optimize rms_norm and soft_max kernels
* mtl : add f16 mat x f32 vec multiplication kernel
* mtl : fix bug in f16 x f32 mul mat + speed-up computation
* mtl : faster mul_mat_q4_0_f32 kernel
* mtl : fix kernel signature + roll inner loop
* mtl : more threads for rms_norm + better timing
* mtl : remove printfs from inner loop
* mtl : simplify implementation
* mtl : add save/load vocab to ggml file
* mtl : plug Metal inference into llama.cpp (very quick-n-dirty)
* mtl : make it work with main example
Lots of hacks but at least now it generates text
* mtl : preparing for merge
* mtl : clean-up ggml mtl interface + suport scratch / inplace
* mtl : remove temp / debug code
* metal : final refactoring and simplification
* Revert "ci : disable temporary"
This reverts commit 98c267fc77fe811082f672538fc91bcfc9072d63.
* metal : add comments
* metal : clean-up stuff, fix typos
* readme : add Metal instructions
* readme : add example for main
2023-06-04 23:34:30 +03:00
Georgi Gerganov
7552ac5863
ggml : sync cgraph import / export API
2023-05-29 19:31:44 +03:00
Georgi Gerganov
5d1830b99d
ggml : fix bug in ggml_alibi
2023-05-29 19:30:49 +03:00
Georgi Gerganov
93618031c7
ggml : add ggml_tensor_overhead()
2023-05-27 16:19:56 +03:00
Georgi Gerganov
bdbda1b17a
ggml : sync ggml core (minor additions, e.g. ggml_get_tensor_by_name())
2023-05-27 12:23:16 +03:00
Georgi Gerganov
265db9834e
ggml : output 3d sizes in ggml_graph_dump_dot()
2023-05-21 11:56:23 +03:00
Georgi Gerganov
fab49c685e
ggml : update WASM SIMD
2023-05-20 20:00:41 +03:00
Georgi Gerganov and GitHub
3de84b2606
ggml : add ggml_clamp() ( #1539 )
...
* ggml : add ggml_clamp()
* ggml : indentation
2023-05-20 15:34:45 +03:00
Georgi Gerganov
ea600071cb
Revert "feature : add blis and other BLAS implementation support ( #1502 )"
...
This reverts commit 07e9ace0f9 .
2023-05-20 12:03:48 +03:00
Georgi Gerganov
ec2e10c444
llama : add llama_init_backend() API ( close #1527 )
2023-05-20 11:06:37 +03:00
Georgi Gerganov
8a203f9fa1
llama : fix compile warnings in llama_set_state_data()
2023-05-20 10:14:43 +03:00
Georgi Gerganov
4fd3e29297
ggml : fix scalar implementation of Q4_1 dot
2023-05-20 10:13:19 +03:00
Georgi Gerganov and GitHub
2d5db48371
ggml : use F16 instead of F32 in Q4_0, Q4_1, Q8_0 ( #1508 )
...
* ggml : use F16 instead of F32 in Q4_0, Q4_1 and Q8_0
* llama : bump LLAMA_FILE_VERSION to 3
* cuda : update Q4 and Q8 dequantize kernels
* ggml : fix AVX dot products
* readme : update performance table + hot topics
2023-05-19 22:17:18 +03:00
Georgi Gerganov
6986c7835a
tests : add missing header
2023-05-19 21:17:28 +03:00
Georgi Gerganov
4b7e245adf
minor : fix compile warnings
2023-05-19 20:14:51 +03:00
Georgi Gerganov and GitHub
13c351ad72
ggml : various fixes ( #1450 )
...
- `ggml_rope()`
- `ggml_diag_mask_inf()` multi-threaded
- compatibility with scratch buffers
2023-05-14 18:22:50 +03:00
Georgi Gerganov
601a033475
ggml : add GGML_QNT_VERSION to track quantization format changes
...
https://github.com/ggerganov/ggml/issues/150#issuecomment-1546625668
2023-05-14 10:20:19 +03:00
Georgi Gerganov
08737ef720
cuda : fix convert function ( #1412 )
2023-05-13 17:40:58 +03:00
Georgi Gerganov
bda4d7c215
make : fix PERF build with cuBLAS
2023-05-13 17:25:09 +03:00
Georgi Gerganov
5a5aeb1e91
llama : fix unused warning
2023-05-13 16:55:14 +03:00
Georgi Gerganov and GitHub
66841fdb0e
ggml : multi-thread mul and diag_mask ops ( #1428 )
2023-05-13 16:48:03 +03:00
Georgi Gerganov
f048af0230
ggml : sync alibi fix from ggml repo
2023-05-13 11:54:33 +03:00
Georgi Gerganov
0cd22e190a
llama : fix various warnings
2023-05-13 11:23:15 +03:00
Georgi Gerganov and GitHub
cdd5350892
readme : update Q4_0 perplexities
...
I think these were affected by the removal of the `round` during quantization
2023-05-13 09:12:44 +03:00
Georgi Gerganov
738ace394a
llama : free ggml context in set / copy state data ( close #1425 )
2023-05-13 09:08:52 +03:00
Georgi Gerganov
fb62f92433
llama : fix --mtest option ( close #1414 )
2023-05-12 21:44:20 +03:00
b9fd7eee57
ggml : remove bit shuffling ( #1405 )
...
* ggml : remove Q4_0 bit shufling (ARM NEON)
* ggml : remove Q4_1 bit shuffling (ARM NEON + reference)
* ggml : nibbles_from_floats() + bytes_from_nibbles() (ARM NEON)
* ggml : remove Q4_2 bit shuffling (WIP, BROKEN)
* ggml : remove Q5_0 bit shuffling (ARM NEON)
* ggml : 2x faster scalar implementations
* ggml : remove Q5_1 bit shuffling (ARM NEON + scalar)
* ggml : simplify scalar dot
* ggml : remove WASM SIMD bit shuffling + remove vzip for ARM 32-bit
* ggml : fix Q4_1 quantization
* ggml : update cuBLAS + normalize variable names
* ggml : remove Q4_2 mode
* ggml : minor formatting
* ggml : fix Q5_0 quantization
* scripts : add script for measuring the time per token
* AVX implementations (#1370 )
* ggml : uniform 5th bit extraction
* llama : produce error upon loading old model files
* llama : fix model magic/version write
* ggml : speed-up Q5_0 + Q5_1 at 4 threads
* ggml : preserve old Q4 and Q5 formats
* ggml : simplify Q8_1 - no need for low / high sums anymore
* ggml : fix Q8_0 and Q8_1 rounding
* Revert "AVX implementations (#1370 )"
This reverts commit 948d124837f9d287d8490f41338e0e4cceb0814f.
* ggml : fix AVX2 implementation
* sha : update hashes for 7B and 13B
* readme : update timings + remove warning banner
* llama : update v2 PR number to 1405
* ggml : fix WASM comments
* ggml : back to original bit order
* readme : add note that Q4 and Q5 have been changed
* llama : fix return for unknown version
---------
Co-authored-by: Stephan Walter <stephan@walter.name >
2023-05-12 00:23:08 +03:00
Georgi Gerganov and GitHub
56551bc11f
readme : add notice about upcoming breaking change
2023-05-08 22:52:18 +03:00
Georgi Gerganov and GitHub
f9a6364912
llama : require first token to be BOS ( #1303 )
...
* llama : require first token to be BOS
* scripts : add ppl-run-all.sh
* perplexity : add BOS for each chunk
* readme : update perplexity values after BOS fix
* perplexity : add clarifying comments
2023-05-08 17:41:54 +03:00
Georgi Gerganov
799fdc1b5d
ggml : vectorize Q8_0 quantization
...
https://github.com/ggerganov/ggml/pull/127#issuecomment-1533648531
2023-05-03 23:24:20 +03:00
Georgi Gerganov and GitHub
bca9ad938a
minor : fix whitespaces ( #1302 )
2023-05-03 20:09:42 +03:00