Georgi Gerganov and GitHub
5921b8f089
llama : cache llama_token_to_piece ( #7587 )
...
* llama : cache llama_token_to_piece
ggml-ci
* llama : use vectors and avoid has_cache
ggml-ci
* llama : throw on unknown tokenizer types
ggml-ci
* llama : print a log of the total cache size
2024-05-31 02:01:41 +10:00
Georgi Gerganov
55d62262a9
metal : remove invalid asserts ( #7617 )
2024-05-29 22:21:20 +03:00
Georgi Gerganov
975ec63ff2
metal : add missing asserts ( #7617 )
2024-05-29 20:45:25 +03:00
Georgi Gerganov and GitHub
fb76ec31a9
ggml : fix YARN + add tests + add asserts ( #7617 )
...
* tests : add rope tests
ggml-ci
* ggml : fixes (hopefully)
ggml-ci
* tests : add non-cont tests
ggml-ci
* cuda : add asserts for rope/norm + fix DS2
ggml-ci
* ggml : assert contiguousness
* tests : reduce RoPE tests
ggml-ci
2024-05-29 20:17:31 +03:00
Georgi Gerganov and GitHub
cce3dcffc5
cuda : non-cont concat support ( #7610 )
...
* tests : add non-cont concat tests
* cuda : non-cont concat support
ggml-ci
2024-05-29 15:38:26 +03:00
Georgi Gerganov
00281b7be3
scripts : remove mpi remnants
2024-05-29 14:31:18 +03:00
Georgi Gerganov
2ab977282b
sync : ggml
2024-05-29 14:29:52 +03:00
Georgi Gerganov
72de268bec
ggml : restore ggml_rope_xpos_inplace (ggml/0)
...
ggml-ci
2024-05-29 14:29:33 +03:00
Georgi Gerganov
6bd12ce409
sycl : fix assert ( #7563 )
2024-05-28 22:22:50 +03:00
Georgi Gerganov
edc29433fa
tests : fix test-tokenizer-0.sh
2024-05-28 15:04:09 +03:00
Georgi Gerganov and GitHub
8b99e2aa66
llama : handle unknown utf8 bytes ( #7588 )
2024-05-28 13:55:35 +03:00
Georgi Gerganov and GitHub
0548a4187f
ggml : generalize GGML_OP_CONCAT ( #7563 )
...
* ggml : generalize GGML_OP_CONCAT (WIP)
ggml-ci
* tests : add dim != 2 tests
* metal : generalize concat kernel
* tests : naming
* cuda : generalize concat kernel
ggml-ci
* sycl : add warning and assert
* ggml : fix op params handling
* metal : bugfix kernel
ggml-ci
* ggml : reimplement CPU and Metal
* cuda : add asserts
ggml-ci
* ggml : fix ptrs
ggml-ci
2024-05-28 11:04:19 +03:00
Georgi Gerganov and GitHub
1d8fca72ae
metal : add GGML_OP_REPEAT kernels ( #7557 )
...
ggml-ci
2024-05-27 12:10:19 +03:00
Georgi Gerganov and GitHub
62bfef5194
metal : disable FA kernel for HS=256 ( #7556 )
...
ggml-ci
2024-05-27 10:38:39 +03:00
Georgi Gerganov and GitHub
eaf6e03174
llama : add comments about experimental flags ( #7544 )
2024-05-27 09:24:13 +03:00
dff451cfa1
flake.lock: Update ( #7540 )
...
Flake lock file updates:
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/4a6b83b05df1a8bd7d99095ec4b4d271f2956b64?narHash=sha256-%2BNpbZRCRisUHKQJZF3CT%2Bxn14ZZQO%2BKjxIIanH3Pvn4%3D' (2024-05-17)
→ 'github:NixOS/nixpkgs/bfb7a882678e518398ce9a31a881538679f6f092?narHash=sha256-4zSIhSRRIoEBwjbPm3YiGtbd8HDWzFxJjw5DYSDy1n8%3D' (2024-05-24)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-05-26 08:54:56 -07:00
Georgi Gerganov
9588f196b1
train : change default FA argument ( #7528 )
2024-05-25 15:22:35 +03:00
d041d2ceaa
flake.lock: Update ( #7232 )
...
Flake lock file updates:
• Updated input 'flake-parts':
'github:hercules-ci/flake-parts/e5d10a24b66c3ea8f150e47dfdb0416ab7c3390e?narHash=sha256-yzcRNDoyVP7%2BSCNX0wmuDju1NUCt8Dz9%2BlyUXEI0dbI%3D' (2024-05-02)
→ 'github:hercules-ci/flake-parts/8dc45382d5206bd292f9c2768b8058a8fd8311d9?narHash=sha256-/GJvTdTpuDjNn84j82cU6bXztE0MSkdnTWClUCRub78%3D' (2024-05-16)
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/63c3a29ca82437c87573e4c6919b09a24ea61b0f?narHash=sha256-4cPymbty65RvF1DWQfc%2BBc8B233A1BWxJnNULJKQ1EY%3D' (2024-05-02)
→ 'github:NixOS/nixpkgs/4a6b83b05df1a8bd7d99095ec4b4d271f2956b64?narHash=sha256-%2BNpbZRCRisUHKQJZF3CT%2Bxn14ZZQO%2BKjxIIanH3Pvn4%3D' (2024-05-17)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-05-24 08:59:06 -07:00
Georgi Gerganov
74f33adf5f
readme : remove trailing space ( #7469 )
2024-05-23 17:43:18 +03:00
Georgi Gerganov
1debe72737
ggml : silence UB sanitizer error during iq2_xxs quantization ( #0 )
2024-05-23 17:25:38 +03:00
Georgi Gerganov and GitHub
55ac3b7aea
ci : use Pythia models instead of OpenLlama ( #7470 )
...
* ci : start using Pythia models over OpenLlama
ggml-ci
* ci : disable q2_k ppl tests
* ci : use convert-hf-to-gguf.py
* ci : update gg_get_model
* ci : fix convert outfile name
ggml-ci
* llama : gptneox arch use F32 attn prec
ggml-ci
2024-05-23 15:28:14 +03:00
Georgi Gerganov
a61a94e543
llama : rename n_ctx -> cache.size, less confusing ( #0 )
2024-05-23 12:38:18 +03:00
Georgi Gerganov and GitHub
d48c88cbd5
ggml : remove ggml_flash_attn and ggml_flash_ff ( #7463 )
...
ggml-ci
2024-05-23 10:00:44 +03:00
Georgi Gerganov and GitHub
e84b71c2c6
ggml : drop support for QK_K=64 ( #7473 )
...
* ggml : drop support for QK_K=64
ggml-ci
* opencl : restore QK_K=256 define
2024-05-23 10:00:21 +03:00
Georgi Gerganov
fbf777d2b9
main : minor ( #7462 )
2024-05-23 09:43:49 +03:00
Georgi Gerganov and GitHub
197ff91462
build : remove zig ( #7471 )
2024-05-22 20:05:38 +03:00
Georgi Gerganov and GitHub
6ff13987ad
common : normalize naming style ( #7462 )
...
* common : normalize naming style
ggml-ci
* common : match declaration / definition order
* zig : try to fix build
2024-05-22 20:04:20 +03:00
Georgi Gerganov
9b3d833189
cuda : fix compile warning ( #7454 )
2024-05-22 12:36:37 +03:00
Georgi Gerganov and GitHub
3e5faa8503
cuda : fix rope + add tests ( #7452 )
...
* cuda : fix rope pos data
ggml-ci
* ggml : drop mode & 1 == 1 support for ggml_rope
ggml-ci
* ggml : support freq_factors for f16 rope (CPU)
ggml-ci
* tests : add rope tests using frequency factors
ggml-ci
2024-05-22 11:01:35 +03:00
Georgi Gerganov and GitHub
6369bf0433
metal : handle F16 inf values, fix FA partial offload ( #7434 )
...
ggml-ci
2024-05-21 23:03:42 +03:00
Georgi Gerganov
c3f8d58356
tests : test-tokenizer-0.sh print more info ( #7402 )
2024-05-21 19:53:48 +03:00
Georgi Gerganov and GitHub
fabf30b4c4
llama : remove Persimmon ( #7408 )
...
* llama : remove Persimmon
* requirements : remove
2024-05-21 02:35:28 +10:00
Georgi Gerganov and GitHub
3bc10cb485
server : fix temperature + disable some tests ( #7409 )
...
* server : fix temperature
* server : disable tests relying on parallel determinism
* ci : change server Debug -> RelWithDebInfo
2024-05-20 22:10:03 +10:00
Georgi Gerganov and GitHub
1cc0155d04
server : tuning tests ( #7388 )
...
* server : don't pass temperature as string
* server : increase timeout
* tests : fix the fix 0.8f -> 0.8
ggml-ci
* tests : set explicit temperature
2024-05-20 10:16:41 +03:00
Georgi Gerganov and GitHub
e932094d58
server : return error on too large embedding input ( #7389 )
2024-05-20 08:56:05 +03:00
Georgi Gerganov
2789baf480
tests : fix --keep_split -> --keep-split ( #7374 )
2024-05-20 08:55:09 +03:00
Georgi Gerganov
854d365aba
cmake : update android comments ( #7341 )
2024-05-19 11:01:01 +03:00
Georgi Gerganov and GitHub
059031b8c4
ci : re-enable sanitizer runs ( #7358 )
...
* Revert "ci : temporary disable sanitizer builds (#6128 )"
This reverts commit 4f6d1337ca .
* ci : trigger
2024-05-18 18:55:54 +03:00
Georgi Gerganov and GitHub
511182eabb
android : use "ci-android" branch for CI ( #7341 )
...
* android : use "ci-android" branch for CI
* ggml : disable SIMD exp and silu for 32-bit ARM
ggml-ci
* android : do not fetch, use add_subdirectory instead
* cmake : provide binary dir
2024-05-18 20:40:39 +10:00
Georgi Gerganov and GitHub
b49a13dd2f
convert : fix set_vocab_sentencepiece ( #6866 )
...
* convert : fix set_vocab_sentencepiece
* Update convert-hf-to-gguf.py
2024-05-18 08:46:20 +03:00
Georgi Gerganov
29499bb593
sync : ggml
2024-05-15 13:23:41 +03:00
Georgi Gerganov
9f773486ab
script : sync ggml-rpc
2024-05-14 19:14:38 +03:00
Georgi Gerganov and GitHub
e8a7fd4fb0
metal : support FA without mask + add asserts ( #7278 )
...
* ggml : fa without mask + add asserts
ggml-ci
* metal : support non-contiguous KV
ggml-ci
2024-05-14 19:09:30 +03:00
Georgi Gerganov
a5e3fde857
sync : ggml
...
ggml-ci
2024-05-14 19:08:09 +03:00
Georgi Gerganov
f308ea7059
metal : tune soft_max number of threads (whisper/0)
2024-05-14 19:08:09 +03:00
Georgi Gerganov
c3c88f296a
ggml : try fix ppc64 (whisper/0)
2024-05-14 19:08:09 +03:00
Georgi Gerganov and GitHub
614d3b914e
llama : less KV padding when FA is off ( #7257 )
...
ggml-ci
2024-05-13 17:15:15 +03:00
Georgi Gerganov
6f1b63606f
cmake : fix version cmp ( #7227 )
2024-05-12 18:30:23 +03:00
Georgi Gerganov
7bd4ffb780
metal : fix warnings (skipme) ( #0 )
2024-05-11 21:38:13 +03:00
Georgi Gerganov
1622ac023f
sync : ggml
2024-05-11 21:35:05 +03:00
Georgi Gerganov
6aeff24f8b
metal : fix indent (ggml/0)
2024-05-11 21:34:21 +03:00
Georgi Gerganov
325756d28d
ggml : resolve merge (ggml/0)
...
ggml-ci
2024-05-11 21:33:08 +03:00
Georgi Gerganov
fae9d234b6
sync : ggml
...
ggml-ci
2024-05-11 15:38:34 +03:00
Georgi Gerganov and GitHub
9cb317f77e
ggml : full ALiBi support ( #7192 )
...
* ggml : full ALiBi support
* ggml : update ggml_soft_max_ext() CUDA, SYCL
* ggml : ggml_flash_attn_ext() support ALiBi (CPU)
* ggml : ggml_flash_attn_ext() support ALiBi (Metal)
* ggml : fix warning
* ggml : ggml_flash_attn_ext() support ALiBi (CUDA)
ggml-ci
* ggml : fix assert message
* vulkan : add dev notes
* ggml : require mask when using ALiBi
ggml-ci
* convert : fix convert for refact models
2024-05-11 10:32:41 +03:00
Georgi Gerganov and GitHub
18e437665c
metal : fix flash attention kernel requirements ( #7169 )
...
* metal : fix flash attention kernel requirements
ggml-ci
* metal : fix ggml_metal_supports_op
ggml-ci
2024-05-10 18:20:10 +03:00
Georgi Gerganov
8c660242d7
convert : print "ignore_merges" field
2024-05-10 17:53:04 +03:00
Georgi Gerganov and GitHub
d46dbc76f8
readme : add scheduled server workflow status badge
2024-05-09 16:40:42 +03:00
Georgi Gerganov
9da243b36a
Revert "llava : add support for moondream vision language model ( #6899 )"
...
This reverts commit 46e12c4692 .
2024-05-08 22:14:39 +03:00
Georgi Gerganov
7e0b6a7b3b
py : also print the normalizers
2024-05-08 12:47:07 +03:00
Georgi Gerganov
c0e6fbf8c3
metal : fix unused warning
2024-05-08 09:14:50 +03:00
Georgi Gerganov and GitHub
53d6c52e22
readme : update hot topics
2024-05-07 21:43:13 +03:00
Georgi Gerganov and GitHub
947d3ad27d
ci : add GG_BUILD_EXTRA_TESTS_0 env ( #7098 )
...
* ci : add GG_BUILD_EXTRA_TESTS_0 env
ggml-ci
* Update run.sh
ggml-ci
2024-05-07 11:08:49 +03:00
b3a995b416
flake.lock: Update ( #7079 )
...
Flake lock file updates:
• Updated input 'flake-parts':
'github:hercules-ci/flake-parts/9126214d0a59633752a136528f5f3b9aa8565b7d?narHash=sha256-sB4SWl2lX95bExY2gMFG5HIzvva5AVMJd4Igm%2BGpZNw%3D' (2024-04-01)
→ 'github:hercules-ci/flake-parts/e5d10a24b66c3ea8f150e47dfdb0416ab7c3390e?narHash=sha256-yzcRNDoyVP7%2BSCNX0wmuDju1NUCt8Dz9%2BlyUXEI0dbI%3D' (2024-05-02)
• Updated input 'flake-parts/nixpkgs-lib':
'github:NixOS/nixpkgs/d8fe5e6c92d0d190646fb9f1056741a229980089?dir=lib&narHash=sha256-iMUFArF0WCatKK6RzfUJknjem0H9m4KgorO/p3Dopkk%3D' (2024-03-29)
→ 'https://github.com/NixOS/nixpkgs/archive/50eb7ecf4cd0a5756d7275c8ba36790e5bd53e33.tar.gz?narHash=sha256-QBx10%2Bk6JWz6u7VsohfSw8g8hjdBZEf8CFzXH1/1Z94%3D ' (2024-05-02)
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/7bb2ccd8cdc44c91edba16c48d2c8f331fb3d856?narHash=sha256-Drmja/f5MRHZCskS6mvzFqxEaZMeciScCTFxWVLqWEY%3D' (2024-04-25)
→ 'github:NixOS/nixpkgs/63c3a29ca82437c87573e4c6919b09a24ea61b0f?narHash=sha256-4cPymbty65RvF1DWQfc%2BBc8B233A1BWxJnNULJKQ1EY%3D' (2024-05-02)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-05-06 08:36:06 -07:00
Georgi Gerganov
bcdee0daa7
minor : fix trailing whitespace
2024-05-06 09:31:30 +03:00
Georgi Gerganov and GitHub
92139b90af
tests : add test-tokenizer-0.sh + fix some tokenizers ( #7036 )
...
* tests : add test-tokenizer-0.sh
* unicode : add all unicode number ranges
* starcoder : fix pre-tokenizer
* tests : add test that fails with DeepSeek tokenizers
* falcon : fix regex
* unicode : regenerate unicode tables
* refact : add tokenizer model
* lint : fix
* tests : disable failing tests
ggml-ci
* refact : add tests files
ggml-ci
* convert : print -> logging
ggml-ci
* lint : fix
* unicode : digit -> number
* phi-3 : update
2024-05-04 08:32:32 +03:00
Georgi Gerganov and GitHub
77e15bec62
metal : remove deprecated error code ( #7008 )
2024-04-30 15:52:21 +03:00
9c67c2773d
ggml : add Flash Attention ( #5021 )
...
* ggml : add ggml_flash_attn_ext API
* ggml : fix GQA support in ggml_flash_attn_ext
* ggml : online attention (CPU)
* metal : initial implementation
* metal : f16 precision
* metal : reduce branches
* metal : specialize for head size
* wip : 8 rows per simd group
* wip : 4 rows per simd group
* wip : template for rows per warp
* metal : parallelize across KV size
* metal : parallel reduce across heads
* metal : efficient flash_attn_f16 implementation
* metal : avoid redundant loads of the attention
* metal : scale and mask in matrix form
* metal : fix comment
* llama : avoid ggml_cast, use F32 query
* metal : add parallel reduce version (disabled)
* metal : move output into local memory + optimize
- the result from each simdgroup now stays in the registers
- significantly reduced SRAM usage
- more efficient skipping of -INF blocks
- avoid simdgroup barrier in hot loop
- add comments
* metal : add tests, fix scaling, support C > 32
* metal : improve precision
* ggml : fix f16 mad
* metal : minor
* metal : support Q > 8
* tests : add ATTN tests
* metal : disable buffer allocation logs
* tests : more
* metal : faster inner loop for C == 32
* metal : fix array initialization
* tests : ifdef
* ggml : switch to padded F16 mask for ggml_soft_max, ggml_flash_attn_ext
* ggml : fix ggml_soft_max mask requirement
* cuda : fix soft_max to use correct mask size
* cuda : add flash_attn kernel (wip)
* metal : optimize softmax for C > 32
* metal : optimize softmax
* tests : minor fix
* cuda : avoid zeroing fragments
* tests : update dims
* cuda : fix __hisinf() result check
* cuda : avoid warp_reduce for smax
* cuda : use int instead of int64_t
Noticeably improves performance (thanks to Johannes)
* cuda : make loops use the same loop values
Thanks Johannes again for the tip
* cuda : unroll some of the loops
* cuda : avoid __hisinf branches
* cuda : use half2 in softmax
* cuda : switch to 1 warp for bs > 16
* cuda : speed-up reduce part of the kernel
* cuda : unroll Q*K^T loop
* cuda : fix -INF block check
* cuda : simplify softmax
* cuda : fix matrix names
* cuda : minor
* llama : adapt to F16 KQ_pos
* llama : adapt new models to F16 KQ_mask
* ggml : fix F16 store (ARM NEON)
* llama : fix type of KQ_mask and KQ_pos
* ggml : fix CPU soft_max
* tests : add hs=256
* cuda : fix build
* metal : improve perf via smaller int registers
* cuda : adapt soft_max to F16 mask and pos
* CUDA: faster FlashAttention, kernel for bs == 1
* 16 cols for Phi-2
* no vec for hs, no hs==256 ncols==32 for Volta
* adjust kernel selection logic
* 4 warps, 256 stride for all D
* no ncols == 64
* Multiple parallel blocks for batch size 1
* fix compile warnings
* fix excessive KQ_b loads
* fix cmake build
* fix KV cache padding, NaN from INFINITY (#6438 )
* llama : flash_attn cparam + fix defrag
* server: support flash_attn param
* server: bench: enable flash_attn param
* CUDA: refactor host code, dyn. par. blocks
* fix flash_attn_vec_f16 race condition
* flush softmax exp below threshold to 0
* store temp KQ in registers
* Calculate KQ as FP32 if KQV has GGML_PREC_F32
* Add __hgt2_mask implementation for CUDA 11
* fix KQ FP32 precision fpr parallel_blocks > 1
* llama-bench : add -fa,--flash-attn arg
* metal : add BS=1 kernel for flash attention (#6508 )
* metal : add BS=1 kernel for flash attention (wip)
* metal : support more than 1 warps
* metal : opts
* metal : opt
* metal : switch to parallel reduce
* metal : reduce registers
* metal : simplify
* metal : initial FA vec kernel
* metal : use F32 attention accumulators
* batched-bench : add fattn arg
* llama : simplify llama_build_kv_store
ggml-ci
* llama : adapt build_olmo to changes
* ggml : fix arm fp16 store on windows
* metal : clean-up
* metal : clean-up kernel code
* metal : minor
* tests : remove benchmarks
ggml-ci
* ggml : fix avx512 const correctness
ggml-ci
* ggml : fix soft_max with bias on CPU
ggml-ci
* common : print --flash-attn in help
* ggml : fix num dimensions in ggml_flash_attn_ext
* llama : force disable flash attention for incompatible models
* ggml : ggml_soft_max support F16/F32 mask/pos
ggml-ci
* cuda : uint -> uint32_t
* cuda : "constexpr dim3" -> "const dim3"
ggml-ci
* cuda : try to fix __hgt2_mask
ggml-ci
* ggml : add TODO's for F16/F32 mask/pos support in other backends
* llama : replace bool need_kq_pos with use_alibi
* llama : prep ALiBi support for BERT models
ggml-ci
* llama : fix n_batch requirements
ggml-ci
* cont
* server : add help for --flash-attn arg
* llama : disable FA for AMD
* tests : remove TMP_ATTN_BENCH
ggml-ci
* llama : support save/load state with FA enabled
ggml-ci
* ci : add CUDA save-load-state tests
ggml-ci
* llama : llama_kv_cache_clear zeroes data + fix save-load seq
ggml-ci
* llama : fix copy-paste errors, add TODO
* llama : disallow incompatible states
* llama : update llama_state_get_size after v_trans field
* metal : remove tmp log
* llama : add static reminder for llama_state_get_size
* metal : fix max nsg
ggml-ci
* ci : fix arg order
ggml-ci
---------
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
Co-authored-by: Pierrick HYMBERT <pierrick.hymbert@gmail.com >
2024-04-30 12:16:08 +03:00
Georgi Gerganov and GitHub
952d03dbea
convert : use utf8 encoding ( #7000 )
...
* convert : use utf8 encoding
* convert : update instructions and warning message
2024-04-30 11:05:25 +03:00
Georgi Gerganov and GitHub
d2c898f746
ci : tmp disable gguf-split ( #6983 )
...
ggml-ci
2024-04-29 18:36:39 +03:00
Georgi Gerganov and GitHub
544f1f10ad
ggml : fix __MSC_VER -> _MSC_VER ( #6977 )
...
ggml-ci
2024-04-29 17:55:02 +03:00
Georgi Gerganov and GitHub
24affa7db3
readme : update hot topics
2024-04-29 17:06:19 +03:00
f4ab2a4147
llama : fix BPE pre-tokenization ( #6920 )
...
* merged the changes from deepseeker models to main branch
* Moved regex patterns to unicode.cpp and updated unicode.h
* Moved header files
* Resolved issues
* added and refactored unicode_regex_split and related functions
* Updated/merged the deepseek coder pr
* Refactored code
* Adding unicode regex mappings
* Adding unicode regex function
* Added needed functionality, testing remains
* Fixed issues
* Fixed issue with gpt2 regex custom preprocessor
* unicode : fix? unicode_wstring_to_utf8
* lint : fix whitespaces
* tests : add tokenizer tests for numbers
* unicode : remove redundant headers
* tests : remove and rename tokenizer test scripts
* tests : add sample usage
* gguf-py : reader prints warnings on duplicate keys
* llama : towards llama3 tokenization support (wip)
* unicode : shot in the dark to fix tests on Windows
* unicode : first try custom implementations
* convert : add "tokenizer.ggml.pre" GGUF KV (wip)
* llama : use new pre-tokenizer type
* convert : fix pre-tokenizer type writing
* lint : fix
* make : add test-tokenizer-0-llama-v3
* wip
* models : add llama v3 vocab file
* llama : adapt punctuation regex + add llama 3 regex
* minor
* unicode : set bomb
* unicode : set bomb
* unicode : always use std::wregex
* unicode : support \p{N}, \p{L} and \p{P} natively
* unicode : try fix windows
* unicode : category support via std::regex
* unicode : clean-up
* unicode : simplify
* convert : add convert-hf-to-gguf-update.py
ggml-ci
* lint : update
* convert : add falcon
ggml-ci
* unicode : normalize signatures
* lint : fix
* lint : fix
* convert : remove unused functions
* convert : add comments
* convert : exercise contractions
ggml-ci
* lint : fix
* cmake : refactor test targets
* tests : refactor vocab tests
ggml-ci
* tests : add more vocabs and tests
ggml-ci
* unicode : cleanup
* scripts : ignore new update script in check-requirements.sh
* models : add phi-3, mpt, gpt-2, starcoder
* tests : disable obsolete
ggml-ci
* tests : use faster bpe test
ggml-ci
* llama : more prominent warning for old BPE models
* tests : disable test-tokenizer-1-bpe due to slowness
ggml-ci
---------
Co-authored-by: Jaggzh <jaggz.h@gmail.com >
Co-authored-by: Kazim Abrar Mahi <kazimabrarmahi135@gmail.com >
2024-04-29 16:58:41 +03:00
83b72cb086
Merge pull request from GHSA-p5mv-gjc5-mwqv
...
* always use calloc
clamp n_kv on failure to read a kv
* ggml : alternative ctx->header.n_kv update
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-04-26 10:41:53 +03:00
Georgi Gerganov
dba497e0c1
cmake : restore LLAMA_LLAMAFILE_DEFAULT
2024-04-25 21:37:27 +03:00
Georgi Gerganov
fa0b4ad252
cmake : remove obsolete ANDROID check
2024-04-25 18:59:51 +03:00
Georgi Gerganov
853d06ffe2
ci : tmp disable slow tests
2024-04-25 17:06:27 +03:00
Georgi Gerganov and GitHub
51543729ff
ggml : fix redefinition of vaddvq_f32 for 32-bit ARM ( #6906 )
2024-04-25 15:48:25 +03:00
Georgi Gerganov and GitHub
54770413c4
ggml : fix MIN / MAX macros ( #6904 )
...
ggml-ci
2024-04-25 15:12:28 +03:00
Georgi Gerganov and GitHub
aa750c1ede
tests : minor bash stuff ( #6902 )
...
* tests : minor bash stuff
ggml-ci
* llama : fix build
ggml-ci
* tests : fix CUR_DIR -> ROOT_DIR
ggml-ci
* tests : fix fname
ggml-ci
2024-04-25 14:27:20 +03:00
Georgi Gerganov and GitHub
c0d1b3e03e
ggml : move 32-bit arm compat in ggml-impl.h ( #6865 )
...
ggml-ci
2024-04-24 12:00:07 +03:00
Georgi Gerganov
8960fe86ae
llama : fix typo in <|im_end|> token text ( #6745 )
2024-04-22 15:41:11 +03:00
Georgi Gerganov and GitHub
40f74e4d73
llama : add option to render special/control tokens ( #6807 )
...
* make : fix common dep on llama.h
* llama : add option to render special tokens
* readme : add API change notice
ggml-ci
* swift : fix build
2024-04-21 18:36:45 +03:00
Georgi Gerganov
b9cc76d87e
ggml : fix ggml_backend_cpu_supports_op() for CPY ( #0 )
2024-04-21 16:48:50 +03:00
Georgi Gerganov and GitHub
aed82f6837
common : try to fix Android CI ( #6780 )
...
* common : disable get_math_cpu_count() until Android CI gets fixed
* common : another try
2024-04-20 13:27:12 +03:00
Georgi Gerganov and GitHub
3b8f1ec4b1
llamafile : tmp disable + build sgemm.o when needed ( #6716 )
...
* build : sgemm.o only when needed
ggml-ci
* llamafile : tmp disable due to MoE bug
ggml-ci
2024-04-17 23:58:26 +03:00
Georgi Gerganov and GitHub
532c1737a1
llama : make general.name optional ( #6709 )
2024-04-16 23:50:38 +03:00
Georgi Gerganov and GitHub
666867b799
ggml : fix llamafile sgemm wdata offsets ( #6710 )
...
ggml-ci
2024-04-16 23:50:22 +03:00
Georgi Gerganov and GitHub
58227ffdeb
perplexity : require positive --ctx-size arg ( #6695 )
2024-04-16 09:28:33 +03:00
Georgi Gerganov and GitHub
f184dd9208
flake.lock: Update ( #6669 )
2024-04-14 06:55:30 -07:00
Georgi Gerganov and GitHub
ef21ce4ccb
imatrix : remove invalid assert ( #6632 )
2024-04-12 11:49:58 +03:00
Georgi Gerganov and GitHub
9ed2737acc
ci : disable Metal for macOS-latest-cmake-x64 ( #6628 )
2024-04-12 11:15:05 +03:00
Georgi Gerganov
c4a3a4ff47
sync : ggml
2024-04-09 20:29:06 +03:00
Georgi Gerganov and GitHub
e11a8999b5
license : update copyright notice + add AUTHORS ( #6405 )
...
* license : add AUTHORS
* authors : update
* scipts : add LICENSE and gen-authors.sh to sync
2024-04-09 09:23:19 +03:00
cc4a95426d
llama : fix attention layer count sanity check ( #6550 )
...
* llama : fix attention layer count sanity check
* llama : fix parentheses in attention layer count sanity check
There was otherwise a warning when compiling.
---------
Co-authored-by: Francis Couture-Harpin <git@compilade.net >
2024-04-08 22:25:49 +03:00
Georgi Gerganov and GitHub
b73e564b16
quantize : fix precedence of cli args ( #6541 )
2024-04-08 16:23:01 +03:00
b909236c0b
flake.lock: Update ( #6517 )
...
Flake lock file updates:
• Updated input 'flake-parts':
'github:hercules-ci/flake-parts/f7b3c975cf067e56e7cda6cb098ebe3fb4d74ca2' (2024-03-01)
→ 'github:hercules-ci/flake-parts/9126214d0a59633752a136528f5f3b9aa8565b7d' (2024-04-01)
• Updated input 'flake-parts/nixpkgs-lib':
'github:NixOS/nixpkgs/1536926ef5621b09bba54035ae2bb6d806d72ac8?dir=lib' (2024-02-29)
→ 'github:NixOS/nixpkgs/d8fe5e6c92d0d190646fb9f1056741a229980089?dir=lib' (2024-03-29)
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/d8fe5e6c92d0d190646fb9f1056741a229980089' (2024-03-29)
→ 'github:NixOS/nixpkgs/fd281bd6b7d3e32ddfa399853946f782553163b5' (2024-04-03)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-04-07 11:25:30 -07:00
Georgi Gerganov
c37247796b
sync : ggml
2024-04-07 17:05:51 +03:00
Georgi Gerganov
43e8995e75
scripts : sync ggml-cuda folder
2024-04-07 16:08:12 +03:00
Georgi Gerganov
54ea0698fb
sync : ggml
2024-04-06 18:27:46 +03:00
Georgi Gerganov
4399f13fb9
server : remove obsolete --memory-f32 option
2024-04-04 09:34:58 +03:00