Georgi Gerganov
9da243b36a
Revert "llava : add support for moondream vision language model ( #6899 )"
...
This reverts commit 46e12c4692 .
2024-05-08 22:14:39 +03:00
Georgi Gerganov
7e0b6a7b3b
py : also print the normalizers
2024-05-08 12:47:07 +03:00
Georgi Gerganov
c0e6fbf8c3
metal : fix unused warning
2024-05-08 09:14:50 +03:00
Georgi Gerganov and GitHub
53d6c52e22
readme : update hot topics
2024-05-07 21:43:13 +03:00
Georgi Gerganov and GitHub
947d3ad27d
ci : add GG_BUILD_EXTRA_TESTS_0 env ( #7098 )
...
* ci : add GG_BUILD_EXTRA_TESTS_0 env
ggml-ci
* Update run.sh
ggml-ci
2024-05-07 11:08:49 +03:00
b3a995b416
flake.lock: Update ( #7079 )
...
Flake lock file updates:
• Updated input 'flake-parts':
'github:hercules-ci/flake-parts/9126214d0a59633752a136528f5f3b9aa8565b7d?narHash=sha256-sB4SWl2lX95bExY2gMFG5HIzvva5AVMJd4Igm%2BGpZNw%3D' (2024-04-01)
→ 'github:hercules-ci/flake-parts/e5d10a24b66c3ea8f150e47dfdb0416ab7c3390e?narHash=sha256-yzcRNDoyVP7%2BSCNX0wmuDju1NUCt8Dz9%2BlyUXEI0dbI%3D' (2024-05-02)
• Updated input 'flake-parts/nixpkgs-lib':
'github:NixOS/nixpkgs/d8fe5e6c92d0d190646fb9f1056741a229980089?dir=lib&narHash=sha256-iMUFArF0WCatKK6RzfUJknjem0H9m4KgorO/p3Dopkk%3D' (2024-03-29)
→ 'https://github.com/NixOS/nixpkgs/archive/50eb7ecf4cd0a5756d7275c8ba36790e5bd53e33.tar.gz?narHash=sha256-QBx10%2Bk6JWz6u7VsohfSw8g8hjdBZEf8CFzXH1/1Z94%3D ' (2024-05-02)
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/7bb2ccd8cdc44c91edba16c48d2c8f331fb3d856?narHash=sha256-Drmja/f5MRHZCskS6mvzFqxEaZMeciScCTFxWVLqWEY%3D' (2024-04-25)
→ 'github:NixOS/nixpkgs/63c3a29ca82437c87573e4c6919b09a24ea61b0f?narHash=sha256-4cPymbty65RvF1DWQfc%2BBc8B233A1BWxJnNULJKQ1EY%3D' (2024-05-02)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-05-06 08:36:06 -07:00
Georgi Gerganov
bcdee0daa7
minor : fix trailing whitespace
2024-05-06 09:31:30 +03:00
Georgi Gerganov and GitHub
92139b90af
tests : add test-tokenizer-0.sh + fix some tokenizers ( #7036 )
...
* tests : add test-tokenizer-0.sh
* unicode : add all unicode number ranges
* starcoder : fix pre-tokenizer
* tests : add test that fails with DeepSeek tokenizers
* falcon : fix regex
* unicode : regenerate unicode tables
* refact : add tokenizer model
* lint : fix
* tests : disable failing tests
ggml-ci
* refact : add tests files
ggml-ci
* convert : print -> logging
ggml-ci
* lint : fix
* unicode : digit -> number
* phi-3 : update
2024-05-04 08:32:32 +03:00
Georgi Gerganov and GitHub
77e15bec62
metal : remove deprecated error code ( #7008 )
2024-04-30 15:52:21 +03:00
9c67c2773d
ggml : add Flash Attention ( #5021 )
...
* ggml : add ggml_flash_attn_ext API
* ggml : fix GQA support in ggml_flash_attn_ext
* ggml : online attention (CPU)
* metal : initial implementation
* metal : f16 precision
* metal : reduce branches
* metal : specialize for head size
* wip : 8 rows per simd group
* wip : 4 rows per simd group
* wip : template for rows per warp
* metal : parallelize across KV size
* metal : parallel reduce across heads
* metal : efficient flash_attn_f16 implementation
* metal : avoid redundant loads of the attention
* metal : scale and mask in matrix form
* metal : fix comment
* llama : avoid ggml_cast, use F32 query
* metal : add parallel reduce version (disabled)
* metal : move output into local memory + optimize
- the result from each simdgroup now stays in the registers
- significantly reduced SRAM usage
- more efficient skipping of -INF blocks
- avoid simdgroup barrier in hot loop
- add comments
* metal : add tests, fix scaling, support C > 32
* metal : improve precision
* ggml : fix f16 mad
* metal : minor
* metal : support Q > 8
* tests : add ATTN tests
* metal : disable buffer allocation logs
* tests : more
* metal : faster inner loop for C == 32
* metal : fix array initialization
* tests : ifdef
* ggml : switch to padded F16 mask for ggml_soft_max, ggml_flash_attn_ext
* ggml : fix ggml_soft_max mask requirement
* cuda : fix soft_max to use correct mask size
* cuda : add flash_attn kernel (wip)
* metal : optimize softmax for C > 32
* metal : optimize softmax
* tests : minor fix
* cuda : avoid zeroing fragments
* tests : update dims
* cuda : fix __hisinf() result check
* cuda : avoid warp_reduce for smax
* cuda : use int instead of int64_t
Noticeably improves performance (thanks to Johannes)
* cuda : make loops use the same loop values
Thanks Johannes again for the tip
* cuda : unroll some of the loops
* cuda : avoid __hisinf branches
* cuda : use half2 in softmax
* cuda : switch to 1 warp for bs > 16
* cuda : speed-up reduce part of the kernel
* cuda : unroll Q*K^T loop
* cuda : fix -INF block check
* cuda : simplify softmax
* cuda : fix matrix names
* cuda : minor
* llama : adapt to F16 KQ_pos
* llama : adapt new models to F16 KQ_mask
* ggml : fix F16 store (ARM NEON)
* llama : fix type of KQ_mask and KQ_pos
* ggml : fix CPU soft_max
* tests : add hs=256
* cuda : fix build
* metal : improve perf via smaller int registers
* cuda : adapt soft_max to F16 mask and pos
* CUDA: faster FlashAttention, kernel for bs == 1
* 16 cols for Phi-2
* no vec for hs, no hs==256 ncols==32 for Volta
* adjust kernel selection logic
* 4 warps, 256 stride for all D
* no ncols == 64
* Multiple parallel blocks for batch size 1
* fix compile warnings
* fix excessive KQ_b loads
* fix cmake build
* fix KV cache padding, NaN from INFINITY (#6438 )
* llama : flash_attn cparam + fix defrag
* server: support flash_attn param
* server: bench: enable flash_attn param
* CUDA: refactor host code, dyn. par. blocks
* fix flash_attn_vec_f16 race condition
* flush softmax exp below threshold to 0
* store temp KQ in registers
* Calculate KQ as FP32 if KQV has GGML_PREC_F32
* Add __hgt2_mask implementation for CUDA 11
* fix KQ FP32 precision fpr parallel_blocks > 1
* llama-bench : add -fa,--flash-attn arg
* metal : add BS=1 kernel for flash attention (#6508 )
* metal : add BS=1 kernel for flash attention (wip)
* metal : support more than 1 warps
* metal : opts
* metal : opt
* metal : switch to parallel reduce
* metal : reduce registers
* metal : simplify
* metal : initial FA vec kernel
* metal : use F32 attention accumulators
* batched-bench : add fattn arg
* llama : simplify llama_build_kv_store
ggml-ci
* llama : adapt build_olmo to changes
* ggml : fix arm fp16 store on windows
* metal : clean-up
* metal : clean-up kernel code
* metal : minor
* tests : remove benchmarks
ggml-ci
* ggml : fix avx512 const correctness
ggml-ci
* ggml : fix soft_max with bias on CPU
ggml-ci
* common : print --flash-attn in help
* ggml : fix num dimensions in ggml_flash_attn_ext
* llama : force disable flash attention for incompatible models
* ggml : ggml_soft_max support F16/F32 mask/pos
ggml-ci
* cuda : uint -> uint32_t
* cuda : "constexpr dim3" -> "const dim3"
ggml-ci
* cuda : try to fix __hgt2_mask
ggml-ci
* ggml : add TODO's for F16/F32 mask/pos support in other backends
* llama : replace bool need_kq_pos with use_alibi
* llama : prep ALiBi support for BERT models
ggml-ci
* llama : fix n_batch requirements
ggml-ci
* cont
* server : add help for --flash-attn arg
* llama : disable FA for AMD
* tests : remove TMP_ATTN_BENCH
ggml-ci
* llama : support save/load state with FA enabled
ggml-ci
* ci : add CUDA save-load-state tests
ggml-ci
* llama : llama_kv_cache_clear zeroes data + fix save-load seq
ggml-ci
* llama : fix copy-paste errors, add TODO
* llama : disallow incompatible states
* llama : update llama_state_get_size after v_trans field
* metal : remove tmp log
* llama : add static reminder for llama_state_get_size
* metal : fix max nsg
ggml-ci
* ci : fix arg order
ggml-ci
---------
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
Co-authored-by: Pierrick HYMBERT <pierrick.hymbert@gmail.com >
2024-04-30 12:16:08 +03:00
Georgi Gerganov and GitHub
952d03dbea
convert : use utf8 encoding ( #7000 )
...
* convert : use utf8 encoding
* convert : update instructions and warning message
2024-04-30 11:05:25 +03:00
Georgi Gerganov and GitHub
d2c898f746
ci : tmp disable gguf-split ( #6983 )
...
ggml-ci
2024-04-29 18:36:39 +03:00
Georgi Gerganov and GitHub
544f1f10ad
ggml : fix __MSC_VER -> _MSC_VER ( #6977 )
...
ggml-ci
2024-04-29 17:55:02 +03:00
Georgi Gerganov and GitHub
24affa7db3
readme : update hot topics
2024-04-29 17:06:19 +03:00
f4ab2a4147
llama : fix BPE pre-tokenization ( #6920 )
...
* merged the changes from deepseeker models to main branch
* Moved regex patterns to unicode.cpp and updated unicode.h
* Moved header files
* Resolved issues
* added and refactored unicode_regex_split and related functions
* Updated/merged the deepseek coder pr
* Refactored code
* Adding unicode regex mappings
* Adding unicode regex function
* Added needed functionality, testing remains
* Fixed issues
* Fixed issue with gpt2 regex custom preprocessor
* unicode : fix? unicode_wstring_to_utf8
* lint : fix whitespaces
* tests : add tokenizer tests for numbers
* unicode : remove redundant headers
* tests : remove and rename tokenizer test scripts
* tests : add sample usage
* gguf-py : reader prints warnings on duplicate keys
* llama : towards llama3 tokenization support (wip)
* unicode : shot in the dark to fix tests on Windows
* unicode : first try custom implementations
* convert : add "tokenizer.ggml.pre" GGUF KV (wip)
* llama : use new pre-tokenizer type
* convert : fix pre-tokenizer type writing
* lint : fix
* make : add test-tokenizer-0-llama-v3
* wip
* models : add llama v3 vocab file
* llama : adapt punctuation regex + add llama 3 regex
* minor
* unicode : set bomb
* unicode : set bomb
* unicode : always use std::wregex
* unicode : support \p{N}, \p{L} and \p{P} natively
* unicode : try fix windows
* unicode : category support via std::regex
* unicode : clean-up
* unicode : simplify
* convert : add convert-hf-to-gguf-update.py
ggml-ci
* lint : update
* convert : add falcon
ggml-ci
* unicode : normalize signatures
* lint : fix
* lint : fix
* convert : remove unused functions
* convert : add comments
* convert : exercise contractions
ggml-ci
* lint : fix
* cmake : refactor test targets
* tests : refactor vocab tests
ggml-ci
* tests : add more vocabs and tests
ggml-ci
* unicode : cleanup
* scripts : ignore new update script in check-requirements.sh
* models : add phi-3, mpt, gpt-2, starcoder
* tests : disable obsolete
ggml-ci
* tests : use faster bpe test
ggml-ci
* llama : more prominent warning for old BPE models
* tests : disable test-tokenizer-1-bpe due to slowness
ggml-ci
---------
Co-authored-by: Jaggzh <jaggz.h@gmail.com >
Co-authored-by: Kazim Abrar Mahi <kazimabrarmahi135@gmail.com >
2024-04-29 16:58:41 +03:00
83b72cb086
Merge pull request from GHSA-p5mv-gjc5-mwqv
...
* always use calloc
clamp n_kv on failure to read a kv
* ggml : alternative ctx->header.n_kv update
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-04-26 10:41:53 +03:00
Georgi Gerganov
dba497e0c1
cmake : restore LLAMA_LLAMAFILE_DEFAULT
2024-04-25 21:37:27 +03:00
Georgi Gerganov
fa0b4ad252
cmake : remove obsolete ANDROID check
2024-04-25 18:59:51 +03:00
Georgi Gerganov
853d06ffe2
ci : tmp disable slow tests
2024-04-25 17:06:27 +03:00
Georgi Gerganov and GitHub
51543729ff
ggml : fix redefinition of vaddvq_f32 for 32-bit ARM ( #6906 )
2024-04-25 15:48:25 +03:00
Georgi Gerganov and GitHub
54770413c4
ggml : fix MIN / MAX macros ( #6904 )
...
ggml-ci
2024-04-25 15:12:28 +03:00
Georgi Gerganov and GitHub
aa750c1ede
tests : minor bash stuff ( #6902 )
...
* tests : minor bash stuff
ggml-ci
* llama : fix build
ggml-ci
* tests : fix CUR_DIR -> ROOT_DIR
ggml-ci
* tests : fix fname
ggml-ci
2024-04-25 14:27:20 +03:00
Georgi Gerganov and GitHub
c0d1b3e03e
ggml : move 32-bit arm compat in ggml-impl.h ( #6865 )
...
ggml-ci
2024-04-24 12:00:07 +03:00
Georgi Gerganov
8960fe86ae
llama : fix typo in <|im_end|> token text ( #6745 )
2024-04-22 15:41:11 +03:00
Georgi Gerganov and GitHub
40f74e4d73
llama : add option to render special/control tokens ( #6807 )
...
* make : fix common dep on llama.h
* llama : add option to render special tokens
* readme : add API change notice
ggml-ci
* swift : fix build
2024-04-21 18:36:45 +03:00
Georgi Gerganov
b9cc76d87e
ggml : fix ggml_backend_cpu_supports_op() for CPY ( #0 )
2024-04-21 16:48:50 +03:00
Georgi Gerganov and GitHub
aed82f6837
common : try to fix Android CI ( #6780 )
...
* common : disable get_math_cpu_count() until Android CI gets fixed
* common : another try
2024-04-20 13:27:12 +03:00
Georgi Gerganov and GitHub
3b8f1ec4b1
llamafile : tmp disable + build sgemm.o when needed ( #6716 )
...
* build : sgemm.o only when needed
ggml-ci
* llamafile : tmp disable due to MoE bug
ggml-ci
2024-04-17 23:58:26 +03:00
Georgi Gerganov and GitHub
532c1737a1
llama : make general.name optional ( #6709 )
2024-04-16 23:50:38 +03:00
Georgi Gerganov and GitHub
666867b799
ggml : fix llamafile sgemm wdata offsets ( #6710 )
...
ggml-ci
2024-04-16 23:50:22 +03:00
Georgi Gerganov and GitHub
58227ffdeb
perplexity : require positive --ctx-size arg ( #6695 )
2024-04-16 09:28:33 +03:00
Georgi Gerganov and GitHub
f184dd9208
flake.lock: Update ( #6669 )
2024-04-14 06:55:30 -07:00
Georgi Gerganov and GitHub
ef21ce4ccb
imatrix : remove invalid assert ( #6632 )
2024-04-12 11:49:58 +03:00
Georgi Gerganov and GitHub
9ed2737acc
ci : disable Metal for macOS-latest-cmake-x64 ( #6628 )
2024-04-12 11:15:05 +03:00
Georgi Gerganov
c4a3a4ff47
sync : ggml
2024-04-09 20:29:06 +03:00
Georgi Gerganov and GitHub
e11a8999b5
license : update copyright notice + add AUTHORS ( #6405 )
...
* license : add AUTHORS
* authors : update
* scipts : add LICENSE and gen-authors.sh to sync
2024-04-09 09:23:19 +03:00
cc4a95426d
llama : fix attention layer count sanity check ( #6550 )
...
* llama : fix attention layer count sanity check
* llama : fix parentheses in attention layer count sanity check
There was otherwise a warning when compiling.
---------
Co-authored-by: Francis Couture-Harpin <git@compilade.net >
2024-04-08 22:25:49 +03:00
Georgi Gerganov and GitHub
b73e564b16
quantize : fix precedence of cli args ( #6541 )
2024-04-08 16:23:01 +03:00
b909236c0b
flake.lock: Update ( #6517 )
...
Flake lock file updates:
• Updated input 'flake-parts':
'github:hercules-ci/flake-parts/f7b3c975cf067e56e7cda6cb098ebe3fb4d74ca2' (2024-03-01)
→ 'github:hercules-ci/flake-parts/9126214d0a59633752a136528f5f3b9aa8565b7d' (2024-04-01)
• Updated input 'flake-parts/nixpkgs-lib':
'github:NixOS/nixpkgs/1536926ef5621b09bba54035ae2bb6d806d72ac8?dir=lib' (2024-02-29)
→ 'github:NixOS/nixpkgs/d8fe5e6c92d0d190646fb9f1056741a229980089?dir=lib' (2024-03-29)
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/d8fe5e6c92d0d190646fb9f1056741a229980089' (2024-03-29)
→ 'github:NixOS/nixpkgs/fd281bd6b7d3e32ddfa399853946f782553163b5' (2024-04-03)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-04-07 11:25:30 -07:00
Georgi Gerganov
c37247796b
sync : ggml
2024-04-07 17:05:51 +03:00
Georgi Gerganov
43e8995e75
scripts : sync ggml-cuda folder
2024-04-07 16:08:12 +03:00
Georgi Gerganov
54ea0698fb
sync : ggml
2024-04-06 18:27:46 +03:00
Georgi Gerganov
4399f13fb9
server : remove obsolete --memory-f32 option
2024-04-04 09:34:58 +03:00
Georgi Gerganov and GitHub
076b08649e
readme : update hot topics
2024-04-03 16:11:15 +03:00
f87f7b8986
flake.lock: Update ( #6402 )
...
Flake lock file updates:
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/44d0940ea560dee511026a53f0e2e2cde489b4d4' (2024-03-23)
→ 'github:NixOS/nixpkgs/d8fe5e6c92d0d190646fb9f1056741a229980089' (2024-03-29)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-04-01 09:05:57 -07:00
Georgi Gerganov and GitHub
c50a82ce0f
readme : update hot topics
2024-03-31 11:56:30 +03:00
d48ccf3ad4
sync : ggml ( #6351 )
...
* sync : ggml
ggml-ci
* cuda : move GGML_CUDA_DMMV constants to dmmv.cuh
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-03-29 17:45:46 +02:00
Georgi Gerganov and GitHub
cfde806eb9
ci : fix BGE wget ( #6383 )
...
ggml-ci
2024-03-29 14:34:28 +02:00
Georgi Gerganov and GitHub
bfe7dafc9c
readme : add notice for UI list
2024-03-28 22:56:03 +02:00
Georgi Gerganov
3a0345970e
make : whitespace
2024-03-27 15:02:49 +02:00
Georgi Gerganov
2ab4f00d25
llama2c : open file as binary ( #6332 )
2024-03-27 09:16:02 +02:00
43139cc528
flake.lock: Update ( #6266 )
...
Flake lock file updates:
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/d691274a972b3165335d261cc4671335f5c67de9' (2024-03-14)
→ 'github:NixOS/nixpkgs/44d0940ea560dee511026a53f0e2e2cde489b4d4' (2024-03-23)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-03-25 08:22:27 -07:00
a0e584defd
imatrix : fix wname for mul_mat_id ops ( #6271 )
...
* imatrix : fix wname for mul_mat_id ops
* also filter tensor names in mul_mat_id ops
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-03-24 16:18:45 +02:00
Georgi Gerganov
95562175f8
gitignore : gguf-split
2024-03-23 21:35:23 +02:00
Georgi Gerganov
56a00f0a2f
common : default --hf-file to --model ( #6234 )
2024-03-22 21:10:39 +02:00
Georgi Gerganov and GitHub
80bd33bc2c
common : add HF arg helpers ( #6234 )
...
* common : add HF arg helpers
* common : remove defaults
2024-03-22 15:33:38 +02:00
Georgi Gerganov and GitHub
68e210b354
server : enable continuous batching by default ( #6231 )
2024-03-22 13:08:28 +02:00
Georgi Gerganov and GitHub
b3e94f26ba
metal : proper assert for mat-mat memory alignment ( #6225 )
...
* metal : proper assert for mat-mat memory alignment
ggml-ci
* readme : add notice about the bug fix
* metal : fix the fix
ggml-ci
2024-03-22 11:35:53 +02:00
Georgi Gerganov and GitHub
95d576b48e
metal : pad n_ctx by 32 ( #6177 )
...
* metal : require ne00 >= 128 for mat-mat kernels
ggml-ci
* llama : pad n_ctx by 32
ggml-ci
2024-03-22 09:36:03 +02:00
Georgi Gerganov and GitHub
924ce1dce7
tests : disable system() calls ( #6198 )
...
ggml-ci
2024-03-21 16:20:05 +02:00
Georgi Gerganov
6b7e76d28c
gitignore : ignore curl-related files
2024-03-20 14:17:34 +02:00
Georgi Gerganov
bc0baab2ea
server : allow to override -ngl in tests ( #6170 )
2024-03-20 14:14:32 +02:00
Georgi Gerganov
d795988d9e
Revert "llava : add a MobileVLM_V2-1.7B backup ( #6152 )"
...
This reverts commit f8c4e745e1 .
2024-03-20 13:29:49 +02:00
Georgi Gerganov and GitHub
b80cf3b2d1
common : disable repeat penalties by default ( #6127 )
2024-03-19 10:21:54 +02:00
Georgi Gerganov and GitHub
ac9ee6a4ad
ci : disable stale issue messages ( #6126 )
2024-03-18 13:45:38 +02:00
Georgi Gerganov and GitHub
4f6d1337ca
ci : temporary disable sanitizer builds ( #6128 )
2024-03-18 13:45:27 +02:00
Georgi Gerganov and GitHub
cd776c37c9
ci : close all stale issues at once ( #6115 )
2024-03-17 18:51:57 +01:00
Georgi Gerganov
131b058409
make : ggml-metal.o depends on ggml.h
2024-03-15 11:38:40 +02:00
Georgi Gerganov and GitHub
4755afd1cb
llama : fix integer overflow during quantization ( #6063 )
2024-03-14 22:58:41 +02:00
Georgi Gerganov
044ec4b2a5
embedding : add EOS token if not present ( #899 )
2024-03-14 15:14:14 +02:00
Georgi Gerganov
77178eedc8
gguf-py : fix dtype check ( #6045 )
2024-03-14 13:32:14 +02:00
Georgi Gerganov
a44bc969e4
llama : fix typo
2024-03-14 13:13:06 +02:00
Georgi Gerganov and GitHub
3fe8d7a17f
ggml : designate enum vals for integer types ( #6050 )
2024-03-14 12:38:37 +02:00
Georgi Gerganov
68265ebfc6
embedding : print all resulting embeddings ( #899 )
2024-03-14 12:37:20 +02:00
Georgi Gerganov and GitHub
381da2d9f0
metal : build metallib + fix embed path ( #6015 )
...
* metal : build metallib + fix embed path
ggml-ci
* metal : fix embed build + update library load logic
ggml-ci
* metal : fix embeded library build
ggml-ci
* ci : fix iOS builds to use embedded library
2024-03-14 11:55:23 +02:00
Georgi Gerganov
0fd6c1f015
embedding : print cosine similarity ( #899 )
2024-03-14 10:12:29 +02:00
Georgi Gerganov and GitHub
76a936c893
readme : update API changes and hot topics
2024-03-13 20:33:56 +02:00
Georgi Gerganov and GitHub
8030da7afe
ggml : reuse quantum structs across backends ( #5943 )
...
* ggml : reuse quant blocks across backends
ggml-ci
* ggml : define helper constants only for CUDA and SYCL
ggml-ci
* ggml : define helper quantum constants for SYCL
ggml-ci
2024-03-12 14:27:20 +02:00
Georgi Gerganov and GitHub
184215e783
ggml : fix UB in IQ2_S and IQ3_S ( #6012 )
2024-03-12 13:49:55 +02:00
Georgi Gerganov and GitHub
48358b2e5b
sycl : update IQ1_S kernels (WIP - not working!) ( #5995 )
...
* sycl : try to fix after IQ1_S changes
* sycl : iq1s_grid -> iq1s_grid_gpu
* sycl : fix grid type
2024-03-12 11:15:05 +02:00
Georgi Gerganov and GitHub
05b06210c9
llama : more consistent names of count variables ( #5994 )
...
* llama : more consistent names of count variables
ggml-ci
* llama : n_parallel -> n_seq_max
* common : fix param name
* examples : fix param name
2024-03-11 17:49:47 +02:00
Georgi Gerganov and GitHub
83796e62bc
llama : refactor unicode stuff ( #5992 )
...
* llama : refactor unicode stuff
ggml-ci
* unicode : names
* make : fix c++ compiler
* unicode : names
* unicode : straighten tables
* zig : fix build
* unicode : put nfd normalization behind API
ggml-ci
* swift : fix build
* unicode : add BOM
* unicode : add <cstdint>
ggml-ci
* unicode : pass as cpts as const ref
2024-03-11 17:47:47 +02:00
Georgi Gerganov
ee35600b90
llama : fix F16/F32 downcast + improve names ( #5980 )
2024-03-11 09:56:47 +02:00
Georgi Gerganov and GitHub
bb6d00bbf9
metal : move mm_id indices to shared mem ( #5982 )
2024-03-10 23:12:48 +02:00
Georgi Gerganov and GitHub
d9f65c97c3
readme : update hot topics
2024-03-10 20:58:26 +02:00
Georgi Gerganov
b838b53ad6
sync : ggml
2024-03-10 20:10:46 +02:00
Georgi Gerganov
df4dc3e7cb
ggml : try fix 32-bit arm compat (whisper/1938)
...
* ggml : try fix 32-bit arm compat
* ggml : fix cont
2024-03-10 20:10:39 +02:00
Georgi Gerganov
bf47a5eefc
ggml : remove __constant__ specifier for CUDA tables ( #5940 )
2024-03-10 20:09:24 +02:00
c78541479c
nix: update flake.lock ( #5969 )
...
Flake lock file updates:
• Updated input 'nixpkgs':
'github:NixOS/nixpkgs/1536926ef5621b09bba54035ae2bb6d806d72ac8' (2024-02-29)
→ 'github:NixOS/nixpkgs/9df3e30ce24fd28c7b3e2de0d986769db5d6225d' (2024-03-06)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2024-03-10 16:43:08 +02:00
Georgi Gerganov
77d1ac7e00
server : print chat template info
2024-03-09 22:04:00 +02:00
Georgi Gerganov and GitHub
098dbaab44
readme : update hot topics
2024-03-09 18:14:13 +02:00
Georgi Gerganov and GitHub
8380ecfb21
ggml : fix unnecessary f32 -> f16 -> f32 casts (mmla) ( #5951 )
2024-03-09 17:36:20 +02:00
Georgi Gerganov and GitHub
58308a0ecc
server : fix metrics init ( #5964 )
2024-03-09 17:34:15 +02:00
Georgi Gerganov and GitHub
5b09797321
ggml : remove old quantization functions ( #5942 )
...
* ggml : remove old quantization functions
ggml-ci
* ggml : simplify ggml_quantize_chunk
ggml-ci
* ggml : restrict correctness
ggml-ci
* ggml : remove hist data from the quantization API
ggml-ci
* tests : remove hist usage in test-backend-ops
ggml-ci
* vulkan : remove hist and fix typo
2024-03-09 15:53:59 +02:00
Georgi Gerganov and GitHub
97c09585d6
server : clarify some items in the readme ( #5957 )
...
* server : clarify some items in the readme
* server : fix typo
2024-03-09 15:47:47 +02:00
Georgi Gerganov
2c4f566c88
tests : gitignore ggml-common.h
2024-03-09 14:17:11 +02:00
Georgi Gerganov and GitHub
8a3012a4ad
ggml : add ggml-common.h to deduplicate shared code ( #5940 )
...
* ggml : add ggml-common.h to shared code
ggml-ci
* scripts : update sync scripts
* sycl : reuse quantum tables
ggml-ci
* ggml : minor
* ggml : minor
* sycl : try to fix build
2024-03-09 12:47:57 +02:00
Georgi Gerganov and GitHub
9674aaf35c
server : simplify logic for empty prompts ( #5953 )
2024-03-09 12:34:18 +02:00
Georgi Gerganov and GitHub
af37fd8b30
server : fix EOS token detection with disabled cache ( #5938 )
2024-03-08 12:40:02 +02:00
6cdabe6526
llama-bench : add embeddings option ( #5924 )
...
* llama-bench : add embeddings option
* llama-bench : do not hard code embd default value
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-03-07 16:32:38 +02:00