dda64fc17c
convert : get general.name from model dir, not its parent ( #5615 )
...
Co-authored-by: Brian <mofosyne@gmail.com >
2024-05-16 16:15:23 +10:00
Jared Van Bortel and GitHub
4426e2987b
cmake : fix typo ( #7151 )
2024-05-08 19:55:32 -04:00
Jared Van Bortel and GitHub
1b67731e18
BERT tokenizer fixes ( #6498 )
...
Key changes:
* BERT conversion: fix abuse of LlamaHfVocab, do not set BOS or EOS
* Nomic Embed conversion: pad vocab instead of slicing embedding tensor
* llama_tokenize: handle added special tokens like HF does
2024-04-09 13:44:08 -04:00
Jared Van Bortel and GitHub
be55134a53
convert : refactor vocab selection logic ( #6355 )
2024-03-28 11:44:36 -04:00
Jared Van Bortel and GitHub
32c8486e1f
wpm : portable unicode tolower ( #6305 )
...
Also use C locale for ispunct/isspace, and split unicode-data.cpp from unicode.cpp.
2024-03-26 17:46:21 -04:00
Jared Van Bortel and GitHub
94d1b3b411
use _wfopen instead of fopen on Windows ( #6248 )
...
also fix missing #defines before windows.h, and BPE LF token on MSVC
2024-03-23 18:48:02 -04:00
bd60d82d0c
server tests : more pythonic process management; fix bare except: ( #6146 )
...
* server tests : remove seemingly redundant newlines in print()
* server tests : use built-in subprocess features, not os.kill and psutil
* server tests : do not catch e.g. SystemExit; use print_exc
* server tests: handle TimeoutExpired exception
* server tests: fix connect on dual-stack systems
* server: tests: add new tokens regex on windows generated following new repeat penalties default changed in (#6127 )
* server: tests: remove the hack on windows since now we get the good socket family
* server: tests: add new tokens regex following new repeat penalties default changed in (#6127 )
* server: tests: add new tokens regex following new repeat penalties default changed in (#6127 )
---------
Co-authored-by: Pierrick HYMBERT <pierrick.hymbert@gmail.com >
2024-03-20 06:33:49 +01:00
Jared Van Bortel and GitHub
d199ca79f2
mpt : implement backwards compatiblity with duped output tensor ( #6139 )
2024-03-18 12:49:02 -04:00
Jared Van Bortel and GitHub
e04e04f8fa
ggml : use SYS_get_cpu if SYS_getcpu is not defined ( #5906 )
...
Fixes #5694
Fixes ggerganov/whisper.cpp#1894
2024-03-06 15:42:23 -05:00
Jared Van Bortel and GitHub
bd836944f8
quants : use MM256_SET_M128I consistently to fix gcc 7 build ( #5889 )
2024-03-05 11:56:37 -05:00
Jared Van Bortel and GitHub
4d4d2366fc
convert : automatically fall back to HfVocab if tokenizer.model doesn't exist ( #5821 )
2024-03-02 12:27:26 -05:00
Jared Van Bortel and GitHub
c7a0ad8ec9
convert-hf : make model class definitions self-contained ( #5825 )
2024-03-02 12:21:47 -05:00
Jared Van Bortel and GitHub
54fbcd2ce6
convert : fix missing ftype for gemma ( #5690 )
2024-02-23 20:39:14 +02:00
Jared Van Bortel and GitHub
15499eb942
mpt : do not duplicate token_embd.weight on disk ( #5670 )
2024-02-22 17:05:23 -05:00
Jared Van Bortel and GitHub
89febfed93
examples : do not assume BOS when shifting context ( #5622 )
2024-02-21 10:33:54 -05:00
Jared Van Bortel and GitHub
f24ed14ee0
make : pass CPPFLAGS directly to nvcc, not via -Xcompiler ( #5598 )
2024-02-19 15:54:12 -05:00
Jared Van Bortel and GitHub
a0c2dad9d4
build : pass all warning flags to nvcc via -Xcompiler ( #5570 )
...
* build : pass all warning flags to nvcc via -Xcompiler
* make : fix apparent mis-merge from #3952
* make : fix incorrect GF_CC_VER for CUDA host compiler
2024-02-18 16:21:52 -05:00
Jared Van Bortel and GitHub
ea9c8e1143
llama : add support for Nomic Embed ( #5468 )
2024-02-13 12:03:53 -05:00
Jared Van Bortel and GitHub
1ec3332ade
YaRN : store rope scaling type as int32_t in memory ( #5285 )
...
* YaRN : store rope scaling type as int32_t in memory
* llama : store mapped names as const char *
2024-02-03 13:22:06 +02:00
Jared Van Bortel and GitHub
e8dc55d006
kompute : llama-bench support and ggml_cpu_has_kompute() ( #5226 )
2024-01-30 19:04:37 -05:00
Jared Van Bortel and GitHub
6daa69ee81
kompute : fix fallback to CPU ( #5201 )
2024-01-29 17:11:27 -05:00
fbf1ddec69
Nomic Vulkan backend ( #4456 )
...
Signed-off-by: Jared Van Bortel <jared@nomic.ai >
Co-authored-by: niansa <anton-sa@web.de >
Co-authored-by: Adam Treat <treat.adam@gmail.com >
Co-authored-by: Aaron Miller <apage43@ninjawhale.com >
Co-authored-by: ToKiNoBug <tokinobug@163.com >
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
Co-authored-by: slaren <slarengh@gmail.com >
2024-01-29 15:50:50 -05:00
Jared Van Bortel and GitHub
bbe7c56c99
cmake : pass CPU architecture flags to nvcc ( #5146 )
2024-01-26 15:34:06 -05:00
Jared Van Bortel and GitHub
d292f4f204
examples : make pydantic scripts pass mypy and support py3.8 ( #5099 )
2024-01-25 14:51:24 -05:00
Jared Van Bortel and GitHub
b43ebde3b0
convert : partially revert PR #4818 ( #5041 )
2024-01-20 18:14:18 -05:00
Jared Van Bortel and GitHub
97c1549808
perplexity : fix MSVC build after #5020 ( #5043 )
...
* perplexity : fix MSVC build after #5020
* try a differerent fix
2024-01-20 17:08:08 +02:00
Jared Van Bortel and GitHub
8fe03ffdda
common : remove incorrect --model-draft default ( #4568 )
2023-12-21 19:55:34 +02:00
Jared Van Bortel and GitHub
2994f0c5a2
decode : fix logits_valid for legacy API ( #4516 )
2023-12-17 19:39:02 -05:00
Jared Van Bortel and GitHub
f7f468a97d
gguf-py : fail fast on nonsensical special token IDs ( #4489 )
2023-12-17 10:45:46 -05:00
8a5be3bd58
llama : sanity checks for access to logits ( #4274 )
...
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2023-12-15 22:16:15 -05:00
Jared Van Bortel and GitHub
70f806b821
build : detect host compiler and cuda compiler separately ( #4414 )
2023-12-13 12:10:10 -05:00
Jared Van Bortel and GitHub
6138963fb2
build : target Windows 8 for standard mingw-w64 ( #4405 )
...
* build : target Windows 8 for standard mingw-w64
* make : fix missing console.o deps
This was causing a link error with `make all` on Windows.
2023-12-12 11:27:26 +02:00
Jared Van Bortel and GitHub
511f52c334
build : enable libstdc++ assertions for debug builds ( #4275 )
2023-12-01 20:18:35 +02:00
Jared Van Bortel and GitHub
15f5d96037
build : fix build info generation and cleanup Makefile ( #3920 )
...
* cmake : fix joining of REAL_GIT_DIR
* fix includes with help from include-what-you-use
* make : remove unneeded deps and add test-rope target
* fix C includes in C++ source files
* Revert "fix includes with help from include-what-you-use"
This reverts commit 635e9fadfd516d4604a0fecf4a854bfb25ad17ae.
2023-12-01 00:23:08 +02:00
Jared Van Bortel and GitHub
64e64aa255
ggml : restore abort() in GGML_ASSERT ( #4242 )
2023-11-28 11:51:11 +02:00
Jared Van Bortel and GitHub
f3b269813f
ggml : fix -Warray-bounds warning with gcc ( #4231 )
2023-11-26 22:58:43 -05:00
Jared Van Bortel and GitHub
a6fc554e26
llama : restore prefix space in llama tokenizer ( #4081 )
2023-11-15 11:34:47 -05:00