Sigbjørn Skjæret and GitHub
4281c7b315
ci : exempt correct research label ( #15825 )
2025-09-06 01:21:15 +02:00
Sigbjørn Skjæret and GitHub
7d3c9f2b21
ci : explicitly set fa off or on ( #15692 )
2025-08-31 15:30:20 +02:00
Sigbjørn Skjæret and GitHub
84ab83cc0b
model : jina-embeddings-v3 support ( #13693 )
...
* initial jina-embeddings-v3 support
* initial jina-embeddings-v3 support
* initial jina-embeddings-v3 support
* fix vocab parsing with only tokenizer.json
* set mask token lstrip attribute
* additional unk_token_id fallback just in case [no ci]
* revert vocab_size() change [no ci]
* merge tensor loading into general bert
* rope
* add lora embedding and loading (non-functional)
* export separate lora ggufs instead
* add adapter metadata api
* use std::string
* convert_hf_to_lora compatibility
* fix assert
* apply suggestions from review
* apply suggestion from review
2025-08-28 15:49:50 +02:00
Sigbjørn Skjæret and GitHub
39842a7f73
gguf-py : remove erroneous FFN_GATE entry ( #15583 )
2025-08-26 09:08:08 +02:00
Sigbjørn Skjæret and GitHub
0fd90db585
metal : remove contiguous assertion for src0 in IM2COL ( #15577 )
...
* remove contiguous assertion for src0 in IM2COL
* add contiguous check in supports_op
2025-08-26 09:51:43 +03:00
Sigbjørn Skjæret and GitHub
baa9255a45
llama : merge conts and reshapes and remove unnecessary cont ( #15380 )
...
* remove unnecessary conts and merge reshapes
* restore necessary conts
* merge more conts and reshapes
* merge even more conts and reshapes
2025-08-18 19:30:17 +02:00
Sigbjørn Skjæret and GitHub
4d196981d4
convert : force patch_embd weights to F16 or F32 to avoid broken GGUFs ( #15367 )
...
* force patch_embd weights to f32
* use MmprojModel base tensor_force_quant instead
2025-08-17 14:47:42 +02:00
Sigbjørn Skjæret and GitHub
b143fbc87a
ci : fix hang in windows-hip build/release ( #15365 )
...
* fix hang in windows-latest-cmake-hip
* apply fix to release as well
2025-08-17 13:30:23 +02:00
Sigbjørn Skjæret and GitHub
d3248d9b65
ci : fix ios-xcode-build ( #15324 )
...
* fix ios-xcode-build
* use xcode-select with fixed version
* switch to macos-15 to get xcode 16.4
2025-08-15 14:02:39 +02:00
Sigbjørn Skjæret and GitHub
4ebd0c125b
cuda : fix GGML_CUDA_GRAPHS=OFF ( #15300 )
...
* fix USE_CUDA_GRAPH=OFF
ggml-ci
* check capture status
* completely disable capturing check instead
2025-08-14 13:22:07 +03:00
Sigbjørn Skjæret and GitHub
b3e16665e1
server : enable -td and -tbd parameters ( #15172 )
2025-08-13 15:43:00 +02:00
Sigbjørn Skjæret and GitHub
07aa869a91
ci : add more python requirements to copilot-setup-steps ( #15289 )
...
* ci : add flake8 and pyright to copilot-setup-steps.yml
* add tools/server/tests/requirements.txt
2025-08-13 11:30:45 +02:00
Sigbjørn Skjæret and GitHub
bc5182272c
ci : add copilot-setup-steps.yml ( #15214 )
2025-08-13 09:07:13 +02:00
Sigbjørn Skjæret and GitHub
50e81bdf5d
convert : fix merge conflicts ( #15229 )
2025-08-11 11:15:44 +02:00
Sigbjørn Skjæret and GitHub
65c797c4fa
chat : fix yandex chat template ( #15116 )
2025-08-06 13:26:49 +02:00
Sigbjørn Skjæret and GitHub
f324a3b715
chat : only remove double bos/eos if added ( #15086 )
...
* only remove double bos/eos if added
* fix tests
2025-08-05 20:43:36 +02:00
Sigbjørn Skjæret and GitHub
e5bebe5251
gguf-py : add --chat-template-file to gguf_new_metadata ( #15075 )
2025-08-04 21:01:48 +02:00
Sigbjørn Skjæret and GitHub
2721257e3e
quantize : fix confusing error message if ftype is invalid ( #15071 )
2025-08-04 18:11:02 +02:00
Sigbjørn Skjæret and GitHub
2bf3fbf0b5
ci : check that pre-tokenizer hashes are up-to-date ( #15032 )
...
* torch is not required for convert_hf_to_gguf_update
* add --check-missing parameter
* check that pre-tokenizer hashes are up-to-date
2025-08-02 14:39:01 +02:00
Sigbjørn Skjæret and GitHub
138b288b59
cuda : add softcap fusion ( #14907 )
2025-07-29 14:22:03 +02:00
Sigbjørn Skjæret and GitHub
221c0e0c58
ci : correct label refactor->refactoring ( #14832 )
2025-07-23 14:27:54 +02:00
Sigbjørn Skjæret and GitHub
e28c0b80c2
cuda : implement bf16 cpy ops and enable bf16 cont ( #14763 )
...
* implement bf16 cpy ops and enable bf16 cont
* deduplicate copy functions
* deduplicate checks
2025-07-22 12:33:10 +02:00
Sigbjørn Skjæret and GitHub
38d3af1b73
opencl: fix im2col when KW!=KH ( #14803 )
2025-07-21 13:55:10 -07:00
Sigbjørn Skjæret and GitHub
1ba45d4982
ci : disable failing vulkan crossbuilds ( #14723 )
2025-07-16 20:52:08 -03:00
Sigbjørn Skjæret and GitHub
19e5943d9e
convert : make hf token optional ( #14717 )
...
* make hf token optional
* fail if we can't get necessary tokenizer config
2025-07-16 23:17:43 +02:00
Sigbjørn Skjæret and GitHub
4b91d6f71f
convert : only check for tokenizer folder if we need it ( #14704 )
2025-07-16 08:52:04 +02:00
Sigbjørn Skjæret and GitHub
cf91f217f1
convert : add pre-computed hashes first to prevent order mishaps ( #14701 )
2025-07-16 08:51:12 +02:00
Sigbjørn Skjæret and GitHub
923e3ea2e3
cuda : add set rows for bf16 ( #14664 )
2025-07-13 15:01:24 +02:00
Sigbjørn Skjæret and GitHub
105554595f
llama : remove unintended whitespace ( #14592 )
2025-07-09 10:19:50 +02:00
Sigbjørn Skjæret and GitHub
e1a7059053
llama : fix incorrect minicpm3 v_states shape ( #14571 )
2025-07-07 23:35:35 +02:00
Sigbjørn Skjæret and GitHub
12f55c302b
llama : remove ggml_cont where possible ( #14568 )
2025-07-07 21:35:08 +02:00
Sigbjørn Skjæret and GitHub
ddef99522d
server : fix assistant prefilling when content is an array ( #14360 )
2025-07-05 09:17:14 +02:00
Sigbjørn Skjæret and GitHub
6681688146
opencl: add GELU_ERF ( #14476 )
2025-07-04 23:24:56 -07:00
Sigbjørn Skjæret and GitHub
28657a8229
ggml : implement GEGLU_ERF and GEGLU_QUICK ops ( #14445 )
2025-07-03 23:07:22 +02:00
Sigbjørn Skjæret and GitHub
e75ba4c043
gguf-py : add support for chat template jinja files ( #14508 )
...
* add support for chat template jinja files
* remove gemma3n hack
2025-07-02 21:02:35 +02:00
Sigbjørn Skjæret and GitHub
611ba4b264
ci : add OpenCL to labeler workflow ( #14496 )
2025-07-02 09:02:51 +02:00
Sigbjørn Skjæret and GitHub
eff5e45443
add GELU_ERF ( #14455 )
2025-07-01 10:14:21 +02:00
Sigbjørn Skjæret and GitHub
a5d1fb6212
ggml : fix unmerged GGML_FPxx_TO_FPxx refactoring ( #14443 )
2025-06-29 14:38:10 +02:00
a0535ffa0d
ggml : implement REGLU/GEGLU/SWIGLU ops ( #14158 )
...
* implement unary REGLU/GEGLU/SWIGLU cpu ops
* relax constraints
* duplicate shape of source
* fix ggml_vec_geglu_f16
* special case gated ops
* implement unary REGLU/GEGLU/SWIGLU cuda ops
* tighten constraints again
* refactor into GGML_GLU_OP
* metal : add glu kernels
ggml-ci
* add CUDA_GLU_BLOCK_SIZE [no ci]
* more constraints and use 64bit ints
ggml-ci
* 64bit multiplication [no ci]
* implement swapped variants (cpu/cuda)
* update comment [no ci]
ggml-ci
* Vulkan: Add GLU ops and shaders
* SYCL: Implement fused kernel GEGLU, SWIGLU and REGLU for single up+gate
* ggml : implement GLU for split up/gate (#14181 )
* implement GLU for split up/gate
* add tests for ggml_glu_split
* Vulkan: Implement glu_split logic and shader support
* add split to logging [no ci]
* SYCL: refactor element_size ops and add split up and gate support to gated kernels
* SYCL: switch GEGLU to use tanh approximation
---------
Co-authored-by: 0cc4m <picard12@live.de >
Co-authored-by: Akarshan <akarshan@menlo.ai >
* GGML: increase OP count in assertion
* Refactor: Optimize SYCL element-wise operations with unary function inlining
This commit refactors the SYCL element-wise operations to improve performance by:
- Inlining unary operations (sgn, abs, elu, gelu, silu, etc.) to reduce kernel launch overhead.
- Introducing helper functions `op_xxx` for each unary operation to encapsulate the logic.
- Replacing direct kernel calls with calls to these inlined functions.
- Using `__dpct_inline__` to encourage compiler inlining.
- Minor code cleanup and consistency improvements.
The changes aim to reduce kernel launch overhead and improve the overall efficiency of element-wise operations on SYCL devices.
* vulkan: Increase workgroup size for GLU, for performance (#14345 )
* vulkan: Increase workgroup size for GLU, for performance
* vulkan: change GLU shaders to do one element per invocation rather than one row per workgroup
* merge fix
* metal : add support for split and swap
ggml-ci
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
Co-authored-by: 0cc4m <picard12@live.de >
Co-authored-by: Akarshan <akarshan@menlo.ai >
Co-authored-by: Jeff Bolz <jbolz@nvidia.com >
2025-06-29 11:04:10 +02:00
Sigbjørn Skjæret and GitHub
6609507a91
ci : fix windows build and release ( #14431 )
2025-06-28 09:57:07 +02:00
Sigbjørn Skjæret and GitHub
f667f1e624
convert : fix broken sentencepiece vocab ( #14416 )
2025-06-27 10:42:19 +02:00
Sigbjørn Skjæret and GitHub
b25346221d
llama : return mistral-v7-tekken as default template only ( #14390 )
2025-06-26 15:01:14 +02:00
Sigbjørn Skjæret and GitHub
b193d53069
ggml : do not output unprintable characters on GGUF load failure ( #14381 )
2025-06-25 23:26:51 +02:00
Sigbjørn Skjæret and GitHub
abf241045d
main : honor --verbose-prompt on interactive prompts ( #14350 )
2025-06-24 09:31:00 +02:00
Sigbjørn Skjæret and GitHub
238005c2dc
gguf-py : fix SpecialVocab parsing when post_processor is null ( #14330 )
2025-06-22 19:46:17 +02:00
Sigbjørn Skjæret and GitHub
40bfa04c95
common : use std::string_view now that we target c++17 ( #14319 )
2025-06-22 08:37:43 +03:00
Sigbjørn Skjæret and GitHub
aa0ef5c578
gguf-py : fix Qwen3-Embedding eos token ( #14314 )
2025-06-21 18:12:05 +02:00
Sigbjørn Skjæret and GitHub
58cba76a9a
gguf-py : fix TemplateProcessing pair when bos/eos is missing ( #14312 )
2025-06-21 07:33:21 +02:00
Sigbjørn Skjæret and GitHub
22015b2092
lint : remove trailing whitepace ( #14304 )
2025-06-20 16:37:44 +02:00
Sigbjørn Skjæret and GitHub
88fc854b4b
llama : improve sep token handling ( #14272 )
2025-06-20 14:04:09 +02:00
Sigbjørn Skjæret and GitHub
3865cff4f5
convert : fix null head_dim AutoConfig regression ( #14248 )
2025-06-18 09:52:07 +02:00
Sigbjørn Skjæret and GitHub
e434e69183
common : suggest --jinja when autodetection fails ( #14222 )
2025-06-16 21:58:42 +02:00
Sigbjørn Skjæret and GitHub
d4e0d95cf5
chore : clean up relative source dir paths ( #14128 )
2025-06-11 19:04:23 +02:00
Sigbjørn Skjæret and GitHub
cc66a7f78f
tests : add test-tokenizers-repo ( #14017 )
2025-06-11 17:16:32 +02:00
Sigbjørn Skjæret and GitHub
55f6b9fa65
convert : fix duplicate key DeepSeek-R1 conversion error ( #14103 )
2025-06-10 23:29:52 +02:00
Sigbjørn Skjæret and GitHub
3678b838bb
llama : support GEGLU for jina-bert-v2 ( #14090 )
2025-06-10 18:02:08 +02:00
Sigbjørn Skjæret and GitHub
0974ad7a7c
llama : fix llama_model_chat_template with template name (LLM_KV with suffix) ( #14050 )
2025-06-07 14:13:12 +02:00
Sigbjørn Skjæret and GitHub
d17a809ef0
llama : support multiple classifier outputs and labels ( #13940 )
2025-06-06 09:03:25 +02:00
Sigbjørn Skjæret and GitHub
1caae7fc6c
gguf-py : add add_classifier_output_labels method to writer ( #14031 )
...
* add add_classifier_output_labels
* use add_classifier_output_labels
2025-06-05 17:42:31 +02:00
Sigbjørn Skjæret and GitHub
9f47fa5792
vocab : warn about missing mask token ( #14022 )
2025-06-05 09:29:18 +02:00
Sigbjørn Skjæret and GitHub
5e1c3aed40
convert : fix nomic-bert-moe mask token ( #13757 )
2025-06-01 18:07:21 +02:00
Sigbjørn Skjæret and GitHub
c496fe0b1d
convert : fix vocab padding code for bert models ( #13954 )
2025-06-01 17:23:11 +02:00
Sigbjørn Skjæret and GitHub
db38704f01
convert : fix rwkv bos/eos token ( #13844 )
2025-05-30 14:50:43 +02:00
Sigbjørn Skjæret and GitHub
e83ba3e460
llama : add support for jina-reranker-v2 ( #13900 )
2025-05-29 21:42:31 +02:00
Sigbjørn Skjæret and GitHub
2b131621e6
gguf-py : add support for sub_type (in arrays) in GGUFWriter add_key_value method ( #13561 )
2025-05-29 15:36:05 +02:00
Sigbjørn Skjæret and GitHub
5ca82fc1d7
convert : workaround for AutoConfig dummy labels ( #13881 )
2025-05-29 10:00:57 +02:00
Sigbjørn Skjæret and GitHub
6385b843a8
llama : add RobertaForSequenceClassification reranker support ( #13875 )
2025-05-29 08:15:01 +02:00
Sigbjørn Skjæret and GitHub
aa50ba462f
tests : improve UGM tokenizer test coverage ( #13773 )
2025-05-25 16:22:29 +02:00
Sigbjørn Skjæret and GitHub
c3a2624339
vocab : fix ugm tokenizer precision ( #13743 )
2025-05-24 12:29:09 +02:00
Sigbjørn Skjæret and GitHub
5be24af73d
gguf-py : correct charsmap parameter typing ( #13701 )
2025-05-22 14:25:05 +02:00
Sigbjørn Skjæret and GitHub
2aa777d86d
examples : switch retrieval to llama_encode ( #13685 )
...
* switch retrieval to llama_encode
* enable --no-warmup for retrieval
2025-05-21 16:57:38 +02:00
Sigbjørn Skjæret and GitHub
759e37b0d8
tests : avoid github urls due to throttling ( #13654 )
2025-05-20 12:03:17 +02:00
Sigbjørn Skjæret and GitHub
7c07ac244d
ci : add ppc64el to build-linux-cross ( #13575 )
2025-05-16 14:54:23 +02:00
Sigbjørn Skjæret and GitHub
f5170c1d7a
editorconfig : fix trailing whitespace from #13542 ( #13546 )
2025-05-14 21:22:49 +03:00
Sigbjørn Skjæret and GitHub
be1d4a13db
scripts : fix compare-llama-bench.py show parameter ( #13514 )
2025-05-14 08:41:01 +02:00
Sigbjørn Skjæret and GitHub
bf79371120
scripts : support arbitrary input file formats in compare-llama-bench.py ( #13455 )
2025-05-13 15:31:12 +02:00
Sigbjørn Skjæret and GitHub
09232370fc
scripts : exit compare-llama-bench.py gracefully when there's nothing to compare ( #13451 )
2025-05-11 16:20:39 +02:00
Sigbjørn Skjæret and GitHub
d2a4ef05c6
vocab : add ByteDance-Seed/Seed-Coder ( #13423 )
2025-05-10 22:08:07 +02:00
Sigbjørn Skjæret and GitHub
43dfd741a5
llguidance : set tokenizer slices to default ( #13424 )
2025-05-10 17:19:52 +02:00
Sigbjørn Skjæret and GitHub
1a844be132
convert : support rope_scaling type and rope_type ( #13349 )
2025-05-08 15:34:29 +02:00
Sigbjørn Skjæret and GitHub
bc4e1128f7
llama : deci : support ffn-free with attention ( #13296 )
2025-05-07 12:49:27 +02:00
764b85627b
convert : qwen2/3moe : set yarn metadata if present ( #13331 )
...
* set yarn metadata if present
* add comment about enabling YaRN
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co >
---------
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co >
2025-05-06 11:12:06 +02:00
Sigbjørn Skjæret and GitHub
ae803bfc3d
convert : bailingmoe : set yarn metadata if present ( #13312 )
2025-05-05 12:34:26 +02:00
Sigbjørn Skjæret and GitHub
cb06a3c363
llama : orion rope type is neox ( #13261 )
2025-05-02 12:44:24 +02:00
Sigbjørn Skjæret and GitHub
626083faf7
llama : plamo rope type is neox ( #13260 )
2025-05-02 12:40:56 +02:00
Sigbjørn Skjæret and GitHub
7d3af70b08
llama : llm_type order by size ( #13177 )
2025-04-29 13:25:53 +02:00
Sigbjørn Skjæret and GitHub
e98b3692be
llama : set qwen3 model type sizes ( #13175 )
2025-04-29 11:00:31 +02:00
Sigbjørn Skjæret and GitHub
fb28f4f80e
gguf-py : fix upload python package workflow ( #13020 )
2025-04-19 16:26:38 +02:00
Sigbjørn Skjæret and GitHub
7538246e7c
cuda : add f32 to bf16 copy op ( #12806 )
...
This allows BF16 KV-cache on CUDA.
2025-04-08 23:21:31 +02:00
Sigbjørn Skjæret and Georgi Gerganov
36ca8b3628
CUDA: don't convert BF16 weights to FP32 (ggml/1174)
...
* add bf16 support
* use convert_from_bf16_cuda instead of convert_unary_cuda for f32
* revert 7ec5085
* move functionality into convert_unary with constexpr
2025-04-07 18:44:17 +03:00
Sigbjørn Skjæret and GitHub
83a88bd6af
vocab : BailingMoE : change possessive quantifiers to greedy ( #12677 )
2025-04-02 11:21:48 +02:00
Sigbjørn Skjæret and GitHub
5936a616e4
convert : BailingMoE : fix qkv split when head_dim is 0 ( #12687 )
...
NOTE: Ling-lite-base is broken, see https://huggingface.co/inclusionAI/Ling-lite-base/discussions/2
2025-04-01 14:37:13 +02:00
Sigbjørn Skjæret and GitHub
35782aeedb
convert : BailingMoE : avoid setting rope_dim to 0 ( #12678 )
2025-03-31 23:09:48 +02:00
Sigbjørn Skjæret and GitHub
403fbacbbc
convert : Qwerky : use lora_rank_tokenshift and lora_rank_decay if present ( #12667 )
2025-03-31 16:36:25 +02:00
Sigbjørn Skjæret and GitHub
1a85949067
llava : proper description fix ( #12668 )
2025-03-31 11:28:30 +02:00
Sigbjørn Skjæret and GitHub
f52d59d771
llava : fix clip loading GGUFs with missing description ( #12660 )
2025-03-31 11:07:07 +02:00
Sigbjørn Skjæret and GitHub
2c3f8b850a
llama : support BailingMoE (Ling) ( #12634 )
2025-03-30 22:21:03 +02:00
Sigbjørn Skjæret and GitHub
3714c3ee1a
llama : fix incorrect Qwen2Moe ffn_moe_out graph callback ( #12631 )
2025-03-28 22:13:02 +01:00
Sigbjørn Skjæret and GitHub
53af4dba42
convert: fix Mistral3/Gemma3 model hparams init ( #12571 )
...
* Fix Mistral3/Gemma3 model hparams init
* set positional args correctly
* use existing hparams if passed
2025-03-25 23:03:10 +01:00
Sigbjørn Skjæret and GitHub
960e726077
chore : cleanup llama_model_loader::TENSOR_ usage ( #12492 )
2025-03-21 10:21:36 +01:00