Georgi Gerganov and GitHub
37964f44f9
mtmd : fix padding of n_tokens ( #19930 )
2026-02-26 18:39:49 +02:00
Georgi Gerganov and GitHub
01cd448b8c
server : fix ctx checkpoint restore logic ( #19924 )
2026-02-26 18:20:16 +02:00
Georgi Gerganov and GitHub
99bd67c9b2
kv-cache : fix can_shift() check to take into account M-RoPE ( #19928 )
2026-02-26 18:08:54 +02:00
Georgi Gerganov and GitHub
1ca3d1de15
gguf : avoid too many file size calls ( #19919 )
2026-02-26 12:46:32 +02:00
Georgi Gerganov and GitHub
f20469d919
server : enable multi-modal prompt caching ( #19877 )
2026-02-25 15:15:42 +02:00
d7d826b3c1
server : support multi-modal context checkpoints ( #19849 )
...
* Modify llama-memory-hybrid-iswa.cpp
* Modify llama-memory-recurrent.cpp
* Modify server-common.cpp
* Modify server-common.h
* Modify server-context.cpp
* Modify server-task.h
* Added comment to llama-memory-hybrid-iswa.cpp
* Remove comment from server-context.cpp
* Stylistic fix server-context.cpp
* Fix an issue when seqrm isn't called in server-context.cpp
* cont : alternative impl
* cont : cleanup
* cont : n_tokens -> int64_t
---------
Co-authored-by: timkhronos <timkhronos@gmail.com >
2026-02-25 15:14:27 +02:00
Georgi Gerganov and GitHub
244641955f
models : fix graph splits ( #19866 )
2026-02-25 00:01:13 +02:00
418dea39ce
ggml/gguf : prevent integer overflows ( #19856 )
...
* gguf : prevent integer overflow for ggml_context mem size
* ggml : fix int overflows in ggml_new_object()
* gguf : prevent string exhaustion
* gguf : prevent array elements exhaustion
* ggml : fix negative tensor type oob
* py : assert that alignment is non-zero power of 2
* ggml : check int overflow in ggml_new_tensor_impl and ggml_new_object
* gguf-py : error on duplicate keys when reading
* py : restore tensor_fields
* enforce proper alignment in add_custom_alignment
* gguf : better name
* gguf : fix ctx size for no_alloc == true
* gguf : minor print fix
* ggml : print values when overflow
* ggml : remove deprecated ggml_type_sizef()
* ggml : relax ggml_type asserts to debug-only
* gguf : add mem_size overflow test
* gguf : add file size check for arrays
* ggml : relax asseerts for ggml_get_type_traits()
* flake8 fix
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2026-02-24 20:17:11 +02:00
Georgi Gerganov and GitHub
da348c9dfb
models : fix qwen3.5 beta/gate shapes ( #19730 )
...
* models : fix qwen3.5 beta/gate shapes
* cont : avoid extra reshapes
2026-02-19 15:19:53 +02:00
Georgi Gerganov and GitHub
27326bfce1
models : dedup qwen35 graphs ( #19660 )
...
* models : dedup qwen35 graphs
* cont : add missing sigmoid
2026-02-19 08:17:49 +02:00
ad8207af77
cuda : enable CUDA graphs for MMID 1 <= BS <= 4 ( #19645 )
...
* cuda : enable CUDA graphs for MMID BS <= 4
* cont : add stream capture check
Co-authored-by: Oliver Simons <osimons@nvidia.com >
* cont : add MMVQ_MMID_MAX_BATCH_SIZE
---------
Co-authored-by: Oliver Simons <osimons@nvidia.com >
2026-02-17 12:31:49 +02:00
Georgi Gerganov and GitHub
cc45f2ada6
models : deduplicate delta-net graphs for Qwen family ( #19597 )
...
* models : add llm_build_delta_net_base
* cont : keep qwen35 and qwen35moe graphs intact
* cont : add comments
2026-02-16 14:35:04 +02:00
Georgi Gerganov and GitHub
d5dfc33027
graph : fix KQ mask, lora, cvec reuse checks ( #19644 )
...
* graph : fix KQ mask reuse condition
* cont : dedup KQ mask build and can_reuse
* cont : fix build
* graph : fix adapter check for reuse
2026-02-16 09:21:11 +02:00
Georgi Gerganov
ff4affb4c1
sync : ggml
2026-02-15 22:24:29 +02:00
Georgi Gerganov
55d58599c8
ggml : bump version to 0.9.7 (ggml/1425)
2026-02-15 22:24:29 +02:00
Georgi Gerganov
1a8c700bfd
ggml : bump version to 0.9.6 (ggml/1423)
2026-02-15 22:24:29 +02:00
Georgi Gerganov and GitHub
341bc7d23c
context : fix output reorder with backend sampling ( #19638 )
2026-02-15 14:57:40 +02:00
Georgi Gerganov and GitHub
08e6d914b8
ggml : avoid UB in gemm ukernel ( #19642 )
2026-02-15 14:56:35 +02:00
Georgi Gerganov and GitHub
1725e316c1
models : optimize qwen3next graph ( #19375 )
...
* models : optimizing qwen3next graph
* cont
* wip
* wip
* wip
* wip
* wip
* wip
* wip
* wip
* wip
* wip
* cont : remove redundant q, g chunking
* minor
* minor
* avoid passing masks around
* avoid concats during chunking
* naming + shapes
* update names and use prefix to disable CUDA graphs
2026-02-14 12:57:36 +02:00
Georgi Gerganov and GitHub
6e473fb384
metal : fix ACC op ( #19427 )
2026-02-14 09:54:03 +02:00
Georgi Gerganov and GitHub
bb96bfd361
memory : fix kv cache size for hybrid models ( #19559 )
2026-02-13 07:36:24 +02:00
Georgi Gerganov and GitHub
0644baefde
metal : improve concurrency ( #19555 )
2026-02-13 07:35:57 +02:00
Georgi Gerganov and GitHub
490eb96b88
metal : support GGML_OP_SET ( #19548 )
2026-02-13 07:34:52 +02:00
Georgi Gerganov and GitHub
338085c69e
args : add -kvu to llama-parallel ( #19577 )
2026-02-12 21:52:41 +02:00
Georgi Gerganov and GitHub
3b3a948134
metal : update sum_rows kernel to support float4 ( #19524 )
2026-02-12 11:35:28 +02:00
Georgi Gerganov and GitHub
914dde72ba
ggml : unary ops support non-cont src0 + metal F16 unary ops ( #19511 )
...
* ggml : unary ops support non-cont src0
* metal : support F16 unary ops + fix ELU
2026-02-11 18:58:43 +02:00
Georgi Gerganov and GitHub
9ab072ebbe
metal : extend l2_norm support for non-cont src0 ( #19502 )
2026-02-11 14:53:19 +02:00
Georgi Gerganov and GitHub
6d95707827
model : fix wavtokenizer embedding notions ( #19479 )
2026-02-11 07:52:20 +02:00
Georgi Gerganov and GitHub
89181c0b6d
ggml : extend bin bcast for permuted src1 ( #19484 )
...
* tests : extend bin bcast for permuted src1
* cont : extend bin support
* cont : s0 is always 1
* tests : simplify
2026-02-11 07:52:00 +02:00
Georgi Gerganov and GitHub
ceaa89b786
metal : consolidate unary ops ( #19490 )
2026-02-11 07:51:12 +02:00
Georgi Gerganov and GitHub
a0d585537c
cuda : extend GGML_OP_PAD to work with non-cont src0 ( #19429 )
...
* cuda : extend GGML_OP_PAD to work with non-cont src0
* tests : add permuted pad
2026-02-10 08:07:16 +02:00
81ddc60cb3
ci : add metal server workflows ( #19293 )
...
* ci : add metal server workflows
* cont : try fix python init
* cont : move to a separate workflow that runs only on master
* cont : fix num jobs
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2026-02-09 15:09:30 +02:00
Georgi Gerganov and GitHub
972f323e73
revert : "[Model] Qwen3.5 dense and MoE support (no vision) ( #19435 )" ( #19453 )
...
This reverts commit 39bf692af1 .
2026-02-09 14:57:51 +02:00
Georgi Gerganov and GitHub
eb449cdfa4
server : improve context checkpoint logic ( #19408 )
2026-02-08 09:40:04 +02:00
Georgi Gerganov and GitHub
96441c955e
ci : use -j param correctly when building with sanitizers ( #19411 )
...
* ci : use less jobs when building with sanitizers
* cont : fix nproc
* cont : fix the fix
* cont : simplify
2026-02-07 23:50:47 +01:00
Georgi Gerganov and GitHub
8872ad2125
metal : consolidate bin kernels ( #19390 )
...
* metal : refactor bin kernels
* cont
* cont : fix cv
2026-02-07 10:35:56 +02:00
Georgi Gerganov and GitHub
34ba7b5a2f
metal : fix event synchronization in cpy_tensor_async ( #19402 )
2026-02-07 07:37:15 +02:00
Georgi Gerganov and GitHub
dfde5993ea
common : add common_speculative_is_compat() ( #19270 )
...
* llama : add llama_memory_can_rm_suffix()
* Revert "llama : add llama_memory_can_rm_suffix()"
This reverts commit d30e59b62a15ef4266a6503e3f4eba770aec001b.
* spec : check if the target context is compatible for spec decoding
2026-02-06 16:47:22 +02:00
Georgi Gerganov and GitHub
7fcf1ef45d
metal : skip loading all-zero mask ( #19337 )
...
* metal : skip loading all-zero mask
* cont : minor
2026-02-06 09:25:11 +02:00
Georgi Gerganov and GitHub
3e21647666
cuda : cuda graphs now compare all node params ( #19383 )
2026-02-06 07:55:06 +02:00
Georgi Gerganov and GitHub
22cae83218
metal : adaptive CPU/GPU interleave based on number of nodes ( #19369 )
2026-02-05 19:07:22 +02:00
Georgi Gerganov and GitHub
3795cc1e89
benches : update models + numbers ( #19359 )
...
* bench : update script
* benches : update numbers
2026-02-05 14:34:07 +02:00
Georgi Gerganov and GitHub
7a4f97d196
metal : add diag ( #19330 )
2026-02-05 10:08:45 +02:00
Georgi Gerganov and GitHub
423bee462b
ci : fix sanitize workflow to enable ggml sanitizers too ( #19323 )
2026-02-04 15:12:03 +02:00
eaba92c3dc
tests : add non-cont, inplace rope tests ( #19296 )
...
* tests : add non-cont, inplace rope tests
* cont : exercise dim 3
Co-authored-by: Jeff Bolz <jbolz@nvidia.com >
* cont : more dim3 exercises
---------
Co-authored-by: Jeff Bolz <jbolz@nvidia.com >
2026-02-04 12:45:21 +02:00
Georgi Gerganov and GitHub
d838c22bb3
spec : fix the check-rate logic of ngram-simple ( #19261 )
...
* spec : fix the check-rate logic of ngram-simple
* cont : refactor + fix checks
2026-02-04 10:39:53 +02:00
Georgi Gerganov and GitHub
44008ce8f9
metal : add solve_tri ( #19302 )
2026-02-03 23:43:14 +02:00
Georgi Gerganov and GitHub
6a9bf2f788
ci : add sanitizer runs for server ( #19291 )
2026-02-03 22:41:20 +02:00
Georgi Gerganov and GitHub
faa1bc26ee
sampling : delegate input allocation to the scheduler ( #19266 )
...
* sampling : delegate input allocation to the scheduler
* graph : compute backend samplers only if needed
2026-02-03 22:16:16 +02:00
Georgi Gerganov and GitHub
c55bce4159
metal : minor cleanup ( #19251 )
2026-02-03 13:43:29 +02:00
Georgi Gerganov and GitHub
aeb827a3cc
spec : simplify time measurement using common_time_meas ( #19262 )
2026-02-03 08:20:15 +02:00
Georgi Gerganov and GitHub
6fdddb4987
metal : support virtual devices ( #18919 )
...
* metal : support virtual devices
* cont : manage buffer type context memory
* metal : add events
* cont : implement cpy_tensor_async
2026-02-02 14:29:44 +02:00
Georgi Gerganov and GitHub
1239267cc4
authors : update ( #19263 )
...
[no ci]
2026-02-02 08:51:25 +02:00
Georgi Gerganov and GitHub
4927795810
ngram-mod : fix build [no ci] ( #19216 )
2026-01-30 21:27:27 +02:00
Georgi Gerganov
d9a2a4bcaa
sync : ggml
2026-01-30 20:09:21 +02:00
Georgi Gerganov
dfd6106c84
cuda : fix compile warnings (whisper/0)
2026-01-30 20:09:21 +02:00
Georgi Gerganov and GitHub
bbada8bfb9
server : wrap around the "id_slot" parameter ( #19207 )
...
* server : wrap around the "id_slot" parameter
* cont : minor
2026-01-30 19:46:10 +02:00
Georgi Gerganov and GitHub
dabaa2e77a
spec : add ngram-mod ( #19164 )
...
* spec : add ngram-mod
* cont : simplify + keep track of occupancy
* cont : cleanup
* cont : move initialization to common/speculative
* cont : cleanup
* cont : cleanup
* cont : fix
2026-01-30 18:21:48 +02:00
Georgi Gerganov and GitHub
c3b87cebff
tests : add GQA=20 FA test ( #19095 )
2026-01-30 13:52:57 +02:00
Georgi Gerganov and GitHub
4fdbc1e4db
cuda : fix nkvo, offload and cuda graph node properties matching ( #19165 )
...
* cuda : fix nkvo
* cont : more robust cuda graph node property matching
* cont : restore pre-leafs implementation
* cont : comments + static_assert
2026-01-29 18:45:30 +02:00
Georgi Gerganov and GitHub
eed25bc6b0
arg : add -kvu to llama-batched-bench ( #19172 )
2026-01-29 08:50:47 +02:00
Georgi Gerganov and GitHub
631cbfcc7a
cuda : fix "V is K view" check for non-unified KV cache ( #19145 )
2026-01-28 09:15:27 +02:00
Georgi Gerganov and GitHub
2eee6c866c
CUDA: tune GLM 4.7 Flash FA kernel selection logic (DGX Spark) ( #19142 )
2026-01-28 09:15:11 +02:00
Georgi Gerganov and GitHub
b931f81b5a
server : adjust spec tests to generate up to 16 tokens ( #19093 )
2026-01-28 09:11:40 +02:00
Georgi Gerganov and GitHub
c5c64f72ac
llama : disable Direct IO by default ( #19109 )
...
* llama : disable Direct IO by default
* cont : override mmap if supported
2026-01-28 09:11:13 +02:00
Georgi Gerganov and GitHub
8f80d1b254
graph : fix nkvo offload with FA ( #19105 )
2026-01-26 20:18:34 +02:00
Georgi Gerganov and GitHub
56f3ebf38e
model : add correct type for GLM 4.7 Flash ( #19106 )
2026-01-26 11:24:30 +02:00
Georgi Gerganov and GitHub
d9c6ce46f7
kv-cache : support V-less cache ( #19067 )
...
* kv-cache : support V-less cache
* cuda : better check for V_is_K_view
* cuda : improve V_is_K_view check
* graph : add comments
* hparams : refactor
2026-01-25 15:48:56 +02:00
Georgi Gerganov and GitHub
080b161995
completion : fix prompt cache for recurrent models ( #19045 )
2026-01-25 09:12:50 +02:00
Georgi Gerganov and GitHub
557515be1e
graph : utilize ggml_build_forward_select() to avoid reallocations ( #18898 )
...
* graph : avoid branches between embedding and token inputs
* models : make deepstack graphs (e.g. Qwen3 VL) have constant topology
* ci : enable -DGGML_SCHED_NO_REALLOC=ON for server CI
* cont : pad token embeddings to n_embd_inp
2026-01-23 18:22:34 +02:00
a5eaa1d6a3
mla : make the V tensor a view of K ( #18986 )
...
* mla : pass V as a view of K to the FA op
* cuda : adjust mla logic to new layout
* kv-cache : fix rope shift
* tests : remove comment
* cuda : fix reusable_cutoff
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
---------
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
2026-01-22 22:09:01 +02:00
Georgi Gerganov and GitHub
0e4ebeb057
quant : manual overrides of tensor types take precedence ( #18952 )
2026-01-22 16:17:06 +02:00
Georgi Gerganov and GitHub
271191906c
metal : enable FA for MLA heads ( #18950 )
2026-01-20 12:21:28 +02:00
Georgi Gerganov and GitHub
365a3e8c31
ggml : add ggml_build_forward_select ( #18550 )
...
* ggml : add ggml_build_forward_select
* cuda : adapt CUDA graph compat to new feature
* vulkan : update logic to handle command buffer closing
* ggml : check compute for fusion
* ggml : add comment
2026-01-19 20:03:19 +02:00
Georgi Gerganov and GitHub
2fbde785bc
kv-cache : optimize KQ mask construction ( #18842 )
...
* kv-cache : optimize KQ mask construction
* cont : add explanation + improve
* cont : fix
2026-01-17 15:42:42 +02:00
Georgi Gerganov and GitHub
6e7fc8a146
cuda : print less debug logs when disabling cuda graphs ( #18868 )
2026-01-15 20:53:01 +02:00
Georgi Gerganov and GitHub
be8e3d9515
context : do not reserve scheduler for warmups ( #18867 )
2026-01-15 19:35:57 +02:00
Georgi Gerganov and GitHub
39173bcacb
context : reserve new scheduler when graph topology changes ( #18547 )
...
* context : reserve new scheduler when graph topology changes
* cont : fix
* cont : fix reserve
* cont : reserve only when changes occur + timing
* context : add comments
* llama : reserve on sampler changes
* common : allow null common_sampler
* server : task declares needs (embd, logits, sampling)
* server : do not init sampler if not needed
* llama : fix need_reserve when unsetting a sampler
* server : consolidate slot reset/clear logic
2026-01-15 16:39:17 +02:00
Georgi Gerganov and GitHub
e4832e3ae4
vocab : fix attribute overrides for harmony ( #18806 )
...
* vocab : fix attribute overrides for harmony
* cont : add warning log
2026-01-13 17:40:13 +02:00
Georgi Gerganov and GitHub
0a57271ab6
CUDA : fix unused argument when USE_CUDA_GRAPH=OFF ( #18800 )
2026-01-13 12:25:53 +02:00
Georgi Gerganov and GitHub
84ae04f163
tests : refactor test-backend-sampler ( #18753 )
...
* tests : use "auto", use std::string
* tests : refactor test-backend-sampler.cpp
* cmake : remove redundant declarations
* ci : use smaller model
* tests : add struct test_params
* tests : reduce logit bias 100.0f -> 10.0f
2026-01-11 17:31:03 +02:00
Georgi Gerganov and GitHub
f307926482
server : adjust unified KV cache tests ( #18716 )
2026-01-10 17:51:56 +02:00
Georgi Gerganov and GitHub
53eb9435da
server : fix timing of prompt/generation ( #18713 )
2026-01-09 12:59:50 +02:00
Georgi Gerganov and GitHub
d3435efc8a
scripts : pr2wt.sh reset to remote head ( #18695 )
...
* scripts : pr2wt.sh reset to remote head
* cont : cleaner
* cont : restore --set-upstream-to
2026-01-09 12:16:40 +02:00
Georgi Gerganov and GitHub
f5f8812f7c
server : use different seeds for child completions ( #18700 )
...
* server : use different seeds for child completions
* cont : handle default seed
* cont : note
2026-01-09 09:33:50 +02:00
Georgi Gerganov and GitHub
f2f6c88067
scripts : support chaining commands in pr2wt.sh ( #18671 )
2026-01-08 13:40:23 +02:00
56426673cb
scripts : add pr2wt.sh ( #18644 )
...
* scripts : add pr2wt.sh
* script : shebang
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2026-01-07 15:16:20 +02:00
Georgi Gerganov and GitHub
2da64a2f8a
models : fix backend assignment for Granite/Nemotron graphs ( #18599 )
...
* models : fix backend assignment for Granite/Nemotron graphs
* cont : add ref
* cont : move call to build_inp_embd()
2026-01-05 12:34:23 +02:00
Georgi Gerganov and GitHub
c69c7ebc90
graph : fix graph reuse logic when n_pos_per_embd > 1 ( #18566 )
2026-01-03 23:59:06 +02:00
Georgi Gerganov and GitHub
a554a1ecc7
context : fix reserve token padding to n_seqs ( #18536 )
2026-01-03 15:45:34 +02:00
Georgi Gerganov and GitHub
f38de16341
metal : adjust extra size for FA buffer to avoid reallocations ( #18545 )
2026-01-02 19:02:18 +02:00
Georgi Gerganov and GitHub
af1e8e1a6c
graph : reduce topology branching ( #18548 )
2026-01-02 19:01:56 +02:00
Georgi Gerganov and GitHub
d84a6a98be
vocab : reduce debug logs about non-EOG control tokens ( #18541 )
...
* vocab : reduce debug logs about non-EOG control tokens
* cont : add comment
2026-01-02 16:17:33 +02:00
Georgi Gerganov
13814eb370
sync : ggml
2025-12-31 18:54:43 +02:00
Georgi Gerganov
54f67b9b66
ggml : bump version to 0.9.5 (ggml/1410)
2025-12-31 18:54:43 +02:00
Georgi Gerganov and GitHub
01ade96e71
metal : remove BF16 x F16 kernels ( #18456 )
2025-12-31 09:53:48 +02:00
Georgi Gerganov and GitHub
2a85f720b8
server : handle closed connection for tasks ( #18459 )
2025-12-29 15:34:41 +02:00
Georgi Gerganov and GitHub
4301e27319
common : restore grammar-based rejection sampling ( #18137 )
...
* common : restart grammar-based rejection sampling
* sampling : allow null samplers
2025-12-17 19:46:00 +02:00
Georgi Gerganov and GitHub
5ba95754ee
security : add collaborator guidance ( #18081 )
2025-12-16 11:17:11 +02:00
Georgi Gerganov and GitHub
c560316440
graph : reuse SSM graphs ( #16490 )
...
* graph : reuse hybrid graphs
* graph : reuse recurrent graphs
* graph : fix reuse check for recurrent inputs
* memory : move the recurrent state into the memory context
* Revert "memory : move the recurrent state into the memory context"
This reverts commit 00f115fe810815d4a22a6dee0acc346131e970e1.
* cont : fix build
2025-12-16 09:36:21 +02:00