Georgi Gerganov and GitHub
7fcf1ef45d
metal : skip loading all-zero mask ( #19337 )
...
* metal : skip loading all-zero mask
* cont : minor
2026-02-06 09:25:11 +02:00
Georgi Gerganov and GitHub
3e21647666
cuda : cuda graphs now compare all node params ( #19383 )
2026-02-06 07:55:06 +02:00
Georgi Gerganov and GitHub
22cae83218
metal : adaptive CPU/GPU interleave based on number of nodes ( #19369 )
2026-02-05 19:07:22 +02:00
Georgi Gerganov and GitHub
3795cc1e89
benches : update models + numbers ( #19359 )
...
* bench : update script
* benches : update numbers
2026-02-05 14:34:07 +02:00
Georgi Gerganov and GitHub
7a4f97d196
metal : add diag ( #19330 )
2026-02-05 10:08:45 +02:00
Georgi Gerganov and GitHub
423bee462b
ci : fix sanitize workflow to enable ggml sanitizers too ( #19323 )
2026-02-04 15:12:03 +02:00
eaba92c3dc
tests : add non-cont, inplace rope tests ( #19296 )
...
* tests : add non-cont, inplace rope tests
* cont : exercise dim 3
Co-authored-by: Jeff Bolz <jbolz@nvidia.com >
* cont : more dim3 exercises
---------
Co-authored-by: Jeff Bolz <jbolz@nvidia.com >
2026-02-04 12:45:21 +02:00
Georgi Gerganov and GitHub
d838c22bb3
spec : fix the check-rate logic of ngram-simple ( #19261 )
...
* spec : fix the check-rate logic of ngram-simple
* cont : refactor + fix checks
2026-02-04 10:39:53 +02:00
Georgi Gerganov and GitHub
44008ce8f9
metal : add solve_tri ( #19302 )
2026-02-03 23:43:14 +02:00
Georgi Gerganov and GitHub
6a9bf2f788
ci : add sanitizer runs for server ( #19291 )
2026-02-03 22:41:20 +02:00
Georgi Gerganov and GitHub
faa1bc26ee
sampling : delegate input allocation to the scheduler ( #19266 )
...
* sampling : delegate input allocation to the scheduler
* graph : compute backend samplers only if needed
2026-02-03 22:16:16 +02:00
Georgi Gerganov and GitHub
c55bce4159
metal : minor cleanup ( #19251 )
2026-02-03 13:43:29 +02:00
Georgi Gerganov and GitHub
aeb827a3cc
spec : simplify time measurement using common_time_meas ( #19262 )
2026-02-03 08:20:15 +02:00
Georgi Gerganov and GitHub
6fdddb4987
metal : support virtual devices ( #18919 )
...
* metal : support virtual devices
* cont : manage buffer type context memory
* metal : add events
* cont : implement cpy_tensor_async
2026-02-02 14:29:44 +02:00
Georgi Gerganov and GitHub
1239267cc4
authors : update ( #19263 )
...
[no ci]
2026-02-02 08:51:25 +02:00
Georgi Gerganov and GitHub
4927795810
ngram-mod : fix build [no ci] ( #19216 )
2026-01-30 21:27:27 +02:00
Georgi Gerganov
d9a2a4bcaa
sync : ggml
2026-01-30 20:09:21 +02:00
Georgi Gerganov
dfd6106c84
cuda : fix compile warnings (whisper/0)
2026-01-30 20:09:21 +02:00
Georgi Gerganov and GitHub
bbada8bfb9
server : wrap around the "id_slot" parameter ( #19207 )
...
* server : wrap around the "id_slot" parameter
* cont : minor
2026-01-30 19:46:10 +02:00
Georgi Gerganov and GitHub
dabaa2e77a
spec : add ngram-mod ( #19164 )
...
* spec : add ngram-mod
* cont : simplify + keep track of occupancy
* cont : cleanup
* cont : move initialization to common/speculative
* cont : cleanup
* cont : cleanup
* cont : fix
2026-01-30 18:21:48 +02:00
Georgi Gerganov and GitHub
c3b87cebff
tests : add GQA=20 FA test ( #19095 )
2026-01-30 13:52:57 +02:00
Georgi Gerganov and GitHub
4fdbc1e4db
cuda : fix nkvo, offload and cuda graph node properties matching ( #19165 )
...
* cuda : fix nkvo
* cont : more robust cuda graph node property matching
* cont : restore pre-leafs implementation
* cont : comments + static_assert
2026-01-29 18:45:30 +02:00
Georgi Gerganov and GitHub
eed25bc6b0
arg : add -kvu to llama-batched-bench ( #19172 )
2026-01-29 08:50:47 +02:00
Georgi Gerganov and GitHub
631cbfcc7a
cuda : fix "V is K view" check for non-unified KV cache ( #19145 )
2026-01-28 09:15:27 +02:00
Georgi Gerganov and GitHub
2eee6c866c
CUDA: tune GLM 4.7 Flash FA kernel selection logic (DGX Spark) ( #19142 )
2026-01-28 09:15:11 +02:00
Georgi Gerganov and GitHub
b931f81b5a
server : adjust spec tests to generate up to 16 tokens ( #19093 )
2026-01-28 09:11:40 +02:00
Georgi Gerganov and GitHub
c5c64f72ac
llama : disable Direct IO by default ( #19109 )
...
* llama : disable Direct IO by default
* cont : override mmap if supported
2026-01-28 09:11:13 +02:00
Georgi Gerganov and GitHub
8f80d1b254
graph : fix nkvo offload with FA ( #19105 )
2026-01-26 20:18:34 +02:00
Georgi Gerganov and GitHub
56f3ebf38e
model : add correct type for GLM 4.7 Flash ( #19106 )
2026-01-26 11:24:30 +02:00
Georgi Gerganov and GitHub
d9c6ce46f7
kv-cache : support V-less cache ( #19067 )
...
* kv-cache : support V-less cache
* cuda : better check for V_is_K_view
* cuda : improve V_is_K_view check
* graph : add comments
* hparams : refactor
2026-01-25 15:48:56 +02:00
Georgi Gerganov and GitHub
080b161995
completion : fix prompt cache for recurrent models ( #19045 )
2026-01-25 09:12:50 +02:00
Georgi Gerganov and GitHub
557515be1e
graph : utilize ggml_build_forward_select() to avoid reallocations ( #18898 )
...
* graph : avoid branches between embedding and token inputs
* models : make deepstack graphs (e.g. Qwen3 VL) have constant topology
* ci : enable -DGGML_SCHED_NO_REALLOC=ON for server CI
* cont : pad token embeddings to n_embd_inp
2026-01-23 18:22:34 +02:00
a5eaa1d6a3
mla : make the V tensor a view of K ( #18986 )
...
* mla : pass V as a view of K to the FA op
* cuda : adjust mla logic to new layout
* kv-cache : fix rope shift
* tests : remove comment
* cuda : fix reusable_cutoff
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
---------
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
2026-01-22 22:09:01 +02:00
Georgi Gerganov and GitHub
0e4ebeb057
quant : manual overrides of tensor types take precedence ( #18952 )
2026-01-22 16:17:06 +02:00
Georgi Gerganov and GitHub
271191906c
metal : enable FA for MLA heads ( #18950 )
2026-01-20 12:21:28 +02:00
Georgi Gerganov and GitHub
365a3e8c31
ggml : add ggml_build_forward_select ( #18550 )
...
* ggml : add ggml_build_forward_select
* cuda : adapt CUDA graph compat to new feature
* vulkan : update logic to handle command buffer closing
* ggml : check compute for fusion
* ggml : add comment
2026-01-19 20:03:19 +02:00
Georgi Gerganov and GitHub
2fbde785bc
kv-cache : optimize KQ mask construction ( #18842 )
...
* kv-cache : optimize KQ mask construction
* cont : add explanation + improve
* cont : fix
2026-01-17 15:42:42 +02:00
Georgi Gerganov and GitHub
6e7fc8a146
cuda : print less debug logs when disabling cuda graphs ( #18868 )
2026-01-15 20:53:01 +02:00
Georgi Gerganov and GitHub
be8e3d9515
context : do not reserve scheduler for warmups ( #18867 )
2026-01-15 19:35:57 +02:00
Georgi Gerganov and GitHub
39173bcacb
context : reserve new scheduler when graph topology changes ( #18547 )
...
* context : reserve new scheduler when graph topology changes
* cont : fix
* cont : fix reserve
* cont : reserve only when changes occur + timing
* context : add comments
* llama : reserve on sampler changes
* common : allow null common_sampler
* server : task declares needs (embd, logits, sampling)
* server : do not init sampler if not needed
* llama : fix need_reserve when unsetting a sampler
* server : consolidate slot reset/clear logic
2026-01-15 16:39:17 +02:00
Georgi Gerganov and GitHub
e4832e3ae4
vocab : fix attribute overrides for harmony ( #18806 )
...
* vocab : fix attribute overrides for harmony
* cont : add warning log
2026-01-13 17:40:13 +02:00
Georgi Gerganov and GitHub
0a57271ab6
CUDA : fix unused argument when USE_CUDA_GRAPH=OFF ( #18800 )
2026-01-13 12:25:53 +02:00
Georgi Gerganov and GitHub
84ae04f163
tests : refactor test-backend-sampler ( #18753 )
...
* tests : use "auto", use std::string
* tests : refactor test-backend-sampler.cpp
* cmake : remove redundant declarations
* ci : use smaller model
* tests : add struct test_params
* tests : reduce logit bias 100.0f -> 10.0f
2026-01-11 17:31:03 +02:00
Georgi Gerganov and GitHub
f307926482
server : adjust unified KV cache tests ( #18716 )
2026-01-10 17:51:56 +02:00
Georgi Gerganov and GitHub
53eb9435da
server : fix timing of prompt/generation ( #18713 )
2026-01-09 12:59:50 +02:00
Georgi Gerganov and GitHub
d3435efc8a
scripts : pr2wt.sh reset to remote head ( #18695 )
...
* scripts : pr2wt.sh reset to remote head
* cont : cleaner
* cont : restore --set-upstream-to
2026-01-09 12:16:40 +02:00
Georgi Gerganov and GitHub
f5f8812f7c
server : use different seeds for child completions ( #18700 )
...
* server : use different seeds for child completions
* cont : handle default seed
* cont : note
2026-01-09 09:33:50 +02:00
Georgi Gerganov and GitHub
f2f6c88067
scripts : support chaining commands in pr2wt.sh ( #18671 )
2026-01-08 13:40:23 +02:00
56426673cb
scripts : add pr2wt.sh ( #18644 )
...
* scripts : add pr2wt.sh
* script : shebang
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2026-01-07 15:16:20 +02:00
Georgi Gerganov and GitHub
2da64a2f8a
models : fix backend assignment for Granite/Nemotron graphs ( #18599 )
...
* models : fix backend assignment for Granite/Nemotron graphs
* cont : add ref
* cont : move call to build_inp_embd()
2026-01-05 12:34:23 +02:00
Georgi Gerganov and GitHub
c69c7ebc90
graph : fix graph reuse logic when n_pos_per_embd > 1 ( #18566 )
2026-01-03 23:59:06 +02:00
Georgi Gerganov and GitHub
a554a1ecc7
context : fix reserve token padding to n_seqs ( #18536 )
2026-01-03 15:45:34 +02:00
Georgi Gerganov and GitHub
f38de16341
metal : adjust extra size for FA buffer to avoid reallocations ( #18545 )
2026-01-02 19:02:18 +02:00
Georgi Gerganov and GitHub
af1e8e1a6c
graph : reduce topology branching ( #18548 )
2026-01-02 19:01:56 +02:00
Georgi Gerganov and GitHub
d84a6a98be
vocab : reduce debug logs about non-EOG control tokens ( #18541 )
...
* vocab : reduce debug logs about non-EOG control tokens
* cont : add comment
2026-01-02 16:17:33 +02:00
Georgi Gerganov
13814eb370
sync : ggml
2025-12-31 18:54:43 +02:00
Georgi Gerganov
54f67b9b66
ggml : bump version to 0.9.5 (ggml/1410)
2025-12-31 18:54:43 +02:00
Georgi Gerganov and GitHub
01ade96e71
metal : remove BF16 x F16 kernels ( #18456 )
2025-12-31 09:53:48 +02:00
Georgi Gerganov and GitHub
2a85f720b8
server : handle closed connection for tasks ( #18459 )
2025-12-29 15:34:41 +02:00
Georgi Gerganov and GitHub
4301e27319
common : restore grammar-based rejection sampling ( #18137 )
...
* common : restart grammar-based rejection sampling
* sampling : allow null samplers
2025-12-17 19:46:00 +02:00
Georgi Gerganov and GitHub
5ba95754ee
security : add collaborator guidance ( #18081 )
2025-12-16 11:17:11 +02:00
Georgi Gerganov and GitHub
c560316440
graph : reuse SSM graphs ( #16490 )
...
* graph : reuse hybrid graphs
* graph : reuse recurrent graphs
* graph : fix reuse check for recurrent inputs
* memory : move the recurrent state into the memory context
* Revert "memory : move the recurrent state into the memory context"
This reverts commit 00f115fe810815d4a22a6dee0acc346131e970e1.
* cont : fix build
2025-12-16 09:36:21 +02:00
Georgi Gerganov and GitHub
254098a279
common : refactor common_sampler + grammar logic changes ( #17937 )
...
* common : refactor common_sampler + grammar logic changes
* tests : increase max_tokens to get needed response
* batched : fix uninitialized samplers
2025-12-14 10:11:13 +02:00
Georgi Gerganov and GitHub
77ad8542bd
model-conversion : cast logits to float32 ( #18009 )
2025-12-14 08:58:13 +02:00
Georgi Gerganov and GitHub
609a2d0268
models : fix YaRN regression + consolidate logic ( #18006 )
...
* models : fix YaRN regression + consolidate logic
* cont : fix the fix
* cont : remove header
* cont : add header
2025-12-14 08:34:56 +02:00
Georgi Gerganov
a63cbafbbc
ggml : arm repack fix build
2025-12-14 08:33:51 +02:00
Georgi Gerganov
0e59224990
sync : ggml
2025-12-14 08:33:51 +02:00
Georgi Gerganov
71fdcf0616
ggml : arm repack fix build (whisper/0)
2025-12-14 08:33:51 +02:00
Georgi Gerganov and GitHub
3c6391e748
speculative-simple : free batch on exit ( #17985 )
2025-12-13 09:48:34 +02:00
Georgi Gerganov and GitHub
7bed317f53
models : fix the attn_factor for mistral3 graphs + improve consistency ( #17945 )
...
* models : fix the attn_factor for mistral3 graphs
* cont : rework attn_factor correction logic
* cont : make deepseek2 consistent
* cont : add TODO
* cont : special-case DSv2
* cont : revert Mistral 3 Large changes
* cont : fix DS2 to use the original attn_factor
* cont : minor comments
2025-12-12 17:12:40 +02:00
Georgi Gerganov and GitHub
c6f6e4f96a
ggml-alloc : fix reuse-parent logic for misaligned sizes ( #17884 )
2025-12-11 14:30:10 +02:00
Georgi Gerganov and GitHub
d9f8f60618
batch : fix sequence id ownership ( #17915 )
...
* batch : fix sequence id ownage
* cont : reduce allocations
2025-12-11 14:29:47 +02:00
Georgi Gerganov and GitHub
4dff236a52
ggml : remove GGML_KQ_MASK_PAD constant ( #17910 )
...
* ggml : remove GGML_KQ_MASK_PAD constant
* cont : remove comment
2025-12-10 20:53:16 +02:00
Georgi Gerganov and GitHub
6b82eb7883
metal : print node names for debugging ( #17882 )
2025-12-09 15:25:49 +02:00
Georgi Gerganov and GitHub
2bc96931d2
server : make cache_reuse configurable per request ( #17858 )
2025-12-08 12:43:12 +02:00
Georgi Gerganov and GitHub
8e5f4987b1
contrib : stale PRs ( #17803 )
2025-12-06 09:34:18 +02:00
Georgi Gerganov and GitHub
8ce774a102
metal : fix build( #17799 )
...
* metal : fix build
* tests : fix context destruction
2025-12-06 09:33:59 +02:00
Georgi Gerganov and GitHub
8160b38a5f
rpc : fix alloc size logic ( #17116 )
...
* rpc : fix alloc size logic
* rpc : bump version
2025-12-05 19:39:04 +02:00
Georgi Gerganov and GitHub
c41bde6fbd
metal : add residency sets keep-alive heartbeat ( #17766 )
...
* examples : add idle
* metal : attach residency sets to queue
* idle : add link
* idle : adjust intervals
* metal : add residency sets keep-alive heartbeat
* cont : adjust default keep-alive time
2025-12-05 19:38:54 +02:00
Georgi Gerganov and GitHub
0d1324856f
metal : use params per pipeline instance ( #17739 )
2025-12-04 10:34:11 +02:00
Georgi Gerganov and GitHub
a67ef0f47f
llama : fix sanity checks during quantization ( #17721 )
2025-12-04 10:33:42 +02:00
Georgi Gerganov and GitHub
190c4838bd
chat : reserve memory in compute_diffs and improve naming ( #17729 )
2025-12-03 17:22:10 +02:00
Georgi Gerganov and GitHub
3d94e967a1
metal : fix data race in pipeline library ( #17731 )
2025-12-03 14:03:40 +02:00
Georgi Gerganov and GitHub
649495c9d9
metal : add FA head size 48 ( #17619 )
2025-12-01 12:49:53 +02:00
Georgi Gerganov and GitHub
90c72a614a
ggml : extend the GGML_SCHED_NO_REALLOC debug logic of the scheduler ( #17617 )
2025-12-01 12:49:33 +02:00
Georgi Gerganov and GitHub
c386114922
arch : add description about LLM_TENSOR_INFOS ( #17550 )
2025-11-27 16:34:13 +02:00
Georgi Gerganov and GitHub
6783b11fb0
models : fix LFM2 tensors ( #17548 )
2025-11-27 16:04:29 +02:00
Georgi Gerganov and GitHub
583cb83416
ggml : add ggml_top_k ( #17365 )
...
* ggml : add ggml_top_k
* cont : add ggml_argsort_top_k
* metal : add top_k support
* ggml : cleanup
* tests : add virtual err() function for test_case
* ggml : add comments
2025-11-25 15:31:43 +02:00
Georgi Gerganov
2d50b9d8cb
sync : ggml
2025-11-24 15:26:31 +02:00
Georgi Gerganov
2286a360ff
sync : ggml
2025-11-20 14:10:44 +02:00
Georgi Gerganov and GitHub
196f5083ef
common : more accurate sampling timing ( #17382 )
...
* common : more accurate sampling timing
* eval-callback : minor fixes
* cont : add time_meas impl
* cont : fix log msg [no ci]
* cont : fix multiple definitions of time_meas
* llama-cli : exclude chat template init from time measurement
* cont : print percentage of unaccounted time
* cont : do not reset timings
2025-11-20 13:40:10 +02:00
Georgi Gerganov and GitHub
f40a2e5f11
gitignore : be more specific about ignored stuff ( #17354 )
2025-11-18 16:44:53 +02:00
Georgi Gerganov and GitHub
7aaeedc098
metal : support I32 -> I32 copy ( #17317 )
2025-11-17 11:52:00 +02:00
Georgi Gerganov and GitHub
3347e6d904
metal : faster argsort ( #17315 )
...
* metal : faster argsort
* cont : keep data in registers
2025-11-17 11:51:48 +02:00
Georgi Gerganov and GitHub
1a139644a8
metal : add cumsum ( #17305 )
2025-11-17 11:51:13 +02:00
Georgi Gerganov and GitHub
416e7c7f47
metal : remove obosolete asserts ( #17295 )
2025-11-16 09:50:26 +02:00
Georgi Gerganov and GitHub
5b2093becc
server : handle context overflow during decode ( #17267 )
...
* server : handle context overflow during decode
* server : minor refactor
2025-11-16 09:23:37 +02:00
Georgi Gerganov and GitHub
d396b43748
server : fix "can batch with" bug ( #17263 )
2025-11-14 14:03:45 +02:00
Georgi Gerganov and GitHub
45c6ef7307
metal : support argsort for ne00 > 1024 ( #17247 )
...
* metal : refactor argsort
* cont : sort chunks
* cont : merge sorted buckets
* cont : cleanup
2025-11-14 09:36:06 +02:00
Georgi Gerganov and GitHub
2606b0adab
metal : make the FA extra sizes consistent ( #17143 )
2025-11-14 09:13:34 +02:00